Skip to main navigation Skip to search Skip to main content

A comparative study of audio features for audio-to-visual conversion in MPEG-4 compliant facial animation

  • Lei Xie
  • , Zhi-Qiang Liu

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

Audio-to-visual conversion is the basic problem of speech-driven facial animation. Since the conversion problem is to predict facial control parameters from the acoustic speech, the informative representation of audio, i.e., the audio feature, is important to get a good prediction. This paper presents a performance comparison on prosodic features, articulatory features, and perceptual features for the audio-to-visual conversion problem on a common test bed. Experimental results show that the Mel frequency cepstral coefficients (MFCCs) produce the best performance, followed by the perceptual linear prediction coefficients (PLPC), the linear predictive cepstral coefficients (LPCCs), and the prosodie feature set (F0) and energy). The combination of three kinds of features can further improve the prediction performance on facial parameters. It unveils that different audio features carry complementary information relevant to facial animation. © 2006 IEEE.
Original languageEnglish
Title of host publicationProceedings of the 2006 International Conference on Machine Learning and Cybernetics
PublisherIEEE Computer Society
Pages4359-4364
Volume2006
ISBN (Print)1424400619, 9781424400614
DOIs
Publication statusPublished - 2006
Event5th International Conference on Machine Learning and Cybernetics, ICMLC 2006 - Dalian, China
Duration: 13 Aug 200616 Aug 2006

Publication series

NameProceedings of the 2006 International Conference on Machine Learning and Cybernetics
Volume2006

Conference

Conference5th International Conference on Machine Learning and Cybernetics, ICMLC 2006
PlaceChina
CityDalian
Period13/08/0616/08/06

Bibliographical note

Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].

Research Keywords

  • Audio features
  • Audio-to-visual conversion
  • Facial animation
  • MPEG-4
  • Talking face

Fingerprint

Dive into the research topics of 'A comparative study of audio features for audio-to-visual conversion in MPEG-4 compliant facial animation'. Together they form a unique fingerprint.

Cite this