Skip to main navigation Skip to search Skip to main content

Multi-stream articulator model with adaptive reliability measure for audio visual speech recognition

  • Lei Xie
  • , Zhi-Qiang Liu

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

We propose a multi-stream articulator model (MSAM) for audio visual speech recognition (AVSR). This model extends the articulator modelling technique recently used in audio-only speech recognition to audio-visual domain. A multiple-stream structure with a shared articulator layer is used in the model to mimic the speech production process. We also present an adaptive reliability measure (ARM) based on two local dispersion indicators, integrating audio and visual streams with local, temporal reliability. Experiments on the AVCONDIG database shows that our model can achieve comparable recognition performance with the multi-stream hidden Markov model (MSHMM) under various noisy conditions. With the help of the ARM, our model even performs the best at some testing SNRs. © Springer-Verlag Berlin Heidelberg 2006.
Original languageEnglish
Title of host publicationAdvances in Machine Learning and Cybernetics - 4th International Conference, ICMLC 2005, Revised Selected Papers
PublisherSpringer Verlag
Pages994-1004
Volume3930 LNAI
ISBN (Print)3540335846, 9783540335849
DOIs
Publication statusPublished - 2006
EventInternational Conference on Machine Learning and Cybernetics, ICMLC 2005 - Guangzhou, China
Duration: 18 Aug 200521 Aug 2005

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume3930 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

ConferenceInternational Conference on Machine Learning and Cybernetics, ICMLC 2005
PlaceChina
CityGuangzhou
Period18/08/0521/08/05

Bibliographical note

Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].

Fingerprint

Dive into the research topics of 'Multi-stream articulator model with adaptive reliability measure for audio visual speech recognition'. Together they form a unique fingerprint.

Cite this