Abstract
The paper overviews recent progress and challenges in a number of audiovisual speech processing technologies with main emphasis on the problem of automatic speech recognition. It is well known that visual channel information can improve automatic speech processing for human-computer interaction. To automatically process and incorporate such information into automatic systems, a number of steps are required that are surprisingly similar accross speech technologies. Crucial above all is the issue of feature representation of visual speech and its robust extraction. In addition, appropriate integration of the audio and visual representations is required, in order to ensure improved performance of the bimodal systems over audio-only baselines. These topics are discussed in detail in the talk, with main emphasis on their application to the speech recognition problem in the challenging environments of automobiles and smart rooms.
Cite
CITATION STYLE
Potamianos, G. (2008). Audiovisual automatic speech recognition: Progress and challenges. The Journal of the Acoustical Society of America, 123(5_Supplement), 3939–3939. https://doi.org/10.1121/1.2936018
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.