An unsupervised method for learning to track tongue position from an acoustic signal.

  • Hogden J
  • Rubin P
  • Saltzman E
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

A procedure for learning to recover the relative positions of the articulators from speech signals is demonstrated. The algorithm learns without supervision, that is, it does not require information about which articulator configurations created the acoustic signals in the training set. The procedure consists of vector quantizing short time windows of a speech signal, then using multidimensional scaling to represent quantization codes that were temporally close in the encoded speech signal by nearby points in a continuitymap. Since temporally close sounds must have been produced by similar articulator configurations, sounds which were produced by similar articulator positions should be represented close to each other in the continuity map. Using an articulatory speech synthesizer to produce acoustic signals from known articulator positions, relative articulator positions were estimated from synthesized acoustic signals and compared to the synthesizer’s actual articulator positions. High rank-order correlations, ranging from 0.92 to 0.99, were found between the estimated and actual articulator positions. Reasonable estimates of relative articulator positions were made using 32 categories of sound, and the accuracy improved when more sound categories were used.

Cite

CITATION STYLE

APA

Hogden, J., Rubin, P., & Saltzman, E. (1992). An unsupervised method for learning to track tongue position from an acoustic signal. The Journal of the Acoustical Society of America, 91(4_Supplement), 2443–2443. https://doi.org/10.1121/1.403129

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free