News Story
New Speech Model Reads Mouth Movement Across Languages
Carol Espy-Wilson, Distinguished University Professor in the Department of Electrical and Computer Engineering and a member of the Institute for Systems Research (ISR), and PhD student Saba Tabatabaee have developed a system that can estimate how a person’s tongue, lips, and vocal tract move during speech using only the audio recording, without ever being trained on the language being spoken.
The Challenge
Understanding how the tongue, lips, and other articulators move during speech is valuable for research and clinical care, from studying accents to diagnosing speech disorders. Capturing that movement directly requires expensive, specialized equipment that is impractical for many patients, including children, and nearly impossible to use outside a lab. Researchers have long known that acoustic recordings of speech contain clues about articulator movement, but building a system that reliably decodes those clues has been difficult, and almost all prior systems were built and tested only in English and are speaker dependent.
What They Did
The team built a speech inversion system that listens to an audio clip and estimates seven oral tract movements, such as tongue, lip positions and velum position, plus three additional voice source characteristics, all from the sound alone. The system uses a pretrained speech model called WavLM combined with Conformer neural network layers to map acoustic patterns to physical articulator movement.
Figure 2: Waveforms and spectrograms of English, Russian and French utterances, with comparisons between ground-truth TBCD, TTCD, and LA, and the corresponding estimates from the SI system.
The key test came after training the system exclusively on English speech. The researchers then asked it to analyze French and Russian speech it had never encountered. The system estimated articulator movements with strong accuracy in both unseen languages, achieving average correlation scores of 0.83 for French and 0.74 for Russian compared to ground truth measurements. This suggests that much of what the system learns about how speech sounds map to physical movement holds true across languages. The team also extended the system to estimate velopharyngeal movement, the opening and closing that produces nasal sounds, scoring 0.89 for French and 0.82 for Russian, even though French and English follow fundamentally different rules for nasal sounds. The Russian nasal result is based on a single speaker, a small sample worth noting.
In Their Words
“The speaker-independent, cross-linguistically generalizable speech inversion system has the potential to transform the field by enabling entirely new directions in research and a broad range of real-world applications," explained Espy-Wilson. “For example, based on the articulatory information provided by the speech inversion system, my lab has developed speech biomarkers for brain, motor and speech-communication function.”
Why It Matters
The findings suggest that speech inversion systems, once thought to require language specific training data, may generalize much more broadly than assumed. That has real implications for studying speech disorders and articulation patterns in under resourced languages where large training datasets do not exist. The work is a collaboration between UMD, Yale, and the University of Cincinnati, and builds on Espy-Wilson’s lab’s ongoing research in speech inversion for clinical assessment applications.
Learn More
Published August 5, 2026