← Technique taxonomy

modality:ultrasound 18 pages 18 reviewed 0 imported

Ultrasound Silent Speech Interfaces

Ultrasound tongue imaging observes tongue movement rather than recording a spoken voice. What happens next depends on the task: recognizing words or phonemes, or reconstructing speech audio. Start with the output you need, then check whether the study actually tested silent articulation.

The list below includes every paper page that currently carries this technique label.

What it is
In tongue-imaging systems, an ultrasound probe under the chin captures changing tongue images. A learned model maps those images—sometimes together with lip video—to linguistic labels or acoustic features for speech synthesis.
Who it’s for
Readers choosing papers on image-based speech recognition, silent articulation-to-audio conversion, or adaptation when a probe is remounted or the user changes.
Verdict
A result on images recorded during voiced speech is not automatically a result on silent speech. Likewise, adapting to a new recording session does not by itself solve the difference between voiced and silent articulation.

From tongue images to text or speech

Recognition predicts linguistic units such as phonemes or words. Speech reconstruction instead predicts acoustic information used to generate a voice. A system can combine recognition and synthesis, but the output and evaluation still need to be stated: a word error rate and an acoustic reconstruction error do not measure the same outcome.

Lip video can supplement the tongue images. The recognition and silent-reconstruction studies linked below use both, so their results should not be described as evidence for ultrasound alone.

The full list below is generated from the database’s ultrasound technique label, not from a claim that every entry demonstrates silent speech communication. The selected reading routes also include a related ultrasound-and-video review; they do not change the underlying tags or imply that all ultrasound sensing uses tongue images.

Two different adaptation problems

Before reusing a result, check what changed between training and evaluation. These are separate questions, not interchangeable definitions of generalization:

Voiced training → silent use
Ribeiro and colleagues train on images captured during normal voiced speech and compare matched voiced and silent test conditions. Silent recognition performs worse; their work investigates this speaking-mode mismatch. Check whether the test articulation was actually silent, not merely whether microphone audio was excluded from model input.
Same user → another session or user
Tóth and colleagues address speaker and session adaptation, including changes after recording equipment is removed and remounted. Their spatial transformer adjusts input images. Check held-out sessions, held-out speakers, and how much adaptation data is available; this is not evidence that an unchanged model works for every user.
Choose evidence for the intended output
For text, inspect recognition errors and the vocabulary tested. For reconstructed audio, distinguish listening-based intelligibility, automatic-recognizer errors, and acoustic errors. Also check required lip video, recording setup, and whether live use was evaluated before treating a research system as a deployable interface.

Which paper answers your question?

Can a recognizer trained on voiced images read silent articulation?

Silent versus modal multi-speaker speech recognition from ultrasound and video compares the two speaking modes and investigates adaptation for recognition. This is a related review using both tongue ultrasound and lip video, not an ultrasound-only result.

Read the expert review · Read the original paper

Can silent tongue and lip movement be turned into speech audio?

Speech Reconstruction from Silent Tongue and Lip Articulation studies audio reconstruction from images recorded in silent mode. It uses pseudo targets and domain-adversarial training; read its evaluation as speech reconstruction, not direct text recognition.

Read the expert review · Read the original paper

What if the probe is removed and put back on?

Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks studies speaker/session adaptation through learned image transformations. It addresses recording alignment and adaptation, a different problem from changing the speaking mode.

Read the expert review · Read the original paper

Papers

reviewedarXiv / imported corpus page2021

Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces

Honarmandi Shandiz Amin, László Tóth, Gosztolya Gábor, Alexandra Markó, Csapó Tamás Gábor

The ultrasound-based x-vector speaker embedding is highly effective for speaker recognition, achieving under 1% error on unseen speakers, but its integration yields only a marginal improvement in multi-speaker ultrasound-to-speech synthesis accuracy.

reviewedarXiv / imported corpus page2021

Improving Neural Silent Speech Interface Models by Adversarial Training

Amin Honarmandi Shandiz, László Tóth, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó

A clean, well-executed incremental advance using GAN loss to modestly improve articulatory-to-acoustic mapping from ultrasound, validated objectively on two single-speaker corpora.

reviewedarXiv / imported corpus page2019

Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder

Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth, Alexandra Markó

The key advancement is continuous F0 tracking via CNNs yielding lower pitch error and slight naturalness improvement over discontinuous F0 pipelines in ultrasound SSI.

reviewedarXiv / imported corpus page2019

Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces

Gábor Gosztolya, Ádám Pintér, László Tóth, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó

The paper advances ultrasound silent speech interfaces by compressing ultrasound images using an autoencoder bottleneck prior to spectral parameter prediction, resulting in improved accuracy and more natural synthesized speech with smaller models.