Ultrasound Silent Speech Interfaces
Ultrasound tongue imaging observes tongue movement rather than recording a spoken voice. What happens next depends on the task: recognizing words or phonemes, or reconstructing speech audio. Start with the output you need, then check whether the study actually tested silent articulation.
The list below includes every paper page that currently carries this technique label.
- What it is
- In tongue-imaging systems, an ultrasound probe under the chin captures changing tongue images. A learned model maps those images—sometimes together with lip video—to linguistic labels or acoustic features for speech synthesis.
- Who it’s for
- Readers choosing papers on image-based speech recognition, silent articulation-to-audio conversion, or adaptation when a probe is remounted or the user changes.
- Verdict
- A result on images recorded during voiced speech is not automatically a result on silent speech. Likewise, adapting to a new recording session does not by itself solve the difference between voiced and silent articulation.
From tongue images to text or speech
Recognition predicts linguistic units such as phonemes or words. Speech reconstruction instead predicts acoustic information used to generate a voice. A system can combine recognition and synthesis, but the output and evaluation still need to be stated: a word error rate and an acoustic reconstruction error do not measure the same outcome.
Lip video can supplement the tongue images. The recognition and silent-reconstruction studies linked below use both, so their results should not be described as evidence for ultrasound alone.
The full list below is generated from the database’s ultrasound technique label, not from a claim that every entry demonstrates silent speech communication. The selected reading routes also include a related ultrasound-and-video review; they do not change the underlying tags or imply that all ultrasound sensing uses tongue images.
Two different adaptation problems
Before reusing a result, check what changed between training and evaluation. These are separate questions, not interchangeable definitions of generalization:
- Voiced training → silent use
- Ribeiro and colleagues train on images captured during normal voiced speech and compare matched voiced and silent test conditions. Silent recognition performs worse; their work investigates this speaking-mode mismatch. Check whether the test articulation was actually silent, not merely whether microphone audio was excluded from model input.
- Same user → another session or user
- Tóth and colleagues address speaker and session adaptation, including changes after recording equipment is removed and remounted. Their spatial transformer adjusts input images. Check held-out sessions, held-out speakers, and how much adaptation data is available; this is not evidence that an unchanged model works for every user.
- Choose evidence for the intended output
- For text, inspect recognition errors and the vocabulary tested. For reconstructed audio, distinguish listening-based intelligibility, automatic-recognizer errors, and acoustic errors. Also check required lip video, recording setup, and whether live use was evaluated before treating a research system as a deployable interface.
Which paper answers your question?
Can a recognizer trained on voiced images read silent articulation?
Silent versus modal multi-speaker speech recognition from ultrasound and video compares the two speaking modes and investigates adaptation for recognition. This is a related review using both tongue ultrasound and lip video, not an ultrasound-only result.
Can silent tongue and lip movement be turned into speech audio?
Speech Reconstruction from Silent Tongue and Lip Articulation studies audio reconstruction from images recorded in silent mode. It uses pseudo targets and domain-adversarial training; read its evaluation as speech reconstruction, not direct text recognition.
What if the probe is removed and put back on?
Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks studies speaker/session adaptation through learned image transformations. It addresses recording alignment and adaptation, a different problem from changing the speaking mode.
Papers
SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces
A useful single-speaker open-vocabulary eyewear dataset: mixed voiced/silent training reaches 26.3% silent WER, but population and everyday-environment generalization remain untested.
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
HEar-ID jointly models ear-based spelling and identity, but all-user Top-1 is 67.3%, not the selected eight-user 90.25%; whisper dependence, participant failures, modified hardware, and untested attack resistance limit deployment claims.
Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks
Strong full-text-backed evidence that most of the gain comes from fast input alignment, not from inventing a new SSI stack.
Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training
Strong SSI paper improving silent speech reconstruction by generating pseudo acoustic targets and using domain adversarial training to address domain mismatch; validated with TaL dataset showing substantial WER and MOS gains over TaLNet.
Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks
An empirically supported, incremental advancement showing that hybrid 3D-CNN plus ConvLSTM models modestly outperform prior ultrasound tongue video SSI architectures in mel-spectrogram regression accuracy and model efficiency on single-speaker data.
Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input
Helpful side information, not standalone SSI.
Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces
The ultrasound-based x-vector speaker embedding is highly effective for speaker recognition, achieving under 1% error on unseen speakers, but its integration yields only a marginal improvement in multi-speaker ultrasound-to-speech synthesis accuracy.
Voice Activity Detection for Ultrasound-based Silent Speech Interfaces using Convolutional Neural Networks
Preprocessing paper, narrow but legitimate.
Improving Neural Silent Speech Interface Models by Adversarial Training
A clean, well-executed incremental advance using GAN loss to modestly improve articulatory-to-acoustic mapping from ultrasound, validated objectively on two single-speaker corpora.
3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces
Temporal context helps, but the evidence is a single-speaker vocoder-parameter study.
Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image
Real signal, wrong target for SSI.
Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images
Strong ultrasound SSI paper with unusually clear quantitative gains.
Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder
The key advancement is continuous F0 tracking via CNNs yielding lower pitch error and slight naturalness improvement over discontinuous F0 pipelines in ultrasound SSI.
Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces
The paper advances ultrasound silent speech interfaces by compressing ultrasound images using an autoencoder bottleneck prior to spectral parameter prediction, resulting in improved accuracy and more natural synthesized speech with smaller models.
Denoising convolutional autoencoder based B-mode ultrasound tongue image feature extraction
DCAE provides cleaner, more robust ultrasound tongue features leading to improved silent speech recognition, outperforming prior feature extraction strategies.
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks
A solid proof of concept that reconstructs speech audio from ultrasound for controlling unmodified smart speakers, showcasing important system design insight despite prototype limitations in latency, hardware bulk, and speaker dependency.
Updating the silent speech challenge benchmark with deep learning
Benchmark update with a real, reproducible WER gain.
Contour-based 3d tongue motion visualization using ultrasound image sequences
Useful tongue-modeling tool, not a recognizer.