SilentSpeller, SottoVoce, and NasoVoce compared
Start with the output you need: SilentSpeller turns unvoiced spelling into text; SottoVoce reconstructs audio for smart-speaker commands; NasoVoce enhances low-audibility speech captured at the nose. These are different tasks, not three interchangeable products.
This is a comparison of three selected research systems, not a survey of every silent speech sensing method. For the wider field, start with the SSI sensing map.
What goes in, and what comes out?
| Research system | User action and sensor | Output | Important boundary |
|---|---|---|---|
| SilentSpeller | Spell words without voicing; a custom dental retainer measures tongue–palate contact (electropalatography, or EPG). | Recognized spelling becomes text. | User-dependent silent spelling, not natural conversational speech or audio reconstruction. |
| SottoVoce | Articulate without voicing; an under-jaw ultrasound probe images internal articulation. | Reconstructed speech audio is sent to an unmodified smart speaker. | Speaker-dependent, small-command proof of concept; continuous real-time, open-vocabulary use is not demonstrated. |
| NasoVoce | Speak softly or whisper; microphone and vibration sensors sit in smart-glasses nose pads. | Enhanced speech audio for downstream recognition and voice interaction. | Low-audibility speech is not zero-acoustic-input silent spelling; fully streaming fusion remains a limitation in the review. |
Choose a review by your goal
Enter text without voicing
Read the live text-entry study, correction interface, and seated-versus-walking evaluation. Check user enrollment and vocabulary conditions before transferring the result to another setting.
Control an existing smart speaker
Read how ultrasound becomes audio and how smart-speaker command success was tested. Check processing delay, probe placement, and adaptation to silent articulation.
Capture low-audibility speech in noise
Read the microphone–vibration fusion results under noise and the speech-quality evaluation. Check audibility and deployment conditions rather than treating wearable placement as proof of unrestricted daily use.
Which one is most accurate?
There is no shared accuracy ranking here. Text-entry errors, smart-speaker command success, and enhanced-speech quality answer different questions. A larger reported score in one paper does not establish that its interface is better for another paper’s task.
- Match the task and output: text entry, command control, or speech enhancement.
- Check the split: known users versus unseen users, sessions, words, or sentences are different tests.
- Separate offline and live use: inspect delay, correction steps, and the complete interaction, not recognition alone.
- Keep physical conditions visible: sensor fit, movement, noise, and whether acoustic speech is allowed.
Use the review rubric to examine the evidence, or browse the full review database for alternatives. None of this comparison establishes clinical suitability or that a research prototype is a purchasable product.
Detailed evidence from the reviews
This page uses only fields already stored in the review database. Missing fields are shown as "Not stated in review" instead of being filled in by guesswork.
| Paper | Evidence strength | Sensing modality | Evaluation setting | Practicality | Open questions |
|---|---|---|---|---|---|
| SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography | Confidence high · 8 evidence records | electropalatography · SmartPalate custom dental retainer with 124 capacitive electrodes sampled at 100 Hz, connected wired or wireless to processing device. | Offline isolated word recognition with 10-fold cross validation; reserve testing on 100 unseen words; seated vs walking phrase recognition; live interactive text entry with push-to-talk interface and edit gestures. | High for privacy-sensitive communication and hands-busy users; useful where speech is socially inappropriate and users can manage oral hardware. | user_independence; comfort; social_acceptability; broader_symbol_input · The system targets discreet text entry, explicitly trading away naturalness of silent speech for reliability; not a conversational silent speech system. · Recognition confusions occur mainly for letters with similar palatograms, especially EE-sound letters (B/P, D/T/Z). Strong user-dependence; user-independent recognition remains poor. |
| SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks | Confidence high · 7 evidence records | ultrasound · 3.5 MHz convex ultrasound probe attached under the jaw, with ultrasound images captured to display monitor and digitized video stored | Quantitative smart speaker success rates, word error rates with Google speech-to-text, and qualitative user adaptation observations. | medium as a systems design contribution and research direction; low-to-medium as a direct deployable interface in reported form | real_time_interaction; open_vocabulary; speaker_independence; wearable_ultrasound · Prototype supports only a fixed small command vocabulary in speaker-dependent training; no demonstration of open vocabulary or continuous real-time interaction. · Speaker-dependent training; latency unsuitable for real-time use (2.61 s per utterance); differences in silent versus voiced articulation require user adaptation; bulky hardware; potential unknown safety issues with continuous ultrasound emission; small vocabulary size. |
| NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction | Confidence high · 4 evidence records | acoustic; vibration; multimodal · MEMS microphone (Syntiant SPH0141LM4H-1) and MEMS vibration sensor (Syntiant V2S200D) integrated in smart glasses nose pads providing synchronized PDM output. | Quantitative ASR accuracy (WER, CER) on held-out data, objective perceptual quality metrics (PESQ, STOI), MUSHRA subjective ratings with 50 evaluators, and qualitative in-the-wild recordings in four real-world environments. | High for wearable AI voice agents by addressing sensor placement, noise robustness, perceptual quality, and practical evaluation in diverse contexts. | continuous_streaming; adaptive_sensor_gating; longitudinal_wearability; physiological_variability · Targets low-audibility whispered speech, not fully silent speech without any acoustic leakage; assumes hand-covering mouth for privacy. · Fusion model not fully streaming; whispered vibration signals remain weak limiting enhancement quality; performance under extreme noise favors vibration sensor input only at very low SNR. |
Source reviews
SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography
SilentSpeller is a strong, rigorously tested SSI system that reframes silent speech as silent spelling, enabling large vocabulary, live text entry, and walking robustness with in-mouth electropalatography sensors.
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks
A solid proof of concept that reconstructs speech audio from ultrasound for controlling unmodified smart speakers, showcasing important system design insight despite prototype limitations in latency, hardware bulk, and speaker dependency.
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
A strong deployment-focused speech interface leveraging a novel nose-pad dual-sensor configuration and multimodal fusion to enable robust low-audibility speech interaction with AI under noise, backed by extensive evaluation.