← SSI archive

Comparison page 3 reviewed papers

SilentSpeller, SottoVoce, and NasoVoce compared

Start with the output you need: SilentSpeller turns unvoiced spelling into text; SottoVoce reconstructs audio for smart-speaker commands; NasoVoce enhances low-audibility speech captured at the nose. These are different tasks, not three interchangeable products.

This is a comparison of three selected research systems, not a survey of every silent speech sensing method. For the wider field, start with the SSI sensing map.

What goes in, and what comes out?

Compare the task before comparing results
Research systemUser action and sensor OutputImportant boundary
SilentSpeller Spell words without voicing; a custom dental retainer measures tongue–palate contact (electropalatography, or EPG).Recognized spelling becomes text.User-dependent silent spelling, not natural conversational speech or audio reconstruction.
SottoVoce Articulate without voicing; an under-jaw ultrasound probe images internal articulation.Reconstructed speech audio is sent to an unmodified smart speaker.Speaker-dependent, small-command proof of concept; continuous real-time, open-vocabulary use is not demonstrated.
NasoVoce Speak softly or whisper; microphone and vibration sensors sit in smart-glasses nose pads.Enhanced speech audio for downstream recognition and voice interaction.Low-audibility speech is not zero-acoustic-input silent spelling; fully streaming fusion remains a limitation in the review.

Choose a review by your goal

Enter text without voicing

Read the live text-entry study, correction interface, and seated-versus-walking evaluation. Check user enrollment and vocabulary conditions before transferring the result to another setting.

Read the SilentSpeller review · CHI 2022 paper

Control an existing smart speaker

Read how ultrasound becomes audio and how smart-speaker command success was tested. Check processing delay, probe placement, and adaptation to silent articulation.

Read the SottoVoce review · CHI 2019 paper

Capture low-audibility speech in noise

Read the microphone–vibration fusion results under noise and the speech-quality evaluation. Check audibility and deployment conditions rather than treating wearable placement as proof of unrestricted daily use.

Read the NasoVoce review · Author preprint

Which one is most accurate?

There is no shared accuracy ranking here. Text-entry errors, smart-speaker command success, and enhanced-speech quality answer different questions. A larger reported score in one paper does not establish that its interface is better for another paper’s task.

  • Match the task and output: text entry, command control, or speech enhancement.
  • Check the split: known users versus unseen users, sessions, words, or sentences are different tests.
  • Separate offline and live use: inspect delay, correction steps, and the complete interaction, not recognition alone.
  • Keep physical conditions visible: sensor fit, movement, noise, and whether acoustic speech is allowed.

Use the review rubric to examine the evidence, or browse the full review database for alternatives. None of this comparison establishes clinical suitability or that a research prototype is a purchasable product.

Detailed evidence from the reviews

This page uses only fields already stored in the review database. Missing fields are shown as "Not stated in review" instead of being filled in by guesswork.

Paper Evidence strength Sensing modality Evaluation setting Practicality Open questions
SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography
CHI '22 · 2022
Confidence high · 8 evidence records electropalatography · SmartPalate custom dental retainer with 124 capacitive electrodes sampled at 100 Hz, connected wired or wireless to processing device. Offline isolated word recognition with 10-fold cross validation; reserve testing on 100 unseen words; seated vs walking phrase recognition; live interactive text entry with push-to-talk interface and edit gestures. High for privacy-sensitive communication and hands-busy users; useful where speech is socially inappropriate and users can manage oral hardware. user_independence; comfort; social_acceptability; broader_symbol_input · The system targets discreet text entry, explicitly trading away naturalness of silent speech for reliability; not a conversational silent speech system. · Recognition confusions occur mainly for letters with similar palatograms, especially EE-sound letters (B/P, D/T/Z). Strong user-dependence; user-independent recognition remains poor.
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks
CHI '19 · 2019
Confidence high · 7 evidence records ultrasound · 3.5 MHz convex ultrasound probe attached under the jaw, with ultrasound images captured to display monitor and digitized video stored Quantitative smart speaker success rates, word error rates with Google speech-to-text, and qualitative user adaptation observations. medium as a systems design contribution and research direction; low-to-medium as a direct deployable interface in reported form real_time_interaction; open_vocabulary; speaker_independence; wearable_ultrasound · Prototype supports only a fixed small command vocabulary in speaker-dependent training; no demonstration of open vocabulary or continuous real-time interaction. · Speaker-dependent training; latency unsuitable for real-time use (2.61 s per utterance); differences in silent versus voiced articulation require user adaptation; bulky hardware; potential unknown safety issues with continuous ultrasound emission; small vocabulary size.
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
CHI '26 / arXiv · 2026
Confidence high · 4 evidence records acoustic; vibration; multimodal · MEMS microphone (Syntiant SPH0141LM4H-1) and MEMS vibration sensor (Syntiant V2S200D) integrated in smart glasses nose pads providing synchronized PDM output. Quantitative ASR accuracy (WER, CER) on held-out data, objective perceptual quality metrics (PESQ, STOI), MUSHRA subjective ratings with 50 evaluators, and qualitative in-the-wild recordings in four real-world environments. High for wearable AI voice agents by addressing sensor placement, noise robustness, perceptual quality, and practical evaluation in diverse contexts. continuous_streaming; adaptive_sensor_gating; longitudinal_wearability; physiological_variability · Targets low-audibility whispered speech, not fully silent speech without any acoustic leakage; assumes hand-covering mouth for privacy. · Fusion model not fully streaming; whispered vibration signals remain weak limiting enhancement quality; performance under extreme noise favors vibration sensor input only at very low SNR.

Source reviews

reviewedconfidence high

SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography

Naoki Kimura, Tan Gemicioglu, Jonathan Womack, Richard Li, Yuhui Zhao, Abdelkareem Bedri, Zixiong Su, Alex Olwal, Jun Rekimoto, Thad Starner

SilentSpeller is a strong, rigorously tested SSI system that reframes silent speech as silent spelling, enabling large vocabulary, live text entry, and walking robustness with in-mouth electropalatography sensors.