← SSI archive · Review rubric

2022 · CHI '22 · Kimura expert review · confidence high

SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography

Naoki Kimura, Tan Gemicioglu, Jonathan Womack, Richard Li, Yuhui Zhao, Abdelkareem Bedri, Zixiong Su, Alex Olwal, Jun Rekimoto, Thad Starner

BibTeX
@misc{silentspeller,
  title = {SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography},
  author = {Naoki Kimura and Tan Gemicioglu and Jonathan Womack and Richard Li and Yuhui Zhao and Abdelkareem Bedri and Zixiong Su and Alex Olwal and Jun Rekimoto and Thad Starner},
  year = {2022},
  note = {CHI '22},
  doi = {10.1145/3491102.3502015},
  url = {https://doi.org/10.1145/3491102.3502015},
}

SilentSpeller is a strong, rigorously tested SSI system that reframes silent speech as silent spelling, enabling large vocabulary, live text entry, and walking robustness with in-mouth electropalatography sensors.

Verdict: full-text draftPriority: highConfidence: highBasis: full text + existing expert seedCoverage: high

Expert review (33-view rubric)

33 view results across 4 families. Reviewed 2026-06-19T10:31:51.764677+00:00. Reviewer: ficr-coverage-serializer.

Conclusion

SilentSpeller is best placed as D02: a near-term, narrow-output silent text-entry route, not a speech-restoration or brain-signal route. Within the mobile silent text/command-entry route, it ranks first on communication practicality because it demonstrates live compositional text entry at 37 wpm and 87% accuracy, while peers are command/phrase recognition or speech-playback systems with stronger output limits.

Why it matters
The focal paper outputs text by silent spelling with an in-mouth electropalatography retainer; it is evaluated as text entry on MacKenzie-Soukoreff phrases, with live WPM/accuracy, dictionary limits, unseen-word tests, and walking/seated robustness. It does not restore natural speech, decode brain signals, or merely amplify low-volume speech.
route: narrow_taskroute_id: D02input: electropalatographyoutput: textcoverage: 29 applicable / 4 N/A / 0 not-evaluated of 33
provisional narrow_task route provisional comparison scope: route

Route ranking not yet recomputed across all papers.

A: Initial Reading Views 12

12 views in the A_initial_views family.

完全無音の前に来る入力

完全無音だけでなく、囁き、低音量、唇コマンド、スマホ音響、鼻パッド、EPG spellingを実用候補として読めるか。

SilentSpeller is best read as fully unvoiced EPG spelling for text entry, not as whisper, low-volume speech, nasal-pad speech, or smartphone acoustic input. Its practical move is to put a closed-dictionary spelling layer before harder open silent-speech recognition: 1164 isolated words offline and 321 live phrase words, with 37 wpm at 87% live text-entry accuracy. The cost is a custom dental retainer and user-dependent enrollment.

Facts
  • Abstract: the system uses a dental retainer with capacitive touch sensors to track tongue movement and lets users type by spelling words without voicing.
  • §1: silent spelling means the user spells words without voicing instead of saying them audibly.
  • Figure 1 and §3.2: SmartPalate has 124 palate electrodes, samples at 100 Hz, and requires a custom dental impression.
  • §7.2: live input is silent spelling word by word, followed by 5-best word candidates and palate gestures for selection or erasure.
Inferences
  • The focal paper is a strong A01 positive case for EPG spelling as an input mode before general silent speech.
  • It should not be used as evidence that whisper or low-audibility speech is silent; it avoids sound by changing the task to spelling with tongue-palate contact.
  • Its ecosystem link is mobile text entry and smartphone keyboard practice, not voice restoration or speech enhancement.
Unknowns
  • The paper does not measure ambient acoustic leakage from silent spelling.
  • The paper does not compare against whispered, low-volume, nasal-pad, or smartphone acoustic input in the same experiment.
  • The focal prototype is not yet a fully everyday in-mouth consumer device.
Warnings
  • Do not discard it because it is not whole-word silent speech; the spelling redesign is the main contribution.
  • Do not treat any whisper or low-volume neighboring result as evidence for SilentSpeller's fully unvoiced EPG input.
Evidence
  • ev_silentspeller_a01 — View A01 (完全無音の前に来る入力) coverage status=applicable; locator: Abstract; §1; §3; §3.2; §7.2; Figure 1; Figure 12; Table 6 | traced to original papers -> rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce is a nose-bridge microphone plus vibration-sensor interface for normal and whispered low-audibility speech, evaluated with Whisper Large-v2, PESQ, STOI, and MUSHRA; it is adjacent low-audibility speech, not fully silent EPG spelling. ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: Smartphone acoustic sensing uses inaudible acoustic signals with a limited 54-sentence vocabulary and reports 8.4% WER speaker/environment independent and 8.1% WER unseen-sentence testing. ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner is smartphone lipreading command input, with 30 custom commands and 81.7% one-shot, 96.5% three-shot, 98.8% five-shot accuracy.

confidence high (0.90)

狭い課題を実用進歩として読む

開語彙会話、文章、コマンド、keyword spotting、spelling、固定句のどれを狙い、狭めたことが実用性にどう効くか。

The paper explicitly narrows open silent speech into spelling-based text entry. This is a real practical advance for text entry, not evidence of open-vocabulary conversation recovery. Offline performance is 97% character accuracy on a 1164-word dictionary after 2328 word examples from 2 users; live performance is 37 wpm at 87% accuracy on 107 MacKenzie phrases with 321 unique words after about 1-2 hours of participant training for the live setup.

Facts
  • Table 1: Experiment 1 uses 2328 collected words, a 1164-word dictionary, and 10-fold cross-validation.
  • Table 3: HMM results are 97% character accuracy for both P1 and P2, with 93% and 91% word accuracy.
  • §5 and Table 4: 100 words, 200 examples, are removed from training; results are P1 93% character/84% word and P2 96% character/87% word, averaging 94.5% character and 85.5% word.
  • §6.1: 107 phrases contain 556 words and 321 unique words.
  • §7.5 and Table 6: live text entry averages 37 wpm and 87% accuracy.
Inferences
  • The intended task is composition from a bounded text-entry dictionary, not free conversation or speech reconstruction.
  • The high offline accuracy depends on user-dependent training and a dictionary/language-model setting.
  • The unseen-word test is important because spelling/triletter modeling generalizes beyond trained word exemplars, but it is still within the same 1164-word dictionary framework.
Unknowns
  • Open-vocabulary OOV entry is not solved.
  • No punctuation, capitalization, emoji, or long-form composition is evaluated.
  • No direct comparison to a live whole-word SSI text-entry system is available in the focal paper.
Warnings
  • Do not read 97% isolated-word accuracy or 100% strict-grammar phrase accuracy as free-conversation recognition.
  • Do not equate 30-command or 54-sentence neighboring results with mobile text-entry progress.
Evidence
  • ev_silentspeller_a02 — View A02 (狭い課題を実用進歩として読む) coverage status=applicable; locator: Table 1; Table 3; §5; Table 4; §6.1; §7.5; Table 6 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner's narrow task is customizable silent command recognition on smartphones: Table 1 reports 30 custom commands and 81.7%/96.5%/98.8% accuracy for 1/3/5 shots. ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: Acoustic-sensing SSI is a limited-vocabulary sentence task: 54 sentences, 13,500 samples, 2.6% domain-dependent WER, 8.4% domain-independent WER, and 8.1% unseen-sentence WER. ; kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoce maps ultrasound images to generated audio for smart-speaker commands; about 500 speech commands were collected per collaborator and WERs of 20.61%, 41.03%, and 33.56% are reported for generated command audio.

confidence high (0.90)

話者依存

speaker-dependent、seen speaker、unseen speaker、cross-subject、cross-session、remountをどう扱っているか。

The focal results are speaker/user-dependent. The strongest numbers are within-user: P1/P2 train and test separately with 10-fold CV, and live participants use personal training data. The paper includes a weak cross-user exploration only in future-work/limitations: leave-one-user-out on first 500 words from five participants gives 55% character and 36% word accuracy, far below the user-dependent results.

Facts
  • §4.3: the authors state they create user-dependent recognizers where only one participant's data is used for training and testing at a time.
  • §4.1: P1 and P2 are both male, ages 25 and 50, and each provides 2328 isolated words.
  • §7.1: the live experiment has seven participants; P3-P7 are male, ages 23 to 45, and P2/P5 are native English speakers.
  • §7.1: because the system is user-dependent, ideally all users would collect 2328 examples and 107 phrases, but this was too much time for volunteers.
  • §8.1.3: leave-one-user-out cross-validation on the first 500 words from five participants averages 55% character and 36% word accuracy.
Inferences
  • The paper should be cited as a user-dependent SSI/text-entry result, not as immediate first-use cross-subject performance.
  • The live study is broader than the two-person tuning study but still relies on per-user enrollment.
  • The cross-user result is valuable as a boundary: it shows current user-independent SilentSpeller is not ready.
Unknowns
  • No robust unseen-speaker model is demonstrated.
  • No remount-specific train/test split is reported.
  • The paper does not isolate cross-session versus cross-subject versus hardware-fit effects.
Warnings
  • Do not project P1/P2's 97% character accuracy to a new first-time user.
  • Do not hide the 55% character/36% word user-independent result.
Evidence
  • ev_silentspeller_a03 — View A03 (話者依存) coverage status=applicable; locator: §4.1; §4.3; §7.1; §8.1.3; Table 3; Table 6 | traced to original papers -> arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: Neighbor acoustic-sensing paper explicitly reports speaker- and environment-independent limited-vocabulary testing with 8.4% WER, a different generalization claim from SilentSpeller. ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner uses few-shot transfer to unseen users/commands, with 11 model-test participants and a 16-participant live user study; Table 1 reports 1/3/5-shot command accuracies.

confidence high (0.90)

発話モードの違い

silent、whispered、normal、vocalized、mimed、post-laryngectomyを別モードとして測っているか。

The focal paper tests silent spelling, not a mode-transfer stack across normal, whispered, vocalized, mimed, or post-laryngectomy speech. The Wizard-of-Oz pilot compares ideal silent speech, silent spelling, and mini-QWERTY for speed/workload, but the built recognizer is trained and evaluated on silent spelling only. Therefore it gives no focal silent-vs-modal transfer number.

Facts
  • §3: silent spelling is contrasted with saying the word; the user spells letters one by one instead of silently mouthing the word.
  • §3.1: the Wizard-of-Oz pilot compares silent speech input, silent spelling input, and mini-QWERTY.
  • §3.1.2: ideal silent speech is fastest at 115 wpm, while silent spelling and mini-QWERTY average 38.7 and 36.6 wpm.
  • §4-§7: recognition experiments are on unvoiced spelling, walking/seated spelling, and live spelling text entry.
  • §7.6 and §8: authors suggest hybrid common spoken words as future work, but do not evaluate it.
Inferences
  • This is not a normal-to-silent transfer paper.
  • The focal contribution avoids the silent-vs-modal gap by using spelled letters and user-dependent training, rather than transferring vocalized speech models.
  • Any claim that normal/whispered data transfers to SilentSpeller would be unsupported by the focal paper.
Unknowns
  • No normal speech, whispered speech, vocalized speech, mimed speech, or post-laryngectomy condition is measured with SmartPalate.
  • No adaptation method across speech modes is tested.
  • No silent-vs-whispered or silent-vs-modal gap is reported for EPG spelling.
Warnings
  • Do not assume voiced/normal-speech training would transfer to SilentSpeller's silent spelling.
  • Do not treat the Wizard-of-Oz ideal silent speech condition as recognizer evidence.
Evidence
  • ev_silentspeller_a04 — View A04 (発話モードの違い) coverage status=applicable; locator: §3; §3.1; §3.1.2; §4; §7.6; §8 | traced to original papers -> arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: Visual-only neighbor directly tests normal/whispered/silent modes; phrase results include normal matched 69.7, normal-to-silent 61.2, and silent matched 64.4 classification rate, with silent generally hardest. ; arxiv_2103.00333_core_silent-versus-modal-multi-speaker-speech-recognition-from-ultrasound-and-video.txt: Ultrasound/video neighbor reports a strong modal-to-silent gap: TaL80 multi-speaker raw WER 39.34 on modal versus 77.79 on silent; TaL1 also includes whispered speech. ; arxiv_2305.14203_core_improving-the-gap-in-visual-speech-recognition-between-normal-and-silent-speech-based-on-m.txt: Metric-learning VSR neighbor explicitly targets normal/silent gap; baseline silent WER 12.91 improves to 11.06 with LKL, while normal WER is 10.21 baseline and 8.12 proposed.

confidence high (0.90)

センサー配置

電極位置、プローブ固定、カメラ角度、鼻パッド、レーダーアンテナ、歯科装置、再装着が性能を決めていないか。

Sensor placement is central. Performance is tied to a custom-fit palate retainer with 124 binary electrodes at 100 Hz. The mouth fit likely helps walking robustness, but it is also the main daily-wear constraint: dental impression, external ribbon/USB in the evaluated system, salivation/comfort issues, and possible fit/calibration differences such as P5's poor result.

Facts
  • Figure 1: SmartPalate retainer has 124 electrodes and senses tongue position at 100 Hz.
  • §3.2: SmartPalate is a dental retainer-type device with 124 binary capacitive sensors lining the palate.
  • §3.2: data goes through a flex ribbon cable to an external module, then USB to a computer or smartphone.
  • §3.2: each user needs a dental impression for custom fit.
  • §6: the electrode array fits snugly in the mouth, which the authors use to explain walking robustness.
  • §8: P5's poor performance may be due to fit, mouth shape, or subtle electrode miscalibration.
Inferences
  • This is not a sensor-agnostic algorithmic result; the fixed palate geometry is part of the result.
  • Walking robustness is credible for this mounting method, but should not be generalized to loose or remounted oral sensors.
  • Daily wearability remains an open hardware problem despite wireless prototypes.
Unknowns
  • No formal remount or day-to-day don/doff test is reported.
  • No quantitative calibration sensitivity test is reported.
  • The current system's final battery life, sterilization, durability, and long-term comfort are not evaluated.
Warnings
  • Do not assume the lab/custom retainer placement can be reproduced casually in everyday use.
  • Do not separate the accuracy numbers from the dental-fit and enrollment burden.
Evidence
  • ev_silentspeller_a05 — View A05 (センサー配置) coverage status=applicable; locator: Figure 1; Figure 2; §3.2; §6; §7.6; §8; §8.1.1 | traced to original papers -> arxiv_2312.09572_core_ir-uwb-radar-based-contactless-silent-speech-recognition-of-vowels-consonants-words-and-ph.txt: IR-UWB radar neighbor shows placement sensitivity explicitly: upper radar 5-10 cm from lips and lower radar 10-15 cm from chin; Table 2 reports upper FERASEC+DNN-HMM 86.47/81.59/88.95/96.88 for vowels/consonants/words/phrases versus lower 70.59/63.57/81.10/94.27. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce neighbor mounts a microphone and MEMS vibration sensor on smart-glasses nose pads; this verifies a contrasting daily-wear placement strategy for low-audibility speech.

confidence high (0.90)

借りた音声知識

HuBERT、AV-HuBERT、wav2vec2、speech units、lip-reader、ASR guidance、TTS/vocoderなど外部知識をどれだけ使っているか。

SilentSpeller uses borrowed classical ASR structure, not large pretrained speech priors. There is no HuBERT, AV-HuBERT, wav2vec2, ASR guidance, TTS, or vocoder in the focal system. The borrowed knowledge is PCA feature reduction, HMM/HTK training, phone/triphone-like letter/triletter modeling, Viterbi/Baum-Welch alignment, dictionaries, and n-gram language models.

Facts
  • §3.3: PCA is fit on training data only, and top 16 components are used.
  • §3.3: GT2K, a wrapper around HTK, is used for HMM training/testing.
  • §3.3: 26 letters are trained first, then triletters, analogous to phones and triphones.
  • §6.2: recognition uses a dictionary of 321 unique words from 107 phrases and a bigram with Laplace smoothing.
  • Table 3: Transformer baselines perform poorly, 37% character/9.1% word for P1 and 34% character/8.8% word for P2.
Inferences
  • The input signal, not a massive speech prior, carries most of the focal claim.
  • The dictionary and bigram contribute substantially in phrase/live settings, so accuracy is not pure sensor decoding.
  • The paper's negative Transformer result supports the choice of low-data user-dependent HMMs.
Unknowns
  • No no-language-model ablation is reported for live text entry.
  • No explicit input-versus-bigram contribution split is reported.
  • No comparison to modern self-supervised speech/audio-visual models is included.
Warnings
  • Do not confuse language-model constrained word choice with unconstrained reading of tongue motion.
  • Do not treat generated/natural speech in neighboring studies as evidence for SilentSpeller input decoding.
Evidence
  • ev_silentspeller_a06 — View A06 (借りた音声知識) coverage status=applicable; locator: §3.3; §6.2; Table 3; Table 5; Figure 9 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner contrasts with focal by using LRW pretraining and a frozen encoder plus logistic-regression classifier for few-shot command learning. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce uses OpenAI Whisper Large-v2 for evaluation and a Whisper-based distillation/training objective, a type of external speech prior absent from SilentSpeller.

confidence high (0.90)

出力タイプの分離

出力は文字、音声、コマンド、spelling、音声強調、同期TTSのどれか。

The output is text entry, not speech audio. SilentSpeller decodes each silent-spelled word into word candidates and edits the transcript. Its communication metric is text-entry speed and accuracy, so it should not be grouped with speech-generation or enhancement papers whose output quality is PESQ/STOI/MOS/MUSHRA.

Facts
  • Title and abstract: the goal is silent speech text entry.
  • §7.2: after input, the interface displays the five best word predictions; the user accepts, selects, or erases.
  • §7.4: metrics are WPM and Total Error Rate adapted to word-level recognition and edit gestures.
  • Table 6: reports WPM and accuracy for SilentSpeller and mini-QWERTY.
  • §8.1.4: punctuation, capitalization, and emoji are not included.
Inferences
  • The user goal is composing text, not restoring a voice or improving audio quality.
  • The proper comparison family is text-entry systems and silent text interfaces.
  • Speech naturalness metrics are not relevant to this focal paper's success.
Unknowns
  • No audio output or voice quality is evaluated.
  • No end-to-end message comprehension study is reported beyond text-entry accuracy.
  • No punctuation/capitalization/emoji output is supported in the study.
Warnings
  • Do not treat speech-synthesis naturalness as comparable to SilentSpeller's text-entry accuracy.
  • Do not infer patient voice restoration from text output alone.
Evidence
  • ev_silentspeller_a07 — View A07 (出力タイプの分離) coverage status=applicable; locator: Title; Abstract; §7.2; §7.4; Table 6; §8.1.4 | traced to original papers -> kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoce outputs generated audio to control smart speakers and reports generated-command WERs, contrasting with SilentSpeller's text output. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce outputs/enhances speech signals and evaluates WER/CER plus PESQ/STOI/MUSHRA, contrasting with text-entry metrics.

confidence high (0.90)

指標の目的

WER/CER/PER、PESQ/STOI/ESTOI/MCD/MOS、成功率、latency、RTF、wpm、EERは何を測っているか。

The primary focal metrics measure text-entry utility, not perceptual audio quality. Offline character/word accuracy measures dictionary recognition; live accuracy is 1-TER and includes user errors, recognizer errors, and corrections. WPM excludes recognizer time, so the 37 wpm number should be read with that timing caveat. The user cost beside the numbers is custom hardware plus roughly 1-2 hours of live-study enrollment, or about 5 hours for the full 2328-word P1/P2 data sets.

Facts
  • Table 3: 97% character with 93%/91% word accuracy for HMMs; Transformers are 37%/34% character and 9.1%/8.8% word.
  • Table 4: held-out unseen-word test gives 93%/84% for P1 and 96%/87% for P2.
  • Table 5: bigram walking/seated results are P1 99%/98% seated, 99%/97% walking; P2 94%/93% seated, 96%/93% walking.
  • §7.4: WPM uses MacKenzie's formula, and recognition time is removed from overall time.
  • §7.5: live average is 37 wpm and 87% accuracy; this accuracy includes user typing failures, recognizer failures, and corrections.
Inferences
  • Offline accuracy is optimistic for live use because it does not include correction behavior and flow interruptions.
  • TER-derived live accuracy is the most user-facing metric in the paper.
  • The WPM timing convention is defensible for measuring spelling speed, but it should not be reported as full system latency-inclusive throughput.
Unknowns
  • No sentence-level semantic success metric is reported.
  • No long-term fatigue-adjusted WPM is reported.
  • No human comprehension or downstream task success study is reported.
Warnings
  • Do not map character accuracy directly to conversation speed or misunderstanding rate.
  • Do not ignore that WPM excludes recognizer latency.
Evidence
  • ev_silentspeller_a08 — View A08 (指標の目的) coverage status=applicable; locator: Table 3; Table 4; Table 5; §7.4; §7.5; Table 6 | traced to original papers -> arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: Neighbor acoustic paper uses WER as primary metric and reports 2.6%, 8.4%, and 8.1% WER for domain-dependent, domain-independent, and unseen-sentence settings. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: Neighbor NasoVoce evaluates ASR with WER/CER and audio quality/intelligibility with PESQ, STOI, and MUSHRA, confirming those metric families are not SilentSpeller's metrics.

confidence high (0.90)

運用条件

リアルタイム、モバイル、装着負担、登録データ量、再学習、ユーザー修正、騒音・照明・動きが現実的か。

Operationally, SilentSpeller is closer to a live prototype than most offline SSI work, but not yet an everyday mobile product. It has a live interactive UI, walking robustness, 5-best correction, and 200 ms optimized word recognition, yet the study system uses a push-to-record key, custom dental retainer, external ribbon/USB hardware, per-user training, and WPM excludes recognition time.

Facts
  • §7.2: users press a push-to-record button, silently spell a word, release, and see five candidates in about a second.
  • §7.2: TAP and STICK gestures select n-best candidates or erase a word.
  • §7.4: recognition time is removed from WPM timing.
  • §7.5: optimized average word recognition time is 200 ms while spelling a word takes about 1 second.
  • §6: walking phrase input is tested indoors at home; Table 5 shows little seated/walking degradation.
  • §8.1.1: USB tether limits mobility; BLE dongle and in-mouth prototype are future/parallel hardware work.
Inferences
  • The paper deserves operational credit for live text entry and walking tests.
  • The operational ceiling is constrained by enrollment, retainer fit, tethering, and correction ergonomics.
  • The live system is not fully hands-free because recording/phrase advancement still uses keyboard buttons.
Unknowns
  • No outdoor/noise/lighting dependency is relevant to EPG and not tested except walking indoors.
  • No long-term daily wear or field deployment is reported.
  • No fully on-device mobile implementation is demonstrated for the focal live recognizer.
Warnings
  • Do not judge practical readiness from benchmark accuracy alone.
  • Do not omit enrollment time, custom fitting, and latency/timing conventions when reporting performance.
Evidence
  • ev_silentspeller_a09 — View A09 (運用条件) coverage status=applicable; locator: §3.2; §4.2; §6; Table 5; §7.2; §7.4; §7.5; §8.1.1 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner demonstrates a different operational profile: commodity iPhone, on-device recognition/fine-tuning, about 422 ms recognition pipeline, and 2217 ms on-device fine-tuning for 30 commands x 5 shots. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce runs as a wearable nose-pad sensor concept and includes in-the-wild trials in a cafe, roadside, walking outdoors, and train car, a field-condition style not present in SilentSpeller.

confidence high (0.90)

患者価値の証拠

喉頭摘出、失声、麻痺、aphasia、臨床ユーザー、長期使用で直接試しているか。

The paper has a clear patient-value motivation but no direct clinical evidence. It motivates low manual dexterity and severe dysphonia use cases, including a muscular-dystrophy wheelchair-user scenario, ALS, cerebral palsy, stroke, MS, Parkinson's, tremor, and arthritis. However, all evaluated participants are healthy volunteers; no patient, long-term clinical user, or restoration metric is reported.

Facts
  • §1: the motivating scenario is a wheelchair user with muscular dystrophy who can no longer type emails/documents and cannot rely on speech in an open office.
  • §1: the paper lists ALS, cerebral palsy, stroke, MS, Parkinson's disease, essential tremor, and arthritis as conditions that can limit manual dexterity.
  • §6.3: for people with low dexterity and severe dysphonia, a limited phrase set with strict grammar could be useful.
  • §6.3: strict grammar matching to one of 500 MacKenzie phrases increases accuracy to 100% for all four walking/seated conditions.
  • §8: future work should directly compare against current methods for participants with MS, Parkinson's, cerebral palsy, and muscular dystrophy.
Inferences
  • Patient value is plausible for quiet hands-free text entry, but not clinically proven.
  • The strongest direct evidence is healthy-user text-entry speed plus a constrained phrase-control result.
  • The patient scenario is design motivation, not patient trial evidence.
Unknowns
  • No actual users with muscular dystrophy, ALS, MS, cerebral palsy, Parkinson's, stroke, tremor, arthritis, or dysphonia are tested.
  • No long-term use, fatigue, saliva/comfort adaptation, or disease-progression robustness is measured.
  • No direct AAC comparison is performed with target users.
Warnings
  • Do not treat healthy silent-spelling performance as clinical restoration performance.
  • Do not overstate severe dysphonia/low-dexterity value without patient data.
Evidence
  • ev_silentspeller_a10 — View A10 (患者価値の証拠) coverage status=applicable; locator: §1; §6.3; §8; Table 5 | traced to original papers -> arxiv_2009.02110_core_silent-speech-interfaces-for-speech-restoration-a-review.txt: SSI restoration review states that, with few exceptions, SSI outcomes have mostly been validated only for healthy users and that clinical validation remains a key challenge.

confidence high (0.90)

隣接研究を部品として読む

Foley、音声分離、音声強調、AV segmentation、環境音分類、music generation、TTS pauseをSSI証拠ではなく部品として読めるか。

SilentSpeller is core SSI/text-entry work, not merely adjacent audio-visual generation or enhancement. Its reusable adjacent components are the HCI text-entry protocol, mini-QWERTY baseline, gesture-keyboard-like 5-best correction, HMM/HTK ASR machinery, n-gram language modeling, and possible BLE/in-mouth hardware packaging. Foley/music/speech-separation evidence is not relevant to its focal SSI claim.

Facts
  • §2.2: the paper grounds comparison in mobile text-entry literature and smartphone mini-QWERTY rates.
  • §3.3: recognizer uses GT2K/HTK HMM machinery and triletter modeling.
  • §6.2: phrase recognition uses unigram, bigram, and trigram grammars.
  • §7.2: UI mimics smartphone gesture-keyboard 5-best candidate interaction.
  • §7.4: TER is adapted from text-entry evaluation.
  • §8.1.1: BLE dongle and full in-mouth prototype are proposed for portability.
Inferences
  • The paper should be mined for text-entry evaluation and oral-sensor design patterns, not for audio generation.
  • Its adjacent value is strongest for correction UI, enrollment workflow, and wearable oral hardware tradeoffs.
  • Because it measures actual silent input to text, it is stronger SSI evidence than adjacent speech enhancement/generation work for the text-entry goal.
Unknowns
  • No reusable public hardware design is fully evaluated as a consumer device.
  • No comparison is made to Foley, separation, AV segmentation, music generation, or TTS pause insertion domains.
  • No component-level ablation isolates UI correction versus recognizer versus language model.
Warnings
  • Do not classify papers as SSI evidence only because they use words like silent, audio-visual, or speech.
  • Do not import speech-enhancement or generation metrics into SilentSpeller's text-entry claim.
Evidence
  • ev_silentspeller_a11 — View A11 (隣接研究を部品として読む) coverage status=applicable; locator: §2.2; §3.3; §6.2; §7.2; §7.4; §8.1.1 | traced to original papers -> kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoce is useful as an adjacent speech-generation/control component comparison: ultrasound-to-audio command control, not text entry. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce is useful as a wearable low-audibility speech/speech-enhancement component comparison, not evidence for unvoiced EPG text entry.

confidence high (0.90)

失敗条件を価値として読む

unseen speaker、cross-session、silent vs modal、field/noise、ablation、negative control、weak resultを出しているか。

The weak and failure results are highly informative. The paper reports poor Transformer baselines, weak user-independent recognition, participant-specific failure, lip-insensitive letter confusions, cognitive/comfort complaints, and live users trading accuracy for speed. These failures define SilentSpeller's boundary: strong user-dependent in-mouth spelling with dictionary support; weak first-use cross-user recognition and unresolved daily-wear hardware.

Facts
  • Table 3: Transformers score only 37%/34% character accuracy and 9.1%/8.8% word accuracy, far below HMMs.
  • §8.1.3: leave-one-user-out on first 500 words from five participants averages 55% character and 36% word accuracy.
  • §7.5: live accuracy includes user failures, recognizer failures, and corrections; users often choose speed over accuracy.
  • Table 6: P5 has low SilentSpeller accuracies of 78.7%, 82.1%, and 74.6%, while P2 reaches 46/53/52 wpm at about 90-91% accuracy.
  • §7.6: participants report retainer drawback, excess salivation, false triggering, tiredness, cognitive demand, and overthinking.
  • §8: EE-sound letters B/C/D/E/P/T/V are often confused because SmartPalate cannot sense lips.
Inferences
  • The failure conditions are not disqualifying; they show exactly where the system needs lip sensors, better enrollment, and user adaptation.
  • P5's failure suggests fit/mouth-shape/calibration effects may dominate for some users.
  • The user-independent result is a strong warning against claiming immediate out-of-box deployment.
Unknowns
  • No formal ablation quantifies how much each failure source contributes.
  • No long-term adaptation study tests whether P5-like failures recover with more data or hardware adjustment.
  • No negative control for non-spelling tongue movement is reported beyond swallow/TAP/STICK thresholds.
Warnings
  • Do not dismiss the paper because of weak user-independent results; they are valuable boundary evidence.
  • Do not hide the weak numbers when reporting the high 97% offline result.
Evidence
  • ev_silentspeller_a12 — View A12 (失敗条件を価値として読む) coverage status=applicable; locator: Table 3; §7.5; Table 6; §7.6; §8; §8.1.2; §8.1.3 | traced to original papers -> arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: Neighbor visual-only paper shows failure-as-map value: normal/whispered/silent mode mismatch produces drops, and silent is consistently hardest. ; arxiv_2103.00333_core_silent-versus-modal-multi-speaker-speech-recognition-from-ultrasound-and-video.txt: Neighbor ultrasound/video paper reports large modal-silent mismatch, including TaL80 raw WER 39.34 modal versus 77.79 silent, showing why mode-specific failure reporting matters. ; arxiv_2009.02110_core_silent-speech-interfaces-for-speech-restoration-a-review.txt: Review paper verifies broader field boundary: most SSI validation is still healthy-user/offline and realistic clinical/longitudinal validation is a known gap.

confidence high (0.90)

B: Field Maps 5

5 views in the B_field_maps family.

入力modality map

唇・顔映像、超音波、sEMG、EPG、スマホ音響、鼻パッド、囁き、レーダー、脳信号などのどこに属するか。

EPG/口腔内capacitive tongue-contact型。SilentSpellerは唇映像・超音波・sEMG・スマホ音響・鼻パッド・囁き・レーダー・脳信号ではなく、SmartPalate retainerで舌-口蓋接触を124 electrodes/100 Hzで読む。強みは歩行耐性で、offlineはwalking 97.5% char accuracy、seated 96.5%。弱みは唇を読めず、B/C/D/E/P/T/V系の混同、個別歯型、口内装着、現行tether/外部回路。

Facts
  • Figure 1 says the SmartPalate retainer has 124 electrodes and samples tongue position at 100 Hz.
  • §3.2 describes SmartPalate as a dental retainer-type device with 124 binary capacitive sensors lining the palate.
  • §6 reports the walking/seated experiment and Table 5 reports bigram results: P1 seated 99%(98%), walking 99%(97%); P2 seated 94%(93%), walking 96%(93%).
  • §8 says EE-sound letters such as B, C, D, E, P, T, and V are often confused because SmartPalate cannot sense lip position.
Inferences
  • Field placement is closest to EPG/oral tongue-contact sensing, not acoustic lip sensing or visual lipreading.
  • The modality evidence is unusually strong for motion robustness among wearable SSI papers, but only under controlled indoor walking.
  • The main sensor limitation is articulatory coverage: palate/tongue contact is strong, lip articulation is missing.
Unknowns
  • No direct outdoor walking, public transit, eating/drinking, long-term wear, or clinical-user modality data.
  • No validated fully in-mouth production hardware result.
  • No strong user-independent modality result; §8.1.3 gives only 55% character and 36% word accuracy.
Warnings
  • 入力名だけで実用性を決めない。配置は口腔内で安定するが、歯型、外部回路、唇非観測、話者差が残る。
  • Table 2の他 modality 数値は単位がphrase/word/characterで混在するため、同一性能表として読まない。
Evidence
  • ev_silentspeller_b01 — View B01 (入力modality map) coverage status=applicable; locator: Figure 1; Figure 2; §3.2; §6; Table 5; §8; §8.1.1; §8.1.3 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: lip/face video smartphone command SSI: 25-command F1 0.8947 one-shot; user study 30 commands, 1-shot 81.7%, 5-shot 98.8%. ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: smartphone acoustic sensing uses phone speaker/microphone and inaudible acoustic signals for lip movement; 54 sentences, WER 2.6% domain-dependent, 8.4% speaker/environment-independent, 8.1% unseen-sentence. ; kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: under-jaw ultrasound modality; output is synthesized audio to control Alexa; 2 users, about 500 commands/user, Network1+2 smart-speaker success avg 65%. ; arxiv_2308.06533_core_knowledge-distilled-ensemble-model-for-semg-based-silent-speech-interface.txt: sEMG-based SSI with facial sensors; 3900 samples, 26 NATO alphabet classes, best distilled ensemble accuracy 85.9%. ; arxiv_2312.09572_core_ir-uwb-radar-based-contactless-silent-speech-recognition-of-vowels-consonants-words-and-ph.txt: contactless IR-UWB radar; upper/lip radar FERASEC+DNN-HMM accuracy 86.47% vowels, 81.59% consonants, 88.95% words, 96.88% phrases. ; ssi_brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings-brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings.txt: ECoG brain-to-text modality; Brain2Char WER 10.6%, 8.5%, 7.0% on 3 participants; mimed speech WER 40% and 67% on 20 trials.

confidence high (0.90)

出力map

文字、音声復元、コマンド、spelling、同期TTS/voice-over/Foleyのどれを出す研究か。

出力は文字/単語候補によるtext entry、具体的にはspelling-to-text。音声復元、command-only、同期TTS/voice-over/Foleyではない。評価指標はWPM、TER由来accuracy、character/word accuracyが合う。37 WPM/87% live accuracyはtext-entry性能であり、PESQ/MCD/MOSやcommand accuracyと横並び比較不可。

Facts
  • §7.2 says the interface displays five best word predictions after input.
  • Figure 2 note says individual letters are not recognized in real-time.
  • §7.2 describes TAP for n-best selection and STICK for erase-word.
  • §7.5/Table 6 report live SilentSpeller average 37 WPM and 87% accuracy; mini-QWERTY control averaged 48 WPM and 93% accuracy.
Inferences
  • SilentSpeller belongs in the spelling/text-entry branch of SSI, not speech restoration.
  • The ecosystem link is smartphone text entry and AAC-like short informal text, because the study uses the MacKenzie-Soukoreff phrase set and mini-QWERTY control.
  • Output granularity is word-level dictionary recognition with edit gestures, not free character streaming.
Unknowns
  • No punctuation, capitalization, emoji, open-vocabulary composition, or out-of-vocabulary text-entry evaluation.
  • No direct ASR-style WER for free dictation.
  • No audio output quality because the system does not synthesize speech.
Warnings
  • 異なる出力を同じ性能表に並べない。37 WPM/87% text-entry accuracyは、speech reconstructionのPESQ/MCD/MOSやcommand F1と同じ意味ではない。
  • Recognition time was removed from WPM in §7.4, so user-perceived end-to-end speed is not exactly the same number.
Evidence
  • ev_silentspeller_b02 — View B02 (出力map) coverage status=applicable; locator: §3; §3.3; §7.2; §7.4; §7.5; Table 6; §8.1.4 | traced to original papers -> arxiv_1710.09798_near_lip2audspec-speech-reconstruction-from-silent-lip-movements-video.txt: Lip2AudSpec output is reconstructed speech/audio from silent lip video; metrics include PESQ, STMI/STOI-like intelligibility and Corr2D, not text-entry WPM. ; ssi_visualtts-tts-with-accurate-lip-speech-synchronization-for-automatic-voice-over-visualtts-tts-with-accurate-lip-speech-synchronization-for-automatic-voice-over.txt: VisualTTS output is automatic voice-over speech synchronized to lip video; metrics include LSE-C 5.87, LSE-D 8.45, FD 5.92, MOS 4.17±0.06. ; kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoce output is synthesized audio for smart-speaker command interaction, not text entry; Network1+2 avg smart-speaker success 65%. ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner output is customizable command classification; user study 30 commands, 1-shot 81.7%, 5-shot 98.8%, not free text.

confidence high (0.90)

評価の強さmap

評価はunseen speaker、cross-session、silent-only、real-time、live task、latency、wpm、field condition、ablationを含むか。

評価はSSI/HCIとして中〜強いが、speaker generalizationは弱い。含む: live task、standard text-entry corpus、mini-QWERTY control、walking vs seated、unseen-word holdout、latency/speed記述、ablation/tuning。含まない: 強いunseen speaker、長期field condition、end-to-end latency込みWPM、臨床ユーザ。主要数値: offline 1164 wordsで97% char/92% word級、unseen 100 wordsで94.5% char/85.5% word、walking/seatedほぼ同等、live 7人で37 WPM/87%。ただしtraining costはP1/P2 2328 words約5h、live参加者500+556 words約2h。

Facts
  • Table 1 lists pilot, isolated-word, unseen-word, walking/seated, and live text-entry experiments.
  • §4.3/Table 3: P1/P2, 2328 words, 1164 unique, 10-fold cross-validation; HMM 97% character and 93%/91% word accuracy.
  • §5/Table 4: 100 words and 200 examples left out; P1 93%(84%), P2 96%(87%), average 94.5% character and 85.5% word accuracy.
  • §7.4 removes recognition time from WPM; §7.5 reports optimized average word recognition time 200 ms and spelling a word around 1 second.
Inferences
  • Split quality is good for user-dependent and unseen-vocabulary claims, but not for speaker-independent SSI.
  • The live task is stronger than offline-only SSI work because users actually entered phrases and corrected candidates.
  • Field strength is limited: walking was indoors at home, and main live text entry was seated.
  • Ablation/control evidence is real: HMM vs Transformer, PCA dimensions, training size, FPS, n-gram comparison, mini-QWERTY control.
Unknowns
  • No powered statistical test for live SilentSpeller vs mini-QWERTY in Table 6.
  • No end-to-end WPM including recognizer wait time.
  • No cross-session quantitative breakdown except statements that later phrase/live data came weeks or months after isolated-word data.
  • No clinical participant evaluation.
Warnings
  • single speaker、closed set、random split、MSEだけ、qualitativeだけ、demoだけを強い評価として扱わない。
  • The 97% offline number is not the same evidence level as the 37 WPM/87% live number.
Evidence
  • ev_silentspeller_b03 — View B03 (評価の強さmap) coverage status=applicable; locator: Table 1; §3.1; Figure 5; §4.3; Table 3; §5; Table 4; §6; Table 5; §7; §7.4; §7.5; Table 6; §8.1.3 | traced to original papers -> arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: acoustic SSI has stronger speaker/environment split: 54 sentences, WER 8.4% speaker- and environment-independent, 8.1% unseen-sentence, 2.6% domain-dependent. ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner mobile command evaluation reports one-shot and five-shot command performance and approximately 250 ms feature/fine-tuning processing on commodity phone. ; arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: visual-only normal/whisper/silent study has 53 participants, 3 camera views, subject-independent split; silent phrases normal-to-silent 61.2 vs normal-to-normal 69.7. ; arxiv_2002.03851_core_continuous-silent-speech-recognition-using-eeg.txt: EEG CTC text study: 4 male subjects, 32 wet electrodes, 30 unique USC-TIMIT sentences; random 20% test WER 74.86%/83.34%, leave-subject WER 92.55%.

confidence high (0.90)

実用への近さmap

近い、条件付きで近い、研究段階、遠いのどこに置けるか。

実用への近さは「条件付きで近い」。研究デモではなくlive text entryまで到達し、37 WPM/87%はSMS/AAC的短文なら現実味がある。ただし製品に近いとは言えない。条件は、個別歯型retainer、1-2h以上のユーザ依存training、push-to-record、現行USB/外部module、候補表示delay、唇非観測、句読点/大文字/emoji未対応、健康成人7人のみ。

Facts
  • §7.1 says live participants collected 556 phrase words plus 500 randomly selected dictionary words; about two hours of training data collection.
  • §4.2 says P1/P2 2328-word datasets required about five hours of input for each participant.
  • §7.2 uses a push-to-record button, five-best word list, TAP, and STICK gestures.
  • §8.1.1 says the current SmartPalate sends data through USB cable and this limits mobility; Figure 14/15 show BLE and in-mouth prototypes.
  • §7.6 reports drawbacks: retainer, excess salivation, tiring mouth posture, and cognitive demand of spelling.
Inferences
  • Near-term practicality is plausible only for niche users who accept a dental retainer and training burden.
  • The walking robustness and live text-entry speed make it closer to practice than offline-only SSI recognizers.
  • The deployment burden is higher than camera/acoustic phone approaches but lower than carefully placed sEMG electrodes or ultrasound probes for daily don/doff.
  • Clinical/daily gap is large because target scenarios mention ALS/MS/Parkinson’s etc., but the study tested healthy participants.
Unknowns
  • No long-term comfort, hygiene, battery, wireless reliability, charging, dental fitting cost, or repair data.
  • No real users with low dexterity, dysphonia, ALS, MS, Parkinson’s, cerebral palsy, or muscular dystrophy.
  • No deployment in public social settings despite open-office and mobile motivation.
Warnings
  • 研究として強いことと製品に近いことを混同しない。
  • Walking robustness is promising, but indoor-home walking is not the same as daily field deployment.
Evidence
  • ev_silentspeller_b04 — View B04 (実用への近さmap) coverage status=applicable; locator: §1; §3.2; §7.1; §7.2; §7.5; §7.6; §8; §8.1.1; §8.1.3; §8.1.4; Figure 14; Figure 15; Table 6 | traced to original papers -> arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: phone-based acoustic sensing has lower hardware burden but limited 54-sentence vocabulary; WER 8.4% speaker/environment-independent and 8.1% unseen-sentence. ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearner runs as a mobile customizable command interface; 30-command user study reached 81.7% one-shot and 98.8% five-shot, but output is command recognition. ; rekimoto2026_nasovoce-nasovoce-a-nose-mounted-low-audibility-speech-interface-for-always-available-spe.txt: NasoVoce traces a later practical wearable direction: nose-mounted mic/vibration sensor, Whisper Large-v2/PESQ/STOI/MUSHRA, 1000 items, -10/0/10 dB noise, cafe/roadside/walking/train recordings. ; arxiv_2308.06533_core_knowledge-distilled-ensemble-model-for-semg-based-silent-speech-interface.txt: sEMG alphabet SSI uses facial sensors and reaches 85.9% on 26 NATO alphabet classes, showing electrode placement burden in a different wearable family.

confidence high (0.90)

研究としての強さmap

中核研究、強い隣接研究、弱いが残す価値がある研究、SSI証拠として弱い研究のどれか。

研究としては中核研究。理由は、EPG/口腔内SSIで、silent spellingという出力設計を立て、offlineだけでなくlive text entry、walking/seated、unseen words、mini-QWERTY controlまで評価したため。実用は条件付きだが、研究価値は高い。失敗地図としても有用で、user-independent 55% char/36% word、Transformer 34-37% char、P5低精度、唇非観測のEE音混同、training/hardware負担を明記している。

Facts
  • §1.1 lists contributions including Wizard-of-Oz study, training-data optimization, unseen-word tests, walking/seated tests, live text-entry experiments, and public database.
  • Table 3 reports HMM 97% character accuracy versus Transformer 37%/34% character accuracy.
  • §8.1.3 reports leave-one-user-out initial exploration: 55% character accuracy and 36% word accuracy.
  • §8.1 states limitations in participant count, training data, text-entry sessions, hardware sensing, and recognition pipeline.
Inferences
  • This is core SSI research, not merely adjacent lipreading or audio-visual enhancement, because the sensing path bypasses acoustic speech and the output is communicative text.
  • The method value is not just the HMM; it is the silent-spelling design that trades speech naturalness for recognizability, compositional vocabulary, and text-entry speed.
  • Evidence value is high for HCI viability but not conclusive for population-scale or product readiness.
  • Negative findings are unusually informative for future SSI: lips matter, user-independent modeling is hard, small-data Transformers failed, and live accuracy differs from offline accuracy.
Unknowns
  • Whether a larger user-independent/adaptive model can close the 55%/36% gap.
  • Whether added lip sensors solve EE-letter confusions without harming comfort.
  • Whether disabled target users achieve similar WPM, accuracy, comfort, and fatigue.
Warnings
  • 実用から遠い研究でも、失敗条件や基礎知見が強ければ残す。
  • Do not downgrade research strength merely because custom hardware and training burden remain.
  • Do not inflate it to product-ready: the focal itself reports major limitations.
Evidence
  • ev_silentspeller_b05 — View B05 (研究としての強さmap) coverage status=applicable; locator: §1.1; Table 1; §3.3; §4.3; Table 3; §5; §6; §7; Table 6; §8; §8.1; §8.1.3; §8.1.4 | traced to original papers -> arxiv_2103.00333_core_silent-versus-modal-multi-speaker-speech-recognition-from-ultrasound-and-video.txt: ultrasound+lip-video TaL/TaL80 study shows strong core SSI evidence on silent/modal mismatch; TaL80 WER modal 39.34 raw, silent 77.79 raw, silent 70.21 with fMLLR. ; arxiv_2010.02960_core_digital-voicing-of-silent-speech.txt: sEMG digital voicing is methodologically strong but practical burden remains: single speaker, nearly 20h EMG; closed 67-word vocab human WER 3.6, open 9828-word vocab automatic WER 68.0/human WER 74.8. ; ssi_brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings-brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings.txt: Brain2Char is strong adjacent/core neural SSI evidence but invasive ECoG: WER 10.6%, 8.5%, 7.0%; mimed WER 40% and 67%. ; arxiv_2312.09572_core_ir-uwb-radar-based-contactless-silent-speech-recognition-of-vowels-consonants-words-and-ph.txt: IR-UWB radar paper contributes failure/feasibility map for contactless sensing: upper/lip placement better than lower/chin; 86.47% vowels, 81.59% consonants, 88.95% words, 96.88% phrases.

confidence high (0.90)

C: Loop Views 11

11 views in the C_loop_views family.

通信できるか

今日、意図を通信できる形になっているか。

通信利用は成立している。ただし成立しているのはopen-vocabulary会話ではなく、SmartPalate装着・個人別学習・321語ライブ辞書・5-best訂正つきの短文テキスト入力である。ライブ7人平均は37 wpm、87% accuracy(1-TER)。認識待ち時間はWPMから除外されている。

Facts
  • Abstract: live text entry speeds for seven participants averaged 37 words per minute at 87% accuracy.
  • Abstract/Table 1: offline isolated word testingは1164-word dictionaryで平均97% character accuracy、unseen 100 wordsで平均94% accuracy。
  • §6/Table 5: walkingは97.5% character accuracy、seatedは96.5% character accuracyと要約される。
  • §7.2: wordごとに5-best候補を表示し、TAPで候補選択、STICKでerase-wordして再入力する。
  • §7.4: recognition time was removed from overall time.
  • §7.5/Table 6: mini-QWERTY平均は48 wpm、93% accuracy。SilentSpeller平均は37 wpm、87% accuracy。
Inferences
  • 今日の短文通信としては成立するが、実験の通信契約は語彙・文体・操作・装置を強く固定している。
  • 実験UIはpush-to-record buttonとright shiftを使うため、論文タイトルのhands-free目標はライブ実験そのものでは完全には満たされていない。
  • 近傍SSIの多くがコマンド入力や音声復元に寄る中で、SilentSpellerの強みはライブtext entryを測った点である。
Unknowns
  • 自由作成のopen vocabulary文での実効速度と誤り率。
  • 句読点・大文字・絵文字・記号を含む通信。
  • 臨床対象者や長期日常利用で同じ37 wpm/87%が出るか。
  • 認識待ち時間を含めた実ユーザー体感WPM。
Warnings
  • 論文ランキングより通信利用の成立を先に見る: 成立はするが、成立条件も同時に読む必要がある。
Evidence
  • ev_silentspeller_c01 — View C01 (通信できるか) coverage status=applicable; locator: Abstract; §7.2; §7.4; §7.5; Table 1; Table 5; Table 6 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearnerはmobile command routeで、30 commandsのuser study、one-shot accuracy 81.7%、5 samples per commandで98.8%、online keyword activationを報告。 ; kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoceはultrasound-to-audioでsmart speaker controlを実証し、WERはGT 20.61%、Network 1 41.03%、Network 2 33.56%、total processing time 2.61 s。 ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: smartphone acoustic sensingは54 sentencesのlimited vocabularyでdomain-independent WER 8.4%、unseen sentence WER 8.1%。

confidence high (0.90)

誰が負担するか

成果のための負担をユーザー、装置、学習データ、環境固定、外部モデル、課題制限の誰が払っているか。

性能の代価は主にユーザー、装置、個人別データ、課題制限が払っている。37 wpm/87%は、custom dental retainer、個人別HMM、約1-2時間以上の学習、321語辞書、5-best訂正、MacKenzie短文という条件の上にある。

Facts
  • §3.2: SmartPalateは124 binary capacitive sensorsを持つdental retainerで、100 Hzでtongue movementsを取得する。
  • §3.2: each user must obtain a dental impression so that the electrode array can be custom fit.
  • §4.2: P1/P2の2328-word datasets required approximately five hours of input for each participant.
  • §7.1: P3-P7は556 phrase wordsと500 dictionary wordsを収集し、about two hoursを要した。
  • §7.1: live text entry experimentを含む参加はtotal four hours。
  • §7.1/§7.2: live testは107 phrases、321 unique words、bigram recognizerを使う。
Inferences
  • device burdenは高い。口腔内センサはmotion artifactには強いが、custom fitと有線prototypeが導入障壁になる。
  • data burdenは中から高。大規模一般事前学習より、個人ごとの録音・再学習に寄せている。
  • task constraint burdenは中から高。文字入力の形は広いが、ライブ評価は321語・句読点なし・短文に限定される。
  • external model burdenは比較的低い。SilentSpeller本体はHMM/PCA/dictionary/bigramで、クラウドASRや巨大LMには依存しない。
Unknowns
  • ユーザーが毎日装着・再学習・修正を続けられるか。
  • 全内蔵wireless mouthpieceで同じ性能が出るか。
  • より大きい語彙で学習負担がどう増えるか。
Warnings
  • 高い数値だけを見ず、その数値を成立させる負担の置き場所を残す: 性能値の横にretainer・training・dictionary・repair costを書く必要がある。
Evidence
  • ev_silentspeller_c02 — View C02 (誰が負担するか) coverage status=applicable; locator: §3.2; §4.2; §7.1; §7.2; §8.1; §8.1.4

confidence high (0.90)

どの段階で不足情報を補うか

取得、発話モード差の吸収、解読自由度、受け渡しと訂正のどこで情報不足を直しているか。

不足情報は一段ではなく、取得、個人別適応、辞書/言語制約、5-best手渡し、ユーザー訂正で分散して補われている。特にlip情報の欠落はモデルとdictionary/triletter contextで補っているが、P5のB/C/D/E/P/T/V混同のように補い切れない。

Facts
  • §3.3: 124 binary electrode valuesをtop 16 principal componentsに射影し、HMMでdecodeする。
  • §3.3: training is provided in the form of words, not individual letters, so that co-articulation can be modeled.
  • §4.3: HMMは平均97% character accuracy、92% word accuracy。Transformerは34-37% character、8.8-9.1% wordに留まった。
  • §6.2: walking/seated比較では321 unique words dictionaryと107 phrasesから作ったbigramを使う。
  • §7.2: input後にfive best word predictionsを表示し、ユーザーがTAPまたはSTICKで処理する。
  • §8: P5ではEE sound letters(B, C, D, E, P, T, V)が混同され、lip-facing electrodes追加が解決案とされる。
Inferences
  • acquisition段階ではtongue-palate contactを高頻度に取り、PCAで低次元化して不足を抑える。
  • speech-mode差はsilent spellingに置き換えることで通常silent speechの発話モード差を回避しているが、個人差はuser-dependent trainingで吸収している。
  • decoding freedomは狭い。自由文生成ではなくdictionary/bigram/5-bestに情報を押し込んでいる。
  • 最終不足はUI handoffでユーザーに渡され、候補選択または再入力で閉じる。
Unknowns
  • phrase-level連続認識にした場合、どの段階で不足を補うのが最適か。
  • lip sensors追加後にP5型エラーがどれだけ減るか。
  • 大語彙・未知語・句読点でdictionary/bigram補完がどこまで保つか。
Warnings
  • SSIの失敗を単一モデル精度だけで説明しない: SilentSpellerの成功も失敗も、sensor、個人学習、辞書、grammar、UI repairの合成で読む。
Evidence
  • ev_silentspeller_c03 — View C03 (どの段階で不足情報を補うか) coverage status=applicable; locator: §3.3; §4.3; §6.2; §7.2; §8; §8.1.2; §8.1.3

confidence high (0.90)

誰が誤りを直すか

誤りはpreprocessing/device、model、external model、user、existing service、conversation partnerの誰がいつ直すか。

誤り修復の主役はモデル単体ではなくユーザーである。モデルはbigramつきHMMで5-bestを出す。ユーザーはwordごとにTAPで候補選択、STICKでerase/retryする。自動訂正は限定的で、失敗を放置すれば87%平均精度にそのまま残る。

Facts
  • §7.2: if the next input is started, the first candidate is assumed correct.
  • §7.2: if the correct answer is in the 5-best list, the user selects with TAP.
  • §7.2: if no correct answer is shown, the candidates are deleted with STICK and the system returns to input state.
  • §7.4: TER includes incorrect-not-fixed, incorrect-but-fixed, and fixes.
  • §7.5: participants mostly chose speed over accuracy, often leaving characters uncorrected, especially P3, who chose not to edit at all.
  • §7.3: mini-QWERTY baseline used autocorrection, but SilentSpeller repair is described through n-best/erase gestures.
Inferences
  • repair contractはword-level rollbackで、小さく閉じられる。
  • ユーザーが修復をしない場合、通信は速くなるが誤りが残る。
  • conversation partnerによるrepairや既存service handoffは評価されていない。
Unknowns
  • 修復操作の時間・疲労・学習曲線を独立に測った値。
  • 重要文書入力でユーザーがどの程度訂正するか。
  • 誤認実行が危険なcommand用途でconfirmationが必要か。
Warnings
  • accuracyだけでなく、失敗後に戻れる契約を読む: SilentSpellerはword単位rollback契約を持つが、コストはユーザーが払う。
Evidence
  • ev_silentspeller_c04 — View C04 (誰が誤りを直すか) coverage status=applicable; locator: §7.2; §7.3; §7.4; §7.5

confidence high (0.90)

何に依存しているか

結果は狭い課題、個人訓練、固定装着、speech/text prior、paired data、低音量入力のどれに依存するか。

依存は強い。主依存は固定装着、個人別訓練、辞書/言語prior、課題制限である。外せる依存もあるが、外すと性能は大きく落ちる。特にuser-independentは初期探索で55% character/36% wordまで落ちている。

Facts
  • §4.3: user dependent recognizers where only one participant’s data is used for training and testing at a time.
  • §4.3/Table 3: HMMは97% character、91-93% word。Transformerは34-37% character、8.8-9.1% word。
  • §4.3/Figure 11: 500 wordsで90% accuracy、98% 4-best accuracyとされる。
  • §5/Table 4: 100 removed unseen wordsではP1 93%(84%)、P2 96%(87%)、平均94.5% character/85.5% word。
  • §6.2: phrase recognition uses a dictionary constructed from 321 unique words and a bigram constructed using the 107 phrases.
  • §8.1.3: leave-one-user-out on first 500 words from first five participants averaged 55% character accuracy and 36% word accuracy.
Inferences
  • 個人訓練依存は最重要。user-independentへ外すと通信品質が崩れる。
  • speech/text prior依存は中から高。dictionary、bigram、triletter context、5-bestが性能を支える。
  • paired data依存は中。word-level sensor sequenceとword labelが必要。
  • 低音量入力ではなく完全無声spellingなので、音響入力への依存は低い。
  • 固定装着は高依存。custom retainerの安定性がmotion robustnessの源泉でもある。
Unknowns
  • 辞書を数千語から数万語へ広げた時の劣化。
  • user-independent/adaptive recognizerが実用域まで上がるか。
  • 個人別データをどこまで減らせるか。
  • OOVを辞書追加なしで扱えるか。
Warnings
  • modality名より、何を固定しないと結果が出ないかを重く見る: EPGという名前より、custom fit + user-dependent + dictionary制約が重要。
Evidence
  • ev_silentspeller_c05 — View C05 (何に依存しているか) coverage status=applicable; locator: §4.3; §5; §6.2; §8.1.3; Table 3; Table 4; Figure 11

confidence high (0.90)

時間が経つと何が壊れるか

明日、来週、再装着後、疲労後、新語彙、新環境で何が壊れるか。

時間経過への証拠は部分的に強いが、長期運用証拠ではない。weeks/months後の同一ユーザー再収録でも良好という証拠はある。一方、再装着劣化、日常疲労、新環境、臨床利用は未検証。壊れそうな箇所はfit/calibration、short words、lip-dependent letters、疲労、語彙拡張である。

Facts
  • §6.1: phrase data was recorded weeks or months after the initial isolated words for each user.
  • §6.1: good results suggest the sensing system is robust, consistent, and reproducible between sessions.
  • §7.5: main live experiment occurred about a week after phrase training and weeks to months after initial isolated words.
  • §7.6: P7 mentioned having to keep their mouth open to avoid false triggering, which was tiring.
  • §7.6: P4 reported excess salivation but found the device surprisingly comfortable during long training sessions.
  • §8: P5 poor result may be due to SmartPalate fit, mouth shape, or subtle electrode miscalibration.
  • §8.1.4: punctuation, capitalization, and emojis are not included.
Inferences
  • 明日/来週レベルの同一装着・同一タスク再現性は一定の証拠がある。
  • 再装着そのものの劣化は直接測っていない。custom retainerなので再装着は安定しやすい推測はできるが、数値はない。
  • 疲労はaccuracy曲線ではなく主観コメントとしてのみ見える。
  • 新語彙は100 unseen wordsで部分的に検証済みだが、真のopen vocabularyではない。
  • SSI近傍ではsession/repositioning劣化が大きいことがあるため、SilentSpellerでも長期remount試験は必要。
Unknowns
  • 毎日再装着を数週間続けた劣化。
  • 食事後、乾燥、唾液、歯列変化、センサ汚れの影響。
  • 疲労後の速度/TER変化。
  • 公共空間や屋外歩行での実効性能。
  • 新語彙を大量追加した時の再学習コスト。
Warnings
  • 初回デモの成功を継続利用の証拠として扱わない: 本論文はweeks/monthsの部分証拠を持つが、長期日常利用は未証明。
Evidence
  • ev_silentspeller_c06 — View C06 (時間が経つと何が壊れるか) coverage status=applicable; locator: §6.1; §7.5; §7.6; §8; §8.1.4; Table 4; Table 5 | traced to original papers -> arxiv_2305.19130_core_adaptation-of-tongue-ultrasound-based-silent-speech-interfaces-using-spatial-transformer-n.txt: UTI SSI adaptation paper reports cross-speaker/cross-session adaptation; STN alone eliminates about 75-76% of the gap, STN+output about 88-92%, showing session/device alignment can be a real SSI maintenance issue. ; arxiv_2509.21964_core_a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckb.txt: wearable EMG neckband removed/repositioned between sessions; leave-one-session-out accuracy drops to 64±18% vocalized and 54±7% silent for 8 words.

confidence high (0.90)

失敗後にどう戻るか

誤り、遅延、装着ずれ、疲労が出た後も通信を続けられるか。

失敗後の復帰経路はある。word単位で5-best選択またはerase/retryできるため、誤りは小さく閉じられる。ただしfallback input、active learning、即時再登録、疲労/装着ずれからの自動復帰は実装評価されていない。

Facts
  • §7.2: correct answerが5-bestにあればTAP gestureで選択する。
  • §7.2: no correct answerが表示されなければSTICKでcandidates are deletedし、input stateへ戻る。
  • §7.4: ERASE-WORDとN-BEST gesturesをTERのfixとして扱う。
  • §7.5: P3 chose not to edit at all, showing repair is optional and user-dependent.
  • §8.1.2: troublesome short wordsに追加例を与えるpilotでreliableになったと述べる。
  • §8.1.2: phrase-level transcription interface could use more language context and correct only when necessary.
Inferences
  • 誤りからの基本復帰はword単位re-entryで、通信の破綻を1語に閉じる設計である。
  • 遅延からの復帰は設計というより将来のrecognizer高速化に依存している。
  • 装着ずれや疲労による性能低下を検出して再登録する仕組みはない。
  • fallbackは実験上mini-QWERTY比較があるだけで、SilentSpeller UI内には統合されていない。
Unknowns
  • 連続使用中にエラー率が上がった時の再学習UI。
  • 誤りが連続した場合にユーザーが通信を継続できる閾値。
  • critical textでconfirmationやreviewをどう入れるか。
  • restart latencyやre-enrollment所要時間。
Warnings
  • 最高精度より、失敗を小さく閉じて戻る運用の有無を見る: SilentSpellerはword単位repairを持つが、自動復帰は薄い。
Evidence
  • ev_silentspeller_c07 — View C07 (失敗後にどう戻るか) coverage status=applicable; locator: §7.2; §7.4; §7.5; §8.1.2

confidence high (0.90)

何が通信を止めるか

open vocabulary、speech-mode mismatch、unseen speaker、patient adaptation、remount、latency、real environment、wearable use、clinical evidenceのどれで止まるか。

通信を止める主条件は、open vocabularyより前に、custom device + user-dependent training + live dictionary/repair contractである。テスト済み停止条件はwalking、unseen trained examples、weeks/months gap、live entry。未テスト停止条件はtrue open vocabulary、user-independent、clinical patients、remount repetition、real public environment、full wearable use、punctuation/caps/emojis。

Facts
  • §1: prior SSI is often limited to around 100 words and requires stationary use; paper positions SilentSpeller against this.
  • §5: 100 words were removed from training and tested as unseen words; average 94.5% character and 85.5% word accuracy.
  • §6/Table 5: walking and seated accuracy are similar under dictionary+bigram recognition.
  • §7/Table 6: live text entry is evaluated with seven participants over three SilentSpeller sessions each.
  • §8.1: limitations include number of participants, amount of training data, number of sessions, hardware sensing, and recognition pipeline.
  • §8.1.3: user-independent leave-one-user-out averaged 55% character and 36% word accuracy.
  • §8.1.4: punctuation, capitalization, and emojis are not included.
Inferences
  • speech-mode mismatchはSilentSpellerの中心停止条件ではない。silent spellingへ問題を変換している。
  • walking/motion artifactはかなり外した停止条件である。
  • unseen vocabularyは部分的に外したが、辞書内unseen training wordsであり、open vocabularyではない。
  • route blockerはproduct化ではhardware/wearability、field robustness、user-independent/adaptive recognition、vocabulary expansionが支配的。
Unknowns
  • ALS/MS/Parkinson’s等の対象ユーザーでの性能。
  • 完全wireless in-mouth prototypeでの同一性能。
  • 辞書外語・固有名詞・長文作成。
  • 日常環境でのfalse triggerとrepair burden。
Warnings
  • global rankingより、どの停止条件と戦っている路線かを先に見る: この論文はmovement/live text entryと戦い、open vocabulary/clinical/productは残している。
Evidence
  • ev_silentspeller_c08 — View C08 (何が通信を止めるか) coverage status=applicable; locator: §1; §5; §6; §7; §8.1; §8.1.3; §8.1.4; Table 5; Table 6

confidence high (0.90)

どの路線か

低音量接続、狭い課題入力、無音再生、患者復元、脳信号基礎研究のどの路線か。

路線は「無音spellingによる狭めたテキスト入力」である。低音量接続でも患者復元でも脳信号基礎でもない。route metricはWPM、TER/accuracy、character/word accuracy。peer setはmobile command lipreading、ultrasound-to-audio smart speaker control、smartphone acoustic limited-sentence SSR、wearable EMG command SSIで、停止条件はそれぞれ違う。

Facts
  • §3: silent spelling means the user spells each letter in the words, one by one.
  • §3: silent spelling is compositional because words never seen in training might still be recognized.
  • §7: paper says no comparable silent speech system was found for mobile text entry because neighbors are offline, small vocabulary, command phrases, or combinations.
  • §7.5: SilentSpeller live average is 37 wpm and 87% accuracy.
  • Table 2: SilentSpeller Live is listed as Capacitive/Tongue, 321 words, C87%, 37wpm, walking yes, unseen yes.
Inferences
  • SilentSpellerは音声復元の自然性ではなく、通信可能なtext entryを直接評価する路線である。
  • 同じSSIでも、LipLearnerのcommand customizationやSottoVoceのsmart speaker audio controlと一列ランキングにしない方がよい。
  • route-specific stop conditionは、大語彙自由入力を保ちながらtraining/retainer/repair負担をどこまで減らせるかである。
Unknowns
  • このrouteが最終的にsilent speech full dictationへ拡張できるか。
  • command routeとtext-entry routeの実ユーザー選好差。
  • 低音量speech routeとの速度/社会的許容性比較。
Warnings
  • 停止条件が違う研究を一列順位で比べない: 37 wpm text entryと98.8% command recognitionと8.4% WER sentence SSRは別路線で読む。
Evidence
  • ev_silentspeller_c09 — View C09 (どの路線か) coverage status=applicable; locator: §3; §7; §7.5; Table 2; Table 6 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearnerはcommand customization route。25-command one-shot F1 0.8947、user study 30 commandsでone-shot 81.7%、5-shot 98.8%。 ; kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoceはultrasound imagesからaudioを生成し、Amazon Alexa commandsをsmart speakerで評価。WER 20.61/41.03/33.56%。 ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: smartphone acoustic SSRは54 Chinese sentences、domain-independent WER 8.4%、unseen sentence WER 8.1%。 ; arxiv_2509.21964_core_a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckb.txt: wearable EMG neckband routeは8 target words、silent 5-fold accuracy 68±3%、LOSO silent 54±7%、power 22.2 mW。 ; arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: visual-only speech-mode study reports matched silent phrase classification 64.4% and normal-trained-to-silent 61.2%, showing speech-mode mismatch is a separate route issue.

confidence high (0.90)

路線ごとの負担

各路線でユーザー、装置、モデル、環境、既存音声系の誰が残る負担を払うか。

SilentSpeller routeの残負担は、既存音声系ではなくユーザーと装置が主に払う。既存speech systemへ投げるSottoVoceや、smartphone camera/acousticへ寄せるLipLearner/acoustic SSRとは負担の置き場所が違う。

Facts
  • §3.2: each user must obtain a dental impression for a custom-fit SmartPalate.
  • §4.2/§7.1: training data collection is about five hours for 2328 words or about two hours for 1056 words.
  • §7.2: users must operate word-level repair through TAP/STICK gestures.
  • §8.1.1: current SmartPalate sends sensing data through a USB cable, limiting mobility.
  • §8.1.1: latest prototype suggests full in-mouth integration is possible but presented as future/prototype work.
  • §8.1.3: user-independent result is 55% character and 36% word.
Inferences
  • user costはtraining、mouthpiece tolerance、spelling mental demand、repair操作。
  • device costはcustom intraoral sensorで高いが、環境costはcamera/lightingより低く、walking robustnessが強い。
  • model costはHMMで軽めだが、個人別データとdictionary/bigramに依存する。
  • existing speech system costはSilentSpeller text-entry routeでは低い。これはSottoVoce型の音声生成routeと対照的。
Unknowns
  • 完全in-mouth BLE版の電力、遅延、耐久性。
  • ユーザーがtraining負担を受け入れる条件。
  • 大語彙化した時のmodel/device/user cost配分。
Warnings
  • 同じ停止条件を外していても、代価の支払い先で実用の意味が変わる: SilentSpellerはmotion artifactを外す代わりにcustom oral deviceと個人訓練を要求する。
Evidence
  • ev_silentspeller_c10 — View C10 (路線ごとの負担) coverage status=applicable; locator: §3.2; §4.2; §7.1; §7.2; §8.1.1; §8.1.3 | traced to original papers -> kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: SottoVoceはgenerated audioをspeakerから出し、unchanged Amazon Echo/Echo Showを制御するroute。processing time 2.61 s、WER 41.03/33.56% for generated networks。 ; arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: LipLearnerはcommodity smartphone、on-device fine-tuning、visual keyword spotting、Voice2Lip登録を使うcommand route。 ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: acoustic sensing routeはsmartphone speaker/microphoneで54 sentencesを認識し、domain-independent WER 8.4%。追加口腔内device負担はないがlimited vocabulary。 ; arxiv_2509.21964_core_a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckb.txt: wearable EMG routeはfully-dry neckband、14 differential channels、22.2 mW、8-word silent accuracy 68±3%、LOSO silent 54±7%。

confidence high (0.90)

個別論文を読む最終手順

路線、出力、外した停止条件、代価の支払い先、外部補助が弱くなった時に残るものを順に確認したか。

最終判定: 強い論文。ただし強さは最高accuracyではなく、SSIで珍しいライブtext-entry通信、walking耐性、unseen trained-word耐性、repair UIまで通した点にある。弱さは、custom retainer、user-dependent training、321語ライブ辞書、recognition time除外、完全hands-free未達、臨床/長期/open vocabulary未証明である。

Facts
  • §1/Table 1: contributions include Wizard of Oz speed/usability, training-data optimization, unseen words, walking vs seated, interactive text entry, live text entry, public database.
  • §7.5/Table 6: live SilentSpeller average is 37 wpm and 87% accuracy; mini-QWERTY average is 48 wpm and 93%.
  • §7.4: recognition time was removed from WPM measurement.
  • §6/Table 5: walking and seated offline phrase recognition differ little.
  • §5/Table 4: 100 words removed from training still yield average 94.5% character and 85.5% word accuracy.
  • §8.1: paper explicitly lists limitations in participants, training data, sessions, hardware sensing, and recognition pipeline.
  • §8.1.3: user-independent recognition is poor at 55% character and 36% word.
Inferences
  • route確認済み: silent spelling text-entry。
  • output確認済み: text words/phrases, not restored audible speech。
  • 外した停止条件: seated-only/offline-only/trained-only vocabularyを部分的に外した。
  • 代価確認済み: user/device/data/task constraintsが払う。
  • 外部補助が弱くなった時に残るもの: tongue-sensing spelling signalとHMMは残るが、dictionary/bigram/user training/repairなしの通信性能は未証明。
  • 現実停止条件を外した度合いは高いが、製品/臨床/自由入力の証拠としてはまだ不足。
Unknowns
  • open vocabulary dictationとしての成立。
  • target clinical populationでの実用性。
  • 完全wearable all-in-mouth systemでの性能。
  • 長期日常利用の保守コスト。
  • 句読点・大文字・絵文字・専門語彙を含む作業入力。
Warnings
  • 強い論文を、数字が高い論文ではなく、現実の停止条件を外し代価も読める論文として判定する: SilentSpellerはここで強いが、代価は大きい。
Evidence
  • ev_silentspeller_c11 — View C11 (個別論文を読む最終手順) coverage status=applicable; locator: §1; §5; §6; §7.4; §7.5; §8.1; §8.1.3; Table 1; Table 4; Table 5; Table 6 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: nearby mobile SSI command paper reports high command accuracy with few-shot samples but evaluates command interaction, not continuous text entry. ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: nearby smartphone acoustic SSR reports low WER on 54 limited sentences, including unseen sentence WER 8.1%, making vocabulary/task restriction explicit. ; arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: visual-only speech-mode paper shows silent speech remains hard even in matched condition, e.g. silent-trained/silent-tested phrase classification 64.4%, so SilentSpeller's spelling route avoids a real speech-mode stop condition. ; arxiv_2305.19130_core_adaptation-of-tongue-ultrasound-based-silent-speech-interfaces-using-spatial-transformer-n.txt: UTI adaptation paper shows session/speaker adaptation burden remains central in articulatory SSI; STN adaptation reduces about 75% of gap and STN+out about 88-92%. ; arxiv_2509.21964_core_a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckb.txt: wearable EMG paper verifies reposition/session burden: LOSO silent accuracy 54±7% despite global silent 68±3% on only 8 words.

confidence high (0.90)

D: Route Catalog 5

5 views in the D_route_catalog family.

低音量路線

低音量入力や囁きを既存ASR、スマートスピーカー、音声強調へつなぐ路線か。

SilentSpellerは低音量・囁き入力を既存ASR/TTSやスマートスピーカーへつなぐ路線ではない。入力は「without voicing」のsilent spellingで、出力はテキスト入力。低音量化でASRに寄せるのではなく、124電極SmartPalateで舌接触を100 Hz取得し、HMMで321/1164語辞書内の語を選ぶ。近い用途は「speech is socially inappropriate」な場面だが、音声強調・囁きASR接続ではない。

Facts
  • Figure 1: SmartPalate retainer has 124 electrodes and senses tongue position at 100 Hz.
  • Abstract: users type by spelling words without voicing.
  • §1: speech recognition could distract colleagues and raise privacy/security concerns.
  • §3.2: SmartPalate sends 100 Hz tongue-movement data; no acoustic ASR/TTS link is used.
  • Table 2 lists Fukumoto 2018 as Audio/Ingressive speech, but SilentSpeller rows are Capacitive/Tongue.
Inferences
  • The focal system preserves silent-input purity more than whisper/low-volume routes because it does not intentionally emit or capture low-volume speech.
  • Its privacy claim is about avoiding audible speech in open/shared contexts, not about making existing ASR robust to quiet audio.
  • Because output is text-entry candidates, D01 is at most a contrast class.
Unknowns
  • The paper does not measure actual acoustic leakage during silent spelling.
  • The paper does not test whisper, low-volume speech, speech enhancement, or existing ASR integration.
  • The paper does not compare privacy against whispered ASR in user studies.
Warnings
  • 完全無音を諦める低音量路線として扱わない。SilentSpellerは低音量ではなく無声spelling路線である。
Evidence
  • ev_silentspeller_d01 — View D01 (低音量路線) coverage status=not_applicable; locator: Abstract; Figure 1; §1; §3.2; Table 2 | traced to original papers -> arxiv_1802.06399_core_visual-only-recognition-of-normal-whispered-and-silent-speech.txt: Neighbor low-volume/whisper contrast: visual-only paper explicitly studies normal, whispered, and silent speech, and reports phrase classification 64.4% for silent-trained/silent-tested and 61.2% for normal-trained/silent-tested.

confidence medium (0.75)

狭い課題路線

コマンド、綴り、短文入力を確実に通す路線か。

SilentSpellerはD02の中核例。自由会話ではなく、無声で単語を綴り、辞書・bigram・5-best編集に載せる狭いテキスト入力路線である。性能はoffline isolated wordで1164語、平均97% character accuracy / 92% word accuracy、unseen 100語で94% character / 86% word、walking/seatedで97.5%/96.5% character accuracy、live 7人で平均37 WPM・87% accuracy。コストは個人別SmartPalate作製、124電極、P1/P2は2328語収集に約5時間、P3-P7は556語+500語で約2時間、live実験込み計4時間程度、さらにpush-to-recordボタン使用で完全hands-freeではない。

Facts
  • §1: silent spelling recognizes spelling instead of full silent speech and uses a 1164-word vocabulary in this work.
  • Table 3: HMM recognizers reached 97% character accuracy and 93%/91% word accuracy for P1/P2 on 2328 isolated words, 1164 unique.
  • §5 / Table 4: removing 100 words from training gave average 94% character accuracy and 86% word accuracy.
  • Table 5: walking/seated bigram results were P1 99%(98%) seated, 99%(97%) walking; P2 94%(93%) seated, 96%(93%) walking.
  • §7.5 / Table 6: live text entry average speed was 37 WPM and average accuracy was 87%; mini-QWERTY average was 48 WPM and 93%.
  • §7.2: interaction includes INPUT, N-BEST-SELECT/TAP, and ERASE-WORD/STICK.
Inferences
  • The paper's win is reliable short-text entry under a constrained dictionary/language-model/edit interface, not unrestricted dictation.
  • Spelling is deliberately narrower than silent speech, trading naturalness and speed for recognizer reliability and mobility.
  • The live result is strong for a prototype SSI route, but cost must be read beside performance: custom retainer, user-dependent data, and hours of training.
Unknowns
  • Open-vocabulary composition beyond the tested dictionary and 100 held-out words remains unproven.
  • Long-term daily learning curves and device comfort are not measured beyond study sessions.
  • Patient or low-dexterity users were not directly tested.
Warnings
  • 狭い課題の勝利を自由会話の勝利として扱わない。37 WPM/87%は321語・107 phrase・bigram・5-best編集付きlive text entryの値である。
Evidence
  • ev_silentspeller_d02 — View D02 (狭い課題路線) coverage status=applicable; locator: §1; §3; Table 3; Table 4; Table 5; §7.2; §7.5; Table 6 | traced to original papers -> arxiv_2302.05907_core_liplearner-customizable-silent-speech-interactions-on-mobile-devices.txt: Neighbor D02 command route: LipLearner reports 25-command one-shot F1 0.8947±0.0530, and user study with 30 commands reports one-shot 81.7% and five-shot 98.8% accuracy. ; arxiv_2011.11315_core_end-to-end-silent-speech-recognition-with-acoustic-sensing.txt: Neighbor narrow sentence route: smartphone acoustic sensing uses 54 sentences and reports WER 8.4% domain-independent and 8.1% unseen-sentence testing. ; arxiv_2509.21964_core_a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckb.txt: Neighbor wearable EMG command route was opened via prior route trace and used as contrast for limited-word command systems; focal paper itself does not rely on this number.

confidence high (0.90)

無音再生路線

無音の映像や調音から音声や文字を再生する路線か。

SilentSpellerは無音映像・調音から音声を再生する路線ではない。入力はSmartPalateの舌接触系列、出力は文字候補・テキストであり、生成音声やTTS品質を評価しない。§8.1.1にSmartPalate間の電気刺激フィードバックや二者通信の想像はあるが、これは将来案で、本文の実験結果ではない。

Facts
  • §3.3: captured SmartPalate data are projected to 16 PCA components and decoded using HMMs into words.
  • Figure 12: interface displays five best word predictions, not generated speech.
  • §7.5: recognition latency and text-entry accuracy are discussed; no speech naturalness metric appears.
  • §8.1.1: electrical stimulation could provide feedback and perhaps two-way communication between two SmartPalates, stated as future possibility.
Inferences
  • The paper is recognition-to-text, not articulatory-to-speech synthesis.
  • No guidance/no-guidance generation and content-fidelity split are irrelevant to the demonstrated system.
Unknowns
  • The paper does not evaluate whether tongue-contact signals can reconstruct acoustic speech.
  • The paper does not test generated speech naturalness, intelligibility, speaker identity, or content accuracy.
Warnings
  • 生成音声の自然さと入力内容の正しさを分ける以前に、SilentSpellerは生成音声を出していない。
Evidence
  • ev_silentspeller_d03 — View D03 (無音再生路線) coverage status=not_applicable; locator: §3.3; Figure 9; Figure 12; §7.5; §8.1.1 | traced to original papers -> kimura2019_sottovoce-sottovoce-an-ultrasound-imaging-based-silent-speech-interaction-using-deep-neura.txt: Neighbor D03 playback route: SottoVoce converts ultrasound images to audio for unchanged smart speakers; processing time 2.61 s and WER 20.61%/41.03%/33.56% for GT/Network 1/Network 2 were verified.

confidence medium (0.75)

患者復元路線

患者や失声者の復元を直接狙う路線か。

SilentSpellerは患者復元を直接狙う臨床路線ではなく、健常参加者でmobile/hands-free silent text entryを示すHCI/SSI prototypeである。患者・低手指巧緻性・dysphoniaは強い動機として出るが、実験参加者はP1-P7の健常成人で、ALS/MS/Parkinson/cerebral palsy/muscular dystrophy等の直接患者結果はない。assistive valueは plausible だが、直接臨床証拠は薄い。

Facts
  • §1 motivating scenario describes muscular dystrophy and low manual dexterity in an open office.
  • §2.3 discusses AlterEgo extension with three users with movement impairments and dysphonia due to MS, but this is related work.
  • §7.1: P3-P7 are male, ages 23 to 45; P2 and P5 native English speakers; others non-native with English ability descriptions.
  • §8: future direction includes comparing against current methods for participants with multiple sclerosis, Parkinson’s, cerebral palsy, and muscular dystrophy.
  • §8: for people with severe dysphonia and low dexterity, a conservative version might be used for home automation or choosing one of N phrases.
Inferences
  • The paper is assistive-motivated but not a direct patient restoration study.
  • The healthy-proxy gap is substantial because device fit, fatigue, dysphonia, motor impairment, and daily AAC needs are not directly measured in target users.
  • The social value could be large, but the focal evidence supports feasibility in healthy users first.
Unknowns
  • No clinical protocol, clinician-supervised trial, or patient outcome is reported.
  • No direct result for ALS, MS, Parkinson’s, cerebral palsy, muscular dystrophy, laryngectomy, or severe dysphonia users is reported.
  • Whether target users can tolerate the retainer, training burden, and spelling cognitive load is unknown.
Warnings
  • 社会的価値は大きいが、直接証拠は薄い。患者復元路線の実証として過大評価しない。
Evidence
  • ev_silentspeller_d04 — View D04 (患者復元路線) coverage status=not_applicable; locator: §1; §2.3; §7.1; §8; §8.1 | traced to original papers -> arxiv_2111.01740_near_personalized-one-shot-lipreading-for-an-als-patient.txt: Neighbor direct-patient route: Personalized One-Shot Lipreading explicitly studies an ALS patient and reports over 83% top-5 accuracy / 83.2% versus 62.6% comparable methods for the patient.

confidence medium (0.75)

脳信号路線

EEG、MEG、ECoGなど脳信号から可能性を探る基礎研究路線か。

SilentSpellerは脳信号路線ではない。EEG/MEG/ECoGを使わず、口腔内SmartPalateのcapacitive tongue sensingを使う。関連研究ではBCI/EEG spellersが20 WPMをめったに超えず、動きに弱いと述べるだけで、脳信号からの基礎研究・通信実用性は評価対象外である。

Facts
  • Figure 1: SmartPalate retainer with 124 electrodes senses tongue position at 100 Hz.
  • §2.1: non-invasive gaze and brain computer interfaces rarely exceed 20 WPM and are highly susceptible to body movements.
  • §2.1: most EEG spellers rely on electrical brain signals from visual stimuli to determine selected key.
  • §3.2: SmartPalate is a dental retainer-type device with 124 binary capacitive sensors lining the palate.
  • Table 2: SilentSpeller modality is Capacitive and proxy is Tongue.
Inferences
  • The paper positions BCI as a slower/movement-sensitive text-entry baseline, not as its own route.
  • Daily-use distance for brain-signal systems cannot be inferred from SilentSpeller except as a contrast motivating a non-brain wearable sensor.
Unknowns
  • No EEG/MEG/ECoG data are collected.
  • No invasive/noninvasive brain equipment burden, neural decoding accuracy, or BCI daily-use protocol is evaluated.
  • No comparison experiment against a real BCI keyboard is performed.
Warnings
  • 科学的野心と近い将来の通信実用性を分ける以前に、SilentSpellerは脳信号研究ではない。BCI比較は背景情報に限る。
Evidence
  • ev_silentspeller_d05 — View D05 (脳信号路線) coverage status=not_applicable; locator: Figure 1; §2.1; §3.2; Table 2 | traced to original papers -> arxiv_2002.03851_core_continuous-silent-speech-recognition-using-eeg.txt: Neighbor D05 EEG route: continuous silent speech recognition using non-invasive EEG, 32 wet EEG electrodes/31 sensors used, 30 unique sentences, WER 83.34% with all sensors plus dimension reduction and 92.55% in cross-subject setting. ; arxiv_2604.16441_core_iphoneme-brain-to-text-communication-for-als-using-conformerxl-decoding.txt: Neighbor invasive brain-to-text route: ALS iEEG/ECoG system reports 92.14% phoneme accuracy, 73.39% word accuracy, 26.61% WER, 180 ms latency, and 256-channel intracranial recording.

confidence medium (0.75)

Rankings 1

provisional · route · narrow_task

Axis
route_level
Condition
Route-scoped comparison only (R02 discipline).
Cost
See per-view burden / L04 (user + device custom fit + enrollment).
Caveat
Route-aware numeric rank not recomputed across all papers; rank left null. Route-scoped rank: 1 on axis 'near-term communication practicality within mobile silent text/command-entry systems: live or deployable interaction, output expressiveness, WPM/accuracy or WER/F1, motion tolerance, and user/device cost stated beside the metric'. Rank is qualitative, not a strict numeric leaderboard, because peer papers report different metric families: SilentSpeller reports live WPM/TER-derived accuracy for compositional text; SottoVoce reports smart-speaker command success and WER for generated audio; acoustic sensing reports WER on 54 fixed sentences; LipLearner reports F1/accuracy for customizable commands. Metric families are therefore kept separate per R03.

Reading guidance

Verdict
full-text draft · priority high · confidence high
Why it matters
The paper's main contribution is the novel problem reframing from silent speech to silent spelling using an electropalatography retainer, which yields practical live text entry over large vocabularies including unseen words, with empirical validation of robustness to walking and comparison to mainstream mobile text input.
What to trust
Basis: full text + existing expert seed. Coverage: high. 8 evidence records back the review.
What is weak
Recognition confusions occur mainly for letters with similar palatograms, especially EE-sound letters (B/P, D/T/Z). Strong user-dependence; user-independent recognition remains poor. Offline experiments rely on data from only two main participants for tuning; live text entry and walking tests include seven users but under constrained phrase tasks; vocabulary is limited to English letters and space without punctuation or capitalization. The system requires a custom-fitted in-mouth SmartPalate retainer with 124 electrodes connected by a wired (now partially wireless prototype) interface. The retainer remains obtrusive, and the user-dependent training required limits scalability and ease of deployment. The system targets discreet text entry, explicitly trading away naturalness of silent speech for reliability; not a conversational silent speech system. Overclaim risk: low-medium.
Read before
SSI review rubric
Read next
SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks

Axes

Task
text-entry using silent spelling
Modality
electropalatography
Hardware
SmartPalate custom dental retainer with 124 capacitive electrodes sampled at 100 Hz, connected wired or wireless to processing device.
Body site
palate; tongue
Output
text
Vocabulary
Dictionary-based silent spelling with triletter HMM decoding, phrase composition, and bigram correction
Metrics
Offline HMM accuracy: ~97% character, 92% word; Unseen word offline test: 94.5% character, 85.5% word accuracy; Walking and seated phrase recognition: ~97.5% vs 96.5% character accuracy; Live text entry average 37 words per minute at 87% accuracy, best participant 53 wpm at 91%.
Evaluation mode
Offline isolated word recognition with 10-fold cross validation; reserve testing on 100 unseen words; seated vs walking phrase recognition; live interactive text entry with push-to-talk interface and edit gestures.
Review confidence
high
Overclaim risk
low-medium

Expert take

SilentSpeller offers a carefully validated alternative to classic silent speech interfaces by changing the recognition task from continuous silent speech to discrete silent spelling. This reframing produces a more structured signal that, together with electropalatography sensors and a user-dependent HMM recognizer, supports a large vocabulary of 1164 words with high offline accuracy. The system uniquely supports live interactive text entry at speeds around 37 wpm with 87% accuracy, including robust performance while walking, demonstrating tolerance to motion artifacts. Limitations include the requirement for custom dental impressions, obtrusive mouth hardware, strong user dependence for training, and struggles with user independence and social acceptability. Overall, the work advances SSI toward viable mobile, hands-free text entry applications in privacy-sensitive or hands-busy scenarios, providing extensive empirical evidence backing claims.

True value

The paper's main contribution is the novel problem reframing from silent speech to silent spelling using an electropalatography retainer, which yields practical live text entry over large vocabularies including unseen words, with empirical validation of robustness to walking and comparison to mainstream mobile text input.

What changed

Canon before

Most silent speech interfaces were limited to small vocabularies (~100 words), stationary use, and offline, non-interactive experiments, with little evidence for practical live text entry.

Delta from canon

The key change is reframing the task from silent speech recognition to silent spelling recognition, allowing a larger vocabulary (1164 offline words), robust unseen-word generalization, tolerance to walking motion, and live hands-free text entry at reasonable speeds (~37 wpm average).

Position in field

One of the clearest practical SSI task-reframing papers to date; less natural than silent speech but more usable for mobile text entry.

Evidence

“ SilentSpeller achieves an average of easily recognized than silent speech, allowing larger vocabularies 94% accuracy (86% word). (1164 words in this work) and on-the-go interaction. ”

author_claim · 1 INTRODUCTION · confidence 1.00

“ Character (word) accuracy Transformer 37% (9.1%) 34% (8.8%) For the SilentSpeller use cases of silent text entry while mobile, or Table 3: Average 10-fold, cross-validation, user-dependent for people with movement disorders, one or two hours of training word accuracy on 2328 isolated words, 1164 unique, using data is quite reasonable, especially since such use cases may often HMMs and deep learning Transformers. use a limited vocabulary [28, 38]. ”

metric · TUNING MODELS · confidence 1.00

“ Figure 1: a) A SilentSpeller user wears the SmartPalate retainer whose 124 electrodes sense the position of the tongue at 100 Hz. ”

fact · 3.2 SMARTPALATE · confidence 1.00

“ Character (word) accuracy Transformer 37% (9.1%) 34% (8.8%) For the SilentSpeller use cases of silent text entry while mobile, or Table 3: Average 10-fold, cross-validation, user-dependent for people with movement disorders, one or two hours of training word accuracy on 2328 isolated words, 1164 unique, using data is quite reasonable, especially since such use cases may often HMMs and deep learning Transformers. use a limited vocabulary [28, 38]. ”

validation_scope · TUNING MODELS · confidence 1.00

“ The current system can be made wearable for testing; Figure 1a Early Identification of Recognizer Success shows such a system constructed using a Vufine head worn dis- P2 P5 play, the Smart Palate, and the support hardware in a backpack. ”

deployment_claim · 3.2 SMARTPALATE · confidence 1.00

“ Character (word) accuracy Transformer 37% (9.1%) 34% (8.8%) For the SilentSpeller use cases of silent text entry while mobile, or Table 3: Average 10-fold, cross-validation, user-dependent for people with movement disorders, one or two hours of training word accuracy on 2328 isolated words, 1164 unique, using data is quite reasonable, especially since such use cases may often HMMs and deep learning Transformers. use a limited vocabulary [28, 38]. ”

deployment_claim · 4.3 TUNING USER DEPENDENT RECOGNIZERS · confidence 1.00

“ We have reason to be optimistic: the words from the 107 phrases collected during the seated condition. tongue is relatively isolated from the mechanical shock of walking Similarly, the recognizer for the seated condition was trained with (otherwise, voiced speech while walking would not be possible) the 2328 dictionary words plus the 556 words from the 107 phrases and SilentSpeller’s electrode array fits snugly in the mouth such collected during the walking condition. ”

deployment_claim · 8 DISCUSSION · confidence 1.00

“ Similarly, of 124 binary electrode values is projected to the top 16 principal SilentSpeller focuses on recognizing the 26 letters of the alphabet, components. ”

fact · 3.3 RECOGNIZER PIPELINE · confidence 1.00

Limits

Technical limits

Recognition confusions occur mainly for letters with similar palatograms, especially EE-sound letters (B/P, D/T/Z). Strong user-dependence; user-independent recognition remains poor.

Evaluation limits

Offline experiments rely on data from only two main participants for tuning; live text entry and walking tests include seven users but under constrained phrase tasks; vocabulary is limited to English letters and space without punctuation or capitalization.

Deployment limits

The system requires a custom-fitted in-mouth SmartPalate retainer with 124 electrodes connected by a wired (now partially wireless prototype) interface. The retainer remains obtrusive, and the user-dependent training required limits scalability and ease of deployment.

Scope limits

The system targets discreet text entry, explicitly trading away naturalness of silent speech for reliability; not a conversational silent speech system.