← SSI archive · Review rubric

2026 · arXiv · Field expert review · confidence high

SilentWear: an Ultra-Low Power Wearable System for EMG-based Silent Speech Recognition

Giusy Spacone, Sebastian Frey, Giovanni Pollo, Alessio Burrello, Daniele Jahier Pagliari, Victor Kartsch, Andrea Cossettini, Luca Benini

BibTeX
@misc{silentwear-an-ultra-low-power-wearable-system-for-emg-based-silent-speech-recognition,
  title = {SilentWear: an Ultra-Low Power Wearable System for EMG-based Silent Speech Recognition},
  author = {Giusy Spacone and Sebastian Frey and Giovanni Pollo and Alessio Burrello and Daniele Jahier Pagliari and Victor Kartsch and Andrea Cossettini and Luca Benini},
  year = {2026},
  note = {arXiv},
  eprint = {2603.02847},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2603.02847v2},
}

A useful dry-neckband and embedded-CNN study: silent balanced accuracy falls from 77.5% across pooled-day batches to 59.3% on a new day; 2.47 ms is compute time, while closed-loop usability remains untested.

Verdict: full-text draftPriority: highConfidence: highBasis: full text + summaryCoverage: high

Reading guidance

Verdict
full-text draft · priority high · confidence high
Why it matters
Makes the cost of day-to-day sensor repositioning visible while providing a compact embedded baseline and a concrete supervised recalibration strategy.
What to trust
Basis: full text + summary. Coverage: high. 11 evidence records back the review.
What is weak
Offline preprocessing uses a zero-phase 20 Hz high-pass filter plus 50 Hz notch; causal deployment equivalence is not established. Trigger timing imperfectly aligns articulation. Quantization changes recognition accuracy by an unreported amount. Whole-window ITR omits actual cue/rest/correction overhead. Four subjects and eight commands plus rest. Global evaluation mixes every recording day into train and test through different batches; only leave-one-session-out isolates a new day. All models remain subject- and mode-specific. Rest is downsampled to class balance, unlike typical idle-heavy use. Online closed-loop validation is explicitly future work. Personalized fixed-command models, cued seated data and sensor-fit sensitivity. No closed-loop user study, walking/activity false-trigger study, patient cohort or measured all-day endurance. Recalibration is trained offline rather than demonstrated on-device. Four seated participants, eight prompted English commands plus rest, three days and separate models for vocalized/silent conditions. No unseen words, sentences, patient communication or naturalistic continuous use. Overclaim risk: Moderate for complete real-time usability, all-day endurance and general recognition accuracy if compute benchmarks, battery estimates and pooled-day results are treated as operational user validation..
Read before
SSI review rubric
Read next
SSI archive

Axes

Task
command-recognition
Modality
14-channel differential surface EMG from a fully dry textile neckband.
Hardware
27 dry SoftPulse electrodes in a textile neckband form 14 differential channels: 10 overlapping central channels and four lateral channels; four shorted electrodes provide ground. BioGAP-Ultra uses two ADS1298 AFEs, GAP9/NE16 and nRF5340 BLE, in 26 x 65 x 13 mm electronics with a 150 mAh Li-Po battery.
Body site
throat
Output
commands;labels
Vocabulary
closed-set isolated commands
Metrics
Table II, 1.4 s windows: SpeechNet balanced accuracy 84.8 +/- 4.6% vocalized and 77.5 +/- 6.6% silent for pooled-day held-out batches; 71.1 +/- 8.3% and 59.3 +/- 2.2% for held-out days. Random Forest silent results are 62.9% and 48.7%, respectively. Table III, 800 ms windows: 2.47 ms / 63.9 microjoules at 240 MHz, or 1.60 ms / 97.1 microjoules at 370 MHz. System power 20.5 mW with 100 ms stride; estimated battery life 27.1 h. Author-reported measurements, not independently reproduced.
Evaluation mode
Subject-specific balanced accuracy using pooled-day held-out batches, leave-one-session-out testing and sequential next-batch fine-tuning; window-length ablation and analytical ITR; embedded latency/energy benchmarking.
Review confidence
high
Overclaim risk
Moderate for complete real-time usability, all-day endurance and general recognition accuracy if compute benchmarks, battery estimates and pooled-day results are treated as operational user validation.

Expert take

SilentWear is strongest as an integrated sensing-and-embedded-computing study. It combines a fully dry 14-channel textile neckband, a 15,489-parameter CNN and explicit multiday tests, while reporting both inference energy and the larger acquisition/communication power budget. The key recognition result is the day-transfer gap: silent balanced accuracy is 77.5% when training and testing use different batches from the same set of days, but only 59.3% when an entire day is held out. That difference is more informative for daily use than the headline accuracy alone. Incremental fine-tuning improves next-batch results with roughly one additional acquisition batch, although this is supervised recalibration and is not demonstrated as on-device learning. The embedded measurements are valuable but must retain their boundaries: 2.47 ms is inference time on an 800 ms window, predictions are scheduled every 100 ms, and the 27.1-hour battery duration is an estimate from the reported power and battery capacity. Neither the analytical information-transfer rate nor near-perfect classification of cued rest segments measures actual communication throughput or accidental activations during daily activity. The paper explicitly says that online closed-loop system validation remains future work. Offline zero-phase filtering also needs reconciliation with a causal deployed implementation, and quantized recognition accuracy is not separately reported. This is a useful reproducible direction for personalized command interfaces, with stronger hardware evidence than evidence for everyday communication readiness.

True value

Makes the cost of day-to-day sensor repositioning visible while providing a compact embedded baseline and a concrete supervised recalibration strategy.

What changed

Canon before

Dry neck EMG sensing and the BioGAP-Ultra hardware were presented in prior work. Wearable SSI reports often use small vocabularies and offline classifiers without testing electrode repositioning or reporting integrated energy.

Delta from canon

Extends the prior neckband study with four-person multiday recordings, a small CNN, next-batch fine-tuning and GAP9 deployment benchmarks.

Position in field

Personalized non-invasive EMG command recognition with on-device inference and dry neck sensing; distinct from open-vocabulary speech transcription or synthesis.

Evidence

“ Four subjects each provide three multiday sessions with neckband repositioning, containing eight prompted commands in vocalized and silent modes. Classification includes a ninth rest class, which is downsampled for balance. ”

validation_scope · Sections III-B/C; PDF pp. 3-4 · confidence 0.99

“ SpeechNet uses temporal then spatial CNN blocks with 15,489 parameters; the dry neckband hardware extends prior conference work. GAP9 deployment uses 8-bit post-training quantization. ”

actual_novelty · Sections III-A/E/G; Table I; PDF pp. 3-5 · confidence 0.99

“ Silent balanced accuracy is 77.5 +/- 6.6% for pooled-day held-out batches and 59.3 +/- 2.2% for held-out sessions, an 18.2 percentage-point gap. Corresponding Random Forest values are 62.9% and 48.7%. ”

metric · Table II; PDF p. 6 · confidence 0.99

“ The Global protocol trains and tests on separate batches from all three days. The Inter-Session protocol holds one day out. Models are explicitly subject-specific and condition-specific. ”

validation_scope · Section III-D; PDF p. 4 · confidence 0.99

“ The text reports that one 1.4-second-window adaptation round improves silent accuracy from about 58% to 72%, with roughly 20 repetitions per word and about 10 minutes of acquisition. The protocol evaluates the following unseen batch. ”

metric · Sections III-D and IV-D; Figures 7/8; PDF pp. 4 and 7-8 · confidence 0.98

“ For an 800 ms input, GAP9 achieves 2.47 ms and 63.9 microjoules per inference at 240 MHz/0.65 V, or 1.60 ms and 97.1 microjoules at 370 MHz/0.8 V. ”

metric · Table III; Section IV-E; PDF pp. 7-8 · confidence 0.99

“ A 100 ms sliding-window stride gives reported total power 20.5 mW: 15.3 mW acquisition, 4.55 mW data handling/result streaming and 0.639 mW compute. Battery life of 27.1 h with 150 mAh is explicitly estimated. ”

metric · Section IV-E; PDF p. 8 · confidence 0.99

“ Online validation under real-time closed-loop user interaction and adaptive recalibration is explicitly left for future work. Reviewer assessment: embedded latency does not establish operational communication accuracy or usability. ”

limitation · Section V; PDF p. 10 · confidence 0.99

“ ITR uses class accuracy and input-window duration; reported silent peak is 69.2 bit/min at 800 ms. Reviewer assessment: this analytical rate omits the cued protocol, rest intervals and error correction and is not measured user throughput. ”

limitation · Sections III-B/F and IV-C; PDF pp. 3-6 · confidence 0.99

“ Offline preprocessing uses a zero-phase Butterworth high-pass filter; acquisition labels follow visual triggers with acknowledged reaction-time mismatch. Reviewer assessment: causal online filter/segmentation parity and post-quantization accuracy need explicit verification. ”

limitation · Sections III-C/G and V; PDF pp. 4-5 and 9-10 · confidence 0.99

“ Figures 7/8 use five tested batches with the first batch preceding adaptation, leaving four sequential updates in the specified protocol; the text nevertheless describes five fine-tuning rounds and Figure 8 repeats a vocalized caption. Reviewer assessment: update counts and condition labeling need correction. ”

limitation · Sections III-D and IV-D; Figures 7/8; PDF pp. 4 and 7-8 · confidence 0.99

Limits

Technical limits

Offline preprocessing uses a zero-phase 20 Hz high-pass filter plus 50 Hz notch; causal deployment equivalence is not established. Trigger timing imperfectly aligns articulation. Quantization changes recognition accuracy by an unreported amount. Whole-window ITR omits actual cue/rest/correction overhead.

Evaluation limits

Four subjects and eight commands plus rest. Global evaluation mixes every recording day into train and test through different batches; only leave-one-session-out isolates a new day. All models remain subject- and mode-specific. Rest is downsampled to class balance, unlike typical idle-heavy use. Online closed-loop validation is explicitly future work.

Deployment limits

Personalized fixed-command models, cued seated data and sensor-fit sensitivity. No closed-loop user study, walking/activity false-trigger study, patient cohort or measured all-day endurance. Recalibration is trained offline rather than demonstrated on-device.

Scope limits

Four seated participants, eight prompted English commands plus rest, three days and separate models for vocalized/silent conditions. No unseen words, sentences, patient communication or naturalistic continuous use.