EMG
This page groups the current SSI review database by the real `modality:` tag `modality:emg`.
The list below includes every paper page that currently carries this technique label.
Papers
AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration
耳周囲cEEGridで25の固定中国語文を認識。テスト日較正なしの別日ホルドアウトで平均92.24%、21日以上後のライブ250試行で98.00%。開語彙会話や患者適用は未実証。
Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding
標準化8ch顔/頸部sEMG・閉じた50文・27人LOSOで、多被験者事前学習+対象微調整は21.7% CER / 31.9% WER。3分キャリブレーションは約13分と有意差なし(20.5%/31.7%)。未見文では78.6% CERまで崩壊。開語彙や臨床完成ではない。
Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition
A strong touch-to-activate wearable EMG prototype with high closed-vocabulary accuracy, but the three-participant, speaker-dependent evaluation does not establish broad generalization or deployment readiness.
CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses
A smaller EMG-to-speech model with a modest reported WER gain; large acoustic-metric gains are confounded by reference-based alignment applied asymmetrically.
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
The paper advances silent speech synthesis by leveraging masked training to robustly fuse electromyography and lipreading, showing improved performance and resilience, but adaptation to laryngectomized users remains challenging.
A 1000-hour EEG-EMG-audio dataset of Japanese speech production
A 1020-hour multimodal EEG-EMG-audio dataset for Japanese overt speech vastly expands data resources, enabling diverse speech decoding and EEG research, though generalization is limited by three participants and no decoding benchmarks are presented.
Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features
Articulatory features better predict aligned muscle envelopes across speech modes; this supports representation choice, while actual EMG-to-speech decoding gains remain untested.
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
Useful evidence that prompted silent articulation carries affect cues, with silent-only AUC 0.829 within a person; weak unseen-speaker transfer and inseparable facial-expression effects limit deployment claims.
SilentWear: an Ultra-Low Power Wearable System for EMG-based Silent Speech Recognition
A useful dry-neckband and embedded-CNN study: silent balanced accuracy falls from 77.5% across pooled-day batches to 59.3% on a new day; 2.47 ms is compute time, while closed-loop usability remains untested.
EMG-to-Speech with Fewer Channels
Exhaustive subset search shows useful channel complementarity, and full-channel pretraining helps smaller EMG inputs. Single-person evaluation, unclear selection independence and a dropout text/figure conflict limit layout recommendations.
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
Combines one users silent EMG with face-conditioned target voices without inference audio. Pitch flattening modestly helps silent word accuracy, but multi-user decoding and faithful personal-voice recovery are not demonstrated.
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding
脳波と筋電を組み合わせ、声を出さずに発音した中国語の四声を分類する研究。学習に含まない人で平均85.10%を報告するが、文章認識ではなく、指標名や分割・チャネル選択手順には確認が必要。
Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication
Held-out sentence reconstruction is demonstrated in personalized EEG/EMG experiments, but the strongest aggregate evidence is overt/whispered phoneme decoding—not unrestricted imagined-speech communication.
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
本人の録音を学習に使わず、顔・首の筋電から音声を作る研究。ALS参加者1人の無声発話を音声化したが、聞き取りの単語誤り率は50.82%で、日常会話の回復や長期安定性は未実証。
A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband
完全乾式の首輪型筋電計測で、健康な1人の無声8語分類は68±3%。装着し直した未学習セッションでは54±7%に下がる。22.2mWは計測・無線通信の値で、自由会話や端末内実時間認識の完成を示すものではない。
From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach
筋電から合成した音声の文字起こしをTransformerとGPT-2で修正し、約100発話で単語誤り率36%→30%を報告する。音声そのものの聞き取りやすさや日常利用の実証ではなく、学習・分割・意味保持の検証は不足している。
A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations
電極配置の違う脳波・筋電データの統合学習は64語分類を改善するが、患者1人の本人別評価であり、別日・自由な会話・臨床効果への隔たりが残る。
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
音声から作った合成筋電を選別して実データと混ぜ、発声時の実筋電で単語誤り率23.30%に対し21.87%を報告する。実筋電は1人で、1532人は合成元の音声話者。無声発話への有効性は未実証で、図と本文の不一致も残る。
Knowledge Distilled Ensemble Model for sEMG-based Silent Speech Interface
This paper delivers a practical spelling-focused sEMG silent speech system by compressing a ResNet ensemble into a lightweight model achieving 85.9% accuracy on the NATO alphabet with portable hardware, but remains limited to 5 young male subjects and speaker-dependent scenarios.
Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language
SSRNet innovatively applies duration-aware Seq2Seq modeling and tonal multitask learning to reconstruct intelligible Mandarin speech from facial sEMG signals, markedly improving performance over prior methods but remains speaker-dependent with limited deployment evaluation.
An Improved Model for Voicing Silent Speech
This paper substantially improves open-vocabulary silent speech voicing using learned convolutional EMG features, Transformer modeling, and phoneme supervision, reducing WER from 68.0% to 42.2% automatic and 32.3% human in a single-speaker lab setting.
Digital Voicing of Silent Speech
Core EMG SSI paper with real gains from target transfer.