← Home

Silent speech research 154 reviewed pages 0 imported corpus pages 113 citation-linked pages

Silent speech papers by publication data.

Silent speech interfaces (SSI) let people communicate without vocalizing, using sensors on the face, throat, or mouth instead of a microphone to recognize intended speech.

サイレントスピーチ(無声発話)インタフェースは、発声せずに口・喉・顔の動きをセンサーで読み取り、 意図した発話を認識する技術です。

Browse the SSI review database by year, citation count, title, author, or page type.

Paper pages are expert evaluations, not abstract reposts. Citation counts come from OpenAlex when available. The latest refresh attempt was 2026-09-12; 104 counts were refreshed this run. Existing verified values are retained when an update fails.

Counts are available for 113 of 154 papers; each available count shows its retrieval date. “Not yet fetched” is different from a confirmed count of zero.

New to the category? Start with the silent speech interfaces hub for sensing routes and Kimura’s own lineage, the question map for short Q→A, or flagship reviews (SilentSpeller, SottoVoce, NasoVoce), then camera VSR AKVSR or CHI When Help Hurts, then return here for the multi-author review database.

Browse papers

154 papers shown

arXiv reviewed not yet fetched

AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration

Xiran Xu, Mochu Dong, Yujie Yan, Chenxi Wang, Yu Jiao, Jing Chen

耳周囲cEEGridで25の固定中国語文を認識。テスト日較正なしの別日ホルドアウトで平均92.24%、21日以上後のライブ250試行で98.00%。開語彙会話や患者適用は未実証。

BibTeX
@misc{aessi-an-around-ear-silent-speech-interface-for-cross-day-online-reuse-without-test-day-calibrat,
  title = {AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration},
  author = {Xiran Xu and Mochu Dong and Yujie Yan and Chenxi Wang and Yu Jiao and Jing Chen},
  year = {2026},
  note = {arXiv},
  eprint = {2609.21436},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.21436v1},
}
arXiv reviewed not yet fetched

Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

Chenqian Le, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Tianyu He, Nikasadat Emami, Adeen Flinker, Yao Wang

標準化8ch顔/頸部sEMG・閉じた50文・27人LOSOで、多被験者事前学習+対象微調整は21.7% CER / 31.9% WER。3分キャリブレーションは約13分と有意差なし(20.5%/31.7%)。未見文では78.6% CERまで崩壊。開語彙や臨床完成ではない。

BibTeX
@misc{multi-subject-pretraining-enables-short-calibration-personalization-for-closed-corpus-surface-emg-speech-decoding,
  title = {Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding},
  author = {Chenqian Le and Beatrice Fumagalli and Yasamin Esmaeili and Xupeng Chen and Tianyu He and Nikasadat Emami and Adeen Flinker and Yao Wang},
  year = {2026},
  note = {arXiv},
  eprint = {2609.21288},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2609.21288v1},
}
arXiv reviewed not yet fetched

AVSRBench: A Multi-Condition AVSR Benchmark

Rishabh Jain, Naomi Harte

LRS3のsub-1% AV WERは放送ドメインの指標。六条件比較では視覚のみが領域外で崩壊し、AV融合の明確な利点は主にLombard。RoomReader-AV(6.49h・10,324発話・118人)は会議会話の厳しさを示す。装着SSIの代替ではない。

BibTeX
@misc{avsrbench-a-multi-condition-avsr-benchmark,
  title = {AVSRBench: A Multi-Condition AVSR Benchmark},
  author = {Rishabh Jain and Naomi Harte},
  year = {2026},
  note = {arXiv},
  eprint = {2609.10366},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2609.10366v1},
}
arXiv reviewed 0 citations · 2026-09-12

Brain2Speech-Net: Intelligible, Real-Time Brain-to-Speech Synthesis Without Text Decoding

Shreeram Suresh Chandra, Zexin Cai, Yu Tsao, Simon King, Berrak Sisman

A useful text-free inference path for attempted-speech synthesis, with faster-than-real-time throughput but 47.2% unseen-utterance WER and no demonstrated causal streaming.

BibTeX
@misc{brain2speech-net-intelligible-real-time-brain-to-speech-synthesis-without-text-decoding,
  title = {Brain2Speech-Net: Intelligible, Real-Time Brain-to-Speech Synthesis Without Text Decoding},
  author = {Shreeram Suresh Chandra and Zexin Cai and Yu Tsao and Simon King and Berrak Sisman},
  year = {2026},
  note = {arXiv},
  eprint = {2609.04455},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2609.04455v1},
}
arXiv reviewed not yet fetched

Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition

Yuta Kurotaki, Shusuke Yamakoshi, Reitaro Yoshida, Yutaka Isoda, Tamami Takano, Yuji Isano, Yusuke Miyake, Kentaro Kuribayashi, Hiroki Ota

A strong touch-to-activate wearable EMG prototype with high closed-vocabulary accuracy, but the three-participant, speaker-dependent evaluation does not establish broad generalization or deployment readiness.

BibTeX
@misc{soft-active-electromyography-interface-for-machine-learning-enabled-silent-speech-recognition,
  title = {Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition},
  author = {Yuta Kurotaki and Shusuke Yamakoshi and Reitaro Yoshida and Yutaka Isoda and Tamami Takano and Yuji Isano and Yusuke Miyake and Kentaro Kuribayashi and Hiroki Ota},
  year = {2026},
  note = {arXiv},
  eprint = {2608.27048},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2608.27048v1},
}
arXiv reviewed not yet fetched

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

Ruidong Zhang, Jiacheng Liu, François Guimbretière, Cheng Zhang

A useful single-speaker open-vocabulary eyewear dataset: mixed voiced/silent training reaches 26.3% silent WER, but population and everyday-environment generalization remain untested.

BibTeX
@misc{sonispeech-a-large-scale-open-vocabulary-tri-modal-dataset-for-wearable-silent-speech-interfaces,
  title = {SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces},
  author = {Ruidong Zhang and Jiacheng Liu and François Guimbretière and Cheng Zhang},
  year = {2026},
  note = {arXiv},
  eprint = {2608.00803},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2608.00803v1},
}
arXiv reviewed not yet fetched

Do EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding

Owais Mujtaba Khanday, Mohamed Baha Ben Ticha, Sanae Belfrouh, Marc Ouellet, Jose A. Gonzalez-Lopez

General EEG pretraining shows no consistent speech-decoding advantage here; word decoding remains weak, and test-subject-informed early stopping limits the unseen-user claim.

BibTeX
@misc{do-eeg-foundation-models-transfer-to-speech-a-benchmark-on-overt-and-imagined-speech-decoding,
  title = {Do EEG Foundation Models Transfer to Speech? A Benchmark on Overt and Imagined Speech Decoding},
  author = {Owais Mujtaba Khanday and Mohamed Baha Ben Ticha and Sanae Belfrouh and Marc Ouellet and Jose A. Gonzalez-Lopez},
  year = {2026},
  note = {arXiv},
  eprint = {2607.27268},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2607.27268v2},
}
arXiv reviewed 0 citations · 2026-09-12

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

Sajid Fardin Dipto, Tarikul Islam Tamiti, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua

A smaller EMG-to-speech model with a modest reported WER gain; large acoustic-metric gains are confounded by reference-based alignment applied asymmetrically.

BibTeX
@misc{cs-ets-chaos-inspired-samba-based-emg-to-speech-synthesis-with-nonlinear-chaotic-losses,
  title = {CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses},
  author = {Sajid Fardin Dipto and Tarikul Islam Tamiti and David Vergano and Luke Baja-Ricketts and Anomadarshi Barua},
  year = {2026},
  note = {arXiv},
  eprint = {2607.18629},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2607.18629v1},
}
arXiv reviewed not yet fetched

Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech

Benjamin Ballyk, Teyun Kwon, Miran Özdogan, Oiwi Parker Jones

Introducing PNA, the paper advances non-invasive brain-to-speech decoding by creating artifact-informed augmentations via ICA, significantly improving imagined speech classification accuracy on MEG data when combined with trial averaging.

BibTeX
@misc{physiological-noise-augmentation-improves-non-invasive-brain-to-speech,
  title = {Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech},
  author = {Benjamin Ballyk and Teyun Kwon and Miran Özdogan and Oiwi Parker Jones},
  year = {2026},
  note = {arXiv},
  eprint = {2607.05165},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2607.05165v1},
}
arXiv reviewed not yet fetched

EEG-Based Imagined Speech Decoding Using a Hybrid CNN-SNN Architecture

Fatima Shalhoub, Mariam Al Mawla, Kabalan Chaccour, Iván López-Espejo, Hoda Fares

Promising five-class EEG classification at a reported 80.13% accuracy; independent replication, a matched spiking ablation, and actual power and online tests remain necessary.

BibTeX
@misc{eeg-based-imagined-speech-decoding-using-a-hybrid-cnn-snn-architecture,
  title = {EEG-Based Imagined Speech Decoding Using a Hybrid CNN-SNN Architecture},
  author = {Fatima Shalhoub and Mariam Al Mawla and Kabalan Chaccour and Iván López-Espejo and Hoda Fares},
  year = {2026},
  note = {arXiv},
  eprint = {2607.03844},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2607.03844v1},
}
2026 reviewed not yet fetched

A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR

Lin Chen, Jingping Fang, Hairui Liu, Chenyang Xu, Junhao Chen, Xiaorui Li, Weidong Cai, Xiaoming Chen

イベントストリームのマルチ話者VSR。DVS-LipでWER 22.3%・VER 19.8%・240 ms。カメラVTPとも装着SSIともセンサが異なる。

BibTeX
@misc{lipsflow-a-first-exploration-of-neuromorphic-ot-cfm-for-multi-speaker-vsr,
  title = {A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR},
  author = {Lin Chen and Jingping Fang and Hairui Liu and Chenyang Xu and Junhao Chen and Xiaorui Li and Weidong Cai and Xiaoming Chen},
  year = {2026},
  note = {ECCV 2026},
  doi = {10.1007/978-3-032-37281-9_8},
  eprint = {2606.31225},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1007/978-3-032-37281-9_8},
}
arXiv reviewed 0 citations · 2026-09-12

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

Eder del Blanco, David Gimeno-Gómez, Eva Navas, Carlos-D Martínez-Hinarejos, Inma Hernáez

The paper advances silent speech synthesis by leveraging masked training to robustly fuse electromyography and lipreading, showing improved performance and resilience, but adaptation to laryngectomized users remains challenging.

BibTeX
@misc{cross-modal-masking-for-robust-silent-speech-synthesis-using-semg-and-lipreading,
  title = {Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading},
  author = {Eder del Blanco and David Gimeno-Gómez and Eva Navas and Carlos-D Martínez-Hinarejos and Inma Hernáez},
  year = {2026},
  note = {arXiv},
  eprint = {2606.09667},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2606.09667v1},
}
arXiv reviewed 0 citations · 2026-09-12

A 1000-hour EEG-EMG-audio dataset of Japanese speech production

Motoshige Sato, Ilya Horiguchi, Masakazu Inoue, Kenichi Tomeoka, Eri Hatakeyama, Yuya Kita, Atsushi Yamamoto, Ippei Fujisawa, Shuntaro Sasai

A 1020-hour multimodal EEG-EMG-audio dataset for Japanese overt speech vastly expands data resources, enabling diverse speech decoding and EEG research, though generalization is limited by three participants and no decoding benchmarks are presented.

BibTeX
@misc{a-1000-hour-eeg-emg-audio-dataset-of-japanese-speech-production,
  title = {A 1000-hour EEG-EMG-audio dataset of Japanese speech production},
  author = {Motoshige Sato and Ilya Horiguchi and Masakazu Inoue and Kenichi Tomeoka and Eri Hatakeyama and Yuya Kita and Atsushi Yamamoto and Ippei Fujisawa and Shuntaro Sasai},
  year = {2026},
  note = {arXiv},
  eprint = {2606.01264},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2606.01264v1},
}
arXiv reviewed not yet fetched

Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping

Maryam Maghsoudi, Shihab Shamma

The study convincingly shows zero-shot imagined speech decoding by mapping MEG imagery to listened responses and decoding with a listened-trained contrastive model, marking a promising data-efficient advance despite limited vocabulary and hardware constraints.

BibTeX
@misc{zero-shot-imagined-speech-decoding-via-imagined-to-listened-meg-mapping,
  title = {Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping},
  author = {Maryam Maghsoudi and Shihab Shamma},
  year = {2026},
  note = {arXiv},
  eprint = {2605.08075},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2605.08075v1},
}
arXiv reviewed 0 citations · 2026-09-12

Articulatory movements influence electromagnetic wave transmission through the vocal tract

Rémi Blandin, Martin Laabs, Rudolf von Bünau, Bryn Lloyd, Silvia Farcito, Denys Nikolayev, Gabriela Hossu, Peter Birkholz, Dirk Plettemeier

A useful two-person physical model of contact-RF articulation sensing; qualitative agreement below about 3 GHz does not establish recognition accuracy or general-user robustness.

BibTeX
@misc{articulatory-movements-influence-electromagnetic-wave-transmission-through-the-vocal-tract,
  title = {Articulatory movements influence electromagnetic wave transmission through the vocal tract},
  author = {Rémi Blandin and Martin Laabs and Rudolf von Bünau and Bryn Lloyd and Silvia Farcito and Denys Nikolayev and Gabriela Hossu and Peter Birkholz and Dirk Plettemeier},
  year = {2026},
  note = {arXiv},
  eprint = {2604.19362},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2604.19362v3},
}
arXiv reviewed not yet fetched

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

Chenqian Le, Ruisi Li, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Amirhossein Khalilian-Gourtani, Tianyu He, Adeen Flinker, Yao Wang

Articulatory features better predict aligned muscle envelopes across speech modes; this supports representation choice, while actual EMG-to-speech decoding gains remain untested.

BibTeX
@misc{comparison-of-semg-encoding-accuracy-across-speech-modes-using-articulatory-and-phoneme-features,
  title = {Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features},
  author = {Chenqian Le and Ruisi Li and Beatrice Fumagalli and Yasamin Esmaeili and Xupeng Chen and Amirhossein Khalilian-Gourtani and Tianyu He and Adeen Flinker and Yao Wang},
  year = {2026},
  note = {arXiv},
  eprint = {2604.18920},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2604.18920v2},
}
arXiv reviewed 0 citations · 2026-09-12

iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding

Yoonmin Cha, Dawit Chun, Sung Goon Park

A promising decoder and gaze-confirmation concept, with reported 7.86% phoneme error but no clearly independent final test, complete latency measurement or user validation of the interface.

BibTeX
@misc{iphoneme-brain-to-text-communication-for-als-using-conformerxl-decoding,
  title = {iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding},
  author = {Yoonmin Cha and Dawit Chun and Sung Goon Park},
  year = {2026},
  note = {arXiv},
  eprint = {2604.16441},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2604.16441v1},
}
arXiv reviewed not yet fetched

Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction

Mohammed Salah Al-Radhi, Géza Németh, Andon Tchechmedjiev, Binbin Xu

Encouraging reported reconstruction metrics, with substantial feature-source and evaluation ambiguities; neither neural prosody recovery nor human intelligibility is firmly established.

BibTeX
@misc{brain-to-speech-prosody-feature-engineering-and-transformer-based-reconstruction,
  title = {Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction},
  author = {Mohammed Salah Al-Radhi and Géza Németh and Andon Tchechmedjiev and Binbin Xu},
  year = {2026},
  note = {arXiv},
  eprint = {2604.05751},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2604.05751v1},
}
arXiv reviewed 1 citations · 2026-09-12

Toward Robust, Reproducible, and Widely Accessible Intracranial Language Brain-Computer Interfaces: A Comprehensive Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions

Dongyi He, Wai Ting Siok, Nizhuan Wang

A useful systems-level agenda for speech neuroprostheses, but its composite benchmark is unvalidated and several source attributions and mechanism claims need correction before reuse.

BibTeX
@misc{toward-robust-reproducible-and-widely-accessible-intracranial-language-brain-computer-interfaces,
  title = {Toward Robust, Reproducible, and Widely Accessible Intracranial Language Brain-Computer Interfaces: A Comprehensive Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions},
  author = {Dongyi He and Wai Ting Siok and Nizhuan Wang},
  year = {2026},
  note = {arXiv},
  eprint = {2603.12279},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2603.12279v2},
}
arXiv reviewed not yet fetched

Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review

Kele Xu, Yifan Wang, Ming Feng, Qisheng Xu, Wuyang Chen, Yutao Dou, Cheng Yang, Huaimin Wang

A broad and useful SSI taxonomy, but missing review-selection detail and concrete citation errors make its benchmark tables unsuitable for unverified rankings or deployment claims.

BibTeX
@misc{silent-speech-interfaces-in-the-era-of-large-language-models-a-comprehensive-taxonomy-and-system,
  title = {Silent Speech Interfaces in the Era of Large Language Models: A Comprehensive Taxonomy and Systematic Review},
  author = {Kele Xu and Yifan Wang and Ming Feng and Qisheng Xu and Wuyang Chen and Yutao Dou and Cheng Yang and Huaimin Wang},
  year = {2026},
  note = {arXiv},
  eprint = {2603.11877},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2603.11877v1},
}
arXiv reviewed 0 citations · 2026-09-12

Affect Decoding in Phonated and Silent Speech Production from Surface EMG

Simon Pistrosch, Kleanthis Avramidis, Zhao Ren, Tiantian Feng, Jihwan Lee, Monica Gonzalez-Machorro, A. Batliner, Tanja Schultz, Shrikanth Narayanan, Björn W. Schuller

Useful evidence that prompted silent articulation carries affect cues, with silent-only AUC 0.829 within a person; weak unseen-speaker transfer and inseparable facial-expression effects limit deployment claims.

BibTeX
@misc{affect-decoding-in-phonated-and-silent-speech-production-from-surface-emg,
  title = {Affect Decoding in Phonated and Silent Speech Production from Surface EMG},
  author = {Simon Pistrosch and Kleanthis Avramidis and Zhao Ren and Tiantian Feng and Jihwan Lee and Monica Gonzalez-Machorro and A. Batliner and Tanja Schultz and Shrikanth Narayanan and Björn W. Schuller},
  year = {2026},
  note = {arXiv},
  eprint = {2603.11715},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2603.11715v2},
}
CHI 2026 reviewed 1 citations · 2026-09-12

NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction

Jun Rekimoto, Yu Nishimura, Bojian Yang

A strong deployment-focused speech interface leveraging a novel nose-pad dual-sensor configuration and multimodal fusion to enable robust low-audibility speech interaction with AI under noise, backed by extensive evaluation.

BibTeX
@misc{nasovoce,
  title = {NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction},
  author = {Jun Rekimoto and Yu Nishimura and Bojian Yang},
  year = {2026},
  note = {CHI '26 / arXiv},
  doi = {10.1145/3772318.3791397},
  eprint = {2603.10324},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3772318.3791397},
}
arXiv reviewed 0 citations · 2026-09-12

SilentWear: an Ultra-Low Power Wearable System for EMG-based Silent Speech Recognition

Giusy Spacone, Sebastian Frey, Giovanni Pollo, Alessio Burrello, Daniele Jahier Pagliari, Victor Kartsch, Andrea Cossettini, Luca Benini

A useful dry-neckband and embedded-CNN study: silent balanced accuracy falls from 77.5% across pooled-day batches to 59.3% on a new day; 2.47 ms is compute time, while closed-loop usability remains untested.

BibTeX
@misc{silentwear-an-ultra-low-power-wearable-system-for-emg-based-silent-speech-recognition,
  title = {SilentWear: an Ultra-Low Power Wearable System for EMG-based Silent Speech Recognition},
  author = {Giusy Spacone and Sebastian Frey and Giovanni Pollo and Alessio Burrello and Daniele Jahier Pagliari and Victor Kartsch and Andrea Cossettini and Luca Benini},
  year = {2026},
  note = {arXiv},
  eprint = {2603.02847},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2603.02847v2},
}
arXiv reviewed 0 citations · 2026-09-12

Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech

Maryam Maghsoudi, Rupesh Chillale, Shihab A. Shamma

A useful one-participant cross-mode analysis showing why envelope correlation needs sentence-discrimination checks; incomplete split/alignment reporting and mixed model results limit the stronger conclusions.

BibTeX
@misc{relating-the-neural-representations-of-vocalized-mimed-and-imagined-speech,
  title = {Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech},
  author = {Maryam Maghsoudi and Rupesh Chillale and Shihab A. Shamma},
  year = {2026},
  note = {arXiv},
  eprint = {2602.22597},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.22597v1},
}
arXiv reviewed not yet fetched

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

Yifan Liang, Andong Li, Kang Yang, Guochen Yu, Fangkun Liu, Lingling Dai, Xiaodong Li, Chengshi Zheng

Improves lip-to-speech naturalness with a structured codec-latent flow model, but word accuracy and speaker similarity trail stronger baselines; reference audio and unmeasured runtime limit broader claims.

BibTeX
@misc{sld-l2s-hierarchical-subspace-latent-diffusion-for-high-fidelity-lip-to-speech-synthesis,
  title = {SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis},
  author = {Yifan Liang and Andong Li and Kang Yang and Guochen Yu and Fangkun Liu and Lingling Dai and Xiaodong Li and Chengshi Zheng},
  year = {2026},
  note = {arXiv},
  eprint = {2602.11477},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.11477v1},
}
arXiv reviewed not yet fetched

EMG-to-Speech with Fewer Channels

Injune Hwang, Jaejun Lee, Kyogu Lee

Exhaustive subset search shows useful channel complementarity, and full-channel pretraining helps smaller EMG inputs. Single-person evaluation, unclear selection independence and a dropout text/figure conflict limit layout recommendations.

BibTeX
@misc{emg-to-speech-with-fewer-channels,
  title = {EMG-to-Speech with Fewer Channels},
  author = {Injune Hwang and Jaejun Lee and Kyogu Lee},
  year = {2026},
  note = {arXiv},
  eprint = {2602.06460},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.06460v1},
}
arXiv reviewed not yet fetched

LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency

Jaejun Lee, Yoori Oh, Kyogu Lee

Improves reference-prosody metrics and earns 54.22% listener preference, but does not improve deployed WER; oracle audio cues, normalization changes and limited perceptual reporting qualify broader claims.

BibTeX
@misc{lipsody-lip-to-speech-synthesis-with-enhanced-prosody-consistency,
  title = {LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency},
  author = {Jaejun Lee and Yoori Oh and Kyogu Lee},
  year = {2026},
  note = {arXiv},
  eprint = {2602.01908},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.01908v1},
}
arXiv reviewed not yet fetched

Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only

Jaejun Lee, Yoori Oh, Kyogu Lee

Combines one users silent EMG with face-conditioned target voices without inference audio. Pitch flattening modestly helps silent word accuracy, but multi-user decoding and faithful personal-voice recovery are not demonstrated.

BibTeX
@misc{speaking-without-sound-multi-speaker-silent-speech-voicing-with-facial-inputs-only,
  title = {Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only},
  author = {Jaejun Lee and Yoori Oh and Kyogu Lee},
  year = {2026},
  note = {arXiv},
  eprint = {2602.01879},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.01879v1},
}
arXiv reviewed not yet fetched

Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes

Maryam Maghsoudi, Ayushi Mishra

Offers useful activation-intervention diagnostics, but donor replacement bypasses recipient information, baseline/patch scores are unresolved, and winner counts exceed the stated layer width; strong causal conclusions are not established.

BibTeX
@misc{mechanistic-interpretability-of-brain-to-speech-models-across-speech-modes,
  title = {Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes},
  author = {Maryam Maghsoudi and Ayushi Mishra},
  year = {2026},
  note = {arXiv},
  eprint = {2602.01247},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2602.01247v1},
}
arXiv reviewed not yet fetched

Lip-Siri: Contactless Open-Sentence Silent Speech with Wi-Fi Backscatter

Ye Tian, Haohua Du, Chao Gu, Junyang Zhang, Shanyue Wang, Hao Zhou, Jiahui Hou, Xiang-Yang Li

A credible contactless Wi-Fi backscatter SSI prototype with 85.61% word accuracy and 36.87% sentence WER, but lexicon-composed sentences are not unseen-word recognition and deployment still requires SDR hardware plus manual boundaries.

BibTeX
@misc{lip-siri-contactless-open-sentence-silent-speech-with-wi-fi-backscatter,
  title = {Lip-Siri: Contactless Open-Sentence Silent Speech with Wi-Fi Backscatter},
  author = {Ye Tian and Haohua Du and Chao Gu and Junyang Zhang and Shanyue Wang and Hao Zhou and Jiahui Hou and Xiang-Yang Li},
  year = {2026},
  note = {arXiv},
  eprint = {2601.18177},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2601.18177v1},
}
arXiv reviewed not yet fetched

Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech

Soufiane Jhilal, Stéphanie Martin, Anne-Lise Giraud

ImageNet-pretrained vision models improve closed-set MEG classification, including held-out-subject tests, but the study decodes task/vowel labels—not sentences—and leaves practical communication undemonstrated.

BibTeX
@misc{transfer-learning-from-imagenet-for-meg-based-decoding-of-imagined-speech,
  title = {Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech},
  author = {Soufiane Jhilal and Stéphanie Martin and Anne-Lise Giraud},
  year = {2026},
  note = {arXiv},
  eprint = {2601.15909},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2601.15909v1},
}
arXiv reviewed not yet fetched

EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG

Hanbeot Park, Yunjeong Cho, Hunhee Kim

Subject-specific EEG reconstructs known cued utterances offline, but imagined-speech WER remains 47.48% at 2 s and 43.46% at 4 s before correction; unseen content, real-time use and fully alignment-free processing are not demonstrated.

BibTeX
@misc{eeg-to-voice-decoding-of-spoken-and-imagined-speech-using-non-invasive-eeg,
  title = {EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG},
  author = {Hanbeot Park and Yunjeong Cho and Hunhee Kim},
  year = {2025},
  note = {arXiv},
  eprint = {2512.22146},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2512.22146v1},
}
arXiv reviewed not yet fetched

Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning

Xuefu Dong, Liqiang Xu, Lixing He, Zengyi Han, Ken Christofferson, Yifei Chen, Akihito Taya, Yuuki Nishiyama, Kaoru Sezaki

HEar-ID jointly models ear-based spelling and identity, but all-user Top-1 is 67.3%, not the selected eight-user 90.25%; whisper dependence, participant failures, modified hardware, and untested attack resistance limit deployment claims.

BibTeX
@misc{poster-recognizing-hidden-in-the-ear-private-key-for-reliable-silent-speech-interface-using-mult,
  title = {Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning},
  author = {Xuefu Dong and Liqiang Xu and Lixing He and Zengyi Han and Ken Christofferson and Yifei Chen and Akihito Taya and Yuuki Nishiyama and Kaoru Sezaki},
  year = {2025},
  note = {arXiv},
  eprint = {2512.16518},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2512.16518v1},
}
arXiv reviewed not yet fetched

A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses

Maryam Maghsoudi, Mohsen Rezaeizadeh, Shihab Shamma

音楽・詩を想像したときの脳磁図(MEG)から、聴取時の脳活動を予測する基礎研究。学習から外した参加者でも本人の一部データで調整すると比較用モデルを上回るが、相関は小さく、音声や文章の復元・実用的な意思伝達は未実証である。

BibTeX
@misc{a-convolutional-framework-for-mapping-imagined-auditory-meg-into-listened-brain-responses,
  title = {A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses},
  author = {Maryam Maghsoudi and Mohsen Rezaeizadeh and Shihab Shamma},
  year = {2025},
  note = {arXiv},
  eprint = {2512.03458},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2512.03458v1},
}
arXiv reviewed not yet fetched

VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task

Yuyue Wang, Xin Cheng, Yihan Wu, Xihua Wang, Jinchuan Tian, Ruihua Song

文章と参考音声を使い、唇に合う音声を作る時間制御の研究。単語誤り率と同期指標は改善するが、唇だけから内容を読む方式ではなく、数値の不整合と無声利用の未検証が残る。

BibTeX
@misc{vspeechlm-a-visual-speech-language-model-for-visual-text-to-speech-task,
  title = {VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task},
  author = {Yuyue Wang and Xin Cheng and Yihan Wu and Xihua Wang and Jinchuan Tian and Ruihua Song},
  year = {2025},
  note = {arXiv},
  eprint = {2511.22229},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.22229v1},
}
arXiv reviewed 1 citations · 2026-09-12

A cross-species neural foundation model for end-to-end speech decoding

Zhang, Yizi, He, Linyang, Fan, Chaofei, Liu, Tingkai, Yu, Han, Le, Trung, Li, Jingyuan, Linderman, Scott, Duncker, Lea, Willett, Francis R, Mesgarani, Nima, Paninski, Liam

Introduces a cross-species pretrained transformer encoder enabling state-of-the-art end-to-end neural speech decoding with audio-LLMs, improving accuracy and enabling imagined speech decoding, but latency and real-time deployment remain challenges.

BibTeX
@misc{a-cross-species-neural-foundation-model-for-end-to-end-speech-decoding,
  title = {A cross-species neural foundation model for end-to-end speech decoding},
  author = {Zhang, Yizi and He, Linyang and Fan, Chaofei and Liu, Tingkai and Yu, Han and Le, Trung and Li, Jingyuan and Linderman, Scott and Duncker, Lea and Willett, Francis R and Mesgarani, Nima and Paninski, Liam},
  year = {2025},
  note = {arXiv},
  eprint = {2511.21740},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.21740v5},
}
arXiv reviewed not yet fetched

MultiDiffNet: A Multi-Objective Diffusion Framework for Generalizable Brain Decoding

Mengchun Zhang, Kateryna Shapovalenko, Yucheng Shao, Eddie Guo, Parusha Pradhan

想像発話11クラスの未学習者正解率は混合あり12.12%でEEGNetの10.61%から小幅改善。ただし本人の較正が必要で、拡散モデルなしの構成も上回るため、較正不要の実用的な発話認識とはいえない。

BibTeX
@misc{multidiffnet-a-multi-objective-diffusion-framework-for-generalizable-brain-decoding,
  title = {MultiDiffNet: A Multi-Objective Diffusion Framework for Generalizable Brain Decoding},
  author = {Mengchun Zhang and Kateryna Shapovalenko and Yucheng Shao and Eddie Guo and Parusha Pradhan},
  year = {2025},
  note = {arXiv},
  eprint = {2511.18294},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.18294v1},
}
arXiv reviewed not yet fetched

Subject-Independent Imagined Speech Detection via Cross-Subject Generalization and Calibration

Byung-Kwan Ko, Soowon Kim, Seo-Hyun Lee

6人の脳波で想像発話と休止を区別し、本人の学習用標本10%で較正すると正解率は67.0%から78.1%へ改善。単語の解読ではなく発話状態の検出であり、較正不要の汎化や実利用は未実証。

BibTeX
@misc{subject-independent-imagined-speech-detection-via-cross-subject-generalization-and-calibration,
  title = {Subject-Independent Imagined Speech Detection via Cross-Subject Generalization and Calibration},
  author = {Byung-Kwan Ko and Soowon Kim and Seo-Hyun Lee},
  year = {2025},
  note = {arXiv},
  eprint = {2511.13739},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.13739v1},
}
arXiv reviewed not yet fetched

CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding

Yifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou, Jiawei Ju

脳波と筋電を組み合わせ、声を出さずに発音した中国語の四声を分類する研究。学習に含まない人で平均85.10%を報告するが、文章認識ではなく、指標名や分割・チャネル選択手順には確認が必要。

BibTeX
@misc{cat-net-a-cross-attention-tone-network-for-cross-subject-eeg-emg-fusion-tone-decoding,
  title = {CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding},
  author = {Yifan Zhuang and Calvin Huang and Zepeng Yu and Yongjie Zou and Jiawei Ju},
  year = {2025},
  note = {arXiv},
  eprint = {2511.10935},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.10935v1},
}
arXiv reviewed not yet fetched

Toward Practical BCI: A Real-time Wireless Imagined Speech EEG Decoding System

Ji-Ha Park, Heon-Gyu Kwak, Gi-Hwan Shin, Yoo-In Jeon, Sun-Min Park, Ji-Yeon Hwang, Seong-Whan Lee

想像した4命令を脳波で分類する試作系を有線・無線で実装し、正解率は62.00%と46.67%。本人の較正が必要で、3人・分割不明の評価から日常利用や自由な文章の解読まで実証したとはいえない。

BibTeX
@misc{toward-practical-bci-a-real-time-wireless-imagined-speech-eeg-decoding-system,
  title = {Toward Practical BCI: A Real-time Wireless Imagined Speech EEG Decoding System},
  author = {Ji-Ha Park and Heon-Gyu Kwak and Gi-Hwan Shin and Yoo-In Jeon and Sun-Min Park and Ji-Yeon Hwang and Seong-Whan Lee},
  year = {2025},
  note = {arXiv},
  eprint = {2511.07936},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.07936v1},
}
arXiv reviewed not yet fetched

Lightweight Diffusion-based Framework for Online Imagined Speech Decoding in Aphasia

Eunyeong Ko, Soowon Kim, Ha-Na Jo

失語症のある1人で、想像する3語と休止の実時間分類を試作。著者報告は第1候補65%・上位2候補70%だが、クラス別値と集計が整合せず、性能の確定には原データの確認が必要。

BibTeX
@misc{lightweight-diffusion-based-framework-for-online-imagined-speech-decoding-in-aphasia,
  title = {Lightweight Diffusion-based Framework for Online Imagined Speech Decoding in Aphasia},
  author = {Eunyeong Ko and Soowon Kim and Ha-Na Jo},
  year = {2025},
  note = {arXiv},
  eprint = {2511.07920},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.07920v3},
}
arXiv reviewed not yet fetched

Distinct Theta Synchrony across Speech Modes: Perceived, Spoken, Whispered, and Imagined

Jung-Sun Lee, Ha-Na Jo, Eunyeong Ko

健康な10人の脳波で、聞く・有声発話・ささやき・想像発話の同期分布を比較する基礎研究。認識精度は測らず、有意差や予測性能を裏付ける詳細も不足している。

BibTeX
@misc{distinct-theta-synchrony-across-speech-modes-perceived-spoken-whispered-and-imagined,
  title = {Distinct Theta Synchrony across Speech Modes: Perceived, Spoken, Whispered, and Imagined},
  author = {Jung-Sun Lee and Ha-Na Jo and Eunyeong Ko},
  year = {2025},
  note = {arXiv},
  eprint = {2511.07918},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2511.07918v1},
}
arXiv reviewed not yet fetched

Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication

Deok-Seon Kim, Seo-Hyun Lee, Kang Yin, Seong-Whan Lee

Held-out sentence reconstruction is demonstrated in personalized EEG/EMG experiments, but the strongest aggregate evidence is overt/whispered phoneme decoding—not unrestricted imagined-speech communication.

BibTeX
@misc{reconstructing-unseen-sentences-from-speech-related-biosignals-for-open-vocabulary-neural-commun,
  title = {Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication},
  author = {Deok-Seon Kim and Seo-Hyun Lee and Kang Yin and Seong-Whan Lee},
  year = {2025},
  note = {arXiv},
  eprint = {2510.27247},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2510.27247v1},
}
arXiv reviewed not yet fetched

emg2speech: Synthesizing speech from electromyography using self-supervised speech models

Harshavardhana T. Gowda, Daniel C. Comstock, Lee M. Miller

本人の録音を学習に使わず、顔・首の筋電から音声を作る研究。ALS参加者1人の無声発話を音声化したが、聞き取りの単語誤り率は50.82%で、日常会話の回復や長期安定性は未実証。

BibTeX
@misc{emg2speech-synthesizing-speech-from-electromyography-using-self-supervised-speech-models,
  title = {emg2speech: Synthesizing speech from electromyography using self-supervised speech models},
  author = {Harshavardhana T. Gowda and Daniel C. Comstock and Lee M. Miller},
  year = {2025},
  note = {arXiv},
  eprint = {2510.23969},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2510.23969v2},
}
arXiv reviewed not yet fetched

IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks

Sunghwa Lee, Jaewon Yu

唇の近くに置いた非接触レーダーで、1人の英単語50語を平均91.1%の正解率で分類し、従来方式の74.0%を上回った。別の人、未知語、自由な姿勢や日常会話での性能は未実証。

BibTeX
@misc{ir-uwb-radar-based-contactless-silent-speech-recognition-with-attention-enhanced-temporal-convol,
  title = {IR-UWB Radar-Based Contactless Silent Speech Recognition with Attention-Enhanced Temporal Convolutional Networks},
  author = {Sunghwa Lee and Jaewon Yu},
  year = {2025},
  note = {arXiv},
  eprint = {2509.26409},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2509.26409v1},
}
arXiv reviewed not yet fetched

NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training

Suli Wang, Yangshen Deng, Zhenghua Bao, Xinyu Zhan, Yiqun Duan

課題別の補助学習とテスト時適応により、本人内の5種類の想像発話分類を改善する研究。CBraModのBalanced Accuracyは58.98%だが、未知被験者への想像発話適用は偶然水準で、自由文や較正不要のSSIではない。

BibTeX
@misc{neurottt-bridging-pretraining-downstream-task-misalignment-in-eeg-foundation-models-via-test-tim,
  title = {NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training},
  author = {Suli Wang and Yangshen Deng and Zhenghua Bao and Xinyu Zhan and Yiqun Duan},
  year = {2025},
  note = {arXiv},
  eprint = {2509.26301},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2509.26301v2},
}
arXiv reviewed not yet fetched

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie, Chengshi Zheng

中国語の口元映像からの音声合成で、英語の事前学習と声調を意識した生成を組み合わせ、比較対象より文字・声調の誤りを減らした。ただし文字誤り率61.2%が残り、実際の無声発話や実時間会話の成功を示したわけではない。

BibTeX
@misc{lta-l2s-lexical-tone-aware-lip-to-speech-synthesis-for-mandarin-with-cross-lingual-transfer-lear,
  title = {LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning},
  author = {Kang Yang and Yifan Liang and Fangkun Liu and Zhenping Xie and Chengshi Zheng},
  year = {2025},
  note = {arXiv},
  eprint = {2509.25670},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2509.25670v1},
}
arXiv reviewed not yet fetched

A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband

Fiona Meier, Giusy Spacone, Sebastian Frey, Luca Benini, Andrea Cossettini

完全乾式の首輪型筋電計測で、健康な1人の無声8語分類は68±3%。装着し直した未学習セッションでは54±7%に下がる。22.2mWは計測・無線通信の値で、自由会話や端末内実時間認識の完成を示すものではない。

BibTeX
@misc{a-parallel-ultra-low-power-silent-speech-interface-based-on-a-wearable-fully-dry-emg-neckband,
  title = {A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband},
  author = {Fiona Meier and Giusy Spacone and Sebastian Frey and Luca Benini and Andrea Cossettini},
  year = {2025},
  note = {arXiv},
  eprint = {2509.21964},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2509.21964v1},
}
arXiv reviewed not yet fetched

From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

Nithyashree Sivasubramaniam

筋電から合成した音声の文字起こしをTransformerとGPT-2で修正し、約100発話で単語誤り率36%→30%を報告する。音声そのものの聞き取りやすさや日常利用の実証ではなく、学習・分割・意味保持の検証は不足している。

BibTeX
@misc{from-silent-signals-to-natural-language-a-dual-stage-transformer-llm-approach,
  title = {From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach},
  author = {Nithyashree Sivasubramaniam},
  year = {2025},
  note = {arXiv},
  eprint = {2509.04507},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2509.04507v1},
}
arXiv reviewed not yet fetched

An Introduction to Silent Paralinguistics

Zhao Ren, Simon Pistrosch, Buket Coşkun, Kevin Scheck, Anton Batliner, Björn W. Schuller, Tanja Schultz

声を出さない会話で、言葉だけでなく感情や話し方も伝えるための総説。研究課題の整理は有用だが、感情の復元精度や実用効果を新たに実証した論文ではない。

BibTeX
@misc{an-introduction-to-silent-paralinguistics,
  title = {An Introduction to Silent Paralinguistics},
  author = {Zhao Ren and Simon Pistrosch and Buket Coşkun and Kevin Scheck and Anton Batliner and Björn W. Schuller and Tanja Schultz},
  year = {2025},
  note = {arXiv},
  eprint = {2508.18127},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2508.18127v1},
}
arXiv reviewed not yet fetched

Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource

Lei Yang, Junshan Jin, Mingyuan Zhang, Yi He, Bofan Chen, Shilin Wang

唇の20点の動きを画像と組み合わせ、学習データ3分の1の500語読唇で正解率を85.75%から86.65%へ改善。ただしモデルは大型化し、実際の無声調音・自由文・実時間動作は未検証。

BibTeX
@misc{landmark-guided-visual-feature-extractor-for-visual-speech-recognition-with-limited-resource,
  title = {Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource},
  author = {Lei Yang and Junshan Jin and Mingyuan Zhang and Yi He and Bofan Chen and Shilin Wang},
  year = {2025},
  note = {arXiv},
  eprint = {2508.07233},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2508.07233v1},
}
arXiv reviewed not yet fetched

A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations

Masakazu Inoue, Motoshige Sato, Kenichi Tomeoka, Nathania Nah, Eri Hatakeyama, Kai Arulkumaran, Ilya Horiguchi, Shuntaro Sasai

電極配置の違う脳波・筋電データの統合学習は64語分類を改善するが、患者1人の本人別評価であり、別日・自由な会話・臨床効果への隔たりが残る。

BibTeX
@misc{a-silent-speech-decoding-system-from-eeg-and-emg-with-heterogenous-electrode-configurations,
  title = {A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations},
  author = {Masakazu Inoue and Motoshige Sato and Kenichi Tomeoka and Nathania Nah and Eri Hatakeyama and Kai Arulkumaran and Ilya Horiguchi and Shuntaro Sasai},
  year = {2025},
  note = {arXiv},
  eprint = {2506.13835},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2506.13835v1},
}
arXiv reviewed not yet fetched

Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling

Xiaodan Chen, Xiaoxue Gao, Mathias Quoy, Alexandre Pitti, Nancy F. Chen

音声から作った合成筋電を選別して実データと混ぜ、発声時の実筋電で単語誤り率23.30%に対し21.87%を報告する。実筋電は1人で、1532人は合成元の音声話者。無声発話への有効性は未実証で、図と本文の不一致も残る。

BibTeX
@misc{confidence-based-self-training-for-emg-to-speech-leveraging-synthetic-emg-for-robust-modeling,
  title = {Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling},
  author = {Xiaodan Chen and Xiaoxue Gao and Mathias Quoy and Alexandre Pitti and Nancy F. Chen},
  year = {2025},
  note = {arXiv},
  eprint = {2506.11862},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2506.11862v2},
}
arXiv reviewed 0 citations · 2026-09-12

SonicVisionLM: Playing Sound with Vision Language Models

Zhifeng Xie, Shengye Yu, Qile He, Mengtian Li

A high-quality video-to-audio generation framework leveraging vision-language models for editable, temporally precise sound effect generation; strong experimental validations but outside standard SSI scope.

BibTeX
@misc{sonicvisionlm-playing-sound-with-vision-language-models,
  title = {SonicVisionLM: Playing Sound with Vision Language Models},
  author = {Zhifeng Xie and Shengye Yu and Qile He and Mengtian Li},
  year = {2024},
  note = {arXiv / imported corpus page},
  eprint = {2401.04394},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2401.04394v1},
}
arXiv reviewed 9 citations · 2026-09-12

IR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrases

Sunghwa Lee, Younghoon Shin, Myungjong Kim, Jiwon Seo

This paper introduces FERASEC, a novel radar feature extraction enabling the first contactless IR-UWB radar phoneme-level silent speech recognition with 86% vowel and 81% consonant accuracy, surpassing raw signal baselines and signifying a key advance in practical silent speech interfaces.

BibTeX
@misc{ir-uwb-radar-based-contactless-silent-speech-recognition-of-vowels-consonants-words-and-phrases,
  title = {IR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrases},
  author = {Sunghwa Lee and Younghoon Shin and Myungjong Kim and Jiwon Seo},
  year = {2023},
  note = {arXiv / imported corpus page},
  doi = {10.1109/ACCESS.2023.3344177},
  eprint = {2312.09572},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/ACCESS.2023.3344177},
}
arXiv reviewed 82 citations · 2026-09-12

Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency

Chenyu Tang, Muzi Xu, Wentian Yi, Zibo Zhang, Edoardo Occhipinti, Chaoqun Dong, Dafydd Ravenscroft, Sung‐Min Jung, Sanghyo Lee, Shuo Gao, Jong Min Kim, Luigi G. Occhipinti

Strong SSI system combining a novel ultrasensitive throat textile strain sensor with an efficient 1D residual CNN, achieving high word classification accuracy with low computational cost and promising few-shot transfer to new users and words on small vocabularies.

BibTeX
@misc{ultrasensitive-textile-strain-sensors-redefine-wearable-silent-speech-interfaces-with-high-machine-learning-efficiency,
  title = {Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency},
  author = {Chenyu Tang and Muzi Xu and Wentian Yi and Zibo Zhang and Edoardo Occhipinti and Chaoqun Dong and Dafydd Ravenscroft and Sung‐Min Jung and Sanghyo Lee and Shuo Gao and Jong Min Kim and Luigi G. Occhipinti},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2311.15683},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2311.15683v1},
}
arXiv reviewed 1 citations · 2026-09-12

Distributed pressure matching strategy using diffusion adaptation

Mengfei Zhang, Junqing Zhang, Jie Chen, Cédric Richard

Distributed rootless pressure matching for personal sound zones is presented and validated in simulation, not an SSI paper.

BibTeX
@misc{distributed-pressure-matching-strategy-using-diffusion-adaptation,
  title = {Distributed pressure matching strategy using diffusion adaptation},
  author = {Mengfei Zhang and Junqing Zhang and Jie Chen and Cédric Richard},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2311.07729},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2311.07729v1},
}
arXiv reviewed 5739 citations · 2026-04-29

Advancing Test-Time Adaptation for Acoustic Foundation Models in Open-World Shifts

Andy Clark

Strong acoustic ASR paper proposing confidence-weighted frame adaptation plus temporal consistency regularization for stable test-time adaptation under wild acoustic conditions, yielding substantial WER improvements across noise, accents, and singing datasets.

BibTeX
@misc{advancing-test-time-adaptation-for-acoustic-foundation-models-in-open-world-shifts,
  title = {Advancing Test-Time Adaptation for Acoustic Foundation Models in Open-World Shifts},
  author = {Andy Clark},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2310.09505},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2310.09505v1},
}
arXiv reviewed 21 citations · 2026-09-12

Sound Source Localization is All about Cross-Modal Alignment

Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung

Provides a novel multi-positive contrastive framework enhancing semantic audio-visual alignment for sound source localization. Strong experimental evidence supports claims. Method is outside the SSI domain.

BibTeX
@misc{sound-source-localization-is-all-about-cross-modal-alignment,
  title = {Sound Source Localization is All about Cross-Modal Alignment},
  author = {Arda Senocak and Hyeonggon Ryu and Junsik Kim and Tae-Hyun Oh and Hanspeter Pfister and Joon Son Chung},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2309.10724},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2309.10724v1},
}
arXiv reviewed 4 citations · 2026-09-12

Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

Ji-Hoon Kim, Jaehun Kim, Joon Son Chung

Strong lip-to-speech system that reduces ambiguity via SSL linguistic conditioning, variance predictors, and flow-based refinement, achieving near-vocoded naturalness and improved intelligibility on standard datasets.

BibTeX
@misc{let-there-be-sound-reconstructing-high-quality-speech-from-silent-videos,
  title = {Let There Be Sound: Reconstructing High Quality Speech from Silent Videos},
  author = {Ji-Hoon Kim and Jaehun Kim and Joon Son Chung},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2308.15256},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2308.15256v2},
}
arXiv reviewed 0 citations · 2026-04-29

An Initial Exploration: Learning to Generate Realistic Audio for Silent Video

Matthew Martel, Jackson Wagner

Honest exploratory comparison showing transformer-based model outperforms deep-fusion CNN and Wavenet for generating low-to-mid frequency audio from silent video in a small curated dataset; not a speech or SSI paper.

BibTeX
@misc{an-initial-exploration-learning-to-generate-realistic-audio-for-silent-video,
  title = {An Initial Exploration: Learning to Generate Realistic Audio for Silent Video},
  author = {Matthew Martel and Jackson Wagner},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2308.12408},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2308.12408v1},
}
arXiv reviewed 12 citations · 2026-04-29

Audio Knowledge Empowered Visual Speech Recognition

Jeong Hun Yeo, Minsu Kim, Jeongsoo Choi, Dae Hoe Kim, Yong Man Ro

The paper advances visual speech recognition by selectively transferring refined linguistic audio knowledge via a learned compact memory and cross-attention injection, improving benchmark WERs over prior audio-assisted methods without requiring audio inputs during inference.

BibTeX
@misc{akvsr-audio-knowledge-empowered-visual-speech-recognition-by-compressing-audio-knowledge-of-a-pretrained-model,
  title = {Audio Knowledge Empowered Visual Speech Recognition},
  author = {Jeong Hun Yeo and Minsu Kim and Jeongsoo Choi and Dae Hoe Kim and Yong Man Ro},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2308.07593},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2308.07593v2},
}
arXiv reviewed 3 citations · 2026-09-12

Knowledge Distilled Ensemble Model for sEMG-based Silent Speech Interface

Wenqiang Lai, Qihan Yang, Mao Ye, Endong Sun, Jiangnan Ye

This paper delivers a practical spelling-focused sEMG silent speech system by compressing a ResNet ensemble into a lightweight model achieving 85.9% accuracy on the NATO alphabet with portable hardware, but remains limited to 5 young male subjects and speaker-dependent scenarios.

BibTeX
@misc{knowledge-distilled-ensemble-model-for-semg-based-silent-speech-interface,
  title = {Knowledge Distilled Ensemble Model for sEMG-based Silent Speech Interface},
  author = {Wenqiang Lai and Qihan Yang and Mao Ye and Endong Sun and Jiangnan Ye},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2308.06533},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2308.06533v1},
}
arXiv reviewed 7 citations · 2026-09-12

Automatically measuring speech fluency in people with aphasia: first achievements using read-speech data

Lionel Fontan, Typhanie Prince, Aleksandra Nowakowska, Halima Sahraoui, Silvia Martínez‐Ferreiro

Strong clinical fluency regression method validated on noisy read speech from aphasia patients; outside core SSI modalities and use-cases.

BibTeX
@misc{automatically-measuring-speech-fluency-in-people-with-aphasia-first-achievements-using-read-speech-data,
  title = {Automatically measuring speech fluency in people with aphasia: first achievements using read-speech data},
  author = {Lionel Fontan and Typhanie Prince and Aleksandra Nowakowska and Halima Sahraoui and Silvia Martínez‐Ferreiro},
  year = {2023},
  note = {arXiv / imported corpus page},
  doi = {10.1080/02687038.2023.2244728},
  eprint = {2308.04763},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1080/02687038.2023.2244728},
}
arXiv reviewed 10 citations · 2026-09-12

Exploring how a Generative AI interprets music

Gabriela Barenboim, Luigi Del Debbio, Johannes Hirn, Verónica Sanz

A thorough interpretability analysis reveals that MusicVAE uses only a few dozen latent dimensions to encode music with pitch and rhythm strongly represented in the first two, but the work has no direct relevance to silent speech interfaces.

BibTeX
@misc{exploring-how-a-generative-ai-interprets-music,
  title = {Exploring how a Generative AI interprets music},
  author = {Gabriela Barenboim and Luigi Del Debbio and Johannes Hirn and Verónica Sanz},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2308.00015},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2308.00015v1},
}
arXiv reviewed 2 citations · 2026-09-12

Audio-visual video-to-speech synthesis with synthesized input audio

Triantafyllos Kefalas, Yannis Panagakis, Maja Pantić

The paper credibly shows that incorporating synthesized audio as an auxiliary input in a second-stage audiovisual synthesis model improves video-to-speech reconstruction quality and intelligibility in benchmarks, though gains depend on model variant and dataset.

BibTeX
@misc{audio-visual-video-to-speech-synthesis-with-synthesized-input-audio,
  title = {Audio-visual video-to-speech synthesis with synthesized input audio},
  author = {Triantafyllos Kefalas and Yannis Panagakis and Maja Pantić},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2307.16584},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2307.16584v1},
}
arXiv reviewed 5 citations · 2026-04-29

Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Jinxiang Liu, Chen Ju, Chaofan Ma, Yanfeng Wang, Yu Wang, Ya Zhang

Strong AVS result, outside SSI: the useful idea is audio-conditioned decoder queries plus dynamic mask prediction.

BibTeX
@misc{audio-aware-query-enhanced-transformer-for-audio-visual-segmentation,
  title = {Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation},
  author = {Jinxiang Liu and Chen Ju and Chaofan Ma and Yanfeng Wang and Yu Wang and Ya Zhang},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2307.13236},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2307.13236v1},
}
arXiv reviewed 3 citations · 2026-09-12

RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations

Neha Sahipjohn, Neil Shah, Vishal Tambrahalli, Vineet Gandhi

Strong modular SSL-based lip-to-speech synthesis paper that innovatively maps lip SSL features to disentangled speech embeddings before vocoder synthesis, demonstrating improved intelligibility and robustness across benchmark datasets.

BibTeX
@misc{robustl2s-speaker-specific-lip-to-speech-synthesis-exploiting-self-supervised-representations,
  title = {RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations},
  author = {Neha Sahipjohn and Neil Shah and Vishal Tambrahalli and Vineet Gandhi},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2307.01233},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2307.01233v1},
}
arXiv reviewed 14 citations · 2026-09-12

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao

The real gain is not 'diffusion' alone but aligned conditioning plus guidance that pushes synchronization very hard.

BibTeX
@misc{diff-foley-synchronized-video-to-audio-synthesis-with-latent-diffusion-models,
  title = {Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models},
  author = {Simian Luo and Chuanhao Yan and Chenxu Hu and Hang Zhao},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2306.17203},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2306.17203v1},
}
arXiv reviewed 7 citations · 2026-09-12

High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

Junchen Lu, Berrak Şişman, Mingyang Zhang, Haizhou Li

This video-conditioned AVO system innovatively supervises alignment by predicting discrete speech units rather than reconstructing acoustic features, leading to better lip-sync and speech quality on a single-speaker dataset; however, it is not an SSI interface paper.

BibTeX
@misc{high-quality-automatic-voice-over-with-accurate-alignment-supervision-through-self-supervised-discrete-speech-units,
  title = {High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units},
  author = {Junchen Lu and Berrak Şişman and Mingyang Zhang and Haizhou Li},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2306.17005},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2306.17005v1},
}
arXiv reviewed 3 citations · 2026-09-12

Large-scale unsupervised audio pre-training for video-to-speech synthesis

Triantafyllos Kefalas, Yannis Panagakis, Maja Pantić

Good decoder-transfer pretraining improves video-to-speech quality on several benchmarks, but WER gains are not consistent. A useful methodological contribution with strong benchmark support, adjacent to SSI rather than a deployable system.

BibTeX
@misc{large-scale-unsupervised-audio-pre-training-for-video-to-speech-synthesis,
  title = {Large-scale unsupervised audio pre-training for video-to-speech synthesis},
  author = {Triantafyllos Kefalas and Yannis Panagakis and Maja Pantić},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2306.15464},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2306.15464v2},
}
arXiv reviewed 1 citations · 2026-09-12

LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading

Yochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot, Ethan Fetaya

Strong full-text paper demonstrating that inference-time text guidance via ASR classifier is key to significantly improved intelligibility in lip-to-speech synthesis on challenging in-the-wild video datasets, outperforming prior baselines.

BibTeX
@misc{lipvoicer-generating-speech-from-silent-videos-guided-by-lip-reading,
  title = {LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading},
  author = {Yochai Yemini and Aviv Shamsian and Lior Bracha and Sharon Gannot and Ethan Fetaya},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2306.03258},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2306.03258v1},
}
arXiv reviewed 20 citations · 2026-09-12

Intelligible Lip-to-Speech Synthesis with Speech Units

Jeongsoo Choi, Minsu Kim, Yong Man Ro

Speech units as a pseudo-text target enable strong content supervision that substantially cuts WER without text labels, and the multi-input vocoder improves speech quality from blurry mel outputs, yielding a state-of-the-art lip-to-speech system on LRS benchmarks.

BibTeX
@misc{intelligible-lip-to-speech-synthesis-with-speech-units,
  title = {Intelligible Lip-to-Speech Synthesis with Speech Units},
  author = {Jeongsoo Choi and Minsu Kim and Yong Man Ro},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2305.19603},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2305.19603v1},
}
INTERSPEECH 2023 reviewed 5 citations · 2026-09-12

Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks

László Tóth, Amin Honarmandi Shandiz, Gábor Gosztolya, Tamás Gábor Csapó

Strong full-text-backed evidence that most of the gain comes from fast input alignment, not from inventing a new SSI stack.

BibTeX
@misc{adaptation-of-tongue-ultrasound-based-silent-speech-interfaces-using-spatial-transformer-networks,
  title = {Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks},
  author = {László Tóth and Amin Honarmandi Shandiz and Gábor Gosztolya and Tamás Gábor Csapó},
  year = {2023},
  note = {the Proceedings of Interspeech 2023},
  doi = {10.21437/Interspeech.2023-1607},
  eprint = {2305.19130},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.21437/Interspeech.2023-1607},
}
arXiv reviewed 6 citations · 2026-09-12

Zero-shot personalized lip-to-speech synthesis with face image based voice control

Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling

Demonstrates effective zero-shot voice control in Lip2Speech by leveraging face image-based speaker embeddings, validated on GRID corpus but constrained by dataset vocabulary and speech naturalness.

BibTeX
@misc{zero-shot-personalized-lip-to-speech-synthesis-with-face-image-based-voice-control,
  title = {Zero-shot personalized lip-to-speech synthesis with face image based voice control},
  author = {Zheng-Yan Sheng and Yang Ai and Zhen-Hua Ling},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2305.14359},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2305.14359v1},
}
arXiv reviewed 0 citations · 2026-09-12

Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning

Sara Kashiwagi, Keitaro Tanaka, Feng Qi, Shigeo Morishima

Strong viseme-level metric learning approach reduces silent speech VSR errors on a small 10-phrase dataset, notably achieving parity with baselines using much less silent data.

BibTeX
@misc{improving-the-gap-in-visual-speech-recognition-between-normal-and-silent-speech-based-on-metric-learning,
  title = {Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning},
  author = {Sara Kashiwagi and Keitaro Tanaka and Feng Qi and Shigeo Morishima},
  year = {2023},
  note = {arXiv / imported corpus page},
  doi = {10.21437/Interspeech.2023-370},
  eprint = {2305.14203},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.21437/Interspeech.2023-370},
}
arXiv reviewed 37 citations · 2026-09-12

Conditional Generation of Audio from Video via Foley Analogies

Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, Andrew Owens

The paper matters because it gives V2A generation a controllable exemplar, not because it beats every timing baseline.

BibTeX
@misc{conditional-generation-of-audio-from-video-via-foley-analogies,
  title = {Conditional Generation of Audio from Video via Foley Analogies},
  author = {Yuexi Du and Ziyang Chen and Justin Salamon and Bryan Russell and Andrew Owens},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2304.08490},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2304.08490v1},
}
arXiv reviewed 5 citations · 2026-09-12

Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training

Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

Strong SSI paper improving silent speech reconstruction by generating pseudo acoustic targets and using domain adversarial training to address domain mismatch; validated with TaL dataset showing substantial WER and MOS gains over TaLNet.

BibTeX
@misc{speech-reconstruction-from-silent-tongue-and-lip-articulation-by-pseudo-target-generation-and-domain-adversarial-training,
  title = {Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training},
  author = {Rui-Chen Zheng and Yang Ai and Zhen-Hua Ling},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2304.05574},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2304.05574v1},
}
CHI 2023 reviewed 23 citations · 2026-09-12

WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions

Jun Rekimoto

Strong whisper-conversion paper, but it remains whisper-based rather than truly silent SSI.

BibTeX
@misc{wesper-zero-shot-and-realtime-whisper-to-normal-voice-conversion-for-whisper-based-speech-interactions,
  title = {WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions},
  author = {Jun Rekimoto},
  year = {2023},
  note = {Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI '23), April 23--28, 2023},
  doi = {10.1145/3544548.3580706},
  eprint = {2303.01639},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3544548.3580706},
}
arXiv reviewed 5 citations · 2026-09-12

Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech

Dong Yang, Tomoki Koriyama, Yuki Saito, Takaaki Saeki, Detai Xin, Hiroshi Saruwatari

The paper presents a strong multi-speaker TTS phrasing approach leveraging speaker-conditioned BERT embeddings and pause duration categories to improve pause insertion precision and synthetic speech rhythm; however, it is out-of-scope for SSI as it focuses on audible speech synthesis only.

BibTeX
@misc{duration-aware-pause-insertion-using-pre-trained-language-model-for-multi-speaker-text-to-speech,
  title = {Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech},
  author = {Dong Yang and Tomoki Koriyama and Yuki Saito and Takaaki Saeki and Detai Xin and Hiroshi Saruwatari},
  year = {2023},
  note = {arXiv / imported corpus page},
  eprint = {2302.13652},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2302.13652v1},
}
arXiv reviewed 37 citations · 2026-09-12

LipLearner: Customizable Silent Speech Interactions on Mobile Devices

Zixiong Su, Shitao Fang, Jun Rekimoto

LipLearner is a strong mobile silent speech system that uniquely closes the loop from few-shot lipreading model design to practical on-device customization and keyword spotting, demonstrated robustly in real-world conditions and a user study.

BibTeX
@misc{liplearner-customizable-silent-speech-interactions-on-mobile-devices,
  title = {LipLearner: Customizable Silent Speech Interactions on Mobile Devices},
  author = {Zixiong Su and Shitao Fang and Jun Rekimoto},
  year = {2023},
  note = {arXiv / imported corpus page},
  doi = {10.1145/3544548.3581465},
  eprint = {2302.05907},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3544548.3581465},
}
arXiv reviewed 1 citations · 2026-04-29

Towards Neural Decoding of Imagined Speech based on Spoken Speech

Seo‐Hyun Lee, Young-Eun Lee, Soo-Won Kim, Byung-Kwan Ko, Seong‐Whan Lee

Transfer of CSP+SVM models trained on spoken speech EEG to imagined speech achieves comparable, though slightly lower, accuracy within a limited 5-class, 7-subject offline EEG setup, with visual imagery control supporting specificity.

BibTeX
@misc{towards-neural-decoding-of-imagined-speech-based-on-spoken-speech,
  title = {Towards Neural Decoding of Imagined Speech based on Spoken Speech},
  author = {Seo‐Hyun Lee and Young-Eun Lee and Soo-Won Kim and Byung-Kwan Ko and Seong‐Whan Lee},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2212.02047},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2212.02047v2},
}
arXiv reviewed 1 citations · 2026-09-12

Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation

Hassan Taherian, Şefik Emre Eskimez, Takuya Yoshioka

Strong causal PSE paper, not SSI. The pVAD-guided loss is the part that holds up under full-text reading.

BibTeX
@misc{breaking-the-trade-off-in-personalized-speech-enhancement-with-cross-task-knowledge-distillation,
  title = {Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation},
  author = {Hassan Taherian and Şefik Emre Eskimez and Takuya Yoshioka},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2211.02944},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2211.02944v1},
}
arXiv reviewed 2 citations · 2026-09-12

Movement Detection of Tongue and Related Body Parts Using IR-UWB Radar

Sunghwa Lee, Younghoon Shin

Good sensing primitive, very small task.

BibTeX
@misc{movement-detection-of-tongue-and-related-body-parts-using-ir-uwb-radar,
  title = {Movement Detection of Tongue and Related Body Parts Using IR-UWB Radar},
  author = {Sunghwa Lee and Younghoon Shin},
  year = {2022},
  note = {arXiv / imported corpus page},
  doi = {10.1109/ICTC55196.2022.9952644},
  eprint = {2209.01762},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/ICTC55196.2022.9952644},
}
arXiv reviewed 14 citations · 2026-09-12

Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild

Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar

The real contribution is not just another VAE-GAN; it is turning lip-to-speech into an arbitrary-speaker problem with credible low-data adaptation.

BibTeX
@misc{lip-to-speech-synthesis-for-arbitrary-speakers-in-the-wild,
  title = {Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild},
  author = {Sindhu B Hegde and K R Prajwal and Rudrabha Mukhopadhyay and Vinay P. Namboodiri and C. V. Jawahar},
  year = {2022},
  note = {arXiv / imported corpus page},
  doi = {10.1145/3503161.3548081},
  eprint = {2209.00642},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3503161.3548081},
}
arXiv reviewed 2 citations · 2026-09-12

An Anchor-Free Detector for Continuous Speech Keyword Spotting

Zhiyuan Zhao, Chuanxin Tang, Chengdong Yao, Chong Luo

Strong CSKWS paper, not SSI. The detection framing and unknown class are the points that hold up in full text.

BibTeX
@misc{an-anchor-free-detector-for-continuous-speech-keyword-spotting,
  title = {An Anchor-Free Detector for Continuous Speech Keyword Spotting},
  author = {Zhiyuan Zhao and Chuanxin Tang and Chengdong Yao and Chong Luo},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2208.04622},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2208.04622v1},
}
arXiv reviewed 9 citations · 2026-09-12

FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis

Yongqi Wang, Zhou Zhao

This paper matters because it makes unconstrained lip-to-speech materially faster without obviously sacrificing quality.

BibTeX
@misc{fastlts-non-autoregressive-end-to-end-unconstrained-lip-to-speech-synthesis,
  title = {FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis},
  author = {Yongqi Wang and Zhou Zhao},
  year = {2022},
  note = {arXiv / imported corpus page},
  doi = {10.1145/3503161.3548194},
  eprint = {2207.03800},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3503161.3548194},
}
arXiv reviewed 7 citations · 2026-09-12

Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks

Amin Honarmandi Shandiz, László Tóth

An empirically supported, incremental advancement showing that hybrid 3D-CNN plus ConvLSTM models modestly outperform prior ultrasound tongue video SSI architectures in mel-spectrogram regression accuracy and model efficiency on single-speaker data.

BibTeX
@misc{improved-processing-of-ultrasound-tongue-videos-by-combining-convlstm-and-3d-convolutional-networks,
  title = {Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks},
  author = {Amin Honarmandi Shandiz and László Tóth},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2206.12947},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2206.12947v1},
}
arXiv reviewed 6 citations · 2026-09-12

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

Joanna Hong, Minsu Kim, Yong Man Ro

The paper is really about disentangling identity, and that is why the unseen-speaker results hold up.

BibTeX
@misc{visagesyntalk-unseen-speaker-video-to-speech-synthesis-via-speech-visage-feature-selection,
  title = {VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection},
  author = {Joanna Hong and Minsu Kim and Yong Man Ro},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2206.07458},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2206.07458v2},
}
arXiv reviewed 6 citations · 2026-09-12

Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information

Chi-Luen Feng, Po‐Chun Hsu, Hung-yi Lee

Strong evidence that silence segments in HuBERT representations uniquely store speaker information, improving SID accuracy when silence is augmented; analytical SSL probing paper outside silent speech interface field.

BibTeX
@misc{silence-is-sweeter-than-speech-self-supervised-model-using-silence-to-store-speaker-information,
  title = {Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information},
  author = {Chi-Luen Feng and Po‐Chun Hsu and Hung-yi Lee},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2205.03759},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2205.03759v1},
}
arXiv reviewed 17 citations · 2026-09-12

SVTS: Scalable Video-to-Speech Synthesis

Rodrigo Mira, Alexandros Haliassos, Stavros Petridis, Björn W. Schuller, Maja Pantić

A key scaling contribution that demonstrates simple spectrogram prediction plus pretrained vocoder pipelines outperform prior complex models on diverse datasets, marking foundational progress in large-scale video-to-speech synthesis.

BibTeX
@misc{svts-scalable-video-to-speech-synthesis,
  title = {SVTS: Scalable Video-to-Speech Synthesis},
  author = {Rodrigo Mira and Alexandros Haliassos and Stavros Petridis and Björn W. Schuller and Maja Pantić},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2205.02058},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2205.02058v2},
}
arXiv reviewed 22 citations · 2026-04-29

Listen only to me! How well can target speech extraction handle false alarms?

Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Kateřina Žmolíková, Hiroshi Satō, Tomohiro Nakatani

Strong paper for false-alarm handling in TSE, wrong domain if someone tries to count it as SSI progress.

BibTeX
@misc{listen-only-to-me-how-well-can-target-speech-extraction-handle-false-alarms,
  title = {Listen only to me! How well can target speech extraction handle false alarms?},
  author = {Marc Delcroix and Keisuke Kinoshita and Tsubasa Ochiai and Kateřina Žmolíková and Hiroshi Satō and Tomohiro Nakatani},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2204.04811},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2204.04811v2},
}
arXiv reviewed 51 citations · 2026-09-12

Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video

Minsu Kim, Joanna Hong, Se Jin Park, Yong Man Ro

The key idea is not generic fusion; it is storing cross-modal correspondences so video-only decoding can recover some audio-side structure later.

BibTeX
@misc{multi-modality-associative-bridging-through-memory-speech-sound-recollected-from-face-video,
  title = {Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video},
  author = {Minsu Kim and Joanna Hong and Se Jin Park and Yong Man Ro},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2204.01265},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2204.01265v1},
}
arXiv reviewed 12 citations · 2026-09-12

VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion

Disong Wang, Shan Yang, Dan Su, Xunying Liu, Dong Yu, Helen Meng

The real move is importing structure from voice conversion, not just adding another speaker embedding.

BibTeX
@misc{vcvts-multi-speaker-video-to-speech-synthesis-via-cross-modal-knowledge-transfer-from-voice-conversion,
  title = {VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion},
  author = {Disong Wang and Shan Yang and Dan Su and Xunying Liu and Dong Yu and Helen Meng},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2202.09081},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2202.09081v1},
}
ICASSP 2022 reviewed 18 citations · 2026-09-12

Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals

Xingyu Chen, Qiushi Zhu, Jie Zhang, Li-Rong Dai

Sound classification paper, not SSI.

BibTeX
@misc{supervised-and-self-supervised-pretraining-based-covid-19-detection-using-acoustic-breathing-cough-speech-signals,
  title = {Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals},
  author = {Xingyu Chen and Qiushi Zhu and Jie Zhang and Li-Rong Dai},
  year = {2022},
  note = {ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 561-565},
  doi = {10.1109/ICASSP43922.2022.9746205},
  eprint = {2201.08934},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/ICASSP43922.2022.9746205},
}
arXiv reviewed 17 citations · 2026-09-12

VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over

Junchen Lu, Berrak Şişman, Rui Liu, Mingyang Zhang, Haizhou Li

VisualTTS effectively improves lip-speech synchronization in scripted voice over by conditioning TTS on lip video, but does not tackle silent speech decoding or unscripted scenarios.

BibTeX
@misc{visualtts-tts-with-accurate-lip-speech-synchronization-for-automatic-voice-over,
  title = {VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over},
  author = {Junchen Lu and Berrak Şişman and Rui Liu and Mingyang Zhang and Haizhou Li},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2110.03342},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2110.03342v3},
}
arXiv reviewed 16 citations · 2026-09-12

Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language

Huiyan Li, Haohong Lin, You Wang, Hengyang Wang, Ming Zhang, Han Gao, Qing Ai, Zhiyuan Luo, Guang Li

SSRNet innovatively applies duration-aware Seq2Seq modeling and tonal multitask learning to reconstruct intelligible Mandarin speech from facial sEMG signals, markedly improving performance over prior methods but remains speaker-dependent with limited deployment evaluation.

BibTeX
@misc{sequence-to-sequence-voice-reconstruction-for-silent-speech-in-a-tonal-language,
  title = {Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language},
  author = {Huiyan Li and Haohong Lin and You Wang and Hengyang Wang and Ming Zhang and Han Gao and Qing Ai and Zhiyuan Luo and Guang Li},
  year = {2022},
  note = {arXiv / imported corpus page},
  eprint = {2108.00190},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2108.00190v3},
}
CHI 2022 reviewed 44 citations · 2026-09-12

SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography

Naoki Kimura, Tan Gemicioglu, Jonathan Womack, Richard Li, Yuhui Zhao, Abdelkareem Bedri, Zixiong Su, Alex Olwal, Jun Rekimoto, Thad Starner

SilentSpeller is a strong, rigorously tested SSI system that reframes silent speech as silent spelling, enabling large vocabulary, live text entry, and walking robustness with in-mouth electropalatography sensors.

BibTeX
@misc{silentspeller,
  title = {SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography},
  author = {Naoki Kimura and Tan Gemicioglu and Jonathan Womack and Richard Li and Yuhui Zhao and Abdelkareem Bedri and Zixiong Su and Alex Olwal and Jun Rekimoto and Thad Starner},
  year = {2022},
  note = {CHI '22},
  doi = {10.1145/3491102.3502015},
  url = {https://doi.org/10.1145/3491102.3502015},
}
arXiv reviewed 21 citations · 2026-09-12

SA-SDR: A novel loss function for separation of meeting style data

Thilo von Neumann, Keisuke Kinoshita, Christoph Boeddeker, Marc Delcroix, Reinhold Haeb‐Umbach

Elegant loss fix, not SSI.

BibTeX
@misc{sa-sdr-a-novel-loss-function-for-separation-of-meeting-style-data,
  title = {SA-SDR: A novel loss function for separation of meeting style data},
  author = {Thilo von Neumann and Keisuke Kinoshita and Christoph Boeddeker and Marc Delcroix and Reinhold Haeb‐Umbach},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2110.15581},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2110.15581v2},
}
arXiv reviewed 2 citations · 2026-04-29

Advances and Challenges in Deep Lip Reading

Marzieh Oghbaie, Arian Sabaghi, Kooshan Hashemifard, Mohammad Kazem Akbari

Good survey, not a model result.

BibTeX
@misc{advances-and-challenges-in-deep-lip-reading,
  title = {Advances and Challenges in Deep Lip Reading},
  author = {Marzieh Oghbaie and Arian Sabaghi and Kooshan Hashemifard and Mohammad Kazem Akbari},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2110.07879},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2110.07879v1},
}
arXiv reviewed 107 citations · 2026-09-12

Sub-word Level Lip Reading With Visual Attention

K R Prajwal, Triantafyllos Afouras, Andrew Zisserman

Major lip-reading gain, adjacent to SSI.

BibTeX
@misc{sub-word-level-lip-reading-with-visual-attention,
  title = {Sub-word Level Lip Reading With Visual Attention},
  author = {K R Prajwal and Triantafyllos Afouras and Andrew Zisserman},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2110.07603},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2110.07603v2},
}
arXiv reviewed 2 citations · 2026-09-12

Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input

Csapó Tamás Gábor, László Tóth, Gosztolya Gábor, Alexandra Markó

Helpful side information, not standalone SSI.

BibTeX
@misc{speech-synthesis-from-text-and-ultrasound-tongue-image-based-articulatory-input,
  title = {Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input},
  author = {Csapó Tamás Gábor and László Tóth and Gosztolya Gábor and Alexandra Markó},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2107.02003},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2107.02003v1},
}
arXiv reviewed 7 citations · 2026-09-12

Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits

Qingjian Lin, Lin Yang, Xuyang Wang, Luyuan Xie, Jia Chen, Junjie Wang

Useful separation engineering, not silent speech.

BibTeX
@misc{sparsely-overlapped-speech-training-in-the-time-domain-joint-learning-of-target-speech-separation-and-personal-vad-benefits,
  title = {Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits},
  author = {Qingjian Lin and Lin Yang and Xuyang Wang and Luyuan Xie and Jia Chen and Junjie Wang},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2106.14371},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2106.14371v2},
}
arXiv reviewed 0 citations · 2026-04-29

Silent Speech and Emotion Recognition from Vocal Tract Shape Dynamics in Real-Time MRI

Laxmi Pandey, Ahmed Sabbir Arif

Strong rtMRI recognition result, weak deployment story.

BibTeX
@misc{silent-speech-and-emotion-recognition-from-vocal-tract-shape-dynamics-in-real-time-mri,
  title = {Silent Speech and Emotion Recognition from Vocal Tract Shape Dynamics in Real-Time MRI},
  author = {Laxmi Pandey and Ahmed Sabbir Arif},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2106.08706},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2106.08706v1},
}
arXiv reviewed 5 citations · 2026-09-12

Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces

Honarmandi Shandiz Amin, László Tóth, Gosztolya Gábor, Alexandra Markó, Csapó Tamás Gábor

The ultrasound-based x-vector speaker embedding is highly effective for speaker recognition, achieving under 1% error on unseen speakers, but its integration yields only a marginal improvement in multi-speaker ultrasound-to-speech synthesis accuracy.

BibTeX
@misc{neural-speaker-embeddings-for-ultrasound-based-silent-speech-interfaces,
  title = {Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces},
  author = {Honarmandi Shandiz Amin and László Tóth and Gosztolya Gábor and Alexandra Markó and Csapó Tamás Gábor},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2106.04552},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2106.04552v2},
}
arXiv reviewed 23 citations · 2026-09-12

An Improved Model for Voicing Silent Speech

David Gaddy, Dan Klein

This paper substantially improves open-vocabulary silent speech voicing using learned convolutional EMG features, Transformer modeling, and phoneme supervision, reducing WER from 68.0% to 42.2% automatic and 32.3% human in a single-speaker lab setting.

BibTeX
@misc{an-improved-model-for-voicing-silent-speech,
  title = {An Improved Model for Voicing Silent Speech},
  author = {David Gaddy and Dan Klein},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2106.01933},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2106.01933v2},
}
arXiv reviewed 2 citations · 2026-09-12

Voice Activity Detection for Ultrasound-based Silent Speech Interfaces using Convolutional Neural Networks

Amin Honarmandi Shandiz, László Tóth

Preprocessing paper, narrow but legitimate.

BibTeX
@misc{voice-activity-detection-for-ultrasound-based-silent-speech-interfaces-using-convolutional-neural-networks,
  title = {Voice Activity Detection for Ultrasound-based Silent Speech Interfaces using Convolutional Neural Networks},
  author = {Amin Honarmandi Shandiz and László Tóth},
  year = {2021},
  note = {arXiv / imported corpus page},
  doi = {10.1007/978-3-030-83527-9_43},
  eprint = {2105.13718},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1007/978-3-030-83527-9_43},
}
arXiv reviewed 9 citations · 2026-09-12

Speaker disentanglement in video-to-speech conversion

Dan Oneaţă, Adriana Stan, Horia Cucu

The paper effectively makes speaker identity a controllable factor in multi-speaker video-to-speech synthesis by disentangling it from content, showing the trade-off between intelligibility and voice control on GRID corpus data.

BibTeX
@misc{speaker-disentanglement-in-video-to-speech-conversion,
  title = {Speaker disentanglement in video-to-speech conversion},
  author = {Dan Oneaţă and Adriana Stan and Horia Cucu},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2105.09652},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2105.09652v1},
}
arXiv reviewed 12 citations · 2026-09-12

Improving Neural Silent Speech Interface Models by Adversarial Training

Amin Honarmandi Shandiz, László Tóth, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó

A clean, well-executed incremental advance using GAN loss to modestly improve articulatory-to-acoustic mapping from ultrasound, validated objectively on two single-speaker corpora.

BibTeX
@misc{improving-neural-silent-speech-interface-models-by-adversarial-training,
  title = {Improving Neural Silent Speech Interface Models by Adversarial Training},
  author = {Amin Honarmandi Shandiz and László Tóth and Gábor Gosztolya and Alexandra Markó and Tamás Gábor Csapó},
  year = {2021},
  note = {arXiv / imported corpus page},
  doi = {10.1007/978-3-030-76346-6_39},
  eprint = {2104.11601},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1007/978-3-030-76346-6_39},
}
arXiv reviewed 12 citations · 2026-09-12

3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

László Tóth, Amin Honarmandi Shandiz

Temporal context helps, but the evidence is a single-speaker vocoder-parameter study.

BibTeX
@misc{3d-convolutional-neural-networks-for-ultrasound-based-silent-speech-interfaces,
  title = {3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces},
  author = {László Tóth and Amin Honarmandi Shandiz},
  year = {2021},
  note = {arXiv / imported corpus page},
  doi = {10.1007/978-3-030-61401-0_16},
  eprint = {2104.11532},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1007/978-3-030-61401-0_16},
}
arXiv reviewed 1 citations · 2026-09-12

HTMD-Net: A Hybrid Masking-Denoising Approach to Time-Domain Monaural Singing Voice Separation

Christos Garoufis, Athanasia Zlatintsi, Petros Maragos

Solid time-domain music vocal separation paper with a novel hybrid masking-denoising design showing improved silent-segment suppression; not relevant to SSI applications.

BibTeX
@misc{htmd-net-a-hybrid-masking-denoising-approach-to-time-domain-monaural-singing-voice-separation,
  title = {HTMD-Net: A Hybrid Masking-Denoising Approach to Time-Domain Monaural Singing Voice Separation},
  author = {Christos Garoufis and Athanasia Zlatintsi and Petros Maragos},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2103.04336},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2103.04336v1},
}
arXiv reviewed 2 citations · 2026-09-12

Silent versus modal multi-speaker speech recognition from ultrasound and video

Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals

Large-corpus baseline with real silent-mode gap.

BibTeX
@misc{silent-versus-modal-multi-speaker-speech-recognition-from-ultrasound-and-video,
  title = {Silent versus modal multi-speaker speech recognition from ultrasound and video},
  author = {Manuel Sam Ribeiro and Aciel Eshky and Korin Richmond and Steve Renals},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2103.00333},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2103.00333v1},
}
arXiv reviewed 8 citations · 2026-09-12

EMA2S: An End-to-End Multimodal Articulatory-to-Speech System

Yu‐Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan H. Sherman, Wen-Chin Huang, Xugang Lu, Yu Tsao

EMA2S achieves consistent quality improvements over prior EMA-to-speech baselines by combining multimodal joint loss training with a neural vocoder, though gains remain confined to lab EMA conditions.

BibTeX
@misc{ema2s-an-end-to-end-multimodal-articulatory-to-speech-system,
  title = {EMA2S: An End-to-End Multimodal Articulatory-to-Speech System},
  author = {Yu‐Wen Chen and Kuo-Hsuan Hung and Shang-Yi Chuang and Jonathan H. Sherman and Wen-Chin Huang and Xugang Lu and Yu Tsao},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2102.03786},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2102.03786v2},
}
arXiv reviewed 2 citations · 2026-09-12

Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image

Kele Xu, Tamás Gábor Csapó, Ming Feng

Real signal, wrong target for SSI.

BibTeX
@misc{convolutional-neural-network-based-age-estimation-using-b-mode-ultrasound-tongue-image,
  title = {Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image},
  author = {Kele Xu and Tamás Gábor Csapó and Ming Feng},
  year = {2021},
  note = {arXiv / imported corpus page},
  eprint = {2101.11245},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2101.11245v1},
}
arXiv reviewed 13 citations · 2026-09-12

End-to-end Silent Speech Recognition with Acoustic Sensing

Jian Luo, Jianzong Wang, Ning Cheng, Guilin Jiang, Jing Xiao

Strong mobile-friendly acoustic SSI paper.

BibTeX
@misc{end-to-end-silent-speech-recognition-with-acoustic-sensing,
  title = {End-to-end Silent Speech Recognition with Acoustic Sensing},
  author = {Jian Luo and Jianzong Wang and Ning Cheng and Guilin Jiang and Jing Xiao},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2011.11315},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2011.11315v1},
}
arXiv reviewed 1 citations · 2026-09-12

Speech Prediction in Silent Videos using Variational Autoencoders

Ravindra Yadav, Ashish Sardana, Vinay P. Namboodiri, Rajesh M. Hegde

Strong video-to-speech paper that models ambiguity explicitly.

BibTeX
@misc{speech-prediction-in-silent-videos-using-variational-autoencoders,
  title = {Speech Prediction in Silent Videos using Variational Autoencoders},
  author = {Ravindra Yadav and Ashish Sardana and Vinay P. Namboodiri and Rajesh M. Hegde},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2011.07340},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2011.07340v1},
}
arXiv reviewed 32 citations · 2026-09-12

X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network

Zining Zhang, Bingsheng He, Zhenjie Zhang

Strong time-domain target-speaker extraction using speaker verification and innovative training; improves robustness to absent target but remains speech extraction, not silent speech.

BibTeX
@misc{x-tasnet-robust-and-accurate-time-domain-speaker-extraction-network,
  title = {X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network},
  author = {Zining Zhang and Bingsheng He and Zhenjie Zhang},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2010.12766},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2010.12766v1},
}
arXiv reviewed 23 citations · 2026-09-12

Listening to Sounds of Silence for Speech Denoising

Ruilin Xu, Rundi Wu, Yuko Ishiwaka, Carl Vondrick, Changxi Zheng

Strong denoising work, not SSI.

BibTeX
@misc{listening-to-sounds-of-silence-for-speech-denoising,
  title = {Listening to Sounds of Silence for Speech Denoising},
  author = {Ruilin Xu and Rundi Wu and Yuko Ishiwaka and Carl Vondrick and Changxi Zheng},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2010.12013},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2010.12013v1},
}
arXiv reviewed 51 citations · 2026-09-12

Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching

Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, Dejing Dou

Technically solid self-supervised class-aware audiovisual sounding object localization, but outside the core SSI domain.

BibTeX
@misc{discriminative-sounding-objects-localization-via-self-supervised-audiovisual-matching,
  title = {Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching},
  author = {Di Hu and Rui Qian and Minyue Jiang and Xiao Tan and Shilei Wen and Errui Ding and Weiyao Lin and Dejing Dou},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2010.05466},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2010.05466v1},
}
arXiv reviewed 56 citations · 2026-09-12

Digital Voicing of Silent Speech

David Gaddy, Dan Klein

Core EMG SSI paper with real gains from target transfer.

BibTeX
@misc{digital-voicing-of-silent-speech,
  title = {Digital Voicing of Silent Speech},
  author = {David Gaddy and Dan Klein},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2010.02960},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2010.02960v1},
}
arXiv reviewed 2 citations · 2026-09-12

End-to-End Speaker-Dependent Voice Activity Detection

Yefei Chen, Shuai Wang, Yanmin Qian, Kai Yu

Strong target-speaker VAD paper, not SSI.

BibTeX
@misc{end-to-end-speaker-dependent-voice-activity-detection,
  title = {End-to-End Speaker-Dependent Voice Activity Detection},
  author = {Yefei Chen and Shuai Wang and Yanmin Qian and Kai Yu},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2009.09906},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2009.09906v1},
}
arXiv reviewed 16 citations · 2026-04-29

A comparison of oscillatory characteristics in covert speech and speech perception

Jae Moon, Silvia Orlandi, Tom Chau

Strong covert-speech EEG analysis, not an SSI system.

BibTeX
@misc{a-comparison-of-oscillatory-characteristics-in-covert-speech-and-speech-perception,
  title = {A comparison of oscillatory characteristics in covert speech and speech perception},
  author = {Jae Moon and Silvia Orlandi and Tom Chau},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2009.02816},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2009.02816v1},
}
arXiv reviewed 142 citations · 2026-09-12

Silent Speech Interfaces for Speech Restoration: A Review

José A. González, Alejandro Gomez-Alanis, Juan M. Martín-Doñas, José L. Pérez-Córdoba, Ángel M. Gómez

Core SSI survey with concrete deployment constraints.

BibTeX
@misc{silent-speech-interfaces-for-speech-restoration-a-review,
  title = {Silent Speech Interfaces for Speech Restoration: A Review},
  author = {José A. González and Alejandro Gomez-Alanis and Juan M. Martín-Doñas and José L. Pérez-Córdoba and Ángel M. Gómez},
  year = {2020},
  note = {arXiv / imported corpus page},
  doi = {10.1109/ACCESS.2020.3026579},
  eprint = {2009.02110},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/ACCESS.2020.3026579},
}
arXiv reviewed 27 citations · 2026-09-12

An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation

Daniel Michelsanti, Zheng‐Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, Jesper Jensen

Strong AV speech survey, not an SSI system paper.

BibTeX
@misc{an-overview-of-deep-learning-based-audio-visual-speech-enhancement-and-separation,
  title = {An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation},
  author = {Daniel Michelsanti and Zheng‐Hua Tan and Shi-Xiong Zhang and Yong Xu and Meng Yu and Dong Yu and Jesper Jensen},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2008.09586},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2008.09586v2},
}
arXiv reviewed 7 citations · 2026-09-12

CITISEN: A Deep Learning-Based Speech Signal-Processing Mobile Application

Yu-Wen Chen, Kuo-Hsuan Hung, You-Jin Li, Alexander Kang, Ya‐Hsin Lai, Kai-Chun Liu, Szu‐Wei Fu, Syu‐Siang Wang, Yu Tsao

Strong mobile speech-processing app paper, not SSI.

BibTeX
@misc{citisen-a-deep-learning-based-speech-signal-processing-mobile-application,
  title = {CITISEN: A Deep Learning-Based Speech Signal-Processing Mobile Application},
  author = {Yu-Wen Chen and Kuo-Hsuan Hung and You-Jin Li and Alexander Kang and Ya‐Hsin Lai and Kai-Chun Liu and Szu‐Wei Fu and Syu‐Siang Wang and Yu Tsao},
  year = {2020},
  note = {arXiv / imported corpus page},
  doi = {10.1109/ACCESS.2022.3153469},
  eprint = {2008.09264},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/ACCESS.2022.3153469},
}
arXiv reviewed 118 citations · 2026-09-12

Foley Music: Learning to Generate Music from Videos

Chuang Gan, Deng Huang, Peihao Chen, Joshua B. Tenenbaum, Antonio Torralba

Strong video-to-music paper, not SSI.

BibTeX
@misc{foley-music-learning-to-generate-music-from-videos,
  title = {Foley Music: Learning to Generate Music from Videos},
  author = {Chuang Gan and Deng Huang and Peihao Chen and Joshua B. Tenenbaum and Antonio Torralba},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2007.10984},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2007.10984v1},
}
arXiv reviewed 3 citations · 2026-09-12

Learning Frame Level Attention for Environmental Sound Classification

Zhichao Zhang, Shugong Xu, Shunqing Zhang, Tianhao Qiao, Shan Cao

Strong ESC paper, but outside SSI.

BibTeX
@misc{learning-frame-level-attention-for-environmental-sound-classification,
  title = {Learning Frame Level Attention for Environmental Sound Classification},
  author = {Zhichao Zhang and Shugong Xu and Shunqing Zhang and Tianhao Qiao and Shan Cao},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2007.07241},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2007.07241v1},
}
arXiv reviewed 3 citations · 2026-09-12

Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images

Pramit Saha, Yadong Liu, Bryan Gick, Sidney Fels

Strong ultrasound SSI paper with unusually clear quantitative gains.

BibTeX
@misc{ultra2speech-a-deep-learning-framework-for-formant-frequency-estimation-and-tracking-from-ultrasound-tongue-images,
  title = {Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images},
  author = {Pramit Saha and Yadong Liu and Bryan Gick and Sidney Fels},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2006.16367},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2006.16367v1},
}
arXiv reviewed 16 citations · 2026-09-12

Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task

Babak Naderi, Sebastian Möller

Strong crowdsourcing methodology paper, not SSI.

BibTeX
@misc{application-of-just-noticeable-difference-in-quality-as-environment-suitability-test-for-crowdsourcing-speech-quality-assessment-task,
  title = {Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task},
  author = {Babak Naderi and Sebastian Möller},
  year = {2020},
  note = {arXiv / imported corpus page},
  doi = {10.1109/QoMEX48832.2020.9123093},
  eprint = {2004.05502},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/QoMEX48832.2020.9123093},
}
arXiv reviewed 1 citations · 2026-09-12

Vocoder-Based Speech Synthesis from Silent Videos

Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emília Gómez, Zheng‐Hua Tan, Jesper Jensen

A notable step forward in lip-to-speech synthesis by predicting full vocoder features and jointly training for recognition, achieving strong speaker-dependent results but lacking unseen speaker generalization.

BibTeX
@misc{vocoder-based-speech-synthesis-from-silent-videos,
  title = {Vocoder-Based Speech Synthesis from Silent Videos},
  author = {Daniel Michelsanti and Olga Slizovskaia and Gloria Haro and Emília Gómez and Zheng‐Hua Tan and Jesper Jensen},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2004.02541},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2004.02541v2},
}
arXiv reviewed 1 citations · 2026-09-12

Continuous Silent Speech Recognition using EEG

Gautam Krishna, Co Tran, Mason Carnahan, Ahmed H. Tewfik

Real EEG sentence-level silent speech recognition is demonstrated but at very high WER, confirming feasibility only and underscoring the immature state of current EEG silent speech technology.

BibTeX
@misc{continuous-silent-speech-recognition-using-eeg,
  title = {Continuous Silent Speech Recognition using EEG},
  author = {Gautam Krishna and Co Tran and Mason Carnahan and Ahmed H. Tewfik},
  year = {2020},
  note = {arXiv / imported corpus page},
  eprint = {2002.03851},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2002.03851v7},
}
arXiv reviewed 84 citations · 2026-09-12

Brain2Char: A Deep Architecture for Decoding Text from Brain Recordings

Pengfei Sun, Gopala K. Anumanchipalli, Edward F. Chang

Brain2Char establishes a new state-of-the-art for continuous character decoding from invasive ECoG with competitive WER on large vocabularies and silent speech, demonstrating feasibility for communication BCIs.

BibTeX
@misc{brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings,
  title = {Brain2Char: A Deep Architecture for Decoding Text from Brain Recordings},
  author = {Pengfei Sun and Gopala K. Anumanchipalli and Edward F. Chang},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1909.01401},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1909.01401v1},
}
arXiv reviewed 58 citations · 2026-09-12

Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Alexandre Défossez, Nicolas Usunier, Léon Bottou, Francis R. Bach

This work delivers an improved waveform source separation model combined with a novel remix-based semi-supervised learning scheme using unlabeled music. Though not related to silent speech, it advances music separation benchmarks by closing gaps to spectrogram methods.

BibTeX
@misc{demucs-deep-extractor-for-music-sources-with-extra-unlabeled-data-remixed,
  title = {Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed},
  author = {Alexandre Défossez and Nicolas Usunier and Léon Bottou and Francis R. Bach},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1909.01174},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1909.01174v1},
}
arXiv reviewed 132 citations · 2026-09-12

Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification

Zhichao Zhang, Shugong Xu, Shunqing Zhang, Tianhao Qiao, Shan Cao

The proposed frame-level attention integrated within a convolutional recurrent network effectively improves environmental sound classification accuracy on ESC benchmarks by focusing on informative temporal frames while suppressing irrelevant or silent ones.

BibTeX
@misc{attention-based-convolutional-recurrent-neural-network-for-environmental-sound-classification,
  title = {Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification},
  author = {Zhichao Zhang and Shugong Xu and Shunqing Zhang and Tianhao Qiao and Shan Cao},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1907.02230},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1907.02230v1},
}
arXiv reviewed 32 citations · 2026-09-12

Lipper: Synthesizing Thy Speech using Multi-View Lipreading

Yaman Kumar, Rohit Jain, Khwaja Mohd. Salik, Rajiv Ratn Shah, Yifang Yin, Roger Zimmermann

Strong multi-view lip-to-speech baseline with honest quality limits.

BibTeX
@misc{lipper-synthesizing-thy-speech-using-multi-view-lipreading,
  title = {Lipper: Synthesizing Thy Speech using Multi-View Lipreading},
  author = {Yaman Kumar and Rohit Jain and Khwaja Mohd. Salik and Rajiv Ratn Shah and Yifang Yin and Roger Zimmermann},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1907.01367},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1907.01367v1},
}
arXiv reviewed 5 citations · 2026-09-12

Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder

Tamás Gábor Csapó, Mohammed Salah Al-Radhi, Géza Németh, Gábor Gosztolya, Tamás Grósz, László Tóth, Alexandra Markó

The key advancement is continuous F0 tracking via CNNs yielding lower pitch error and slight naturalness improvement over discontinuous F0 pipelines in ultrasound SSI.

BibTeX
@misc{ultrasound-based-silent-speech-interface-built-on-a-continuous-vocoder,
  title = {Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder},
  author = {Tamás Gábor Csapó and Mohammed Salah Al-Radhi and Géza Németh and Gábor Gosztolya and Tamás Grósz and László Tóth and Alexandra Markó},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1906.09885},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1906.09885v1},
}
arXiv reviewed 1 citations · 2026-09-12

Video-Driven Speech Reconstruction using Generative Adversarial Networks

Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, Maja Pantić

Foundational direct video-to-audio result with clear generalization limits.

BibTeX
@misc{video-driven-speech-reconstruction-using-generative-adversarial-networks,
  title = {Video-Driven Speech Reconstruction using Generative Adversarial Networks},
  author = {Konstantinos Vougioukas and Pingchuan Ma and Stavros Petridis and Maja Pantić},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1906.06301},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1906.06301v1},
}
arXiv reviewed 1 citations · 2026-09-12

A Novel Task-Oriented Text Corpus in Silent Speech Recognition and its Natural Language Generation Construction Method

Dong Cao, Dongdong Zhang, Haibo Chen

Useful EEG-SSR corpus framing paper, but evidence is lighter than a full benchmark paper.

BibTeX
@misc{a-novel-task-oriented-text-corpus-in-silent-speech-recognition-and-its-natural-language-generation-construction-method,
  title = {A Novel Task-Oriented Text Corpus in Silent Speech Recognition and its Natural Language Generation Construction Method},
  author = {Dong Cao and Dongdong Zhang and Haibo Chen},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1905.01974},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1905.01974v1},
}
arXiv reviewed 3 citations · 2026-09-12

Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces

Gábor Gosztolya, Ádám Pintér, László Tóth, Tamás Grósz, Alexandra Markó, Tamás Gábor Csapó

The paper advances ultrasound silent speech interfaces by compressing ultrasound images using an autoencoder bottleneck prior to spectral parameter prediction, resulting in improved accuracy and more natural synthesized speech with smaller models.

BibTeX
@misc{autoencoder-based-articulatory-to-acoustic-mapping-for-ultrasound-silent-speech-interfaces,
  title = {Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces},
  author = {Gábor Gosztolya and Ádám Pintér and László Tóth and Tamás Grósz and Alexandra Markó and Tamás Gábor Csapó},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1904.05259},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1904.05259v1},
}
arXiv reviewed 27 citations · 2026-09-12

Denoising convolutional autoencoder based B-mode ultrasound tongue image feature extraction

Bo Li, Kele Xu, Dawei Feng, Haibo Mi, Huaimin Wang, Jian Zhu

DCAE provides cleaner, more robust ultrasound tongue features leading to improved silent speech recognition, outperforming prior feature extraction strategies.

BibTeX
@misc{denoising-convolutional-autoencoder-based-b-mode-ultrasound-tongue-image-feature-extraction,
  title = {Denoising convolutional autoencoder based B-mode ultrasound tongue image feature extraction},
  author = {Bo Li and Kele Xu and Dawei Feng and Haibo Mi and Huaimin Wang and Jian Zhu},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1903.00888},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1903.00888v1},
}
arXiv reviewed 13 citations · 2026-09-12

All-neural online source separation, counting, and diarization for meeting analysis

Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani, Reinhold Haeb‐Umbach

Strong online diarization/separation paper, but outside SSI.

BibTeX
@misc{all-neural-online-source-separation-counting-and-diarization-for-meeting-analysis,
  title = {All-neural online source separation, counting, and diarization for meeting analysis},
  author = {Thilo von Neumann and Keisuke Kinoshita and Marc Delcroix and Shoko Araki and Tomohiro Nakatani and Reinhold Haeb‐Umbach},
  year = {2019},
  note = {arXiv / imported corpus page},
  eprint = {1902.07881},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1902.07881v1},
}
CHI 2019 reviewed 123 citations · 2026-09-12

SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks

Naoki Kimura, Michinari Kono, Jun Rekimoto

A solid proof of concept that reconstructs speech audio from ultrasound for controlling unmodified smart speakers, showcasing important system design insight despite prototype limitations in latency, hardware bulk, and speaker dependency.

BibTeX
@misc{sottovoce,
  title = {SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks},
  author = {Naoki Kimura and Michinari Kono and Jun Rekimoto},
  year = {2019},
  note = {CHI '19},
  doi = {10.1145/3290605.3300376},
  url = {https://doi.org/10.1145/3290605.3300376},
}
arXiv reviewed 0 citations · 2026-09-12

Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold

Iroro Orife, Shane Walker, Jason Flaks

Strong telephony anti-SPAM paper, not SSI.

BibTeX
@misc{audio-spectrogram-factorization-for-classification-of-telephony-signals-below-the-auditory-threshold,
  title = {Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold},
  author = {Iroro Orife and Shane Walker and Jason Flaks},
  year = {2018},
  note = {arXiv / imported corpus page},
  eprint = {1811.04139},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1811.04139v1},
}
arXiv reviewed 1 citations · 2026-09-12

Proactive Security: Embedded AI Solution for Violent and Abusive Speech Recognition

Christopher Shulby, Leonardo Pombal, Vitor Jordão, Guilherme Ziolle, Bruno Martho, Antônio Postal, Thiago Prochnow

An embedded smartphone NLP classifier detects violent speech with ~87.5% accuracy using known methods but is unrelated to silent speech interfaces; strong practical application in safety alerting.

BibTeX
@misc{proactive-security-embedded-ai-solution-for-violent-and-abusive-speech-recognition,
  title = {Proactive Security: Embedded AI Solution for Violent and Abusive Speech Recognition},
  author = {Christopher Shulby and Leonardo Pombal and Vitor Jordão and Guilherme Ziolle and Bruno Martho and Antônio Postal and Thiago Prochnow},
  year = {2018},
  note = {arXiv / imported corpus page},
  eprint = {1810.09431},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1810.09431v1},
}
arXiv reviewed 16 citations · 2026-09-12

Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed

Yaman Kumar, Mayank Aggarwal, Pratham Nawal, Shin'ichi Satoh, Rajiv Ratn Shah, Roger Zimmermann

Multi-view silent video combined with CNN-LSTM models significantly improves speech audio reconstruction quality over single-view, highlighting the importance of optimal camera placement to address pose variance.

BibTeX
@misc{harnessing-ai-for-speech-reconstruction-using-multi-view-silent-video-feed,
  title = {Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed},
  author = {Yaman Kumar and Mayank Aggarwal and Pratham Nawal and Shin'ichi Satoh and Rajiv Ratn Shah and Roger Zimmermann},
  year = {2018},
  note = {arXiv / imported corpus page},
  doi = {10.1145/3240508.3241911},
  eprint = {1807.00619},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1145/3240508.3241911},
}
arXiv reviewed 1 citations · 2026-09-12

Visual-Only Recognition of Normal, Whispered and Silent Speech

Stavros Petridis, Jie Shen, Doruk Cetin, Maja Pantić

Strong evidence that silent lipreading needs dedicated training.

BibTeX
@misc{visual-only-recognition-of-normal-whispered-and-silent-speech,
  title = {Visual-Only Recognition of Normal, Whispered and Silent Speech},
  author = {Stavros Petridis and Jie Shen and Doruk Cetin and Maja Pantić},
  year = {2018},
  note = {arXiv / imported corpus page},
  eprint = {1802.06399},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1802.06399v1},
}
arXiv reviewed 43 citations · 2026-09-12

Cross-modal Embeddings for Video and Audio Retrieval

Dídac Surís, Amanda Duarte, Amaia Salvador, Jordi Torres, Giró Nieto, Xavier

Useful multimodal retrieval baseline, not SSI.

BibTeX
@misc{cross-modal-embeddings-for-video-and-audio-retrieval,
  title = {Cross-modal Embeddings for Video and Audio Retrieval},
  author = {Dídac Surís and Amanda Duarte and Amaia Salvador and Jordi Torres and Giró Nieto, Xavier},
  year = {2018},
  note = {arXiv / imported corpus page},
  eprint = {1801.02200},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1801.02200v1},
}
arXiv reviewed 9 citations · 2026-09-12

Lip2AudSpec: Speech reconstruction from silent lip movements video

Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani

The paper's auditory spectrogram autoencoder bottleneck target is a key innovation that produces more intelligible, natural reconstructed speech from lip videos than prior methods, as confirmed by objective and human evaluations.

BibTeX
@misc{lip2audspec-speech-reconstruction-from-silent-lip-movements-video,
  title = {Lip2AudSpec: Speech reconstruction from silent lip movements video},
  author = {Hassan Akbari and Himani Arora and Liangliang Cao and Nima Mesgarani},
  year = {2017},
  note = {arXiv / imported corpus page},
  eprint = {1710.09798},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1710.09798v1},
}
arXiv reviewed 59 citations · 2026-09-12

Updating the silent speech challenge benchmark with deep learning

Yan Ji, Licheng Liu, Hongcui Wang, Zhilei Liu, Zhibin Niu, B. Denby

Benchmark update with a real, reproducible WER gain.

BibTeX
@misc{updating-the-silent-speech-challenge-benchmark-with-deep-learning,
  title = {Updating the silent speech challenge benchmark with deep learning},
  author = {Yan Ji and Licheng Liu and Hongcui Wang and Zhilei Liu and Zhibin Niu and B. Denby},
  year = {2017},
  note = {arXiv / imported corpus page},
  eprint = {1709.06818},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1709.06818v1},
}
arXiv reviewed 6 citations · 2026-09-12

Seeing Through Noise: Visually Driven Speaker Separation and Enhancement

Aviv Gabbay, Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Strong audiovisual speech separation and enhancement leveraging face video for speaker-dependent masking; not a silent speech interface paper.

BibTeX
@misc{seeing-through-noise-visually-driven-speaker-separation-and-enhancement,
  title = {Seeing Through Noise: Visually Driven Speaker Separation and Enhancement},
  author = {Aviv Gabbay and Ariel Ephrat and Tavi Halperin and Shmuel Peleg},
  year = {2017},
  note = {arXiv / imported corpus page},
  eprint = {1708.06767},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1708.06767v3},
}
arXiv reviewed 8 citations · 2026-09-12

Improved Speech Reconstruction from Silent Video

Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Strong, benchmark-setting speaker-dependent video-to-speech system that advances speech reconstruction from silent face video but remains limited to per-speaker training and constrained conditions.

BibTeX
@misc{improved-speech-reconstruction-from-silent-video,
  title = {Improved Speech Reconstruction from Silent Video},
  author = {Ariel Ephrat and Tavi Halperin and Shmuel Peleg},
  year = {2017},
  note = {arXiv / imported corpus page},
  eprint = {1708.01204},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1708.01204v3},
}
arXiv reviewed 115 citations · 2026-09-12

Vid2speech: Speech Reconstruction from Silent Video

Ariel Ephrat, Shmuel Peleg

Real lip-to-speech progress, still tightly benchmark-bounded.

BibTeX
@misc{vid2speech-speech-reconstruction-from-silent-video,
  title = {Vid2speech: Speech Reconstruction from Silent Video},
  author = {Ariel Ephrat and Shmuel Peleg},
  year = {2017},
  note = {arXiv / imported corpus page},
  eprint = {1701.00495},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1701.00495v2},
}
arXiv reviewed 6 citations · 2026-09-12

Contour-based 3d tongue motion visualization using ultrasound image sequences

Kele Xu, Yin Yang, Clémence Leboullenger, Pierre Roussel, B. Denby

Useful tongue-modeling tool, not a recognizer.

BibTeX
@misc{contour-based-3d-tongue-motion-visualization-using-ultrasound-image-sequences,
  title = {Contour-based 3d tongue motion visualization using ultrasound image sequences},
  author = {Kele Xu and Yin Yang and Clémence Leboullenger and Pierre Roussel and B. Denby},
  year = {2016},
  note = {arXiv / imported corpus page},
  eprint = {1605.05967},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/1605.05967v1},
}
arXiv reviewed 1 citations · 2026-09-12

Optimal Power Control for Analog Bidirectional Relaying with Long-Term Relay Power Constraint

Zoran Hadzi-Velkov, Nikola Zlatanov, Robert Schober

A rigorous relay power control theory paper optimizing outage under long-term average power constraints for bidirectional AF relaying; solid mathematical contribution but outside SSI relevance.

BibTeX
@misc{optimal-power-control-for-analog-bidirectional-relaying-with-long-term-relay-power-constraint,
  title = {Optimal Power Control for Analog Bidirectional Relaying with Long-Term Relay Power Constraint},
  author = {Zoran Hadzi-Velkov and Nikola Zlatanov and Robert Schober},
  year = {2014},
  note = {arXiv / imported corpus page},
  doi = {10.1109/GLOCOM.2013.6831710},
  eprint = {1404.0906},
  archivePrefix = {arXiv},
  url = {https://doi.org/10.1109/GLOCOM.2013.6831710},
}

Approach comparison

Compare SilentSpeller, SottoVoce, and NasoVoce side by side without treating them as a required reading order.

Research agenda

The reviewed papers keep recurring on wearability, vocabulary, latency, and generalization. The page below turns that into a short, grounded agenda.

agenda4 recurring gaps

Open problems and research agenda

Wearability, open vocabulary, real-time use, and generalization keep reappearing in the current review set.

Technique taxonomy

These pages group the current database by real `modality:` tags from the expert records.

modality:video49 pages

Video

49 reviewed pages · 0 imported pages

modality:acoustic34 pages

Acoustic

34 reviewed pages · 0 imported pages

modality:multimodal24 pages

Multimodal

24 reviewed pages · 0 imported pages

modality:emg22 pages

EMG

22 reviewed pages · 0 imported pages

modality:eeg20 pages

EEG

20 reviewed pages · 0 imported pages

modality:ultrasound18 pages

Ultrasound

18 reviewed pages · 0 imported pages

modality:microphone9 pages

Microphone

9 reviewed pages · 0 imported pages

modality:magnetic6 pages

Magnetic

6 reviewed pages · 0 imported pages

modality:camera4 pages

Camera

4 reviewed pages · 0 imported pages

modality:radar4 pages

Radar

4 reviewed pages · 0 imported pages

modality:vibration2 pages

Vibration

2 reviewed pages · 0 imported pages

Machine-readable exports

These files are generated from repository inputs during build.

JSON exportmachine-readable

SSI review export

Snapshot JSON of the current SSI review records built from repository inputs.

JSON feedmachine-readable

SSI review feed

Snapshot feed of the current SSI review records with source-updated timestamps.

Reference and citation

Use the canonical citation page when you need the database name, maintainer, or last-updated date.

referencecanonical citation

How to cite this database

Canonical citation page for the SSI review database. Last updated 2026-10-05.

Datasets and code resources

Verified links are grouped on a dedicated page so the current corpus can point to code, datasets, and paper pages without inventing any new metadata.

resources6 code-linked6 dataset-linked

Datasets and code resources

Verified links already present in repository data, with paper pages attached wherever the archive has a local review page.

FAQ

Answers grounded in the site's data policy and review methodology.

Silent speechとは何ですか?

Silent speech interface (SSI) は、発声せずに口・喉・顔などの動きをセンサーで読み取り、意図した発話をテキストや音声に変換する技術です。マイクを使わないため、周囲に声を出せない場面や発話が難しい場面でも使えます。

このデータベースの論文はどう選ばれていますか?

リポジトリのarXiv/引用データを基に収集し、対応する論文をNaoki Kimuraが実際に読んで評価しています。評価が済んだものは「reviewed」、まだ評価本文が無いものは「imported corpus」として区別しています。

「reviewed」と「imported corpus」ページの違いは何ですか?

reviewed pageはNaoki Kimuraによる評価本文(強み・弱み・エビデンス付き)を含みます。imported corpus pageは書誌情報とベンチマーク値のみを保持し、評価本文はまだ付いていません。

citation countはどこから来ていますか?

OpenAlexから取得しています。実際に0件と確認できた論文は「0 citations」、まだ取得できていない論文は「not yet fetched」と表示し、混同しません。

レビューは誰が書いていますか?

Naoki Kimura本人による専門家評価です。レビューの見方は /papers/rubric にまとめています。

データポリシーは何ですか?

実データのみを使用し、捏造した指標・主張・受賞歴・リンクは掲載しません。詳細は /papers/cite に記載しています。

JSON形式でデータを取得できますか?

できます。/exports/ssi-review.json(スナップショット)と /feeds/ssi-review.json(フィード)で機械可読形式を提供しています。

このデータベースを引用する際の書式は?

/papers/cite に、データベース名・メンテナー・推奨citation文字列を掲載しています。