Multimodal Speech Recognition Techniques in Silent Speech
A bounded evidence map of systems that combine video, ultrasound, EEG, EMG, acoustic, or vibration signals for speech-related tasks. The list also contains adjacent audiovisual papers, so each review—not the multimodal tag alone—determines relevance to silent-speech interfaces.
The list below includes every paper page that currently carries this technique label.
- What it is
- A tag-based collection from the expert-reviewed silent-speech corpus. It brings together multimodal speech recognition, reconstruction, enhancement, and adjacent audiovisual work while keeping their tasks and evidence boundaries visible.
- Who it’s for
- Speech, HCI, and multimodal-learning readers comparing sensing combinations, outputs, user dependence, and evaluation limits before opening individual papers.
- Verdict
- Use this page as an evidence map, not a leaderboard. Ultrasound-plus-video or EEG-plus-EMG speech studies are not directly comparable with whispered acoustic-vibration systems, video-to-speech, or general audiovisual generation.
What counts as multimodal here?
For this database, multimodal means that a reviewed record carries the `modality:multimodal` tag. In core SSI work, that can mean combining silent articulatory or physiological signals—for example ultrasound tongue imaging with lip video, or EEG with EMG—for recognition or reconstruction.
The same tag also covers SSI-adjacent systems, such as low-audibility acoustic-plus-vibration input, visual speech recognition, video-to-speech, and audiovisual speech enhancement. These papers can inform technique choices without proving performance for a fully silent interface.
Some tagged records study broader audiovisual generation, localization, retrieval, or segmentation rather than speech communication. They remain listed for corpus transparency, but their individual expert reviews mark the boundary from core SSI evidence.
How to compare the techniques
Open the individual reviews and compare like with like. Five fields prevent a broad ‘multimodal’ label from hiding important differences:
- Task and output
- Recognition to text or labels, reconstructed speech, enhancement, and general audio generation answer different questions.
- Inputs and body site
- Video, ultrasound, EEG, EMG, microphones, and vibration sensors observe different signals and impose different hardware constraints.
- Role of each modality
- Check whether a modality is required at inference, used only for training, or supplies complementary evidence during fusion.
- User dependence
- Separate person-specific systems from evaluations that hold out speakers or participants.
- Evaluation boundary
- Dataset, vocabulary, recording condition, real-time status, and expert-noted limits determine how far a result can be generalized.
Papers
AESSI: An Around-Ear Silent Speech Interface for Cross-Day Online Reuse without Test-Day Calibration
耳周囲cEEGridで25の固定中国語文を認識。テスト日較正なしの別日ホルドアウトで平均92.24%、21日以上後のライブ250試行で98.00%。開語彙会話や患者適用は未実証。
AVSRBench: A Multi-Condition AVSR Benchmark
LRS3のsub-1% AV WERは放送ドメインの指標。六条件比較では視覚のみが領域外で崩壊し、AV融合の明確な利点は主にLombard。RoomReader-AV(6.49h・10,324発話・118人)は会議会話の厳しさを示す。装着SSIの代替ではない。
A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
イベントストリームのマルチ話者VSR。DVS-LipでWER 22.3%・VER 19.8%・240 ms。カメラVTPとも装着SSIともセンサが異なる。
NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
A strong deployment-focused speech interface leveraging a novel nose-pad dual-sensor configuration and multimodal fusion to enable robust low-audibility speech interaction with AI under noise, backed by extensive evaluation.
Poster: Recognizing Hidden-in-the-Ear Private Key for Reliable Silent Speech Interface Using Multi-Task Learning
HEar-ID jointly models ear-based spelling and identity, but all-user Top-1 is 67.3%, not the selected eight-user 90.25%; whisper dependence, participant failures, modified hardware, and untested attack resistance limit deployment claims.
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
文章と参考音声を使い、唇に合う音声を作る時間制御の研究。単語誤り率と同期指標は改善するが、唇だけから内容を読む方式ではなく、数値の不整合と無声利用の未検証が残る。
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding
脳波と筋電を組み合わせ、声を出さずに発音した中国語の四声を分類する研究。学習に含まない人で平均85.10%を報告するが、文章認識ではなく、指標名や分割・チャネル選択手順には確認が必要。
Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication
Held-out sentence reconstruction is demonstrated in personalized EEG/EMG experiments, but the strongest aggregate evidence is overt/whispered phoneme decoding—not unrestricted imagined-speech communication.
An Introduction to Silent Paralinguistics
声を出さない会話で、言葉だけでなく感情や話し方も伝えるための総説。研究課題の整理は有用だが、感情の復元精度や実用効果を新たに実証した論文ではない。
A Silent Speech Decoding System from EEG and EMG with Heterogenous Electrode Configurations
電極配置の違う脳波・筋電データの統合学習は64語分類を改善するが、患者1人の本人別評価であり、別日・自由な会話・臨床効果への隔たりが残る。
SonicVisionLM: Playing Sound with Vision Language Models
A high-quality video-to-audio generation framework leveraging vision-language models for editable, temporally precise sound effect generation; strong experimental validations but outside standard SSI scope.
Sound Source Localization is All about Cross-Modal Alignment
Provides a novel multi-positive contrastive framework enhancing semantic audio-visual alignment for sound source localization. Strong experimental evidence supports claims. Method is outside the SSI domain.
Audio-visual video-to-speech synthesis with synthesized input audio
The paper credibly shows that incorporating synthesized audio as an auxiliary input in a second-stage audiovisual synthesis model improves video-to-speech reconstruction quality and intelligibility in benchmarks, though gains depend on model variant and dataset.
Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation
Strong AVS result, outside SSI: the useful idea is audio-conditioned decoder queries plus dynamic mask prediction.
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
The real gain is not 'diffusion' alone but aligned conditioning plus guidance that pushes synchronization very hard.
Conditional Generation of Audio from Video via Foley Analogies
The paper matters because it gives V2A generation a controllable exemplar, not because it beats every timing baseline.
Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training
Strong SSI paper improving silent speech reconstruction by generating pseudo acoustic targets and using domain adversarial training to address domain mismatch; validated with TaL dataset showing substantial WER and MOS gains over TaLNet.
Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
The key idea is not generic fusion; it is storing cross-modal correspondences so video-only decoding can recover some audio-side structure later.
Silent versus modal multi-speaker speech recognition from ultrasound and video
Large-corpus baseline with real silent-mode gap.
Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Technically solid self-supervised class-aware audiovisual sounding object localization, but outside the core SSI domain.
Silent Speech Interfaces for Speech Restoration: A Review
Core SSI survey with concrete deployment constraints.
An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
Strong AV speech survey, not an SSI system paper.
Foley Music: Learning to Generate Music from Videos
Strong video-to-music paper, not SSI.
Cross-modal Embeddings for Video and Audio Retrieval
Useful multimodal retrieval baseline, not SSI.