← Technique taxonomy

modality:multimodal 24 pages 24 reviewed 0 imported

Multimodal Speech Recognition Techniques in Silent Speech

A bounded evidence map of systems that combine video, ultrasound, EEG, EMG, acoustic, or vibration signals for speech-related tasks. The list also contains adjacent audiovisual papers, so each review—not the multimodal tag alone—determines relevance to silent-speech interfaces.

The list below includes every paper page that currently carries this technique label.

What it is
A tag-based collection from the expert-reviewed silent-speech corpus. It brings together multimodal speech recognition, reconstruction, enhancement, and adjacent audiovisual work while keeping their tasks and evidence boundaries visible.
Who it’s for
Speech, HCI, and multimodal-learning readers comparing sensing combinations, outputs, user dependence, and evaluation limits before opening individual papers.
Verdict
Use this page as an evidence map, not a leaderboard. Ultrasound-plus-video or EEG-plus-EMG speech studies are not directly comparable with whispered acoustic-vibration systems, video-to-speech, or general audiovisual generation.

What counts as multimodal here?

For this database, multimodal means that a reviewed record carries the `modality:multimodal` tag. In core SSI work, that can mean combining silent articulatory or physiological signals—for example ultrasound tongue imaging with lip video, or EEG with EMG—for recognition or reconstruction.

The same tag also covers SSI-adjacent systems, such as low-audibility acoustic-plus-vibration input, visual speech recognition, video-to-speech, and audiovisual speech enhancement. These papers can inform technique choices without proving performance for a fully silent interface.

Some tagged records study broader audiovisual generation, localization, retrieval, or segmentation rather than speech communication. They remain listed for corpus transparency, but their individual expert reviews mark the boundary from core SSI evidence.

How to compare the techniques

Open the individual reviews and compare like with like. Five fields prevent a broad ‘multimodal’ label from hiding important differences:

Task and output
Recognition to text or labels, reconstructed speech, enhancement, and general audio generation answer different questions.
Inputs and body site
Video, ultrasound, EEG, EMG, microphones, and vibration sensors observe different signals and impose different hardware constraints.
Role of each modality
Check whether a modality is required at inference, used only for training, or supplies complementary evidence during fusion.
User dependence
Separate person-specific systems from evaluations that hold out speakers or participants.
Evaluation boundary
Dataset, vocabulary, recording condition, real-time status, and expert-noted limits determine how far a result can be generalized.

Papers

reviewedarXiv2026

AVSRBench: A Multi-Condition AVSR Benchmark

Rishabh Jain, Naomi Harte

LRS3のsub-1% AV WERは放送ドメインの指標。六条件比較では視覚のみが領域外で崩壊し、AV融合の明確な利点は主にLombard。RoomReader-AV(6.49h・10,324発話・118人)は会議会話の厳しさを示す。装着SSIの代替ではない。

reviewedECCV 20262026

A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR

Lin Chen, Jingping Fang, Hairui Liu, Chenyang Xu, Junhao Chen, Xiaorui Li, Weidong Cai, Xiaoming Chen

イベントストリームのマルチ話者VSR。DVS-LipでWER 22.3%・VER 19.8%・240 ms。カメラVTPとも装着SSIともセンサが異なる。

reviewedarXiv2025

VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task

Yuyue Wang, Xin Cheng, Yihan Wu, Xihua Wang, Jinchuan Tian, Ruihua Song

文章と参考音声を使い、唇に合う音声を作る時間制御の研究。単語誤り率と同期指標は改善するが、唇だけから内容を読む方式ではなく、数値の不整合と無声利用の未検証が残る。

reviewedarXiv2025

CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding

Yifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou, Jiawei Ju

脳波と筋電を組み合わせ、声を出さずに発音した中国語の四声を分類する研究。学習に含まない人で平均85.10%を報告するが、文章認識ではなく、指標名や分割・チャネル選択手順には確認が必要。

reviewedarXiv2025

An Introduction to Silent Paralinguistics

Zhao Ren, Simon Pistrosch, Buket Coşkun, Kevin Scheck, Anton Batliner, Björn W. Schuller, Tanja Schultz

声を出さない会話で、言葉だけでなく感情や話し方も伝えるための総説。研究課題の整理は有用だが、感情の復元精度や実用効果を新たに実証した論文ではない。

reviewedarXiv / imported corpus page2024

SonicVisionLM: Playing Sound with Vision Language Models

Zhifeng Xie, Shengye Yu, Qile He, Mengtian Li

A high-quality video-to-audio generation framework leveraging vision-language models for editable, temporally precise sound effect generation; strong experimental validations but outside standard SSI scope.

reviewedarXiv / imported corpus page2023

Sound Source Localization is All about Cross-Modal Alignment

Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung

Provides a novel multi-positive contrastive framework enhancing semantic audio-visual alignment for sound source localization. Strong experimental evidence supports claims. Method is outside the SSI domain.

reviewedarXiv / imported corpus page2023

Audio-visual video-to-speech synthesis with synthesized input audio

Triantafyllos Kefalas, Yannis Panagakis, Maja Pantić

The paper credibly shows that incorporating synthesized audio as an auxiliary input in a second-stage audiovisual synthesis model improves video-to-speech reconstruction quality and intelligibility in benchmarks, though gains depend on model variant and dataset.