Camera
This page groups the current SSI review database by the real `modality:` tag `modality:camera`.
The list below includes every paper page that currently carries this technique label.
Papers
AVSRBench: A Multi-Condition AVSR Benchmark
LRS3のsub-1% AV WERは放送ドメインの指標。六条件比較では視覚のみが領域外で崩壊し、AV融合の明確な利点は主にLombard。RoomReader-AV(6.49h・10,324発話・118人)は会議会話の厳しさを示す。装着SSIの代替ではない。
A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
イベントストリームのマルチ話者VSR。DVS-LipでWER 22.3%・VER 19.8%・240 ms。カメラVTPとも装着SSIともセンサが異なる。
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
Combines one users silent EMG with face-conditioned target voices without inference audio. Pitch flattening modestly helps silent word accuracy, but multi-user decoding and faithful personal-voice recovery are not demonstrated.
Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed
Multi-view silent video combined with CNN-LSTM models significantly improves speech audio reconstruction quality over single-view, highlighting the importance of optimal camera placement to address pose variance.