Silent speech papers by publication data.
Silent speech interfaces (SSI) let people communicate without vocalizing, using sensors on the face, throat, or mouth instead of a microphone to recognize intended speech.
サイレントスピーチ(無声発話)インタフェースは、発声せずに口・喉・顔の動きをセンサーで読み取り、 意図した発話を認識する技術です。
Browse the SSI review database by year, citation count, title, author, or page type.
Paper pages are expert evaluations, not abstract reposts. Citation counts come from OpenAlex when available.
Browse papers
107 papers shown
Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech
Introducing PNA, the paper advances non-invasive brain-to-speech decoding by creating artifact-informed augmentations via ICA, significantly improving imagined speech classification accuracy on MEG data when combined with trial averaging.
BibTeX
@misc{physiological-noise-augmentation-improves-non-invasive-brain-to-speech,
title = {Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech},
author = {Benjamin Ballyk and Teyun Kwon and Miran Özdogan and Oiwi Parker Jones},
year = {2026},
note = {arXiv},
eprint = {2607.05165},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2607.05165v1},
} Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
The paper advances silent speech synthesis by leveraging masked training to robustly fuse electromyography and lipreading, showing improved performance and resilience, but adaptation to laryngectomized users remains challenging.
BibTeX
@misc{cross-modal-masking-for-robust-silent-speech-synthesis-using-semg-and-lipreading,
title = {Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading},
author = {Eder del Blanco and David Gimeno-Gómez and Eva Navas and Carlos-D. Martínez-Hinarejos and Inma Hernáez},
year = {2026},
note = {arXiv},
eprint = {2606.09667},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2606.09667v1},
} A 1000-hour EEG-EMG-audio dataset of Japanese speech production
A 1020-hour multimodal EEG-EMG-audio dataset for Japanese overt speech vastly expands data resources, enabling diverse speech decoding and EEG research, though generalization is limited by three participants and no decoding benchmarks are presented.
BibTeX
@misc{a-1000-hour-eeg-emg-audio-dataset-of-japanese-speech-production,
title = {A 1000-hour EEG-EMG-audio dataset of Japanese speech production},
author = {Motoshige Sato and Ilya Horiguchi and Masakazu Inoue and Kenichi Tomeoka and Eri Hatakeyama and Yuya Kita and Atsushi Yamamoto and Ippei Fujisawa and Shuntaro Sasai},
year = {2026},
note = {arXiv},
eprint = {2606.01264},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2606.01264v1},
} Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping
The study convincingly shows zero-shot imagined speech decoding by mapping MEG imagery to listened responses and decoding with a listened-trained contrastive model, marking a promising data-efficient advance despite limited vocabulary and hardware constraints.
BibTeX
@misc{zero-shot-imagined-speech-decoding-via-imagined-to-listened-meg-mapping,
title = {Zero-Shot Imagined Speech Decoding via Imagined-to-Listened MEG Mapping},
author = {Maryam Maghsoudi and Shihab Shamma},
year = {2026},
note = {arXiv},
eprint = {2605.08075},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2605.08075v1},
} NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction
A strong deployment-focused speech interface leveraging a novel nose-pad dual-sensor configuration and multimodal fusion to enable robust low-audibility speech interaction with AI under noise, backed by extensive evaluation.
BibTeX
@misc{nasovoce,
title = {NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech Interaction},
author = {Jun Rekimoto and Yu Nishimura and Bojian Yang},
year = {2026},
note = {CHI '26 / arXiv},
doi = {10.1145/3772318.3791397},
eprint = {2603.10324},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3772318.3791397},
} A cross-species neural foundation model for end-to-end speech decoding
Introduces a cross-species pretrained transformer encoder enabling state-of-the-art end-to-end neural speech decoding with audio-LLMs, improving accuracy and enabling imagined speech decoding, but latency and real-time deployment remain challenges.
BibTeX
@misc{a-cross-species-neural-foundation-model-for-end-to-end-speech-decoding,
title = {A cross-species neural foundation model for end-to-end speech decoding},
author = {Yizi Zhang and Linyang He and Chaofei Fan and Tingkai Liu and Han Yu and Trung Le and Jingyuan Li and Scott Linderman and Lea Duncker and Francis R Willett and Nima Mesgarani and Liam Paninski},
year = {2025},
note = {arXiv},
eprint = {2511.21740},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2511.21740v5},
} SonicVisionLM: Playing Sound with Vision Language Models
A high-quality video-to-audio generation framework leveraging vision-language models for editable, temporally precise sound effect generation; strong experimental validations but outside standard SSI scope.
BibTeX
@misc{sonicvisionlm-playing-sound-with-vision-language-models,
title = {SonicVisionLM: Playing Sound with Vision Language Models},
author = {Zhifeng Xie and Shengye Yu and Qile He and Mengtian Li},
year = {2024},
note = {arXiv / imported corpus page},
eprint = {2401.04394},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2401.04394v1},
} IR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrases
This paper introduces FERASEC, a novel radar feature extraction enabling the first contactless IR-UWB radar phoneme-level silent speech recognition with 86% vowel and 81% consonant accuracy, surpassing raw signal baselines and signifying a key advance in practical silent speech interfaces.
BibTeX
@misc{ir-uwb-radar-based-contactless-silent-speech-recognition-of-vowels-consonants-words-and-phrases,
title = {IR-UWB Radar-Based Contactless Silent Speech Recognition of Vowels, Consonants, Words, and Phrases},
author = {Sunghwa Lee and Younghoon Shin and Myungjong Kim and Jiwon Seo},
year = {2023},
note = {arXiv / imported corpus page},
doi = {10.1109/ACCESS.2023.3344177},
eprint = {2312.09572},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/ACCESS.2023.3344177},
} Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
Strong SSI system combining a novel ultrasensitive throat textile strain sensor with an efficient 1D residual CNN, achieving high word classification accuracy with low computational cost and promising few-shot transfer to new users and words on small vocabularies.
BibTeX
@misc{ultrasensitive-textile-strain-sensors-redefine-wearable-silent-speech-interfaces-with-high-machine-learning-efficiency,
title = {Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency},
author = {Chenyu Tang and Muzi Xu and Wentian Yi and Zibo Zhang and Edoardo Occhipinti and Chaoqun Dong and Dafydd Ravenscroft and Sung‐Min Jung and Sanghyo Lee and Shuo Gao and Jong Min Kim and Luigi G. Occhipinti},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2311.15683},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2311.15683v1},
} Distributed pressure matching strategy using diffusion adaptation
Distributed rootless pressure matching for personal sound zones is presented and validated in simulation, not an SSI paper.
BibTeX
@misc{distributed-pressure-matching-strategy-using-diffusion-adaptation,
title = {Distributed pressure matching strategy using diffusion adaptation},
author = {Mengfei Zhang and Junqing Zhang and Jie Chen and Cédric Richard},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2311.07729},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2311.07729v1},
} Advancing Test-Time Adaptation for Acoustic Foundation Models in Open-World Shifts
Strong acoustic ASR paper proposing confidence-weighted frame adaptation plus temporal consistency regularization for stable test-time adaptation under wild acoustic conditions, yielding substantial WER improvements across noise, accents, and singing datasets.
BibTeX
@misc{advancing-test-time-adaptation-for-acoustic-foundation-models-in-open-world-shifts,
title = {Advancing Test-Time Adaptation for Acoustic Foundation Models in Open-World Shifts},
author = {Andy Clark},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2310.09505},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2310.09505v1},
} Sound Source Localization is All about Cross-Modal Alignment
Provides a novel multi-positive contrastive framework enhancing semantic audio-visual alignment for sound source localization. Strong experimental evidence supports claims. Method is outside the SSI domain.
BibTeX
@misc{sound-source-localization-is-all-about-cross-modal-alignment,
title = {Sound Source Localization is All about Cross-Modal Alignment},
author = {Arda Senocak and Hyeonggon Ryu and Junsik Kim and Tae-Hyun Oh and Hanspeter Pfister and Joon Son Chung},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2309.10724},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2309.10724v1},
} Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
Strong lip-to-speech system that reduces ambiguity via SSL linguistic conditioning, variance predictors, and flow-based refinement, achieving near-vocoded naturalness and improved intelligibility on standard datasets.
BibTeX
@misc{let-there-be-sound-reconstructing-high-quality-speech-from-silent-videos,
title = {Let There Be Sound: Reconstructing High Quality Speech from Silent Videos},
author = {Ji-Hoon Kim and Jaehun Kim and Joon Son Chung},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2308.15256},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2308.15256v2},
} An Initial Exploration: Learning to Generate Realistic Audio for Silent Video
Honest exploratory comparison showing transformer-based model outperforms deep-fusion CNN and Wavenet for generating low-to-mid frequency audio from silent video in a small curated dataset; not a speech or SSI paper.
BibTeX
@misc{an-initial-exploration-learning-to-generate-realistic-audio-for-silent-video,
title = {An Initial Exploration: Learning to Generate Realistic Audio for Silent Video},
author = {Matthew Martel and Jackson Wagner},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2308.12408},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2308.12408v1},
} Audio Knowledge Empowered Visual Speech Recognition
The paper advances visual speech recognition by selectively transferring refined linguistic audio knowledge via a learned compact memory and cross-attention injection, improving benchmark WERs over prior audio-assisted methods without requiring audio inputs during inference.
BibTeX
@misc{akvsr-audio-knowledge-empowered-visual-speech-recognition-by-compressing-audio-knowledge-of-a-pretrained-model,
title = {Audio Knowledge Empowered Visual Speech Recognition},
author = {Jeong Hun Yeo and Minsu Kim and Jeongsoo Choi and Dae Hoe Kim and Yong Man Ro},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2308.07593},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2308.07593v2},
} Knowledge Distilled Ensemble Model for sEMG-based Silent Speech Interface
This paper delivers a practical spelling-focused sEMG silent speech system by compressing a ResNet ensemble into a lightweight model achieving 85.9% accuracy on the NATO alphabet with portable hardware, but remains limited to 5 young male subjects and speaker-dependent scenarios.
BibTeX
@misc{knowledge-distilled-ensemble-model-for-semg-based-silent-speech-interface,
title = {Knowledge Distilled Ensemble Model for sEMG-based Silent Speech Interface},
author = {Wenqiang Lai and Qihan Yang and Mao Ye and Endong Sun and Jiangnan Ye},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2308.06533},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2308.06533v1},
} Automatically measuring speech fluency in people with aphasia: first achievements using read-speech data
Strong clinical fluency regression method validated on noisy read speech from aphasia patients; outside core SSI modalities and use-cases.
BibTeX
@misc{automatically-measuring-speech-fluency-in-people-with-aphasia-first-achievements-using-read-speech-data,
title = {Automatically measuring speech fluency in people with aphasia: first achievements using read-speech data},
author = {Lionel Fontan and Typhanie Prince and Aleksandra Nowakowska and Halima Sahraoui and Silvia Martínez‐Ferreiro},
year = {2023},
note = {arXiv / imported corpus page},
doi = {10.1080/02687038.2023.2244728},
eprint = {2308.04763},
archivePrefix = {arXiv},
url = {https://doi.org/10.1080/02687038.2023.2244728},
} Exploring how a Generative AI interprets music
A thorough interpretability analysis reveals that MusicVAE uses only a few dozen latent dimensions to encode music with pitch and rhythm strongly represented in the first two, but the work has no direct relevance to silent speech interfaces.
BibTeX
@misc{exploring-how-a-generative-ai-interprets-music,
title = {Exploring how a Generative AI interprets music},
author = {Gabriela Barenboim and Luigi Del Debbio and Johannes Hirn and Verónica Sanz},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2308.00015},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2308.00015v1},
} Audio-visual video-to-speech synthesis with synthesized input audio
The paper credibly shows that incorporating synthesized audio as an auxiliary input in a second-stage audiovisual synthesis model improves video-to-speech reconstruction quality and intelligibility in benchmarks, though gains depend on model variant and dataset.
BibTeX
@misc{audio-visual-video-to-speech-synthesis-with-synthesized-input-audio,
title = {Audio-visual video-to-speech synthesis with synthesized input audio},
author = {Triantafyllos Kefalas and Yannis Panagakis and Maja Pantić},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2307.16584},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2307.16584v1},
} Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation
Strong AVS result, outside SSI: the useful idea is audio-conditioned decoder queries plus dynamic mask prediction.
BibTeX
@misc{audio-aware-query-enhanced-transformer-for-audio-visual-segmentation,
title = {Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation},
author = {Jinxiang Liu and Chen Ju and Chaofan Ma and Yanfeng Wang and Yu Wang and Ya Zhang},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2307.13236},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2307.13236v1},
} RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations
Strong modular SSL-based lip-to-speech synthesis paper that innovatively maps lip SSL features to disentangled speech embeddings before vocoder synthesis, demonstrating improved intelligibility and robustness across benchmark datasets.
BibTeX
@misc{robustl2s-speaker-specific-lip-to-speech-synthesis-exploiting-self-supervised-representations,
title = {RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations},
author = {Neha Sahipjohn and Neil Shah and Vishal Tambrahalli and Vineet Gandhi},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2307.01233},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2307.01233v1},
} Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
The real gain is not 'diffusion' alone but aligned conditioning plus guidance that pushes synchronization very hard.
BibTeX
@misc{diff-foley-synchronized-video-to-audio-synthesis-with-latent-diffusion-models,
title = {Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models},
author = {Simian Luo and Chuanhao Yan and Chenxu Hu and Hang Zhao},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2306.17203},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2306.17203v1},
} High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units
This video-conditioned AVO system innovatively supervises alignment by predicting discrete speech units rather than reconstructing acoustic features, leading to better lip-sync and speech quality on a single-speaker dataset; however, it is not an SSI interface paper.
BibTeX
@misc{high-quality-automatic-voice-over-with-accurate-alignment-supervision-through-self-supervised-discrete-speech-units,
title = {High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units},
author = {Junchen Lu and Berrak Şişman and Mingyang Zhang and Haizhou Li},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2306.17005},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2306.17005v1},
} Large-scale unsupervised audio pre-training for video-to-speech synthesis
Good decoder-transfer pretraining improves video-to-speech quality on several benchmarks, but WER gains are not consistent. A useful methodological contribution with strong benchmark support, adjacent to SSI rather than a deployable system.
BibTeX
@misc{large-scale-unsupervised-audio-pre-training-for-video-to-speech-synthesis,
title = {Large-scale unsupervised audio pre-training for video-to-speech synthesis},
author = {Triantafyllos Kefalas and Yannis Panagakis and Maja Pantić},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2306.15464},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2306.15464v2},
} LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
Strong full-text paper demonstrating that inference-time text guidance via ASR classifier is key to significantly improved intelligibility in lip-to-speech synthesis on challenging in-the-wild video datasets, outperforming prior baselines.
BibTeX
@misc{lipvoicer-generating-speech-from-silent-videos-guided-by-lip-reading,
title = {LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading},
author = {Yochai Yemini and Aviv Shamsian and Lior Bracha and Sharon Gannot and Ethan Fetaya},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2306.03258},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2306.03258v1},
} Intelligible Lip-to-Speech Synthesis with Speech Units
Speech units as a pseudo-text target enable strong content supervision that substantially cuts WER without text labels, and the multi-input vocoder improves speech quality from blurry mel outputs, yielding a state-of-the-art lip-to-speech system on LRS benchmarks.
BibTeX
@misc{intelligible-lip-to-speech-synthesis-with-speech-units,
title = {Intelligible Lip-to-Speech Synthesis with Speech Units},
author = {Jeongsoo Choi and Minsu Kim and Yong Man Ro},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2305.19603},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2305.19603v1},
} Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks
Strong full-text-backed evidence that most of the gain comes from fast input alignment, not from inventing a new SSI stack.
BibTeX
@misc{adaptation-of-tongue-ultrasound-based-silent-speech-interfaces-using-spatial-transformer-networks,
title = {Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks},
author = {László Tóth and Amin Honarmandi Shandiz and Gábor Gosztolya and Tamás Gábor Csapó},
year = {2023},
note = {the Proceedings of Interspeech 2023},
doi = {10.21437/Interspeech.2023-1607},
eprint = {2305.19130},
archivePrefix = {arXiv},
url = {https://doi.org/10.21437/Interspeech.2023-1607},
} Zero-shot personalized lip-to-speech synthesis with face image based voice control
Demonstrates effective zero-shot voice control in Lip2Speech by leveraging face image-based speaker embeddings, validated on GRID corpus but constrained by dataset vocabulary and speech naturalness.
BibTeX
@misc{zero-shot-personalized-lip-to-speech-synthesis-with-face-image-based-voice-control,
title = {Zero-shot personalized lip-to-speech synthesis with face image based voice control},
author = {Zheng-Yan Sheng and Yang Ai and Zhen-Hua Ling},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2305.14359},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2305.14359v1},
} Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning
Strong viseme-level metric learning approach reduces silent speech VSR errors on a small 10-phrase dataset, notably achieving parity with baselines using much less silent data.
BibTeX
@misc{improving-the-gap-in-visual-speech-recognition-between-normal-and-silent-speech-based-on-metric-learning,
title = {Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning},
author = {Sara Kashiwagi and Keitaro Tanaka and Feng Qi and Shigeo Morishima},
year = {2023},
note = {arXiv / imported corpus page},
doi = {10.21437/Interspeech.2023-370},
eprint = {2305.14203},
archivePrefix = {arXiv},
url = {https://doi.org/10.21437/Interspeech.2023-370},
} Conditional Generation of Audio from Video via Foley Analogies
The paper matters because it gives V2A generation a controllable exemplar, not because it beats every timing baseline.
BibTeX
@misc{conditional-generation-of-audio-from-video-via-foley-analogies,
title = {Conditional Generation of Audio from Video via Foley Analogies},
author = {Yuexi Du and Ziyang Chen and Justin Salamon and Bryan Russell and Andrew Owens},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2304.08490},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2304.08490v1},
} Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training
Strong SSI paper improving silent speech reconstruction by generating pseudo acoustic targets and using domain adversarial training to address domain mismatch; validated with TaL dataset showing substantial WER and MOS gains over TaLNet.
BibTeX
@misc{speech-reconstruction-from-silent-tongue-and-lip-articulation-by-pseudo-target-generation-and-domain-adversarial-training,
title = {Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training},
author = {Rui-Chen Zheng and Yang Ai and Zhen-Hua Ling},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2304.05574},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2304.05574v1},
} WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions
Strong whisper-conversion paper, but it remains whisper-based rather than truly silent SSI.
BibTeX
@misc{wesper-zero-shot-and-realtime-whisper-to-normal-voice-conversion-for-whisper-based-speech-interactions,
title = {WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech Interactions},
author = {Jun Rekimoto},
year = {2023},
note = {Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI '23), April 23--28, 2023},
doi = {10.1145/3544548.3580706},
eprint = {2303.01639},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3544548.3580706},
} Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech
The paper presents a strong multi-speaker TTS phrasing approach leveraging speaker-conditioned BERT embeddings and pause duration categories to improve pause insertion precision and synthetic speech rhythm; however, it is out-of-scope for SSI as it focuses on audible speech synthesis only.
BibTeX
@misc{duration-aware-pause-insertion-using-pre-trained-language-model-for-multi-speaker-text-to-speech,
title = {Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech},
author = {Dong Yang and Tomoki Koriyama and Yuki Saito and Takaaki Saeki and Detai Xin and Hiroshi Saruwatari},
year = {2023},
note = {arXiv / imported corpus page},
eprint = {2302.13652},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2302.13652v1},
} LipLearner: Customizable Silent Speech Interactions on Mobile Devices
LipLearner is a strong mobile silent speech system that uniquely closes the loop from few-shot lipreading model design to practical on-device customization and keyword spotting, demonstrated robustly in real-world conditions and a user study.
BibTeX
@misc{liplearner-customizable-silent-speech-interactions-on-mobile-devices,
title = {LipLearner: Customizable Silent Speech Interactions on Mobile Devices},
author = {Zixiong Su and Shitao Fang and Jun Rekimoto},
year = {2023},
note = {arXiv / imported corpus page},
doi = {10.1145/3544548.3581465},
eprint = {2302.05907},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3544548.3581465},
} Towards Neural Decoding of Imagined Speech based on Spoken Speech
Transfer of CSP+SVM models trained on spoken speech EEG to imagined speech achieves comparable, though slightly lower, accuracy within a limited 5-class, 7-subject offline EEG setup, with visual imagery control supporting specificity.
BibTeX
@misc{towards-neural-decoding-of-imagined-speech-based-on-spoken-speech,
title = {Towards Neural Decoding of Imagined Speech based on Spoken Speech},
author = {Seo‐Hyun Lee and Young-Eun Lee and Soo-Won Kim and Byung-Kwan Ko and Seong‐Whan Lee},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2212.02047},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2212.02047v2},
} Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation
Strong causal PSE paper, not SSI. The pVAD-guided loss is the part that holds up under full-text reading.
BibTeX
@misc{breaking-the-trade-off-in-personalized-speech-enhancement-with-cross-task-knowledge-distillation,
title = {Breaking the trade-off in personalized speech enhancement with cross-task knowledge distillation},
author = {Hassan Taherian and Şefik Emre Eskimez and Takuya Yoshioka},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2211.02944},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2211.02944v1},
} Movement Detection of Tongue and Related Body Parts Using IR-UWB Radar
Good sensing primitive, very small task.
BibTeX
@misc{movement-detection-of-tongue-and-related-body-parts-using-ir-uwb-radar,
title = {Movement Detection of Tongue and Related Body Parts Using IR-UWB Radar},
author = {Sunghwa Lee and Younghoon Shin},
year = {2022},
note = {arXiv / imported corpus page},
doi = {10.1109/ICTC55196.2022.9952644},
eprint = {2209.01762},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/ICTC55196.2022.9952644},
} Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
The real contribution is not just another VAE-GAN; it is turning lip-to-speech into an arbitrary-speaker problem with credible low-data adaptation.
BibTeX
@misc{lip-to-speech-synthesis-for-arbitrary-speakers-in-the-wild,
title = {Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild},
author = {Sindhu B Hegde and K R Prajwal and Rudrabha Mukhopadhyay and Vinay P. Namboodiri and C. V. Jawahar},
year = {2022},
note = {arXiv / imported corpus page},
doi = {10.1145/3503161.3548081},
eprint = {2209.00642},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3503161.3548081},
} An Anchor-Free Detector for Continuous Speech Keyword Spotting
Strong CSKWS paper, not SSI. The detection framing and unknown class are the points that hold up in full text.
BibTeX
@misc{an-anchor-free-detector-for-continuous-speech-keyword-spotting,
title = {An Anchor-Free Detector for Continuous Speech Keyword Spotting},
author = {Zhiyuan Zhao and Chuanxin Tang and Chengdong Yao and Chong Luo},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2208.04622},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2208.04622v1},
} FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
This paper matters because it makes unconstrained lip-to-speech materially faster without obviously sacrificing quality.
BibTeX
@misc{fastlts-non-autoregressive-end-to-end-unconstrained-lip-to-speech-synthesis,
title = {FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis},
author = {Yongqi Wang and Zhou Zhao},
year = {2022},
note = {arXiv / imported corpus page},
doi = {10.1145/3503161.3548194},
eprint = {2207.03800},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3503161.3548194},
} Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks
An empirically supported, incremental advancement showing that hybrid 3D-CNN plus ConvLSTM models modestly outperform prior ultrasound tongue video SSI architectures in mel-spectrogram regression accuracy and model efficiency on single-speaker data.
BibTeX
@misc{improved-processing-of-ultrasound-tongue-videos-by-combining-convlstm-and-3d-convolutional-networks,
title = {Improved Processing of Ultrasound Tongue Videos by Combining ConvLSTM and 3D Convolutional Networks},
author = {Amin Honarmandi Shandiz and László Tóth},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2206.12947},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2206.12947v1},
} VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection
The paper is really about disentangling identity, and that is why the unseen-speaker results hold up.
BibTeX
@misc{visagesyntalk-unseen-speaker-video-to-speech-synthesis-via-speech-visage-feature-selection,
title = {VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection},
author = {Joanna Hong and Minsu Kim and Yong Man Ro},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2206.07458},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2206.07458v2},
} Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information
Strong evidence that silence segments in HuBERT representations uniquely store speaker information, improving SID accuracy when silence is augmented; analytical SSL probing paper outside silent speech interface field.
BibTeX
@misc{silence-is-sweeter-than-speech-self-supervised-model-using-silence-to-store-speaker-information,
title = {Silence is Sweeter Than Speech: Self-Supervised Model Using Silence to Store Speaker Information},
author = {Chi-Luen Feng and Po‐Chun Hsu and Hung-yi Lee},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2205.03759},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2205.03759v1},
} SVTS: Scalable Video-to-Speech Synthesis
A key scaling contribution that demonstrates simple spectrogram prediction plus pretrained vocoder pipelines outperform prior complex models on diverse datasets, marking foundational progress in large-scale video-to-speech synthesis.
BibTeX
@misc{svts-scalable-video-to-speech-synthesis,
title = {SVTS: Scalable Video-to-Speech Synthesis},
author = {Rodrigo Mira and Alexandros Haliassos and Stavros Petridis and Björn W. Schuller and Maja Pantić},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2205.02058},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2205.02058v2},
} Listen only to me! How well can target speech extraction handle false alarms?
Strong paper for false-alarm handling in TSE, wrong domain if someone tries to count it as SSI progress.
BibTeX
@misc{listen-only-to-me-how-well-can-target-speech-extraction-handle-false-alarms,
title = {Listen only to me! How well can target speech extraction handle false alarms?},
author = {Marc Delcroix and Keisuke Kinoshita and Tsubasa Ochiai and Kateřina Žmolíková and Hiroshi Satō and Tomohiro Nakatani},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2204.04811},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2204.04811v2},
} Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video
The key idea is not generic fusion; it is storing cross-modal correspondences so video-only decoding can recover some audio-side structure later.
BibTeX
@misc{multi-modality-associative-bridging-through-memory-speech-sound-recollected-from-face-video,
title = {Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face Video},
author = {Minsu Kim and Joanna Hong and Se Jin Park and Yong Man Ro},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2204.01265},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2204.01265v1},
} VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion
The real move is importing structure from voice conversion, not just adding another speaker embedding.
BibTeX
@misc{vcvts-multi-speaker-video-to-speech-synthesis-via-cross-modal-knowledge-transfer-from-voice-conversion,
title = {VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion},
author = {Disong Wang and Shan Yang and Dan Su and Xunying Liu and Dong Yu and Helen Meng},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2202.09081},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2202.09081v1},
} Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals
Sound classification paper, not SSI.
BibTeX
@misc{supervised-and-self-supervised-pretraining-based-covid-19-detection-using-acoustic-breathing-cough-speech-signals,
title = {Supervised and Self-supervised Pretraining Based COVID-19 Detection Using Acoustic Breathing/Cough/Speech Signals},
author = {Xingyu Chen and Qiushi Zhu and Jie Zhang and Li-Rong Dai},
year = {2022},
note = {ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 561-565},
doi = {10.1109/ICASSP43922.2022.9746205},
eprint = {2201.08934},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/ICASSP43922.2022.9746205},
} VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over
VisualTTS effectively improves lip-speech synchronization in scripted voice over by conditioning TTS on lip video, but does not tackle silent speech decoding or unscripted scenarios.
BibTeX
@misc{visualtts-tts-with-accurate-lip-speech-synchronization-for-automatic-voice-over,
title = {VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over},
author = {Junchen Lu and Berrak Şişman and Rui Liu and Mingyang Zhang and Haizhou Li},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2110.03342},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2110.03342v3},
} Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language
SSRNet innovatively applies duration-aware Seq2Seq modeling and tonal multitask learning to reconstruct intelligible Mandarin speech from facial sEMG signals, markedly improving performance over prior methods but remains speaker-dependent with limited deployment evaluation.
BibTeX
@misc{sequence-to-sequence-voice-reconstruction-for-silent-speech-in-a-tonal-language,
title = {Sequence-to-Sequence Voice Reconstruction for Silent Speech in a Tonal Language},
author = {Huiyan Li and Haohong Lin and You Wang and Hengyang Wang and Ming Zhang and Han Gao and Qing Ai and Zhiyuan Luo and Guang Li},
year = {2022},
note = {arXiv / imported corpus page},
eprint = {2108.00190},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2108.00190v3},
} SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography
SilentSpeller is a strong, rigorously tested SSI system that reframes silent speech as silent spelling, enabling large vocabulary, live text entry, and walking robustness with in-mouth electropalatography sensors.
BibTeX
@misc{silentspeller,
title = {SilentSpeller: Towards mobile, hands-free, silent speech text entry using electropalatography},
author = {Naoki Kimura and Tan Gemicioglu and Jonathan Womack and Richard Li and Yuhui Zhao and Abdelkareem Bedri and Zixiong Su and Alex Olwal and Jun Rekimoto and Thad Starner},
year = {2022},
note = {CHI '22},
doi = {10.1145/3491102.3502015},
url = {https://doi.org/10.1145/3491102.3502015},
} SA-SDR: A novel loss function for separation of meeting style data
Elegant loss fix, not SSI.
BibTeX
@misc{sa-sdr-a-novel-loss-function-for-separation-of-meeting-style-data,
title = {SA-SDR: A novel loss function for separation of meeting style data},
author = {Thilo von Neumann and Keisuke Kinoshita and Christoph Boeddeker and Marc Delcroix and Reinhold Haeb‐Umbach},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2110.15581},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2110.15581v2},
} Advances and Challenges in Deep Lip Reading
Good survey, not a model result.
BibTeX
@misc{advances-and-challenges-in-deep-lip-reading,
title = {Advances and Challenges in Deep Lip Reading},
author = {Marzieh Oghbaie and Arian Sabaghi and Kooshan Hashemifard and Mohammad Kazem Akbari},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2110.07879},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2110.07879v1},
} Sub-word Level Lip Reading With Visual Attention
Major lip-reading gain, adjacent to SSI.
BibTeX
@misc{sub-word-level-lip-reading-with-visual-attention,
title = {Sub-word Level Lip Reading With Visual Attention},
author = {K R Prajwal and Triantafyllos Afouras and Andrew Zisserman},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2110.07603},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2110.07603v2},
} Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input
Helpful side information, not standalone SSI.
BibTeX
@misc{speech-synthesis-from-text-and-ultrasound-tongue-image-based-articulatory-input,
title = {Speech Synthesis from Text and Ultrasound Tongue Image-based Articulatory Input},
author = {Csapó Tamás Gábor and László Tóth and Gosztolya Gábor and Alexandra Markó},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2107.02003},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2107.02003v1},
} Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits
Useful separation engineering, not silent speech.
BibTeX
@misc{sparsely-overlapped-speech-training-in-the-time-domain-joint-learning-of-target-speech-separation-and-personal-vad-benefits,
title = {Sparsely Overlapped Speech Training in the Time Domain: Joint Learning of Target Speech Separation and Personal VAD Benefits},
author = {Qingjian Lin and Lin Yang and Xuyang Wang and Luyuan Xie and Jia Chen and Junjie Wang},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2106.14371},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2106.14371v2},
} Silent Speech and Emotion Recognition from Vocal Tract Shape Dynamics in Real-Time MRI
Strong rtMRI recognition result, weak deployment story.
BibTeX
@misc{silent-speech-and-emotion-recognition-from-vocal-tract-shape-dynamics-in-real-time-mri,
title = {Silent Speech and Emotion Recognition from Vocal Tract Shape Dynamics in Real-Time MRI},
author = {Laxmi Pandey and Ahmed Sabbir Arif},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2106.08706},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2106.08706v1},
} Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces
The ultrasound-based x-vector speaker embedding is highly effective for speaker recognition, achieving under 1% error on unseen speakers, but its integration yields only a marginal improvement in multi-speaker ultrasound-to-speech synthesis accuracy.
BibTeX
@misc{neural-speaker-embeddings-for-ultrasound-based-silent-speech-interfaces,
title = {Neural Speaker Embeddings for Ultrasound-based Silent Speech Interfaces},
author = {Honarmandi Shandiz Amin and László Tóth and Gosztolya Gábor and Alexandra Markó and Csapó Tamás Gábor},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2106.04552},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2106.04552v2},
} An Improved Model for Voicing Silent Speech
This paper substantially improves open-vocabulary silent speech voicing using learned convolutional EMG features, Transformer modeling, and phoneme supervision, reducing WER from 68.0% to 42.2% automatic and 32.3% human in a single-speaker lab setting.
BibTeX
@misc{an-improved-model-for-voicing-silent-speech,
title = {An Improved Model for Voicing Silent Speech},
author = {David Gaddy and Dan Klein},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2106.01933},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2106.01933v2},
} Voice Activity Detection for Ultrasound-based Silent Speech Interfaces using Convolutional Neural Networks
Preprocessing paper, narrow but legitimate.
BibTeX
@misc{voice-activity-detection-for-ultrasound-based-silent-speech-interfaces-using-convolutional-neural-networks,
title = {Voice Activity Detection for Ultrasound-based Silent Speech Interfaces using Convolutional Neural Networks},
author = {Amin Honarmandi Shandiz and László Tóth},
year = {2021},
note = {arXiv / imported corpus page},
doi = {10.1007/978-3-030-83527-9_43},
eprint = {2105.13718},
archivePrefix = {arXiv},
url = {https://doi.org/10.1007/978-3-030-83527-9_43},
} Speaker disentanglement in video-to-speech conversion
The paper effectively makes speaker identity a controllable factor in multi-speaker video-to-speech synthesis by disentangling it from content, showing the trade-off between intelligibility and voice control on GRID corpus data.
BibTeX
@misc{speaker-disentanglement-in-video-to-speech-conversion,
title = {Speaker disentanglement in video-to-speech conversion},
author = {Dan Oneaţă and Adriana Stan and Horia Cucu},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2105.09652},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2105.09652v1},
} Improving Neural Silent Speech Interface Models by Adversarial Training
A clean, well-executed incremental advance using GAN loss to modestly improve articulatory-to-acoustic mapping from ultrasound, validated objectively on two single-speaker corpora.
BibTeX
@misc{improving-neural-silent-speech-interface-models-by-adversarial-training,
title = {Improving Neural Silent Speech Interface Models by Adversarial Training},
author = {Amin Honarmandi Shandiz and László Tóth and Gábor Gosztolya and Alexandra Markó and Tamás Gábor Csapó},
year = {2021},
note = {arXiv / imported corpus page},
doi = {10.1007/978-3-030-76346-6_39},
eprint = {2104.11601},
archivePrefix = {arXiv},
url = {https://doi.org/10.1007/978-3-030-76346-6_39},
} 3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces
Temporal context helps, but the evidence is a single-speaker vocoder-parameter study.
BibTeX
@misc{3d-convolutional-neural-networks-for-ultrasound-based-silent-speech-interfaces,
title = {3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces},
author = {László Tóth and Amin Honarmandi Shandiz},
year = {2021},
note = {arXiv / imported corpus page},
doi = {10.1007/978-3-030-61401-0_16},
eprint = {2104.11532},
archivePrefix = {arXiv},
url = {https://doi.org/10.1007/978-3-030-61401-0_16},
} HTMD-Net: A Hybrid Masking-Denoising Approach to Time-Domain Monaural Singing Voice Separation
Solid time-domain music vocal separation paper with a novel hybrid masking-denoising design showing improved silent-segment suppression; not relevant to SSI applications.
BibTeX
@misc{htmd-net-a-hybrid-masking-denoising-approach-to-time-domain-monaural-singing-voice-separation,
title = {HTMD-Net: A Hybrid Masking-Denoising Approach to Time-Domain Monaural Singing Voice Separation},
author = {Christos Garoufis and Athanasia Zlatintsi and Petros Maragos},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2103.04336},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2103.04336v1},
} Silent versus modal multi-speaker speech recognition from ultrasound and video
Large-corpus baseline with real silent-mode gap.
BibTeX
@misc{silent-versus-modal-multi-speaker-speech-recognition-from-ultrasound-and-video,
title = {Silent versus modal multi-speaker speech recognition from ultrasound and video},
author = {Manuel Sam Ribeiro and Aciel Eshky and Korin Richmond and Steve Renals},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2103.00333},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2103.00333v1},
} EMA2S: An End-to-End Multimodal Articulatory-to-Speech System
EMA2S achieves consistent quality improvements over prior EMA-to-speech baselines by combining multimodal joint loss training with a neural vocoder, though gains remain confined to lab EMA conditions.
BibTeX
@misc{ema2s-an-end-to-end-multimodal-articulatory-to-speech-system,
title = {EMA2S: An End-to-End Multimodal Articulatory-to-Speech System},
author = {Yu‐Wen Chen and Kuo-Hsuan Hung and Shang-Yi Chuang and Jonathan H. Sherman and Wen-Chin Huang and Xugang Lu and Yu Tsao},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2102.03786},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2102.03786v2},
} Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image
Real signal, wrong target for SSI.
BibTeX
@misc{convolutional-neural-network-based-age-estimation-using-b-mode-ultrasound-tongue-image,
title = {Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image},
author = {Kele Xu and Tamás Gábor Csapó and Ming Feng},
year = {2021},
note = {arXiv / imported corpus page},
eprint = {2101.11245},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2101.11245v1},
} End-to-end Silent Speech Recognition with Acoustic Sensing
Strong mobile-friendly acoustic SSI paper.
BibTeX
@misc{end-to-end-silent-speech-recognition-with-acoustic-sensing,
title = {End-to-end Silent Speech Recognition with Acoustic Sensing},
author = {Jian Luo and Jianzong Wang and Ning Cheng and Guilin Jiang and Jing Xiao},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2011.11315},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2011.11315v1},
} Speech Prediction in Silent Videos using Variational Autoencoders
Strong video-to-speech paper that models ambiguity explicitly.
BibTeX
@misc{speech-prediction-in-silent-videos-using-variational-autoencoders,
title = {Speech Prediction in Silent Videos using Variational Autoencoders},
author = {Ravindra Yadav and Ashish Sardana and Vinay P. Namboodiri and Rajesh M. Hegde},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2011.07340},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2011.07340v1},
} X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network
Strong time-domain target-speaker extraction using speaker verification and innovative training; improves robustness to absent target but remains speech extraction, not silent speech.
BibTeX
@misc{x-tasnet-robust-and-accurate-time-domain-speaker-extraction-network,
title = {X-TaSNet: Robust and Accurate Time-Domain Speaker Extraction Network},
author = {Zining Zhang and Bingsheng He and Zhenjie Zhang},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2010.12766},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2010.12766v1},
} Listening to Sounds of Silence for Speech Denoising
Strong denoising work, not SSI.
BibTeX
@misc{listening-to-sounds-of-silence-for-speech-denoising,
title = {Listening to Sounds of Silence for Speech Denoising},
author = {Ruilin Xu and Rundi Wu and Yuko Ishiwaka and Carl Vondrick and Changxi Zheng},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2010.12013},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2010.12013v1},
} Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Technically solid self-supervised class-aware audiovisual sounding object localization, but outside the core SSI domain.
BibTeX
@misc{discriminative-sounding-objects-localization-via-self-supervised-audiovisual-matching,
title = {Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching},
author = {Di Hu and Rui Qian and Minyue Jiang and Xiao Tan and Shilei Wen and Errui Ding and Weiyao Lin and Dejing Dou},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2010.05466},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2010.05466v1},
} Digital Voicing of Silent Speech
Core EMG SSI paper with real gains from target transfer.
BibTeX
@misc{digital-voicing-of-silent-speech,
title = {Digital Voicing of Silent Speech},
author = {David Gaddy and Dan Klein},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2010.02960},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2010.02960v1},
} End-to-End Speaker-Dependent Voice Activity Detection
Strong target-speaker VAD paper, not SSI.
BibTeX
@misc{end-to-end-speaker-dependent-voice-activity-detection,
title = {End-to-End Speaker-Dependent Voice Activity Detection},
author = {Yefei Chen and Shuai Wang and Yanmin Qian and Kai Yu},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2009.09906},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2009.09906v1},
} A comparison of oscillatory characteristics in covert speech and speech perception
Strong covert-speech EEG analysis, not an SSI system.
BibTeX
@misc{a-comparison-of-oscillatory-characteristics-in-covert-speech-and-speech-perception,
title = {A comparison of oscillatory characteristics in covert speech and speech perception},
author = {Jae Moon and Silvia Orlandi and Tom Chau},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2009.02816},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2009.02816v1},
} Silent Speech Interfaces for Speech Restoration: A Review
Core SSI survey with concrete deployment constraints.
BibTeX
@misc{silent-speech-interfaces-for-speech-restoration-a-review,
title = {Silent Speech Interfaces for Speech Restoration: A Review},
author = {José A. González and Alejandro Gomez-Alanis and Juan M. Martín-Doñas and José L. Pérez-Córdoba and Ángel M. Gómez},
year = {2020},
note = {arXiv / imported corpus page},
doi = {10.1109/ACCESS.2020.3026579},
eprint = {2009.02110},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/ACCESS.2020.3026579},
} An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation
Strong AV speech survey, not an SSI system paper.
BibTeX
@misc{an-overview-of-deep-learning-based-audio-visual-speech-enhancement-and-separation,
title = {An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation},
author = {Daniel Michelsanti and Zheng‐Hua Tan and Shi-Xiong Zhang and Yong Xu and Meng Yu and Dong Yu and Jesper Jensen},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2008.09586},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2008.09586v2},
} CITISEN: A Deep Learning-Based Speech Signal-Processing Mobile Application
Strong mobile speech-processing app paper, not SSI.
BibTeX
@misc{citisen-a-deep-learning-based-speech-signal-processing-mobile-application,
title = {CITISEN: A Deep Learning-Based Speech Signal-Processing Mobile Application},
author = {Yu-Wen Chen and Kuo-Hsuan Hung and You-Jin Li and Alexander Kang and Ya‐Hsin Lai and Kai-Chun Liu and Szu‐Wei Fu and Syu‐Siang Wang and Yu Tsao},
year = {2020},
note = {arXiv / imported corpus page},
doi = {10.1109/ACCESS.2022.3153469},
eprint = {2008.09264},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/ACCESS.2022.3153469},
} Foley Music: Learning to Generate Music from Videos
Strong video-to-music paper, not SSI.
BibTeX
@misc{foley-music-learning-to-generate-music-from-videos,
title = {Foley Music: Learning to Generate Music from Videos},
author = {Chuang Gan and Deng Huang and Peihao Chen and Joshua B. Tenenbaum and Antonio Torralba},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2007.10984},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2007.10984v1},
} Learning Frame Level Attention for Environmental Sound Classification
Strong ESC paper, but outside SSI.
BibTeX
@misc{learning-frame-level-attention-for-environmental-sound-classification,
title = {Learning Frame Level Attention for Environmental Sound Classification},
author = {Zhichao Zhang and Shugong Xu and Shunqing Zhang and Tianhao Qiao and Shan Cao},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2007.07241},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2007.07241v1},
} Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images
Strong ultrasound SSI paper with unusually clear quantitative gains.
BibTeX
@misc{ultra2speech-a-deep-learning-framework-for-formant-frequency-estimation-and-tracking-from-ultrasound-tongue-images,
title = {Ultra2Speech -- A Deep Learning Framework for Formant Frequency Estimation and Tracking from Ultrasound Tongue Images},
author = {Pramit Saha and Yadong Liu and Bryan Gick and Sidney Fels},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2006.16367},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2006.16367v1},
} Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task
Strong crowdsourcing methodology paper, not SSI.
BibTeX
@misc{application-of-just-noticeable-difference-in-quality-as-environment-suitability-test-for-crowdsourcing-speech-quality-assessment-task,
title = {Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task},
author = {Babak Naderi and Sebastian Möller},
year = {2020},
note = {arXiv / imported corpus page},
doi = {10.1109/QoMEX48832.2020.9123093},
eprint = {2004.05502},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/QoMEX48832.2020.9123093},
} Vocoder-Based Speech Synthesis from Silent Videos
A notable step forward in lip-to-speech synthesis by predicting full vocoder features and jointly training for recognition, achieving strong speaker-dependent results but lacking unseen speaker generalization.
BibTeX
@misc{vocoder-based-speech-synthesis-from-silent-videos,
title = {Vocoder-Based Speech Synthesis from Silent Videos},
author = {Daniel Michelsanti and Olga Slizovskaia and Gloria Haro and Emília Gómez and Zheng‐Hua Tan and Jesper Jensen},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2004.02541},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2004.02541v2},
} Continuous Silent Speech Recognition using EEG
Real EEG sentence-level silent speech recognition is demonstrated but at very high WER, confirming feasibility only and underscoring the immature state of current EEG silent speech technology.
BibTeX
@misc{continuous-silent-speech-recognition-using-eeg,
title = {Continuous Silent Speech Recognition using EEG},
author = {Gautam Krishna and Co Tran and Mason Carnahan and Ahmed H. Tewfik},
year = {2020},
note = {arXiv / imported corpus page},
eprint = {2002.03851},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/2002.03851v7},
} Brain2Char: A Deep Architecture for Decoding Text from Brain Recordings
Brain2Char establishes a new state-of-the-art for continuous character decoding from invasive ECoG with competitive WER on large vocabularies and silent speech, demonstrating feasibility for communication BCIs.
BibTeX
@misc{brain2char-a-deep-architecture-for-decoding-text-from-brain-recordings,
title = {Brain2Char: A Deep Architecture for Decoding Text from Brain Recordings},
author = {Pengfei Sun and Gopala K. Anumanchipalli and Edward F. Chang},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1909.01401},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1909.01401v1},
} Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed
This work delivers an improved waveform source separation model combined with a novel remix-based semi-supervised learning scheme using unlabeled music. Though not related to silent speech, it advances music separation benchmarks by closing gaps to spectrogram methods.
BibTeX
@misc{demucs-deep-extractor-for-music-sources-with-extra-unlabeled-data-remixed,
title = {Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed},
author = {Alexandre Défossez and Nicolas Usunier and Léon Bottou and Francis R. Bach},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1909.01174},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1909.01174v1},
} Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification
The proposed frame-level attention integrated within a convolutional recurrent network effectively improves environmental sound classification accuracy on ESC benchmarks by focusing on informative temporal frames while suppressing irrelevant or silent ones.
BibTeX
@misc{attention-based-convolutional-recurrent-neural-network-for-environmental-sound-classification,
title = {Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification},
author = {Zhichao Zhang and Shugong Xu and Shunqing Zhang and Tianhao Qiao and Shan Cao},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1907.02230},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1907.02230v1},
} Lipper: Synthesizing Thy Speech using Multi-View Lipreading
Strong multi-view lip-to-speech baseline with honest quality limits.
BibTeX
@misc{lipper-synthesizing-thy-speech-using-multi-view-lipreading,
title = {Lipper: Synthesizing Thy Speech using Multi-View Lipreading},
author = {Yaman Kumar and Rohit Jain and Khwaja Mohd. Salik and Rajiv Ratn Shah and Yifang Yin and Roger Zimmermann},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1907.01367},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1907.01367v1},
} Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder
The key advancement is continuous F0 tracking via CNNs yielding lower pitch error and slight naturalness improvement over discontinuous F0 pipelines in ultrasound SSI.
BibTeX
@misc{ultrasound-based-silent-speech-interface-built-on-a-continuous-vocoder,
title = {Ultrasound-based Silent Speech Interface Built on a Continuous Vocoder},
author = {Tamás Gábor Csapó and Mohammed Salah Al-Radhi and Géza Németh and Gábor Gosztolya and Tamás Grósz and László Tóth and Alexandra Markó},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1906.09885},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1906.09885v1},
} Video-Driven Speech Reconstruction using Generative Adversarial Networks
Foundational direct video-to-audio result with clear generalization limits.
BibTeX
@misc{video-driven-speech-reconstruction-using-generative-adversarial-networks,
title = {Video-Driven Speech Reconstruction using Generative Adversarial Networks},
author = {Konstantinos Vougioukas and Pingchuan Ma and Stavros Petridis and Maja Pantić},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1906.06301},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1906.06301v1},
} A Novel Task-Oriented Text Corpus in Silent Speech Recognition and its Natural Language Generation Construction Method
Useful EEG-SSR corpus framing paper, but evidence is lighter than a full benchmark paper.
BibTeX
@misc{a-novel-task-oriented-text-corpus-in-silent-speech-recognition-and-its-natural-language-generation-construction-method,
title = {A Novel Task-Oriented Text Corpus in Silent Speech Recognition and its Natural Language Generation Construction Method},
author = {Dong Cao and Dongdong Zhang and Haibo Chen},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1905.01974},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1905.01974v1},
} Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces
The paper advances ultrasound silent speech interfaces by compressing ultrasound images using an autoencoder bottleneck prior to spectral parameter prediction, resulting in improved accuracy and more natural synthesized speech with smaller models.
BibTeX
@misc{autoencoder-based-articulatory-to-acoustic-mapping-for-ultrasound-silent-speech-interfaces,
title = {Autoencoder-Based Articulatory-to-Acoustic Mapping for Ultrasound Silent Speech Interfaces},
author = {Gábor Gosztolya and Ádám Pintér and László Tóth and Tamás Grósz and Alexandra Markó and Tamás Gábor Csapó},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1904.05259},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1904.05259v1},
} Denoising convolutional autoencoder based B-mode ultrasound tongue image feature extraction
DCAE provides cleaner, more robust ultrasound tongue features leading to improved silent speech recognition, outperforming prior feature extraction strategies.
BibTeX
@misc{denoising-convolutional-autoencoder-based-b-mode-ultrasound-tongue-image-feature-extraction,
title = {Denoising convolutional autoencoder based B-mode ultrasound tongue image feature extraction},
author = {Bo Li and Kele Xu and Dawei Feng and Haibo Mi and Huaimin Wang and Jian Zhu},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1903.00888},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1903.00888v1},
} All-neural online source separation, counting, and diarization for meeting analysis
Strong online diarization/separation paper, but outside SSI.
BibTeX
@misc{all-neural-online-source-separation-counting-and-diarization-for-meeting-analysis,
title = {All-neural online source separation, counting, and diarization for meeting analysis},
author = {Thilo von Neumann and Keisuke Kinoshita and Marc Delcroix and Shoko Araki and Tomohiro Nakatani and Reinhold Haeb‐Umbach},
year = {2019},
note = {arXiv / imported corpus page},
eprint = {1902.07881},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1902.07881v1},
} SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks
A solid proof of concept that reconstructs speech audio from ultrasound for controlling unmodified smart speakers, showcasing important system design insight despite prototype limitations in latency, hardware bulk, and speaker dependency.
BibTeX
@misc{sottovoce,
title = {SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks},
author = {Naoki Kimura and Michinari Kono and Jun Rekimoto},
year = {2019},
note = {CHI '19},
doi = {10.1145/3290605.3300376},
url = {https://doi.org/10.1145/3290605.3300376},
} Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold
Strong telephony anti-SPAM paper, not SSI.
BibTeX
@misc{audio-spectrogram-factorization-for-classification-of-telephony-signals-below-the-auditory-threshold,
title = {Audio Spectrogram Factorization for Classification of Telephony Signals below the Auditory Threshold},
author = {Iroro Orife and Shane Walker and Jason Flaks},
year = {2018},
note = {arXiv / imported corpus page},
eprint = {1811.04139},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1811.04139v1},
} Proactive Security: Embedded AI Solution for Violent and Abusive Speech Recognition
An embedded smartphone NLP classifier detects violent speech with ~87.5% accuracy using known methods but is unrelated to silent speech interfaces; strong practical application in safety alerting.
BibTeX
@misc{proactive-security-embedded-ai-solution-for-violent-and-abusive-speech-recognition,
title = {Proactive Security: Embedded AI Solution for Violent and Abusive Speech Recognition},
author = {Christopher Shulby and Leonardo Pombal and Vitor Jordão and Guilherme Ziolle and Bruno Martho and Antônio Postal and Thiago Prochnow},
year = {2018},
note = {arXiv / imported corpus page},
eprint = {1810.09431},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1810.09431v1},
} Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed
Multi-view silent video combined with CNN-LSTM models significantly improves speech audio reconstruction quality over single-view, highlighting the importance of optimal camera placement to address pose variance.
BibTeX
@misc{harnessing-ai-for-speech-reconstruction-using-multi-view-silent-video-feed,
title = {Harnessing AI for Speech Reconstruction using Multi-view Silent Video Feed},
author = {Yaman Kumar and Mayank Aggarwal and Pratham Nawal and Shin'ichi Satoh and Rajiv Ratn Shah and Roger Zimmermann},
year = {2018},
note = {arXiv / imported corpus page},
doi = {10.1145/3240508.3241911},
eprint = {1807.00619},
archivePrefix = {arXiv},
url = {https://doi.org/10.1145/3240508.3241911},
} Visual-Only Recognition of Normal, Whispered and Silent Speech
Strong evidence that silent lipreading needs dedicated training.
BibTeX
@misc{visual-only-recognition-of-normal-whispered-and-silent-speech,
title = {Visual-Only Recognition of Normal, Whispered and Silent Speech},
author = {Stavros Petridis and Jie Shen and Doruk Cetin and Maja Pantić},
year = {2018},
note = {arXiv / imported corpus page},
eprint = {1802.06399},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1802.06399v1},
} Cross-modal Embeddings for Video and Audio Retrieval
Useful multimodal retrieval baseline, not SSI.
BibTeX
@misc{cross-modal-embeddings-for-video-and-audio-retrieval,
title = {Cross-modal Embeddings for Video and Audio Retrieval},
author = {Dídac Surís and Amanda Duarte and Amaia Salvador and Jordi Torres and Giró Nieto, Xavier},
year = {2018},
note = {arXiv / imported corpus page},
eprint = {1801.02200},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1801.02200v1},
} Lip2AudSpec: Speech reconstruction from silent lip movements video
The paper's auditory spectrogram autoencoder bottleneck target is a key innovation that produces more intelligible, natural reconstructed speech from lip videos than prior methods, as confirmed by objective and human evaluations.
BibTeX
@misc{lip2audspec-speech-reconstruction-from-silent-lip-movements-video,
title = {Lip2AudSpec: Speech reconstruction from silent lip movements video},
author = {Hassan Akbari and Himani Arora and Liangliang Cao and Nima Mesgarani},
year = {2017},
note = {arXiv / imported corpus page},
eprint = {1710.09798},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1710.09798v1},
} Updating the silent speech challenge benchmark with deep learning
Benchmark update with a real, reproducible WER gain.
BibTeX
@misc{updating-the-silent-speech-challenge-benchmark-with-deep-learning,
title = {Updating the silent speech challenge benchmark with deep learning},
author = {Yan Ji and Licheng Liu and Hongcui Wang and Zhilei Liu and Zhibin Niu and B. Denby},
year = {2017},
note = {arXiv / imported corpus page},
eprint = {1709.06818},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1709.06818v1},
} Seeing Through Noise: Visually Driven Speaker Separation and Enhancement
Strong audiovisual speech separation and enhancement leveraging face video for speaker-dependent masking; not a silent speech interface paper.
BibTeX
@misc{seeing-through-noise-visually-driven-speaker-separation-and-enhancement,
title = {Seeing Through Noise: Visually Driven Speaker Separation and Enhancement},
author = {Aviv Gabbay and Ariel Ephrat and Tavi Halperin and Shmuel Peleg},
year = {2017},
note = {arXiv / imported corpus page},
eprint = {1708.06767},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1708.06767v3},
} Improved Speech Reconstruction from Silent Video
Strong, benchmark-setting speaker-dependent video-to-speech system that advances speech reconstruction from silent face video but remains limited to per-speaker training and constrained conditions.
BibTeX
@misc{improved-speech-reconstruction-from-silent-video,
title = {Improved Speech Reconstruction from Silent Video},
author = {Ariel Ephrat and Tavi Halperin and Shmuel Peleg},
year = {2017},
note = {arXiv / imported corpus page},
eprint = {1708.01204},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1708.01204v3},
} Vid2speech: Speech Reconstruction from Silent Video
Real lip-to-speech progress, still tightly benchmark-bounded.
BibTeX
@misc{vid2speech-speech-reconstruction-from-silent-video,
title = {Vid2speech: Speech Reconstruction from Silent Video},
author = {Ariel Ephrat and Shmuel Peleg},
year = {2017},
note = {arXiv / imported corpus page},
eprint = {1701.00495},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1701.00495v2},
} Contour-based 3d tongue motion visualization using ultrasound image sequences
Useful tongue-modeling tool, not a recognizer.
BibTeX
@misc{contour-based-3d-tongue-motion-visualization-using-ultrasound-image-sequences,
title = {Contour-based 3d tongue motion visualization using ultrasound image sequences},
author = {Kele Xu and Yin Yang and Clémence Leboullenger and Pierre Roussel and B. Denby},
year = {2016},
note = {arXiv / imported corpus page},
eprint = {1605.05967},
archivePrefix = {arXiv},
url = {http://arxiv.org/abs/1605.05967v1},
} Optimal Power Control for Analog Bidirectional Relaying with Long-Term Relay Power Constraint
A rigorous relay power control theory paper optimizing outage under long-term average power constraints for bidirectional AF relaying; solid mathematical contribution but outside SSI relevance.
BibTeX
@misc{optimal-power-control-for-analog-bidirectional-relaying-with-long-term-relay-power-constraint,
title = {Optimal Power Control for Analog Bidirectional Relaying with Long-Term Relay Power Constraint},
author = {Zoran Hadzi-Velkov and Nikola Zlatanov and Robert Schober},
year = {2014},
note = {arXiv / imported corpus page},
doi = {10.1109/GLOCOM.2013.6831710},
eprint = {1404.0906},
archivePrefix = {arXiv},
url = {https://doi.org/10.1109/GLOCOM.2013.6831710},
} Approach comparison
Compare SilentSpeller, SottoVoce, and NasoVoce side by side without treating them as a required reading order.
Major silent speech approaches compared
SilentSpeller, SottoVoce, and NasoVoce compared on sensing, evaluation, practicality, and open questions.
Research agenda
The reviewed papers keep recurring on wearability, vocabulary, latency, and generalization. The page below turns that into a short, grounded agenda.
Open problems and research agenda
Wearability, open vocabulary, real-time use, and generalization keep reappearing in the current review set.
Technique taxonomy
These pages group the current database by real `modality:` tags from the expert records.
Video
42 reviewed pages · 0 imported pages
Acoustic
30 reviewed pages · 0 imported pages
Ultrasound
16 reviewed pages · 0 imported pages
Multimodal
15 reviewed pages · 0 imported pages
Microphone
7 reviewed pages · 0 imported pages
EEG
6 reviewed pages · 0 imported pages
EMG
6 reviewed pages · 0 imported pages
Magnetic
4 reviewed pages · 0 imported pages
Radar
2 reviewed pages · 0 imported pages
Vibration
2 reviewed pages · 0 imported pages
Camera
1 reviewed pages · 0 imported pages
Electropalatography
1 reviewed pages · 0 imported pages
Machine-readable exports
These files are generated from repository inputs during build.
SSI review export
Snapshot JSON of the current SSI review records built from repository inputs.
SSI review feed
Snapshot feed of the current SSI review records with source-updated timestamps.
Reference and citation
Use the canonical citation page when you need the database name, maintainer, or last-updated date.
How to cite this database
Canonical citation page for the SSI review database. Last updated 2026-07-07.
Datasets and code resources
Verified links are grouped on a dedicated page so the current corpus can point to code, datasets, and paper pages without inventing any new metadata.
Datasets and code resources
Verified links already present in repository data, with paper pages attached wherever the archive has a local review page.
FAQ
Answers grounded in the site's data policy and review methodology.
Silent speechとは何ですか?
Silent speech interface (SSI) は、発声せずに口・喉・顔などの動きをセンサーで読み取り、意図した発話をテキストや音声に変換する技術です。マイクを使わないため、周囲に声を出せない場面や発話が難しい場面でも使えます。
このデータベースの論文はどう選ばれていますか?
リポジトリのarXiv/引用データを基に収集し、対応する論文をNaoki Kimuraが実際に読んで評価しています。評価が済んだものは「reviewed」、まだ評価本文が無いものは「imported corpus」として区別しています。
「reviewed」と「imported corpus」ページの違いは何ですか?
reviewed pageはNaoki Kimuraによる評価本文(強み・弱み・エビデンス付き)を含みます。imported corpus pageは書誌情報とベンチマーク値のみを保持し、評価本文はまだ付いていません。
citation countはどこから来ていますか?
OpenAlexから取得しています。取得できていない論文は「unknown citations」と表示しています。
レビューは誰が書いていますか?
Naoki Kimura本人による専門家評価です。レビューの見方は /papers/rubric にまとめています。
データポリシーは何ですか?
実データのみを使用し、捏造した指標・主張・受賞歴・リンクは掲載しません。詳細は /papers/cite に記載しています。
JSON形式でデータを取得できますか?
できます。/exports/ssi-review.json(スナップショット)と /feeds/ssi-review.json(フィード)で機械可読形式を提供しています。
このデータベースを引用する際の書式は?
/papers/cite に、データベース名・メンテナー・推奨citation文字列を掲載しています。