Silent speech datasets and code resources
Start with the signal you need: facial muscle activity, or images of the tongue and lips. The two official data sources below include silent and voiced material, but they are not interchangeable benchmarks. Follow the source documentation before choosing an experiment.
Official documentation checked on 5 October 2026. This is a selection guide, not a claim that we downloaded or reproduced the datasets. Additional links from the silent-speech review corpus remain below; a paper, demo, or project link does not guarantee a downloadable dataset or model.
- What it is
- Two starting points for data-driven silent-speech experiments, followed by related paper and project links.
- Who it’s for
- Researchers and students looking for SSI datasets or code after a review, and readers who landed here from search and need to know this is an artifact list—not a new paper.
- Verdict
- Choose by input, speaking mode, and evaluation split—not by the largest headline score. Check each source’s access and reuse terms separately from the availability of a paper or demo.
Which silent speech dataset should I start with?
Facial muscle signals: Silent Speech EMG
Choose this starting point for electromyography (EMG): electrical activity recorded from facial muscles
during silent and vocalized speech. The release includes raw eight-channel signals, audio, and prompt metadata.
Reference samples marked with sentence_index: -1 are not ordinary utterance examples.
Dataset files and description on Zenodo · Author’s code and setup instructions
For the research task and its limits, read our reviews of Digital Voicing of Silent Speech and An Improved Model for Voicing Silent Speech. The official code has separate paths for speech-audio synthesis and direct text recognition; select the intended output first.
Tongue ultrasound and lip video: TaL
Choose the Tongue and Lips corpus for synchronized ultrasound, lip video, and audio.
TaL1 records one speaker across six sessions; TaL80 records 81 speakers, each in one session.
Its prompt tags distinguish audible speech (aud), silent speech (sil),
and whispered speech (whi, TaL1 only). Shared prompts carry an x prefix.
Do not assume every recording is silent.
Official TaL documentation, samples, and download instructions · Review: voiced versus silent recognition
Before downloading or reporting a result
- Identify the speaking mode. Removing microphone audio from a model’s input does not make a voiced recording an example of silent articulation.
- Choose a split that answers your question. Keep speaker, session, and prompt overlap explicit. TaL documentation warns that prompts overlap between TaL1 and TaL80; pooling them without checking can change what a held-out test means.
- Match the implementation to the paper. The EMG repository’s current model differs from earlier paper versions, and its default validation set is larger than the original EMNLP 2020 set. Record the commit and split rather than comparing unmatched scores.
- Inspect samples and terms first. TaL offers sample directories before its much larger core downloads. Check the current access, license, citation, and storage instructions at each official source. Code, trained weights, and data can have different conditions.
Other paper and project links
This existing list includes papers and demonstrations, not only downloadable datasets. Links are retained for discovery; their presence does not verify a reusable implementation.