← SSI archive · Review rubric

2026 · arXiv · Field expert review · confidence high

Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features

Chenqian Le, Ruisi Li, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Amirhossein Khalilian-Gourtani, Tianyu He, Adeen Flinker, Yao Wang

BibTeX
@misc{comparison-of-semg-encoding-accuracy-across-speech-modes-using-articulatory-and-phoneme-features,
  title = {Comparison of sEMG Encoding Accuracy Across Speech Modes Using Articulatory and Phoneme Features},
  author = {Chenqian Le and Ruisi Li and Beatrice Fumagalli and Yasamin Esmaeili and Xupeng Chen and Amirhossein Khalilian-Gourtani and Tianyu He and Adeen Flinker and Yao Wang},
  year = {2026},
  note = {arXiv},
  eprint = {2604.18920},
  archivePrefix = {arXiv},
  url = {http://arxiv.org/abs/2604.18920v2},
}

Articulatory features better predict aligned muscle envelopes across speech modes; this supports representation choice, while actual EMG-to-speech decoding gains remain untested.

Verdict: full-text draftPriority: highConfidence: highBasis: full text + summaryCoverage: high

Reading guidance

Verdict
full-text draft · priority high · confidence high
Why it matters
A comparatively broad within-subject physiological analysis that motivates articulatory intermediate targets without pretending to have demonstrated an end-to-end silent speech decoder.
What to trust
Basis: full text + summary. Coverage: high. 10 evidence records back the review.
What is weak
SPARC kinematics are inferred from voiced audio rather than measured during silent production. Aloud uses 14 features including pitch/loudness, silent modes use 12. DTW may inflate absolute correlations. Squared Pearson correlation is not automatically calibrated predictive R-squared; correlated predictors and normalized absolute weights limit causal anatomical interpretation. All models are evaluated within subjects using sentence-level folds, not held-out users or cross-mode transfer. Silent evaluation depends on paired voiced reference and response-derived DTW alignment. Fold count and exact repeat grouping are not detailed. Primary numeric results are plotted rather than tabulated. The Gaddy replication has one participant and incompletely specified uncertainty aggregation. Offline forward encoding uses features derived from paired aloud audio and silently produced EMG warped to the paired aloud envelope. No autonomous silent decoder, latency measurement, wearable implementation, walking test, or patient communication study is provided. Speech-typical participants and prompted sentences. Subvocal means attempted speech with an occluded vocal tract and no phonation; it should not be treated as unconstrained imagined speech. No clinical communication or end-to-end silent recognition study. Overclaim risk: Low for the explicitly bounded encoding contribution; high if correlation is recast as recognition accuracy, cross-mode transfer, or causal articulator recovery..
Read before
SSI review rubric
Read next
SSI archive

Axes

Task
Forward encoding of sEMG from speech representations
Modality
Surface electromyography with paired aloud audio used to derive predictors and timing references
Hardware
Eight surface EMG channels on lower face and neck; band-pass 10-450 Hz and 60 Hz harmonic notches, Hilbert envelopes at 2 kHz downsampled to 50 Hz. Aloud audio is processed at 16 kHz.
Body site
face; jaw; lip; throat
Output
Predicted sEMG envelopes and representation/electrode analyses; no decoded text or speech audio.
Vocabulary
Prompted sentence encoding; no evaluated recognition vocabulary
Metrics
Primary 24-participant results report Pearson r by channel/mode in Figure 2, averaged across folds using Fisher z and summarized with SEM across subjects; precise per-channel means are not tabulated. SPARC outperforms phoneme features on most channels, and subvocal prediction is above the permutation threshold. Separate single-participant Gaddy results: voiced r 0.443 ± 0.017 (phoneme) versus 0.455 ± 0.021 (SPARC); mimed 0.346 ± 0.029 versus 0.364 ± 0.032, with gains on 7/8 electrodes in each mode. No WER, decoding accuracy or synthesis-quality result is reported.
Evaluation mode
Within-subject sentence-level cross-validation of forward feature-to-EMG envelope regression, nested regularization selection, Pearson correlation, permutation nulls, paired FDR-controlled comparisons, variance partitioning and normalized weight maps.
Review confidence
high
Overclaim risk
Low for the explicitly bounded encoding contribution; high if correlation is recast as recognition accuracy, cross-mode transfer, or causal articulator recovery.

Expert take

This study offers useful evidence about which intermediate representations align with facial and neck muscle activity. In 24 speech-typical participants, a regularized linear model predicts sEMG envelopes from either audio-derived SPARC articulatory features or phoneme labels. SPARC performs better on most channels across aloud, mimed and subvocal production. This is forward encoding: the inputs are speech-derived features and the output is a muscle envelope. It is not evidence that EMG has been decoded into words or synthesized speech, a distinction the authors make explicitly. Silent trials are time-warped to paired aloud EMG, and both feature sets are extracted from the aloud audio. Applying the same warping path to both predictors makes their relative comparison more credible, but absolute correlations remain conditional on this reference-assisted alignment. They do not establish independent silent operation or equivalent natural timing across modes. The separate one-person Gaddy comparison shows modest correlation gains, from 0.443 to 0.455 for voiced speech and from 0.346 to 0.364 for mimed speech; these values must not be assigned to the main 24-person cohort. Shared predictive variance dominates, with additional unique SPARC contribution, so the findings do not imply that phonemic information is absent from EMG. Squared correlation and normalized absolute regression weights also require caution: neither establishes calibrated predictive variance explained or causal muscle specificity. The work supports testing articulatory targets in future decoders, while cross-user transfer, actual recognition improvements, and an unaligned real-time silent interface remain open.

True value

A comparatively broad within-subject physiological analysis that motivates articulatory intermediate targets without pretending to have demonstrated an end-to-end silent speech decoder.

What changed

Canon before

The paper contrasts phoneme labels with continuous articulatory targets for sEMG modeling. SPARC estimates vocal-tract kinematics from audio, while mTRF methods characterize time-lagged relationships between stimulus features and measured responses.

Delta from canon

Tests continuous, audio-inferred articulatory coordinates against phoneme one-hot features as forward predictors of muscle envelopes, including low-amplitude subvocal production and shared/unique predictive components.

Position in field

Non-invasive articulatory SSI representation research, using forward muscle-signal encoding rather than a complete speech decoder.

Evidence

“ Twenty-four speech-typical participants each produced 50 TIMIT sentences three times per mode. Aloud, mimed and subvocal modes are defined separately; subvocal involves an occluded vocal tract with no phonation. Eight facial/neck EMG channels were recorded. ”

validation_scope · Section II-A; Figure 1; PDF p. 2 · confidence 0.99

“ The analysis compares forward prediction of EMG envelopes from SPARC and phoneme features; it explicitly does not demonstrate end-to-end decoding gains. ”

actual_novelty · Introduction and Conclusion; PDF pp. 1 and 4 · confidence 0.99

“ Features for every mode are extracted from paired aloud audio: 14 SPARC features for aloud, 12 kinematic features for silent modes, and 40 phoneme/silence one-hot indices at 50 Hz. ”

validation_scope · Section II-B; PDF p. 2 · confidence 0.99

“ Silent envelopes are warped to paired aloud EMG with FastDTW. Both predictors use the same trial-specific path. Authors acknowledge possible inflation of absolute correlations; reviewer assessment: results remain conditional on paired-reference alignment. ”

limitation · Section II-A and Section III-B; PDF pp. 2 and 4 · confidence 0.99

“ Evaluation is within subjects using sentence-level outer folds and training-only inner regularization selection; lags span -300 to +300 ms. Pearson correlations are averaged with Fisher z, and comparisons use FDR-controlled Wilcoxon tests. ”

validation_scope · Section II-C; PDF pp. 2-3 · confidence 0.99

“ Figure 2 reports higher SPARC prediction correlation than phoneme features on most channels across modes; subvocal performance remains above the permutation threshold, and upper-lip channel 6 is strongest. Exact primary per-channel means are not tabulated. ”

metric · Section III-A and Figure 2; PDF p. 3 · confidence 0.98

“ On the separate single-participant Gaddy dataset, voiced correlation increases from 0.443 ± 0.017 to 0.455 ± 0.021, and mimed correlation from 0.346 ± 0.029 to 0.364 ± 0.032. SPARC improves on 7 of 8 electrodes in each mode. ”

metric · Section III-A; PDF p. 3 · confidence 0.99

“ Variance partitioning uses squared correlations from articulatory, phoneme and combined models. Shared contribution dominates, and unique articulatory contribution exceeds unique phoneme contribution; this does not imply absence of shared phonemic information. ”

actual_novelty · Figure 3; Equations 3-5 and accompanying text; PDF pp. 3-4 · confidence 0.99

“ Weight maps sum absolute coefficients over lags, average across participants and normalize within each channel. Anatomical relationships are descriptive; a formal quantitative consistency test is deferred. ”

limitation · Section III-B and Figure 4; PDF p. 4 · confidence 0.98

“ Downstream decoding evaluation is future work. All reported comparisons are within-subject, and above-chance forward envelope prediction does not establish word recognition or new-user communication performance. ”

limitation · Section II-C and Conclusion; PDF pp. 3-4 · confidence 0.99

Limits

Technical limits

SPARC kinematics are inferred from voiced audio rather than measured during silent production. Aloud uses 14 features including pitch/loudness, silent modes use 12. DTW may inflate absolute correlations. Squared Pearson correlation is not automatically calibrated predictive R-squared; correlated predictors and normalized absolute weights limit causal anatomical interpretation.

Evaluation limits

All models are evaluated within subjects using sentence-level folds, not held-out users or cross-mode transfer. Silent evaluation depends on paired voiced reference and response-derived DTW alignment. Fold count and exact repeat grouping are not detailed. Primary numeric results are plotted rather than tabulated. The Gaddy replication has one participant and incompletely specified uncertainty aggregation.

Deployment limits

Offline forward encoding uses features derived from paired aloud audio and silently produced EMG warped to the paired aloud envelope. No autonomous silent decoder, latency measurement, wearable implementation, walking test, or patient communication study is provided.

Scope limits

Speech-typical participants and prompted sentences. Subvocal means attempted speech with an occluded vocal tract and no phonation; it should not be treated as unconstrained imagined speech. No clinical communication or end-to-end silent recognition study.