¹Chung-Ang University, South Korea · ²University of Birmingham, United Kingdom
* Work done during a research visit to the University of Birmingham. † Corresponding author.
Abstract
Existing EEG-to-speech approaches formulate the task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves semantic content. We argue that EEG signals carry semantic as well as acoustic information, and that two limitations follow from the acoustic-only framing.
First, prior encoders treat the 128 electrodes as unordered feature dimensions, ignoring their spatial relationships. Second, they use the N400 corpus in a paradigm-agnostic manner, although the corpus is built around a congruency distinction designed to isolate semantic processing.
SENSE addresses both. A graph-based encoder models the electrodes over their scalp geometry, and EEG Semantic Conditioning (ESC) aligns the EEG latent with a frozen CLIP text embedding using congruent trials only. On the N400 dataset, SENSE outperforms prior methods on both acoustic and semantic metrics, with the largest gains on unseen participants.
Audio samples
All samples come from the unseen-subject split. Generation is conditioned on EEG only, without text or guided sampling. Transcriptions are produced by Whisper-large, with words matching the stimulus shown in bold.
These are best-scoring examples.
“The old man walked with a cane.”
Whisper on SENSE output: “The old man walked with mutting.”
“On one hand I have five fingers.”
Whisper on SENSE output: “On one hand, I have five ears.”
“He liked to sit on the park bench.”
Whisper on SENSE output: “He lies to sit on the hard vans.”
“The angry boys got in a fight.”
Whisper on SENSE output: “The Angry Boys got in the fight.”
“He shut the door.”
Whisper on SENSE output: “He shust the poor.”
“He wore a cast because his arm broke.”
Whisper on SENSE output: “He wore a bass, he does his arm up”
Method
Results
Three test conditions on the N400 corpus: held-out sentences (Audio), held-out participants (Subject), and held-out both. Semantic metrics are computed on congruent trials only. All results are averaged over five seeds.
| Model | Split | MCD ↓ | Mel-Corr ↑ | STOI ↑ | WER ↓ | BERT-R ↑ | CLIP-Sim ↑ |
|---|---|---|---|---|---|---|---|
| FESDE | Both | 11.80 ±0.12 | 19.87 ±1.97 | 0.2785 | 1.2186 | 0.0381 | 0.7557 |
| FE-Phoneme | Both | 10.52 ±0.08 | 28.29 ±1.34 | 0.3979 | 1.1610 | 0.0428 | 0.7509 |
| SENSE | Both | 10.44 ±0.10 | 31.18 ±0.60 | 0.3996 | 1.0186 | 0.0852 | 0.7654 |
| FESDE | Audio | 11.71 ±0.04 | 19.14 ±0.47 | 0.2792 | 1.2125 | 0.0289 | 0.7601 |
| FE-Phoneme | Audio | 10.73 ±0.06 | 26.53 ±0.30 | 0.3918 | 1.1360 | 0.0602 | 0.7554 |
| SENSE | Audio | 10.44 ±0.01 | 28.50 ±0.24 | 0.4039 | 1.0375 | 0.0959 | 0.7601 |
| FESDE | Subject | 11.73 ±0.04 | 18.70 ±0.29 | 0.2744 | 1.2171 | 0.0269 | 0.7587 |
| FE-Phoneme | Subject | 10.65 ±0.02 | 26.90 ±0.08 | 0.3806 | 1.1231 | 0.0564 | 0.7555 |
| SENSE | Subject | 9.928 ±0.03 | 32.07 ±0.14 | 0.4126 | 0.9948 | 0.1088 | 0.7647 |
SENSE outperforms both baselines across all splits and metrics. The margin is largest on the unseen-subject split, indicating that the learned representations generalize across participants rather than relying on subject-specific cues. Absolute WER remains near 1.0 for all methods, so intelligibility-level reconstruction from EEG remains an open challenge.
Cross-subject generalization
Models are trained on k participants and evaluated on a fixed held-out pair. The dashed line marks the strongest baseline trained on all 18 subjects. SENSE exceeds that reference on WER at k=2, and matches it on MCD and Mel-Correlation by k=8.
Ablation
Each row removes one component from the full model. Removing ESC degrades the semantic metrics, while removing CTC produces the largest increase in WER, indicating that phoneme-level supervision also preserves intelligibility.
| Model | Ch.Att | GNN | Skip | CTC | ESC | MCD ↓ | Mel-Corr ↑ | STOI ↑ | WER ↓ | BERT-R ↑ | CLIP-Sim ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| SENSE | ✓ | ✓ | ✓ | ✓ | ✓ | 9.93 ±0.03 | 32.07 ±0.14 | 0.4126 | 0.9948 | 0.1088 | 0.7647 |
| w/o ESC | ✓ | ✓ | ✓ | ✓ | ✗ | 10.34 ±0.03 | 29.43 ±0.27 | 0.4016 | 1.0323 | 0.0763 | 0.7601 |
| w/ shuffled ESC | ✓ | ✓ | ✓ | ✓ | shuf. | 10.38 ±0.04 | 29.25 ±0.28 | 0.4002 | 1.0385 | 0.0721 | 0.7595 |
| w/o CTC | ✓ | ✓ | ✓ | ✗ | ✓ | 10.50 ±0.01 | 27.24 ±0.27 | 0.3721 | 1.0488 | 0.0588 | 0.7624 |
| w/o CNN Skip | ✓ | ✓ | ✗ | ✓ | ✓ | 10.40 ±0.01 | 29.64 ±0.25 | 0.3970 | 1.0306 | 0.0724 | 0.7579 |
| w/o GNN | ✓ | ✗ | ✗ | ✓ | ✓ | 10.11 ±0.02 | 29.54 ±0.22 | 0.3951 | 1.0355 | 0.0788 | 0.7577 |
| w/o Ch.Att | ✗ | ✓ | ✓ | ✓ | ✓ | 10.24 ±0.01 | 29.08 ±0.16 | 0.4046 | 1.0170 | 0.0926 | 0.7639 |
Analysis
The channel attention gate is initialized with a positive bias on auditory-related electrodes and zero elsewhere. No prior is placed on sensorimotor or centro-parietal regions. After training, the auditory-prior channels are modestly down-weighted while central electrodes (C3, C4) gain importance, so the resulting pattern emerges from training rather than design.
The prominence of C3 and C4 is consistent with prior findings that speech perception engages sensorimotor cortex during passive listening. We treat these maps as model-internal diagnostics rather than evidence of neural sources, given the known limitations of saliency-based attribution.