NeurIPS 2026

SENSE: Semantic Neural Speech Synthesis
from Brain Dynamics via Spatial Graph Encoding

Jisoo Park¹*, Seonghak Lee¹, Hyojin Park², Junseok Kwon¹†

¹Chung-Ang University, South Korea  ·  ²University of Birmingham, United Kingdom

susiehome@cau.ac.kr

* Work done during a research visit to the University of Birmingham.  † Corresponding author.

SENSE reads EEG recorded while a person listens, encodes it over the electrode geometry of the scalp, aligns that representation with a frozen CLIP text embedding on congruent trials only, and synthesizes the speech waveform. At inference only EEG is used.
Channel importance across the scalp for prior work and SENSE
fig1.jpg
Figure 1. EEG recorded during speech perception carries acoustic, sensorimotor, and lexical-semantic information distributed across the scalp. Prior work yields channel patterns with no clear anatomical structure (left). SENSE recovers attribution aligned with all three functional regions (right).

Abstract

Existing EEG-to-speech approaches formulate the task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves semantic content. We argue that EEG signals carry semantic as well as acoustic information, and that two limitations follow from the acoustic-only framing.

First, prior encoders treat the 128 electrodes as unordered feature dimensions, ignoring their spatial relationships. Second, they use the N400 corpus in a paradigm-agnostic manner, although the corpus is built around a congruency distinction designed to isolate semantic processing.

SENSE addresses both. A graph-based encoder models the electrodes over their scalp geometry, and EEG Semantic Conditioning (ESC) aligns the EEG latent with a frozen CLIP text embedding using congruent trials only. On the N400 dataset, SENSE outperforms prior methods on both acoustic and semantic metrics, with the largest gains on unseen participants.

Audio samples

Speech generated from EEG alone

All samples come from the unseen-subject split. Generation is conditioned on EEG only, without text or guided sampling. Transcriptions are produced by Whisper-large, with words matching the stimulus shown in bold.

These are best-scoring examples.

“The old man walked with a cane.”

sub-24 · unseen subject · WER 29%

GT
SENSE

Whisper on SENSE output: “The old man walked with mutting.”

“On one hand I have five fingers.”

sub-23 · unseen subject · WER 29%

GT
SENSE

Whisper on SENSE output: “On one hand, I have five ears.”

“He liked to sit on the park bench.”

sub-24 · unseen subject · WER 38%

GT
SENSE

Whisper on SENSE output: “He lies to sit on the hard vans.”

“The angry boys got in a fight.”

sub-23 · unseen subject · WER 43%

GT
SENSE

Whisper on SENSE output: “The Angry Boys got in the fight.”

“He shut the door.”

sub-23 · unseen subject · WER 50%

GT
SENSE

Whisper on SENSE output: “He shust the poor.”

“He wore a cast because his arm broke.”

sub-24 · unseen subject · WER 50%

GT
SENSE

Whisper on SENSE output: “He wore a bass, he does his arm up”

Method

Structure-aware encoding and semantic conditioning

SENSE training pipeline
fig2.png
Figure 2. EEG is encoded into a latent representation and mapped to the prior of a VITS speech decoder. Solid arrows indicate the forward data flow, dashed arrows indicate loss supervision. At inference the EEG decoder, phoneme predictor, and CLIP encoder are discarded, and only EEG is required.

Results

Comparison with prior methods

Three test conditions on the N400 corpus: held-out sentences (Audio), held-out participants (Subject), and held-out both. Semantic metrics are computed on congruent trials only. All results are averaged over five seeds.

Model Split MCD ↓ Mel-Corr ↑ STOI ↑ WER ↓ BERT-R ↑ CLIP-Sim ↑
FESDEBoth11.80 ±0.1219.87 ±1.970.27851.21860.03810.7557
FE-PhonemeBoth10.52 ±0.0828.29 ±1.340.39791.16100.04280.7509
SENSEBoth10.44 ±0.1031.18 ±0.600.39961.01860.08520.7654
FESDEAudio11.71 ±0.0419.14 ±0.470.27921.21250.02890.7601
FE-PhonemeAudio10.73 ±0.0626.53 ±0.300.39181.13600.06020.7554
SENSEAudio10.44 ±0.0128.50 ±0.240.40391.03750.09590.7601
FESDESubject11.73 ±0.0418.70 ±0.290.27441.21710.02690.7587
FE-PhonemeSubject10.65 ±0.0226.90 ±0.080.38061.12310.05640.7555
SENSESubject9.928 ±0.0332.07 ±0.140.41260.99480.10880.7647

SENSE outperforms both baselines across all splits and metrics. The margin is largest on the unseen-subject split, indicating that the learned representations generalize across participants rather than relying on subject-specific cues. Absolute WER remains near 1.0 for all methods, so intelligibility-level reconstruction from EEG remains an open challenge.

Qualitative comparison of generated speech
fig3.jpg
Figure 3. Mel-spectrograms and Whisper-large transcriptions for ground truth, SENSE, FESDE, and FE-Phoneme. Words matching the ground truth are shown in bold, semantically or phonetically similar words are underlined.

Cross-subject generalization

Scaling the number of training subjects

Models are trained on k participants and evaluated on a fixed held-out pair. The dashed line marks the strongest baseline trained on all 18 subjects. SENSE exceeds that reference on WER at k=2, and matches it on MCD and Mel-Correlation by k=8.

Scaling with the number of training subjects
fig4.png
Figure 4. Performance for k in {2, 4, 8, 18}, evaluated on the held-out pair (sub-23, sub-24). Shaded regions show ±1 std over 5 seeds. Orange circles mark the smallest k at which SENSE first matches or exceeds the 18-subject baseline. Since evaluation is on unseen participants, these results reflect cross-subject generalization rather than in-distribution scaling.

Ablation

Component ablation

Each row removes one component from the full model. Removing ESC degrades the semantic metrics, while removing CTC produces the largest increase in WER, indicating that phoneme-level supervision also preserves intelligibility.

Model Ch.AttGNNSkipCTCESC MCD ↓Mel-Corr ↑STOI ↑WER ↓BERT-R ↑CLIP-Sim ↑
SENSE✓✓✓✓✓9.93 ±0.0332.07 ±0.140.41260.99480.10880.7647
w/o ESC✓✓✓✓✗10.34 ±0.0329.43 ±0.270.40161.03230.07630.7601
w/ shuffled ESC✓✓✓✓shuf.10.38 ±0.0429.25 ±0.280.40021.03850.07210.7595
w/o CTC✓✓✓✗✓10.50 ±0.0127.24 ±0.270.37211.04880.05880.7624
w/o CNN Skip✓✓✗✓✓10.40 ±0.0129.64 ±0.250.39701.03060.07240.7579
w/o GNN✓✗✗✓✓10.11 ±0.0229.54 ±0.220.39511.03550.07880.7577
w/o Ch.Att✗✓✓✓✓10.24 ±0.0129.08 ±0.160.40461.01700.09260.7639
EEG-text similarity matrices across semantic encoders
fig6.png
Figure 6. Cosine similarity between mean EEG embeddings (rows) and text embeddings (columns) for matched words. Stronger word-specific alignment appears as a brighter diagonal, and Δ is the gap between diagonal and off-diagonal means. CLIP shows the brightest diagonal and the largest gap.

Analysis

Channel attribution

The channel attention gate is initialized with a positive bias on auditory-related electrodes and zero elsewhere. No prior is placed on sensorimotor or centro-parietal regions. After training, the auditory-prior channels are modestly down-weighted while central electrodes (C3, C4) gain importance, so the resulting pattern emerges from training rather than design.

ESC attribution across encoder variants
fig7.jpg
Figure 7. Per-channel gradient magnitude of the EEG-CLIP alignment, normalized per model. (b) Without GNN: centro-parietal concentration consistent with the N400 response. (c) GNN without skip connection: emphasis shifts to central hub electrodes, a topology artifact of dense connectivity. (d) Full SENSE: a distributed pattern over temporal-auditory (TP7, TP8), motor (C3, C4), and centro-parietal (CPz) regions.

The prominence of C3 and C4 is consistent with prior findings that speech perception engages sensorimotor cortex during passive listening. We treat these maps as model-internal diagnostics rather than evidence of neural sources, given the known limitations of saliency-based attribution.