Saved in:
Bibliographic Details
Main Authors: Zhang, Yizi, He, Linyang, Fan, Chaofei, Liu, Tingkai, Yu, Han, Le, Trung, Li, Jingyuan, Linderman, Scott, Duncker, Lea, Willett, Francis R, Mesgarani, Nima, Paninski, Liam
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.21740
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909040789946368
author Zhang, Yizi
He, Linyang
Fan, Chaofei
Liu, Tingkai
Yu, Han
Le, Trung
Li, Jingyuan
Linderman, Scott
Duncker, Lea
Willett, Francis R
Mesgarani, Nima
Paninski, Liam
author_facet Zhang, Yizi
He, Linyang
Fan, Chaofei
Liu, Tingkai
Yu, Han
Le, Trung
Li, Jingyuan
Linderman, Scott
Duncker, Lea
Willett, Francis R
Mesgarani, Nima
Paninski, Liam
contents Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously. Here, we introduce an end-to-end BraIn-to-Text (BIT) framework that translates neural activity into coherent sentences using a single differentiable neural network. Central to our approach is a cross-task, cross-species pretrained neural encoder, whose representations transfer to both attempted and imagined speech. In a cascaded setting with an n-gram LM, the pretrained encoder establishes a new state-of-the-art (SOTA) on the Brain-to-Text '24 and '25 benchmarks. Integrated end-to-end with audio large language models (LLMs) and trained with contrastive learning for cross-modal alignment, BIT reduces the word error rate (WER) of the prior end-to-end method from 24.69% to 10.22%. Notably, we find that small-scale audio LLMs markedly improve end-to-end decoding. Beyond record-setting performance, BIT aligns attempted and imagined speech embeddings to enable cross-task generalization. Altogether, our approach advances the integration of large, diverse neural datasets, paving the way for an end-to-end decoding framework that supports seamless, differentiable optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A cross-species neural foundation model for end-to-end speech decoding
Zhang, Yizi
He, Linyang
Fan, Chaofei
Liu, Tingkai
Yu, Han
Le, Trung
Li, Jingyuan
Linderman, Scott
Duncker, Lea
Willett, Francis R
Mesgarani, Nima
Paninski, Liam
Computation and Language
Artificial Intelligence
Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram language model (LM), preventing joint optimization of all stages simultaneously. Here, we introduce an end-to-end BraIn-to-Text (BIT) framework that translates neural activity into coherent sentences using a single differentiable neural network. Central to our approach is a cross-task, cross-species pretrained neural encoder, whose representations transfer to both attempted and imagined speech. In a cascaded setting with an n-gram LM, the pretrained encoder establishes a new state-of-the-art (SOTA) on the Brain-to-Text '24 and '25 benchmarks. Integrated end-to-end with audio large language models (LLMs) and trained with contrastive learning for cross-modal alignment, BIT reduces the word error rate (WER) of the prior end-to-end method from 24.69% to 10.22%. Notably, we find that small-scale audio LLMs markedly improve end-to-end decoding. Beyond record-setting performance, BIT aligns attempted and imagined speech embeddings to enable cross-task generalization. Altogether, our approach advances the integration of large, diverse neural datasets, paving the way for an end-to-end decoding framework that supports seamless, differentiable optimization.
title A cross-species neural foundation model for end-to-end speech decoding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.21740