Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moussa, Omer, Toneva, Mariya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913875040927744
author Moussa, Omer
Toneva, Mariya
author_facet Moussa, Omer
Toneva, Mariya
contents Pretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain
Moussa, Omer
Toneva, Mariya
Computation and Language
Sound
Audio and Speech Processing
Neurons and Cognition
Pretrained self-supervised speech models excel in speech tasks but do not reflect the hierarchy of human speech processing, as they encode rich semantics in middle layers and poor semantics in late layers. Recent work showed that brain-tuning (fine-tuning models using human brain recordings) improves speech models' semantic understanding. Here, we examine how well brain-tuned models further reflect the brain's intermediate stages of speech processing. We find that late layers of brain-tuned models substantially improve over pretrained models in their alignment with semantic language regions. Further layer-wise probing reveals that early layers remain dedicated to low-level acoustic features, while late layers become the best at complex high-level tasks. These findings show that brain-tuned models not only perform better but also exhibit a well-defined hierarchical processing going from acoustic to semantic representations, making them better model organisms for human speech processing.
title Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain
topic Computation and Language
Sound
Audio and Speech Processing
Neurons and Cognition
url https://arxiv.org/abs/2506.03832