Speech language models lack important brain-relevant semantics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oota, Subba Reddy, Çelik, Emin, Deniz, Fatma, Toneva, Mariya
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916289067352064
author Oota, Subba Reddy
Çelik, Emin
Deniz, Fatma
Toneva, Mariya
author_facet Oota, Subba Reddy
Çelik, Emin
Deniz, Fatma
Toneva, Mariya
contents Despite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree. This poses the question of what types of information language models truly predict in the brain. We investigate this question via a direct approach, in which we systematically remove specific low-level stimulus features (textual, speech, and visual) from language model representations to assess their impact on alignment with fMRI brain recordings during reading and listening. Comparing these findings with speech-based language models reveals starkly different effects of low-level features on brain alignment. While text-based models show reduced alignment in early sensory regions post-removal, they retain significant predictive power in late language regions. In contrast, speech-based models maintain strong alignment in early auditory regions even after feature removal but lose all predictive power in late language regions. These results suggest that speech-based models provide insights into additional information processed by early auditory regions, but caution is needed when using them to model processing in late language regions. We make our code publicly available. [https://github.com/subbareddy248/speech-llm-brain]
format Preprint
id arxiv_https___arxiv_org_abs_2311_04664
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Speech language models lack important brain-relevant semantics
Oota, Subba Reddy
Çelik, Emin
Deniz, Fatma
Toneva, Mariya
Computation and Language
Machine Learning
Audio and Speech Processing
Neurons and Cognition
Despite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree. This poses the question of what types of information language models truly predict in the brain. We investigate this question via a direct approach, in which we systematically remove specific low-level stimulus features (textual, speech, and visual) from language model representations to assess their impact on alignment with fMRI brain recordings during reading and listening. Comparing these findings with speech-based language models reveals starkly different effects of low-level features on brain alignment. While text-based models show reduced alignment in early sensory regions post-removal, they retain significant predictive power in late language regions. In contrast, speech-based models maintain strong alignment in early auditory regions even after feature removal but lose all predictive power in late language regions. These results suggest that speech-based models provide insights into additional information processed by early auditory regions, but caution is needed when using them to model processing in late language regions. We make our code publicly available. [https://github.com/subbareddy248/speech-llm-brain]
title Speech language models lack important brain-relevant semantics
topic Computation and Language
Machine Learning
Audio and Speech Processing
Neurons and Cognition
url https://arxiv.org/abs/2311.04664