Do self-supervised speech and language models extract similar representations as human brain?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Peili, He, Linyang, Fu, Li, Fan, Lu, Chang, Edward F., Li, Yuanning
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929229902381056
author Chen, Peili
He, Linyang
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
author_facet Chen, Peili
He, Linyang
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
contents Speech and language models trained through self-supervised learning (SSL) demonstrate strong alignment with brain activity during speech and language perception. However, given their distinct training modalities, it remains unclear whether they correlate with the same neural aspects. We directly address this question by evaluating the brain prediction performance of two representative SSL models, Wav2Vec2.0 and GPT-2, designed for speech and language tasks. Our findings reveal that both models accurately predict speech responses in the auditory cortex, with a significant correlation between their brain predictions. Notably, shared speech contextual information between Wav2Vec2.0 and GPT-2 accounts for the majority of explained variance in brain activity, surpassing static semantic and lower-level acoustic-phonetic information. These results underscore the convergence of speech contextual representations in SSL models and their alignment with the neural network underlying speech perception, offering valuable insights into both SSL models and the neural basis of speech and language processing.
format Preprint
id arxiv_https___arxiv_org_abs_2310_04645
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Do self-supervised speech and language models extract similar representations as human brain?
Chen, Peili
He, Linyang
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
Neurons and Cognition
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Speech and language models trained through self-supervised learning (SSL) demonstrate strong alignment with brain activity during speech and language perception. However, given their distinct training modalities, it remains unclear whether they correlate with the same neural aspects. We directly address this question by evaluating the brain prediction performance of two representative SSL models, Wav2Vec2.0 and GPT-2, designed for speech and language tasks. Our findings reveal that both models accurately predict speech responses in the auditory cortex, with a significant correlation between their brain predictions. Notably, shared speech contextual information between Wav2Vec2.0 and GPT-2 accounts for the majority of explained variance in brain activity, surpassing static semantic and lower-level acoustic-phonetic information. These results underscore the convergence of speech contextual representations in SSL models and their alignment with the neural network underlying speech perception, offering valuable insights into both SSL models and the neural basis of speech and language processing.
title Do self-supervised speech and language models extract similar representations as human brain?
topic Neurons and Cognition
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2310.04645