Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Christopher, Yaari, Adam Uri, Singh, Aaditya K, Subramaniam, Vighnesh, Rosenfarb, Dana, DeWitt, Jan, Misra, Pranav, Madsen, Joseph R., Stone, Scellig, Kreiman, Gabriel, Katz, Boris, Cases, Ignacio, Barbu, Andrei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910696349892608
author Wang, Christopher
Yaari, Adam Uri
Singh, Aaditya K
Subramaniam, Vighnesh
Rosenfarb, Dana
DeWitt, Jan
Misra, Pranav
Madsen, Joseph R.
Stone, Scellig
Kreiman, Gabriel
Katz, Boris
Cases, Ignacio
Barbu, Andrei
author_facet Wang, Christopher
Yaari, Adam Uri
Singh, Aaditya K
Subramaniam, Vighnesh
Rosenfarb, Dana
DeWitt, Jan
Misra, Pranav
Madsen, Joseph R.
Stone, Scellig
Kreiman, Gabriel
Katz, Boris
Cases, Ignacio
Barbu, Andrei
contents We present the Brain Treebank, a large-scale dataset of electrophysiological neural responses, recorded from intracranial probes while 10 subjects watched one or more Hollywood movies. Subjects watched on average 2.6 Hollywood movies, for an average viewing time of 4.3 hours, and a total of 43 hours. The audio track for each movie was transcribed with manual corrections. Word onsets were manually annotated on spectrograms of the audio track for each movie. Each transcript was automatically parsed and manually corrected into the universal dependencies (UD) formalism, assigning a part of speech to every word and a dependency parse to every sentence. In total, subjects heard over 38,000 sentences (223,000 words), while they had on average 168 electrodes implanted. This is the largest dataset of intracranial recordings featuring grounded naturalistic language, one of the largest English UD treebanks in general, and one of only a few UD treebanks aligned to multimodal features. We hope that this dataset serves as a bridge between linguistic concepts, perception, and their neural representations. To that end, we present an analysis of which electrodes are sensitive to language features while also mapping out a rough time course of language processing across these electrodes. The Brain Treebank is available at https://BrainTreebank.dev/
format Preprint
id arxiv_https___arxiv_org_abs_2411_08343
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli
Wang, Christopher
Yaari, Adam Uri
Singh, Aaditya K
Subramaniam, Vighnesh
Rosenfarb, Dana
DeWitt, Jan
Misra, Pranav
Madsen, Joseph R.
Stone, Scellig
Kreiman, Gabriel
Katz, Boris
Cases, Ignacio
Barbu, Andrei
Neurons and Cognition
We present the Brain Treebank, a large-scale dataset of electrophysiological neural responses, recorded from intracranial probes while 10 subjects watched one or more Hollywood movies. Subjects watched on average 2.6 Hollywood movies, for an average viewing time of 4.3 hours, and a total of 43 hours. The audio track for each movie was transcribed with manual corrections. Word onsets were manually annotated on spectrograms of the audio track for each movie. Each transcript was automatically parsed and manually corrected into the universal dependencies (UD) formalism, assigning a part of speech to every word and a dependency parse to every sentence. In total, subjects heard over 38,000 sentences (223,000 words), while they had on average 168 electrodes implanted. This is the largest dataset of intracranial recordings featuring grounded naturalistic language, one of the largest English UD treebanks in general, and one of only a few UD treebanks aligned to multimodal features. We hope that this dataset serves as a bridge between linguistic concepts, perception, and their neural representations. To that end, we present an analysis of which electrodes are sensitive to language features while also mapping out a rough time course of language processing across these electrodes. The Brain Treebank is available at https://BrainTreebank.dev/
title Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli
topic Neurons and Cognition
url https://arxiv.org/abs/2411.08343