Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tuckute, Greta, Kotar, Klemen, Fedorenko, Evelina, Yamins, Daniel L. K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909738097180672
author Tuckute, Greta
Kotar, Klemen
Fedorenko, Evelina
Yamins, Daniel L. K.
author_facet Tuckute, Greta
Kotar, Klemen
Fedorenko, Evelina
Yamins, Daniel L. K.
contents We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete \textbf{cochlear tokens}. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11598
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Representing Speech Through Autoregressive Prediction of Cochlear Tokens
Tuckute, Greta
Kotar, Klemen
Fedorenko, Evelina
Yamins, Daniel L. K.
Computation and Language
Sound
Audio and Speech Processing
We introduce AuriStream, a biologically inspired model for encoding speech via a two-stage framework inspired by the human auditory processing hierarchy. The first stage transforms raw audio into a time-frequency representation based on the human cochlea, from which we extract discrete \textbf{cochlear tokens}. The second stage applies an autoregressive sequence model over the cochlear tokens. AuriStream learns meaningful phoneme and word representations, and state-of-the-art lexical semantics. AuriStream shows competitive performance on diverse downstream SUPERB speech tasks. Complementing AuriStream's strong representational capabilities, it generates continuations of audio which can be visualized in a spectrogram space and decoded back into audio, providing insights into the model's predictions. In summary, we present a two-stage framework for speech representation learning to advance the development of more human-like models that efficiently handle a range of speech-based tasks.
title Representing Speech Through Autoregressive Prediction of Cochlear Tokens
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.11598