Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Xuan, Zeng, Chang, Feng, Tiantian, Wang, Shih-Heng, Ma, Jianbo, Narayanan, Shrikanth
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2603.10371
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912960195067904
author Shi, Xuan
Zeng, Chang
Feng, Tiantian
Wang, Shih-Heng
Ma, Jianbo
Narayanan, Shrikanth
author_facet Shi, Xuan
Zeng, Chang
Feng, Tiantian
Wang, Shih-Heng
Ma, Jianbo
Narayanan, Shrikanth
contents Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. These tokenizers are expected to preserve both semantic and acoustic information for downstream understanding and generation. However, emerging evidence suggests that what is termed "semantic" in speech representations does not align with text-derived semantics: a mismatch that can degrade multimodal LLM performance. In this paper, we systematically analyze the information encoded by several widely used speech tokenizers, disentangling their semantic and phonetic content through word-level probing tasks, layerwise representation analysis, and cross-modal alignment metrics such as CKA. Our results show that current tokenizers primarily capture phonetic rather than lexical-semantic structure, and we derive practical implications for the design of next-generation speech tokenization methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10371
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Speech Codec Probing from Semantic and Phonetic Perspectives
Shi, Xuan
Zeng, Chang
Feng, Tiantian
Wang, Shih-Heng
Ma, Jianbo
Narayanan, Shrikanth
Audio and Speech Processing
Computation and Language
Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. These tokenizers are expected to preserve both semantic and acoustic information for downstream understanding and generation. However, emerging evidence suggests that what is termed "semantic" in speech representations does not align with text-derived semantics: a mismatch that can degrade multimodal LLM performance. In this paper, we systematically analyze the information encoded by several widely used speech tokenizers, disentangling their semantic and phonetic content through word-level probing tasks, layerwise representation analysis, and cross-modal alignment metrics such as CKA. Our results show that current tokenizers primarily capture phonetic rather than lexical-semantic structure, and we derive practical implications for the design of next-generation speech tokenization methods.
title Speech Codec Probing from Semantic and Phonetic Perspectives
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2603.10371