Guardado en:
Detalles Bibliográficos
Autores principales: Modica, Luca, Landin, Filip, Farahani, Mehrdad, Qian, Livia, Skantze, Gabriel, Johansson, Richard
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2605.22170
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917519491596288
author Modica, Luca
Landin, Filip
Farahani, Mehrdad
Qian, Livia
Skantze, Gabriel
Johansson, Richard
author_facet Modica, Luca
Landin, Filip
Farahani, Mehrdad
Qian, Livia
Skantze, Gabriel
Johansson, Richard
contents In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when operating in the two modalities. We focus on how these systems encode, store, and retrieve factual knowledge, which has previously been investigated for text-only models. To investigate mechanisms behind the storage and recall of factual association in SLMs, we leverage Causal Mediation Analysis, a technique previously applied to text-based models. Initial results using SpiritLM, a multimodal model integrating discrete speech tokens reveal discrepancies between text-to-text and speech-to-text results, suggesting that the emergent mechanisms for factual recall are only partially carried over from the text to the speech modality. These results advance our understanding of how internal mechanisms encode factual associations in SLMs while contributing insights for improving speech-enabled AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22170
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
Modica, Luca
Landin, Filip
Farahani, Mehrdad
Qian, Livia
Skantze, Gabriel
Johansson, Richard
Computation and Language
In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when operating in the two modalities. We focus on how these systems encode, store, and retrieve factual knowledge, which has previously been investigated for text-only models. To investigate mechanisms behind the storage and recall of factual association in SLMs, we leverage Causal Mediation Analysis, a technique previously applied to text-based models. Initial results using SpiritLM, a multimodal model integrating discrete speech tokens reveal discrepancies between text-to-text and speech-to-text results, suggesting that the emergent mechanisms for factual recall are only partially carried over from the text to the speech modality. These results advance our understanding of how internal mechanisms encode factual associations in SLMs while contributing insights for improving speech-enabled AI systems.
title Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
topic Computation and Language
url https://arxiv.org/abs/2605.22170