Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nakai, Toshiki, Suresh, Varsha, Demberg, Vera
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912997321998336
author Nakai, Toshiki
Suresh, Varsha
Demberg, Vera
author_facet Nakai, Toshiki
Suresh, Varsha
Demberg, Vera
contents Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-specific processing. We introduce a generation-step-aware framework for evaluating cross-modal computation that (i) identifies language-selective neurons for each modality at different decoding steps, (ii) decomposes them into language-representation and language-control roles, and (iii) enables cross-modal comparison via overlap measures and causal intervention, including cross-modal steering of output language. Applying our framework to SeamlessM4T v2, we find that cross-modal language alignment is strongest at the first decoding step, where language-representation neurons are shared across modalities, but weakens as generation proceeds, indicating a shift toward modality-specific autoregressive processing. In contrast, language-control neurons identified from speech transfer causally to text generation, revealing partially shared circuitry for output-language control that strengthens at later decoding steps. These results show that cross-modal processing is both time- and function-dependent, providing a more nuanced view of multilingual computation in speech-text models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17387
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
Nakai, Toshiki
Suresh, Varsha
Demberg, Vera
Computation and Language
Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether this reflects shared computation for the same language or modality-specific processing. We introduce a generation-step-aware framework for evaluating cross-modal computation that (i) identifies language-selective neurons for each modality at different decoding steps, (ii) decomposes them into language-representation and language-control roles, and (iii) enables cross-modal comparison via overlap measures and causal intervention, including cross-modal steering of output language. Applying our framework to SeamlessM4T v2, we find that cross-modal language alignment is strongest at the first decoding step, where language-representation neurons are shared across modalities, but weakens as generation proceeds, indicating a shift toward modality-specific autoregressive processing. In contrast, language-control neurons identified from speech transfer causally to text generation, revealing partially shared circuitry for output-language control that strengthens at later decoding steps. These results show that cross-modal processing is both time- and function-dependent, providing a more nuanced view of multilingual computation in speech-text models.
title Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
topic Computation and Language
url https://arxiv.org/abs/2601.17387