Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shim, Ryan Soh-Eun, Choi, Kwanghee, Chang, Kalvin, Hsu, Ming-Hao, Eichin, Florian, Wu, Zhizheng, Suhr, Alane, Hedderich, Michael A., Harwath, David, Mortensen, David R., Plank, Barbara
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918273890648064
author Shim, Ryan Soh-Eun
Choi, Kwanghee
Chang, Kalvin
Hsu, Ming-Hao
Eichin, Florian
Wu, Zhizheng
Suhr, Alane
Hedderich, Michael A.
Harwath, David
Mortensen, David R.
Plank, Barbara
author_facet Shim, Ryan Soh-Eun
Choi, Kwanghee
Chang, Kalvin
Hsu, Ming-Hao
Eichin, Florian
Wu, Zhizheng
Suhr, Alane
Hedderich, Michael A.
Harwath, David
Mortensen, David R.
Plank, Barbara
contents Multilingual speech foundation models such as Whisper are trained on web-scale data, where data for each language consists of a myriad of regional varieties. However, different regional varieties often employ different scripts to write the same language, rendering speech recognition output also subject to non-determinism in the output script. To mitigate this problem, we show that script is linearly encoded in the activation space of multilingual speech models, and that modifying activations at inference time enables direct control over output script. We find the addition of such script vectors to activations at test time can induce script change even in unconventional language-script pairings (e.g. Italian in Cyrillic and Japanese in Latin script). We apply this approach to inducing post-hoc control over the script of speech recognition output, where we observe competitive performance across all model sizes of Whisper.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02906
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
Shim, Ryan Soh-Eun
Choi, Kwanghee
Chang, Kalvin
Hsu, Ming-Hao
Eichin, Florian
Wu, Zhizheng
Suhr, Alane
Hedderich, Michael A.
Harwath, David
Mortensen, David R.
Plank, Barbara
Computation and Language
Multilingual speech foundation models such as Whisper are trained on web-scale data, where data for each language consists of a myriad of regional varieties. However, different regional varieties often employ different scripts to write the same language, rendering speech recognition output also subject to non-determinism in the output script. To mitigate this problem, we show that script is linearly encoded in the activation space of multilingual speech models, and that modifying activations at inference time enables direct control over output script. We find the addition of such script vectors to activations at test time can induce script change even in unconventional language-script pairings (e.g. Italian in Cyrillic and Japanese in Latin script). We apply this approach to inducing post-hoc control over the script of speech recognition output, where we observe competitive performance across all model sizes of Whisper.
title Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
topic Computation and Language
url https://arxiv.org/abs/2601.02906