Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
Fuente:
arXiv
Saved in:
| Main Authors: | Shim, Ryan Soh-Eun, Choi, Kwanghee, Chang, Kalvin, Hsu, Ming-Hao, Eichin, Florian, Wu, Zhizheng, Suhr, Alane, Hedderich, Michael A., Harwath, David, Mortensen, David R., Plank, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phonotactic Complexity across Dialects
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
by: Choi, Kwanghee, et al.
Published: (2026)
by: Choi, Kwanghee, et al.
Published: (2026)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
by: Li, Chin-Jou, et al.
Published: (2025)
by: Li, Chin-Jou, et al.
Published: (2025)
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
by: Eichin, Florian, et al.
Published: (2025)
by: Eichin, Florian, et al.
Published: (2025)
Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining
by: Körner, Felicia, et al.
Published: (2026)
by: Körner, Felicia, et al.
Published: (2026)
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns
by: Hedderich, Michael A., et al.
Published: (2025)
by: Hedderich, Michael A., et al.
Published: (2025)
Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning
by: Bafna, Niyati, et al.
Published: (2026)
by: Bafna, Niyati, et al.
Published: (2026)
Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
by: Shim, Ryan Soh-Eun, et al.
Published: (2025)
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
by: Casola, Silvia, et al.
Published: (2026)
by: Casola, Silvia, et al.
Published: (2026)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
by: Peng, Puyuan, et al.
Published: (2024)
by: Peng, Puyuan, et al.
Published: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
by: Zhu, Han, et al.
Published: (2024)
by: Zhu, Han, et al.
Published: (2024)
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling
by: Eichin, Florian, et al.
Published: (2024)
by: Eichin, Florian, et al.
Published: (2024)
PRiSM: Benchmarking Phone Realization in Speech Models
by: Bharadwaj, Shikhar, et al.
Published: (2026)
by: Bharadwaj, Shikhar, et al.
Published: (2026)
RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech
by: Zheng, Zhisheng, et al.
Published: (2025)
by: Zheng, Zhisheng, et al.
Published: (2025)
Evaluating Model Perception of Color Illusions in Photorealistic Scenes
by: Mao, Lingjun, et al.
Published: (2024)
by: Mao, Lingjun, et al.
Published: (2024)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
by: Peng, Puyuan, et al.
Published: (2025)
by: Peng, Puyuan, et al.
Published: (2025)
Cross-Lingual IPA Contrastive Learning for Zero-Shot NER
by: Sohn, Jimin, et al.
Published: (2025)
by: Sohn, Jimin, et al.
Published: (2025)
Probing the Robustness Properties of Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Interface Design for Self-Supervised Speech Models
by: Shih, Yi-Jen, et al.
Published: (2024)
by: Shih, Yi-Jen, et al.
Published: (2024)
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods
by: Jung, Haeji, et al.
Published: (2025)
by: Jung, Haeji, et al.
Published: (2025)
Lost in Transliteration: Bridging the Script Gap in Neural IR
by: Chari, Andreas, et al.
Published: (2025)
by: Chari, Andreas, et al.
Published: (2025)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
by: Diwan, Anuj, et al.
Published: (2026)
by: Diwan, Anuj, et al.
Published: (2026)
Multimodal Contextualized Semantic Parsing from Speech
by: Voas, Jordan, et al.
Published: (2024)
by: Voas, Jordan, et al.
Published: (2024)
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
by: Huang, Yiqiao, et al.
Published: (2024)
by: Huang, Yiqiao, et al.
Published: (2024)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Using Language Models to Disambiguate Lexical Choices in Translation
by: Barua, Josh, et al.
Published: (2024)
by: Barua, Josh, et al.
Published: (2024)
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
by: Sclar, Melanie, et al.
Published: (2023)
by: Sclar, Melanie, et al.
Published: (2023)
Long Chain-of-Thought Reasoning Across Languages
by: Barua, Josh, et al.
Published: (2025)
by: Barua, Josh, et al.
Published: (2025)
AyutthayaAlpha: A Thai-Latin Script Transliteration Transformer
by: Lauc, Davor, et al.
Published: (2024)
by: Lauc, Davor, et al.
Published: (2024)
Textless Speech-to-Speech Translation With Limited Parallel Data
by: Diwan, Anuj, et al.
Published: (2023)
by: Diwan, Anuj, et al.
Published: (2023)
SyllableLM: Learning Coarse Semantic Units for Speech Language Models
by: Baade, Alan, et al.
Published: (2024)
by: Baade, Alan, et al.
Published: (2024)
A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic
by: Gonzalez, Juan Moreno, et al.
Published: (2025)
by: Gonzalez, Juan Moreno, et al.
Published: (2025)
Romanized to Native Malayalam Script Transliteration Using an Encoder-Decoder Framework
by: Baiju, Bajiyo, et al.
Published: (2024)
by: Baiju, Bajiyo, et al.
Published: (2024)
Similar Items
-
Phonotactic Complexity across Dialects
by: Shim, Ryan Soh-Eun, et al.
Published: (2024) -
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024) -
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
by: Eichin, Florian, et al.
Published: (2025) -
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025) -
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
by: Choi, Kwanghee, et al.
Published: (2026)