The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Siyu, Ballier, Nicolas, Levow, Gina-Anne, Wright, Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels
by: Liang, Siyu, et al.
Published: (2025)
by: Liang, Siyu, et al.
Published: (2025)
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
by: Liang, Siyu, et al.
Published: (2025)
by: Liang, Siyu, et al.
Published: (2025)
A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
by: Scott, Michael, et al.
Published: (2025)
by: Scott, Michael, et al.
Published: (2025)
Hybrid Neural-LLM Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan
by: Liang, Siyu, et al.
Published: (2026)
by: Liang, Siyu, et al.
Published: (2026)
TEII: Think, Explain, Interact and Iterate with Large Language Models to Solve Cross-lingual Emotion Detection
by: Cheng, Long, et al.
Published: (2024)
by: Cheng, Long, et al.
Published: (2024)
Anatomy of Industrial Scale Multilingual ASR
by: Ramirez, Francis McCann, et al.
Published: (2024)
by: Ramirez, Francis McCann, et al.
Published: (2024)
"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias
by: Liang, Siyu, et al.
Published: (2026)
by: Liang, Siyu, et al.
Published: (2026)
SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding
by: Hou, Shuyang, et al.
Published: (2026)
by: Hou, Shuyang, et al.
Published: (2026)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
by: Chen, Qian, et al.
Published: (2023)
by: Chen, Qian, et al.
Published: (2023)
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
by: Liu, Dancheng, et al.
Published: (2025)
by: Liu, Dancheng, et al.
Published: (2025)
Dynamic Acoustic Model Architecture Optimization in Training for ASR
by: Xu, Jingjing, et al.
Published: (2025)
by: Xu, Jingjing, et al.
Published: (2025)
Romanization Encoding For Multilingual ASR
by: Ding, Wen, et al.
Published: (2024)
by: Ding, Wen, et al.
Published: (2024)
Linguistically Informed Evaluation of Multilingual ASR for African Languages
by: Chen, Fei-Yueh, et al.
Published: (2026)
by: Chen, Fei-Yueh, et al.
Published: (2026)
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
by: Dang, Quy-Anh, et al.
Published: (2026)
by: Dang, Quy-Anh, et al.
Published: (2026)
MLMA: Towards Multilingual ASR With Mamba-based Architectures
by: Ali, Mohamed Nabih, et al.
Published: (2025)
by: Ali, Mohamed Nabih, et al.
Published: (2025)
Revisiting Acoustic Features for Robust ASR
by: Shah, Muhammad A., et al.
Published: (2024)
by: Shah, Muhammad A., et al.
Published: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs
by: Chen, Haoyang, et al.
Published: (2025)
by: Chen, Haoyang, et al.
Published: (2025)
Building Robust and Scalable Multilingual ASR for Indian Languages
by: Gangwar, Arjun, et al.
Published: (2025)
by: Gangwar, Arjun, et al.
Published: (2025)
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025)
by: Omnilingual ASR team, et al.
Published: (2025)
Exploring SSL Discrete Tokens for Multilingual ASR
by: Cui, Mingyu, et al.
Published: (2024)
by: Cui, Mingyu, et al.
Published: (2024)
Configurable Multilingual ASR with Speech Summary Representations
by: Zhu, Harrison, et al.
Published: (2024)
by: Zhu, Harrison, et al.
Published: (2024)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
MultiLingPoT: Enhancing Mathematical Reasoning with Multilingual Program Fine-tuning
by: Li, Nianqi, et al.
Published: (2024)
by: Li, Nianqi, et al.
Published: (2024)
Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning
by: Zhang, Yuchen, et al.
Published: (2026)
by: Zhang, Yuchen, et al.
Published: (2026)
Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages
by: Abdullah, Badr M., et al.
Published: (2026)
by: Abdullah, Badr M., et al.
Published: (2026)
Advocating Character Error Rate for Multilingual ASR Evaluation
by: K, Thennal D, et al.
Published: (2024)
by: K, Thennal D, et al.
Published: (2024)
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
by: Gopal, Shreyas, et al.
Published: (2026)
by: Gopal, Shreyas, et al.
Published: (2026)
A Parameter-efficient Language Extension Framework for Multilingual ASR
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Efficient Adaptation of Multilingual Models for Japanese ASR
by: Bajo, Mark, et al.
Published: (2024)
by: Bajo, Mark, et al.
Published: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Jacobian Scopes: token-level causal attributions in LLMs
by: Liu, Toni J. B., et al.
Published: (2026)
by: Liu, Toni J. B., et al.
Published: (2026)
Assessing the validity of new paradigmatic complexity measures as criterial features for proficiency in L2 writings in English
by: Mallart, Cyriel, et al.
Published: (2025)
by: Mallart, Cyriel, et al.
Published: (2025)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
by: Li, Ruihang, et al.
Published: (2024)
by: Li, Ruihang, et al.
Published: (2024)
Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy
by: Plum, Alistair, et al.
Published: (2024)
by: Plum, Alistair, et al.
Published: (2024)
Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
by: Fucci, Dennis, et al.
Published: (2025)
by: Fucci, Dennis, et al.
Published: (2025)
New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR
by: Lu, Xugang, et al.
Published: (2025)
by: Lu, Xugang, et al.
Published: (2025)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
by: Huang, Shao-Syuan, et al.
Published: (2024)
by: Huang, Shao-Syuan, et al.
Published: (2024)
Scaled and Inter-token Relation Enhanced Transformer for Sample-restricted Residential NILM
by: Rahman, Minhajur, et al.
Published: (2024)
by: Rahman, Minhajur, et al.
Published: (2024)
Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation
by: Rahman, Hanif
Published: (2026)
by: Rahman, Hanif
Published: (2026)
Similar Items
-
Beyond WER: Probing Whisper's Sub-token Decoder Across Diverse Language Resource Levels
by: Liang, Siyu, et al.
Published: (2025) -
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
by: Liang, Siyu, et al.
Published: (2025) -
A Sociophonetic Analysis of Racial Bias in Commercial ASR Systems Using the Pacific Northwest English Corpus
by: Scott, Michael, et al.
Published: (2025) -
Hybrid Neural-LLM Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan
by: Liang, Siyu, et al.
Published: (2026) -
TEII: Think, Explain, Interact and Iterate with Large Language Models to Solve Cross-lingual Emotion Detection
by: Cheng, Long, et al.
Published: (2024)