Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Diarra, Yacouba, Leventhal, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
Dealing with the Hard Facts of Low-Resource African NLP
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025)
Where Are We At with Automatic Speech Recognition for the Bambara Language?
von: Diallo, Seydou, et al.
Veröffentlicht: (2026)
von: Diallo, Seydou, et al.
Veröffentlicht: (2026)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
Deep CLAS: Deep Contextual Listen, Attend and Spell
von: Wang, Mengzhi, et al.
Veröffentlicht: (2024)
von: Wang, Mengzhi, et al.
Veröffentlicht: (2024)
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
von: Li, Zinuo, et al.
Veröffentlicht: (2025)
von: Li, Zinuo, et al.
Veröffentlicht: (2025)
The Serendipity of Claude AI: Case of the 13 Low-Resource National Languages of Mali
von: Dembele, Alou, et al.
Veröffentlicht: (2025)
von: Dembele, Alou, et al.
Veröffentlicht: (2025)
Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task
von: Jung, JungHo, et al.
Veröffentlicht: (2025)
von: Jung, JungHo, et al.
Veröffentlicht: (2025)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Gradient-Informed Training for Low-Resource Multilingual Speech Translation
von: Sun, Ruiyan, et al.
Veröffentlicht: (2026)
von: Sun, Ruiyan, et al.
Veröffentlicht: (2026)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Prediction of Translation Techniques for the Translation Process
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
R2T: Rule-Encoded Loss Functions for Low-Resource Sequence Tagging
von: Keita, Mamadou K., et al.
Veröffentlicht: (2025)
von: Keita, Mamadou K., et al.
Veröffentlicht: (2025)
Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali
von: Coulibaly, Nouhoum, et al.
Veröffentlicht: (2026)
von: Coulibaly, Nouhoum, et al.
Veröffentlicht: (2026)
RVPO: Risk-Sensitive Alignment via Variance Regularization
von: Montero, Ivan, et al.
Veröffentlicht: (2026)
von: Montero, Ivan, et al.
Veröffentlicht: (2026)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
Can Speech LLMs Think while Listening?
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
SpeechQE: Estimating the Quality of Direct Speech Translation
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
SpeechT: Findings of the First Mentorship in Speech Translation
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
von: Moslem, Yasmin, et al.
Veröffentlicht: (2025)
Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension
von: Ai, Lin, et al.
Veröffentlicht: (2024)
von: Ai, Lin, et al.
Veröffentlicht: (2024)
Listen and Chant Before You Read: The Ladder of Beauty in LM Pre-Training
von: Nomura, Yoshinori
Veröffentlicht: (2026)
von: Nomura, Yoshinori
Veröffentlicht: (2026)
REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2025)
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2025)
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
von: Popescu-Belis, Andrei, et al.
Veröffentlicht: (2025)
von: Popescu-Belis, Andrei, et al.
Veröffentlicht: (2025)
Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text
von: Issam, Abderrahmane, et al.
Veröffentlicht: (2026)
von: Issam, Abderrahmane, et al.
Veröffentlicht: (2026)
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024)
von: Hu, Ke, et al.
Veröffentlicht: (2024)
High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers
von: Wong, Brian, et al.
Veröffentlicht: (2025)
von: Wong, Brian, et al.
Veröffentlicht: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
Different Speech Translation Models Encode and Translate Speaker Gender Differently
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
von: Fucci, Dennis, et al.
Veröffentlicht: (2025)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
R-BI: Regularized Batched Inputs enhance Incremental Decoding Framework for Low-Latency Simultaneous Speech Translation
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2024)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
von: Li, Zhaolin, et al.
Veröffentlicht: (2025)
von: Li, Zhaolin, et al.
Veröffentlicht: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Learning When Not to Attend Globally
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
von: Luu, Nam, et al.
Veröffentlicht: (2025)
von: Luu, Nam, et al.
Veröffentlicht: (2025)
FFSTC: Fongbe to French Speech Translation Corpus
von: Kponou, D. Fortune, et al.
Veröffentlicht: (2024)
von: Kponou, D. Fortune, et al.
Veröffentlicht: (2024)
Cross-Lingual Transfer Learning for Speech Translation
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Unveiling the Role of Pretraining in Direct Speech Translation
von: Alastruey, Belen, et al.
Veröffentlicht: (2024)
von: Alastruey, Belen, et al.
Veröffentlicht: (2024)
Contrastive Feedback Mechanism for Simultaneous Speech Translation
von: Tan, Haotian, et al.
Veröffentlicht: (2024)
von: Tan, Haotian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025) -
Kunnafonidilaw ka Cadeau: an ASR dataset of present-day Bambara
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025) -
Dealing with the Hard Facts of Low-Resource African NLP
von: Diarra, Yacouba, et al.
Veröffentlicht: (2025) -
Where Are We At with Automatic Speech Recognition for the Bambara Language?
von: Diallo, Seydou, et al.
Veröffentlicht: (2026) -
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)