Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Yexing, Pan, Youcheng, Ma, Ziyang, Yang, Bo, Yang, Yifan, Deng, Keqi, Chen, Xie, Xiang, Yang, Liu, Ming, Qin, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
by: Du, Yexing, et al.
Published: (2026)
by: Du, Yexing, et al.
Published: (2026)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
by: Deng, Keqi, et al.
Published: (2024)
by: Deng, Keqi, et al.
Published: (2024)
Speech LLMs are Contextual Reasoning Transcribers
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
by: Yang, Sen, et al.
Published: (2025)
by: Yang, Sen, et al.
Published: (2025)
Forget Many, Forget Right: Scalable and Precise Concept Unlearning in Diffusion Models
by: Deng, Kaiyuan, et al.
Published: (2026)
by: Deng, Kaiyuan, et al.
Published: (2026)
Position: Towards Responsible Evaluation for Text-to-Speech
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Toucan: Many-to-Many Translation for 150 African Language Pairs
by: Elmadany, AbdelRahim, et al.
Published: (2024)
by: Elmadany, AbdelRahim, et al.
Published: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
by: Pareras, Oriol, et al.
Published: (2025)
by: Pareras, Oriol, et al.
Published: (2025)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
by: Yan, Jianzhi, et al.
Published: (2025)
by: Yan, Jianzhi, et al.
Published: (2025)
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
by: An, Yanjie, et al.
Published: (2026)
by: An, Yanjie, et al.
Published: (2026)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
by: Yang, Guanrou, et al.
Published: (2025)
by: Yang, Guanrou, et al.
Published: (2025)
Modeling the One-to-Many Property in Open-Domain Dialogue with LLMs
by: Lee, Jing Yang, et al.
Published: (2025)
by: Lee, Jing Yang, et al.
Published: (2025)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
by: Guo, Yiwei, et al.
Published: (2023)
by: Guo, Yiwei, et al.
Published: (2023)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
by: Chen, Yushen, et al.
Published: (2024)
by: Chen, Yushen, et al.
Published: (2024)
Many-Turn Jailbreaking
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
by: Gao, Pengzhi, et al.
Published: (2024)
by: Gao, Pengzhi, et al.
Published: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Primer‐Disk‐Enabled DNA Data Storage System with Index and Record‐Many‐Read‐Many Features
by: Jiaxiang Ma, et al.
Published: (2025)
by: Jiaxiang Ma, et al.
Published: (2025)
Chain of Strategy Optimization Makes Large Language Models Better Emotional Supporter
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Self-Correction Makes LLMs Better Parsers
by: Zhang, Ziyan, et al.
Published: (2025)
by: Zhang, Ziyan, et al.
Published: (2025)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Winning in the Limit: Average-Case Committee Selection with Many Candidates
by: Lin, Yifan, et al.
Published: (2026)
by: Lin, Yifan, et al.
Published: (2026)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
by: Le-Duc, Khai, et al.
Published: (2025)
by: Le-Duc, Khai, et al.
Published: (2025)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
by: Wang, Peidong, et al.
Published: (2024)
by: Wang, Peidong, et al.
Published: (2024)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
by: Deng, Keqi, et al.
Published: (2023)
by: Deng, Keqi, et al.
Published: (2023)
Many-body multipole indices revealed by the real-space dynamical mean-field theory
by: Yang, Guoao, et al.
Published: (2024)
by: Yang, Guoao, et al.
Published: (2024)
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
by: Wei, Wang, et al.
Published: (2025)
by: Wei, Wang, et al.
Published: (2025)
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
by: Perera, David, et al.
Published: (2024)
by: Perera, David, et al.
Published: (2024)
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
by: Li, Ruibin, et al.
Published: (2025)
by: Li, Ruibin, et al.
Published: (2025)
Invariance under Structure Translation as the Origin of Host Immune Capacity Conservation from Noether's Theorem
by: Chen, Yexing, et al.
Published: (2025)
by: Chen, Yexing, et al.
Published: (2025)
Similar Items
-
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
by: Du, Yexing, et al.
Published: (2025) -
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
by: Du, Yexing, et al.
Published: (2026) -
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
by: Du, Yexing, et al.
Published: (2026) -
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
by: Du, Yexing, et al.
Published: (2025) -
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)