MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | Du, Yexing, Liu, Kaiyuan, Pan, Youcheng, Yang, Bo, Deng, Keqi, Chen, Xie, Xiang, Yang, Liu, Ming, Qin, Bing, Wang, YaoWei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
di: Du, Yexing, et al.
Pubblicazione: (2024)
di: Du, Yexing, et al.
Pubblicazione: (2024)
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
di: Du, Yexing, et al.
Pubblicazione: (2026)
di: Du, Yexing, et al.
Pubblicazione: (2026)
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
di: Du, Yexing, et al.
Pubblicazione: (2026)
di: Du, Yexing, et al.
Pubblicazione: (2026)
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
di: Du, Yexing, et al.
Pubblicazione: (2025)
di: Du, Yexing, et al.
Pubblicazione: (2025)
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
di: Du, Yexing, et al.
Pubblicazione: (2026)
di: Du, Yexing, et al.
Pubblicazione: (2026)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
di: Kim, Minsu, et al.
Pubblicazione: (2023)
di: Kim, Minsu, et al.
Pubblicazione: (2023)
Forget Many, Forget Right: Scalable and Precise Concept Unlearning in Diffusion Models
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
di: Deng, Kaiyuan, et al.
Pubblicazione: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025)
di: Deng, Keqi, et al.
Pubblicazione: (2025)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
EnAnchored-X2X: English-Anchored Optimization for Many-to-Many Translation
di: Yang, Sen, et al.
Pubblicazione: (2025)
di: Yang, Sen, et al.
Pubblicazione: (2025)
Toucan: Many-to-Many Translation for 150 African Language Pairs
di: Elmadany, AbdelRahim, et al.
Pubblicazione: (2024)
di: Elmadany, AbdelRahim, et al.
Pubblicazione: (2024)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
di: Jiang, Shixin, et al.
Pubblicazione: (2024)
di: Jiang, Shixin, et al.
Pubblicazione: (2024)
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
di: An, Yanjie, et al.
Pubblicazione: (2026)
di: An, Yanjie, et al.
Pubblicazione: (2026)
Position: Towards Responsible Evaluation for Text-to-Speech
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
di: Gao, Pengzhi, et al.
Pubblicazione: (2024)
di: Gao, Pengzhi, et al.
Pubblicazione: (2024)
Many-body delocalization with a two-dimensional 70-qubit superconducting quantum simulator
di: Li, Tian-Ming, et al.
Pubblicazione: (2025)
di: Li, Tian-Ming, et al.
Pubblicazione: (2025)
Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation
di: Salim, Luis Frentzen, et al.
Pubblicazione: (2026)
di: Salim, Luis Frentzen, et al.
Pubblicazione: (2026)
TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm
di: Zhang, Bingqing, et al.
Pubblicazione: (2024)
di: Zhang, Bingqing, et al.
Pubblicazione: (2024)
Many-body computing on Field Programmable Gate Arrays
di: Lv, Songtai, et al.
Pubblicazione: (2024)
di: Lv, Songtai, et al.
Pubblicazione: (2024)
A Numerical Perspective on Moiré Superlattices: From Single-Particle Properties to Many-Body Physics
di: Lu, Xin, et al.
Pubblicazione: (2025)
di: Lu, Xin, et al.
Pubblicazione: (2025)
Superconducting Quantum Simulation for Many-Body Physics beyond Equilibrium
di: Yao, Yunyan, et al.
Pubblicazione: (2024)
di: Yao, Yunyan, et al.
Pubblicazione: (2024)
Few for Many: Tchebycheff Set Scalarization for Many-Objective Optimization
di: Lin, Xi, et al.
Pubblicazione: (2024)
di: Lin, Xi, et al.
Pubblicazione: (2024)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
di: Le-Duc, Khai, et al.
Pubblicazione: (2025)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
di: Wang, Peidong, et al.
Pubblicazione: (2024)
di: Wang, Peidong, et al.
Pubblicazione: (2024)
A Unifying View of OTFS and Its Many Variants
di: Deng, Qinwen, et al.
Pubblicazione: (2025)
di: Deng, Qinwen, et al.
Pubblicazione: (2025)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2023)
di: Deng, Keqi, et al.
Pubblicazione: (2023)
Speech LLMs are Contextual Reasoning Transcribers
di: Deng, Keqi, et al.
Pubblicazione: (2026)
di: Deng, Keqi, et al.
Pubblicazione: (2026)
Invariance under Structure Translation as the Origin of Host Immune Capacity Conservation from Noether's Theorem
di: Chen, Yexing, et al.
Pubblicazione: (2025)
di: Chen, Yexing, et al.
Pubblicazione: (2025)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
di: Li, Zhaoqing, et al.
Pubblicazione: (2025)
Many-Turn Jailbreaking
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
di: Yang, Xianjun, et al.
Pubblicazione: (2025)
Many-to-Many Matching via Sparsity Controlled Optimal Transport
di: Liu, Weijie, et al.
Pubblicazione: (2025)
di: Liu, Weijie, et al.
Pubblicazione: (2025)
Many Worlds, Many Theories?
di: Cristina Inoue
Pubblicazione: (2016)
di: Cristina Inoue
Pubblicazione: (2016)
Partial Conditioning for Inference of Many-Normal-Means with Hölder Constraints
di: Yang, Jiasen, et al.
Pubblicazione: (2023)
di: Yang, Jiasen, et al.
Pubblicazione: (2023)
General Many-Body Perturbation Framework for Moiré Systems
di: Lu, Xin, et al.
Pubblicazione: (2025)
di: Lu, Xin, et al.
Pubblicazione: (2025)
Work Statistics and Adiabatic Assumption in Nonequilibrium Many-Body Theory
di: Zuo, Yi, et al.
Pubblicazione: (2023)
di: Zuo, Yi, et al.
Pubblicazione: (2023)
ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics
di: Yu, Qiaojun, et al.
Pubblicazione: (2024)
di: Yu, Qiaojun, et al.
Pubblicazione: (2024)
Translation in the Hands of Many:Centering Lay Users in Machine Translation Interactions
di: Savoldi, Beatrice, et al.
Pubblicazione: (2025)
di: Savoldi, Beatrice, et al.
Pubblicazione: (2025)
Many-body interferometry of one-dimensional integrable systems
di: Arzamasovs, Maksims, et al.
Pubblicazione: (2024)
di: Arzamasovs, Maksims, et al.
Pubblicazione: (2024)
Deploy DINO with Many-to-Many Association
di: Jiang, Haodong, et al.
Pubblicazione: (2026)
di: Jiang, Haodong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
di: Du, Yexing, et al.
Pubblicazione: (2024) -
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
di: Du, Yexing, et al.
Pubblicazione: (2026) -
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
di: Du, Yexing, et al.
Pubblicazione: (2026) -
CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation
di: Du, Yexing, et al.
Pubblicazione: (2025) -
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
di: Du, Yexing, et al.
Pubblicazione: (2026)