Omnilingual MT: Machine Translation for 1,600 Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Omnilingual MT Team, Alastruey, Belen, Bafna, Niyati, Caciolai, Andrea, Heffernan, Kevin, Kozhevnikov, Artyom, Ropers, Christophe, Sánchez, Eduardo, Saint-James, Charles-Eric, Tsiamas, Ioannis, Cao, Xiang "Tony", Cheng, Chierh, Chuang, Joe, Duquenne, Paul-Ambroise, Duppenthaler, Mark, Ekberg, Nate, Gao, Cynthia, Cabot, Pere Lluís Huguet, Janeiro, João Maria, Maillard, Jean, Gonzalez, Gabriel Mejia, Schwenk, Holger, Toledo, Edan, Turkatenko, Arina, Ventayol-Boada, Albert, Moritz, Rashel, Mourachko, Alexandre, Parimi, Surya, Williamson, Mary, Yates, Shireen, Dale, David, Costa-jussà, Marta R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
by: Omnilingual SONAR Team, et al.
Published: (2026)
by: Omnilingual SONAR Team, et al.
Published: (2026)
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025)
by: Omnilingual ASR team, et al.
Published: (2025)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025)
by: The Omnilingual MT Team, et al.
Published: (2025)
Large Concept Models: Language Modeling in a Sentence Representation Space
by: LCM team, et al.
Published: (2024)
by: LCM team, et al.
Published: (2024)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
by: Chang, Emily, et al.
Published: (2025)
by: Chang, Emily, et al.
Published: (2025)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
by: Bafna, Niyati, et al.
Published: (2024)
by: Bafna, Niyati, et al.
Published: (2024)
Evaluating Large Language Models along Dimensions of Language Variation: A Systematik Invesdigatiom uv Cross-lingual Generalization
by: Bafna, Niyati, et al.
Published: (2024)
by: Bafna, Niyati, et al.
Published: (2024)
Improving Language and Modality Transfer in Translation by Character-level Modeling
by: Tsiamas, Ioannis, et al.
Published: (2025)
by: Tsiamas, Ioannis, et al.
Published: (2025)
Unified Vision-Language Modeling via Concept Space Alignment
by: Qiu, Yifu, et al.
Published: (2026)
by: Qiu, Yifu, et al.
Published: (2026)
Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders
by: Alastruey, Belen, et al.
Published: (2025)
by: Alastruey, Belen, et al.
Published: (2025)
TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection
by: Sander, Tom, et al.
Published: (2026)
by: Sander, Tom, et al.
Published: (2026)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Unveiling the Role of Pretraining in Direct Speech Translation
by: Alastruey, Belen, et al.
Published: (2024)
by: Alastruey, Belen, et al.
Published: (2024)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Pushing the Limits of Zero-shot End-to-End Speech Translation
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
by: Alastruey, Belen, et al.
Published: (2023)
by: Alastruey, Belen, et al.
Published: (2023)
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
by: Bafna, Niyati, et al.
Published: (2023)
by: Bafna, Niyati, et al.
Published: (2023)
Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning
by: Bafna, Niyati, et al.
Published: (2026)
by: Bafna, Niyati, et al.
Published: (2026)
Spirit LM: Interleaved Spoken and Written Language Model
by: Nguyen, Tu Anh, et al.
Published: (2024)
by: Nguyen, Tu Anh, et al.
Published: (2024)
Research Synthesis 1: Action, Continuity and Value
by: Bafna, Ekta
Published: (2026)
by: Bafna, Ekta
Published: (2026)
How Important is `Perfect' English for Machine Translation Prompts?
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
Measuring the precise photometric period of the probable intermediate polar 1RXS J014549.6+514314 based on extensive photometry
by: Kozhevnikov, V. P.
Published: (2025)
by: Kozhevnikov, V. P.
Published: (2025)
Discovery of eclipses in the cataclysmic variable LAMOST J035913.61+405035.0
by: Kozhevnikov, V. P.
Published: (2024)
by: Kozhevnikov, V. P.
Published: (2024)
Physical nature of 'anomalous' electrons in high-current vacuum diodes
by: Vasily Y. Kozhevnikov
Published: (2021)
by: Vasily Y. Kozhevnikov
Published: (2021)
Kinetic simulation of vacuum plasma expansion beyond the "plasma approximation"
by: Vasily Y. Kozhevnikov
Published: (2022)
by: Vasily Y. Kozhevnikov
Published: (2022)
Detection of Eclipses in the Cataclysmic Variable LAMOST J035913.61 + 405035.0
by: V. P. Kozhevnikov
Published: (2025)
by: V. P. Kozhevnikov
Published: (2025)
Measuring the Precise Photometric Period of the Probable Intermediate Polar 1RXS J014549.6+514314 Based on Extensive Photometry
by: V. P. Kozhevnikov
Published: (2025)
by: V. P. Kozhevnikov
Published: (2025)
SUBMICROSECOND ATMOSPHERIC ELECTRIC DISCHARGE FROM THE NON-UNIFORM ELECTRODE (TIP) TOWARDS THE PLANE ELECTRODE
by: Vasily Y. Kozhevnikov
Published: (2019)
by: Vasily Y. Kozhevnikov
Published: (2019)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
by: Zhu, Han, et al.
Published: (2026)
by: Zhu, Han, et al.
Published: (2026)
Towards Massive Multilingual Holistic Bias
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
CCCE: A Continuous Code Calibration Engine for Autonomous Enterprise Codebase Maintenance via Knowledge Graph Traversal and Adaptive Decision Gating
by: Parimi, Santhosh Kusuma Kumar
Published: (2026)
by: Parimi, Santhosh Kusuma Kumar
Published: (2026)
Ultrahigh Poisson's ratio glasses
by: Lerner, Edan
Published: (2022)
by: Lerner, Edan
Published: (2022)
Effects of coordination and stiffness scale-separation in disordered elastic networks
by: Lerner, Edan
Published: (2023)
by: Lerner, Edan
Published: (2023)
Computation of Graph Polynomials via Tree Decomposition: Theory, Algorithms, and Python Implementation
by: Bafna, Mehul, et al.
Published: (2025)
by: Bafna, Mehul, et al.
Published: (2025)
Similar Items
-
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
by: Omnilingual SONAR Team, et al.
Published: (2026) -
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
by: Omnilingual ASR team, et al.
Published: (2025) -
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025) -
Large Concept Models: Language Modeling in a Sentence Representation Space
by: LCM team, et al.
Published: (2024) -
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
by: Bafna, Niyati, et al.
Published: (2025)