Many-to-English Machine Translation Tools, Data, and Pretrained Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gowda, Thamme, Zhang, Zhao, Mattmann, Chris A, May, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020)
by: Li, Yaoyiran, et al.
Published: (2020)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?
by: Chiaranaipanich, Jirat, et al.
Published: (2024)
by: Chiaranaipanich, Jirat, et al.
Published: (2024)
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
by: Pengpun, Parinthapat, et al.
Published: (2024)
by: Pengpun, Parinthapat, et al.
Published: (2024)
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
by: Bordt, Sebastian, et al.
Published: (2025)
by: Bordt, Sebastian, et al.
Published: (2025)
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
by: Arif, Arwa
Published: (2025)
by: Arif, Arwa
Published: (2025)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Rare but Severe Neural Machine Translation Errors Induced by Minimal Deletion: An Empirical Study on Chinese and English
by: Shi, Ruikang, et al.
Published: (2022)
by: Shi, Ruikang, et al.
Published: (2022)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
by: Zhang, Hengxiang, et al.
Published: (2024)
by: Zhang, Hengxiang, et al.
Published: (2024)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
by: Rashidi, Sina, et al.
Published: (2025)
by: Rashidi, Sina, et al.
Published: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
by: Nadas, Mihai, et al.
Published: (2025)
by: Nadas, Mihai, et al.
Published: (2025)
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
by: Chen, Zhixun, et al.
Published: (2025)
by: Chen, Zhixun, et al.
Published: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
Steering Large Language Models for Machine Translation Personalization
by: Scalena, Daniel, et al.
Published: (2025)
by: Scalena, Daniel, et al.
Published: (2025)
On Translating Technical Terminology: A Translation Workflow for Machine-Translated Acronyms
by: Yue, Richard, et al.
Published: (2024)
by: Yue, Richard, et al.
Published: (2024)
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2025)
by: Mąka, Paweł, et al.
Published: (2025)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
Edeflip: Supervised Word Translation between English and Yoruba
by: Abioye, Ikeoluwa, et al.
Published: (2025)
by: Abioye, Ikeoluwa, et al.
Published: (2025)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
by: Vincent, Sebastian, et al.
Published: (2023)
by: Vincent, Sebastian, et al.
Published: (2023)
A Case Study on Contextual Machine Translation in a Professional Scenario of Subtitling
by: Vincent, Sebastian, et al.
Published: (2024)
by: Vincent, Sebastian, et al.
Published: (2024)
Integrating Pre-trained Language Model into Neural Machine Translation
by: Hwang, Soon-Jae, et al.
Published: (2023)
by: Hwang, Soon-Jae, et al.
Published: (2023)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
by: Ge, Ce, et al.
Published: (2024)
by: Ge, Ce, et al.
Published: (2024)
Investigating Multi-Pivot Ensembling with Massively Multilingual Machine Translation Models
by: Mohammadshahi, Alireza, et al.
Published: (2023)
by: Mohammadshahi, Alireza, et al.
Published: (2023)
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
Does Differential Privacy Impact Bias in Pretrained NLP Models?
by: Islam, Md. Khairul, et al.
Published: (2024)
by: Islam, Md. Khairul, et al.
Published: (2024)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
by: Zhang, Kangning, et al.
Published: (2025)
by: Zhang, Kangning, et al.
Published: (2025)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
by: Tong, Ziyi, et al.
Published: (2026)
by: Tong, Ziyi, et al.
Published: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
Sequence Shortening for Context-Aware Machine Translation
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
by: Yang, Yazheng, et al.
Published: (2023)
by: Yang, Yazheng, et al.
Published: (2023)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Pretraining Large Language Models with NVFP4
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
In-context Pretraining: Language Modeling Beyond Document Boundaries
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
by: Li, Yixiao, et al.
Published: (2025)
by: Li, Yixiao, et al.
Published: (2025)
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods
by: Yue, Richard, et al.
Published: (2024)
by: Yue, Richard, et al.
Published: (2024)
Similar Items
-
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020) -
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024) -
Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?
by: Chiaranaipanich, Jirat, et al.
Published: (2024) -
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
by: Pengpun, Parinthapat, et al.
Published: (2024) -
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)