Many-to-English Machine Translation Tools, Data, and Pretrained Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gowda, Thamme, Zhang, Zhao, Mattmann, Chris A, May, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Emergent Communication Pretraining for Few-Shot Machine Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
A Little Human Data Goes A Long Way
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024)
Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?
von: Chiaranaipanich, Jirat, et al.
Veröffentlicht: (2024)
von: Chiaranaipanich, Jirat, et al.
Veröffentlicht: (2024)
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
von: Pengpun, Parinthapat, et al.
Veröffentlicht: (2024)
von: Pengpun, Parinthapat, et al.
Veröffentlicht: (2024)
Language Models Can Predict Their Own Behavior
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
von: Bordt, Sebastian, et al.
Veröffentlicht: (2025)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2025)
The Saturation Point of Backtranslation in High Quality Low Resource English Gujarati Machine Translation
von: Arif, Arwa
Veröffentlicht: (2025)
von: Arif, Arwa
Veröffentlicht: (2025)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
von: Pelofske, Elijah, et al.
Veröffentlicht: (2024)
von: Pelofske, Elijah, et al.
Veröffentlicht: (2024)
Rare but Severe Neural Machine Translation Errors Induced by Minimal Deletion: An Empirical Study on Chinese and English
von: Shi, Ruikang, et al.
Veröffentlicht: (2022)
von: Shi, Ruikang, et al.
Veröffentlicht: (2022)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
von: Foroutan, Negar, et al.
Veröffentlicht: (2025)
Fine-tuning can Help Detect Pretraining Data from Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
von: Rashidi, Sina, et al.
Veröffentlicht: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
von: Chen, Zhixun, et al.
Veröffentlicht: (2025)
von: Chen, Zhixun, et al.
Veröffentlicht: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
von: Wang, Xinyi, et al.
Veröffentlicht: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
von: Tran, Hieu, et al.
Veröffentlicht: (2025)
von: Tran, Hieu, et al.
Veröffentlicht: (2025)
Steering Large Language Models for Machine Translation Personalization
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
von: Scalena, Daniel, et al.
Veröffentlicht: (2025)
On Translating Technical Terminology: A Translation Workflow for Machine-Translated Acronyms
von: Yue, Richard, et al.
Veröffentlicht: (2024)
von: Yue, Richard, et al.
Veröffentlicht: (2024)
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
von: Mąka, Paweł, et al.
Veröffentlicht: (2025)
von: Mąka, Paweł, et al.
Veröffentlicht: (2025)
Translating Hanja Historical Documents to Contemporary Korean and English
von: Son, Juhee, et al.
Veröffentlicht: (2022)
von: Son, Juhee, et al.
Veröffentlicht: (2022)
Edeflip: Supervised Word Translation between English and Yoruba
von: Abioye, Ikeoluwa, et al.
Veröffentlicht: (2025)
von: Abioye, Ikeoluwa, et al.
Veröffentlicht: (2025)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
von: Vincent, Sebastian, et al.
Veröffentlicht: (2023)
von: Vincent, Sebastian, et al.
Veröffentlicht: (2023)
A Case Study on Contextual Machine Translation in a Professional Scenario of Subtitling
von: Vincent, Sebastian, et al.
Veröffentlicht: (2024)
von: Vincent, Sebastian, et al.
Veröffentlicht: (2024)
Integrating Pre-trained Language Model into Neural Machine Translation
von: Hwang, Soon-Jae, et al.
Veröffentlicht: (2023)
von: Hwang, Soon-Jae, et al.
Veröffentlicht: (2023)
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
von: Ge, Ce, et al.
Veröffentlicht: (2024)
von: Ge, Ce, et al.
Veröffentlicht: (2024)
Investigating Multi-Pivot Ensembling with Massively Multilingual Machine Translation Models
von: Mohammadshahi, Alireza, et al.
Veröffentlicht: (2023)
von: Mohammadshahi, Alireza, et al.
Veröffentlicht: (2023)
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2024)
Does Differential Privacy Impact Bias in Pretrained NLP Models?
von: Islam, Md. Khairul, et al.
Veröffentlicht: (2024)
von: Islam, Md. Khairul, et al.
Veröffentlicht: (2024)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
von: Yu, Zichun, et al.
Veröffentlicht: (2026)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
Sequence Shortening for Context-Aware Machine Translation
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
von: Yang, Yazheng, et al.
Veröffentlicht: (2023)
von: Yang, Yazheng, et al.
Veröffentlicht: (2023)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
Pretraining Large Language Models with NVFP4
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
In-context Pretraining: Language Modeling Beyond Document Boundaries
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
von: Li, Yixiao, et al.
Veröffentlicht: (2025)
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods
von: Yue, Richard, et al.
Veröffentlicht: (2024)
von: Yue, Richard, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Emergent Communication Pretraining for Few-Shot Machine Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020) -
A Little Human Data Goes A Long Way
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2024) -
Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?
von: Chiaranaipanich, Jirat, et al.
Veröffentlicht: (2024) -
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
von: Pengpun, Parinthapat, et al.
Veröffentlicht: (2024) -
Language Models Can Predict Their Own Behavior
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)