Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Yukhymenko, Hanna, Alexandrov, Anton, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024)
by: Yukhymenko, Hanna, et al.
Published: (2024)
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024)
by: Alexandrov, Anton, et al.
Published: (2024)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
On Translating Technical Terminology: A Translation Workflow for Machine-Translated Acronyms
by: Yue, Richard, et al.
Published: (2024)
by: Yue, Richard, et al.
Published: (2024)
Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
by: Zuo, Max, et al.
Published: (2024)
by: Zuo, Max, et al.
Published: (2024)
Automated Multi-Language to English Machine Translation Using Generative Pre-Trained Transformers
by: Pelofske, Elijah, et al.
Published: (2024)
by: Pelofske, Elijah, et al.
Published: (2024)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025)
by: Farinhas, António, et al.
Published: (2025)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
by: Mündler, Niels, et al.
Published: (2023)
by: Mündler, Niels, et al.
Published: (2023)
Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering
by: Varzaneh, Mostafa, et al.
Published: (2024)
by: Varzaneh, Mostafa, et al.
Published: (2024)
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods
by: Yue, Richard, et al.
Published: (2024)
by: Yue, Richard, et al.
Published: (2024)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
by: Liu, Zuxin, et al.
Published: (2024)
by: Liu, Zuxin, et al.
Published: (2024)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
by: Wang, Wenxuan, et al.
Published: (2022)
by: Wang, Wenxuan, et al.
Published: (2022)
Better Alignment with Instruction Back-and-Forth Translation
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Adapting Language Models via Token Translation
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
Sequence Shortening for Context-Aware Machine Translation
by: Mąka, Paweł, et al.
Published: (2024)
by: Mąka, Paweł, et al.
Published: (2024)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
by: Jung, Kyudan, et al.
Published: (2024)
by: Jung, Kyudan, et al.
Published: (2024)
Self-Speculative Biased Decoding for Faster Re-Translation
by: Zeng, Linxiao, et al.
Published: (2025)
by: Zeng, Linxiao, et al.
Published: (2025)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
Accelerating Transformer Inference for Translation via Parallel Decoding
by: Santilli, Andrea, et al.
Published: (2023)
by: Santilli, Andrea, et al.
Published: (2023)
Steering Large Language Models for Machine Translation Personalization
by: Scalena, Daniel, et al.
Published: (2025)
by: Scalena, Daniel, et al.
Published: (2025)
Aligning Pre-trained Models for Spoken Language Translation
by: Sedláček, Šimon, et al.
Published: (2024)
by: Sedláček, Šimon, et al.
Published: (2024)
Structured Document Translation via Format Reinforcement Learning
by: Song, Haiyue, et al.
Published: (2025)
by: Song, Haiyue, et al.
Published: (2025)
Edeflip: Supervised Word Translation between English and Yoruba
by: Abioye, Ikeoluwa, et al.
Published: (2025)
by: Abioye, Ikeoluwa, et al.
Published: (2025)
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020)
by: Li, Yaoyiran, et al.
Published: (2020)
Self-generated Replay Memories for Continual Neural Machine Translation
by: Resta, Michele, et al.
Published: (2024)
by: Resta, Michele, et al.
Published: (2024)
Integrating Pre-trained Language Model into Neural Machine Translation
by: Hwang, Soon-Jae, et al.
Published: (2023)
by: Hwang, Soon-Jae, et al.
Published: (2023)
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
by: Niklaus, Joel, et al.
Published: (2025)
by: Niklaus, Joel, et al.
Published: (2025)
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
Domain-Specific Quality Estimation for Machine Translation in Low-Resource Scenarios
by: Gurav, Namrata Patil, et al.
Published: (2026)
by: Gurav, Namrata Patil, et al.
Published: (2026)
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
by: Dragomir, Alexandra, et al.
Published: (2026)
by: Dragomir, Alexandra, et al.
Published: (2026)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
by: Yang, Wenjie, et al.
Published: (2025)
by: Yang, Wenjie, et al.
Published: (2025)
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
by: Pengpun, Parinthapat, et al.
Published: (2024)
by: Pengpun, Parinthapat, et al.
Published: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
by: Park, Jin Hyun, et al.
Published: (2025)
by: Park, Jin Hyun, et al.
Published: (2025)
Disentangling the Roles of Target-Side Transfer and Regularization in Multilingual Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
by: Vincent, Sebastian, et al.
Published: (2023)
by: Vincent, Sebastian, et al.
Published: (2023)
Shortcomings of LLMs for Low-Resource Translation: Retrieval and Understanding are Both the Problem
by: Court, Sara, et al.
Published: (2024)
by: Court, Sara, et al.
Published: (2024)
Multiple Noises in Diffusion Model for Semi-Supervised Multi-Domain Translation
by: Mayet, Tsiry, et al.
Published: (2023)
by: Mayet, Tsiry, et al.
Published: (2023)
Investigating Multi-Pivot Ensembling with Massively Multilingual Machine Translation Models
by: Mohammadshahi, Alireza, et al.
Published: (2023)
by: Mohammadshahi, Alireza, et al.
Published: (2023)
Similar Items
-
A Synthetic Dataset for Personal Attribute Inference
by: Yukhymenko, Hanna, et al.
Published: (2024) -
BgGPT 1.0: Extending English-centric LLMs to other languages
by: Alexandrov, Anton, et al.
Published: (2024) -
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025) -
On Translating Technical Terminology: A Translation Workflow for Machine-Translated Acronyms
by: Yue, Richard, et al.
Published: (2024) -
Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
by: Zuo, Max, et al.
Published: (2024)