Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Shaomu, Mitani, Ryosuke, Choudhary, Ritvik, Wu, Qiyu, Sekiya, Toshiyuki, Monz, Christof |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating Test-Time Scaling with Reranking for Machine Translation
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Neuron Specialization: Leveraging intrinsic task modularity for multilingual machine translation
by: Tan, Shaomu, et al.
Published: (2024)
by: Tan, Shaomu, et al.
Published: (2024)
How Far Can 100 Samples Go? Unlocking Overall Zero-Shot Multilingual Translation via Tiny Multi-Parallel Data
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation
by: Wu, Di, et al.
Published: (2023)
by: Wu, Di, et al.
Published: (2023)
Disentangling the Roles of Target-Side Transfer and Regularization in Multilingual Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
How to Learn in a Noisy World? Self-Correcting the Real-World Data Noise in Machine Translation
by: Meng, Yan, et al.
Published: (2024)
by: Meng, Yan, et al.
Published: (2024)
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
Calibrating Translation Decoding with Quality Estimation on LLMs
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
The Effect of Language Diversity When Fine-Tuning Large Language Models for Translation
by: Stap, David, et al.
Published: (2025)
by: Stap, David, et al.
Published: (2025)
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Analyzing the Evaluation of Cross-Lingual Knowledge Transfer in Multilingual Language Models
by: Rajaee, Sara, et al.
Published: (2024)
by: Rajaee, Sara, et al.
Published: (2024)
Do Language Models Reason Across Languages?
by: Meng, Yan, et al.
Published: (2026)
by: Meng, Yan, et al.
Published: (2026)
Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?
by: Aycock, Seth, et al.
Published: (2024)
by: Aycock, Seth, et al.
Published: (2024)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Is It a Free Lunch for Removing Outliers during Pretraining?
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM Abilities
by: Stap, David, et al.
Published: (2024)
by: Stap, David, et al.
Published: (2024)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
by: Rajaee, Sara, et al.
Published: (2025)
by: Rajaee, Sara, et al.
Published: (2025)
3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and Composability
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
by: Liu, Hannah, et al.
Published: (2025)
by: Liu, Hannah, et al.
Published: (2025)
Word Alignment as Preference for Machine Translation
by: Wu, Qiyu, et al.
Published: (2024)
by: Wu, Qiyu, et al.
Published: (2024)
GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression
by: Miao, Zhongtao, et al.
Published: (2026)
by: Miao, Zhongtao, et al.
Published: (2026)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
by: Tan, Shaomu, et al.
Published: (2026)
by: Tan, Shaomu, et al.
Published: (2026)
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
by: Bila, Natalia, et al.
Published: (2026)
by: Bila, Natalia, et al.
Published: (2026)
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
by: Yan, Xinlan, et al.
Published: (2025)
by: Yan, Xinlan, et al.
Published: (2025)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Communicating with Speakers and Listeners of Different Pragmatic Levels
by: Naszadi, Kata, et al.
Published: (2024)
by: Naszadi, Kata, et al.
Published: (2024)
Should We be Pedantic About Reasoning Errors in Machine Translation?
by: Bao, Calvin, et al.
Published: (2026)
by: Bao, Calvin, et al.
Published: (2026)
Lost at the Beginning of Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
ApiQ: Finetuning of 2-Bit Quantized Large Language Model
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
On the Limits of Model Merging for Multilinguality in Pre-Training
by: Aycock, Seth, et al.
Published: (2026)
by: Aycock, Seth, et al.
Published: (2026)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
by: Lu, Qingyu, et al.
Published: (2024)
by: Lu, Qingyu, et al.
Published: (2024)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Evaluating Structural Generalization in Neural Machine Translation
by: Kumon, Ryoma, et al.
Published: (2024)
by: Kumon, Ryoma, et al.
Published: (2024)
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
by: Pugachev, Alexander, et al.
Published: (2025)
by: Pugachev, Alexander, et al.
Published: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Similar Items
-
Investigating Test-Time Scaling with Reranking for Machine Translation
by: Tan, Shaomu, et al.
Published: (2025) -
Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
by: Tan, Shaomu, et al.
Published: (2025) -
Neuron Specialization: Leveraging intrinsic task modularity for multilingual machine translation
by: Tan, Shaomu, et al.
Published: (2024) -
How Far Can 100 Samples Go? Unlocking Overall Zero-Shot Multilingual Translation via Tiny Multi-Parallel Data
by: Wu, Di, et al.
Published: (2024) -
Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation
by: Wu, Di, et al.
Published: (2025)