MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Qingyu, Ding, Liang, Zhang, Kanjian, Zhang, Jinxia, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
von: Lu, Qingyu, et al.
Veröffentlicht: (2026)
von: Lu, Qingyu, et al.
Veröffentlicht: (2026)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
von: Lu, Qingyu, et al.
Veröffentlicht: (2025)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
von: Riley, Parker, et al.
Veröffentlicht: (2025)
von: Riley, Parker, et al.
Veröffentlicht: (2025)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
von: Lommel, Arle, et al.
Veröffentlicht: (2024)
von: Lommel, Arle, et al.
Veröffentlicht: (2024)
Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
von: Wang, George, et al.
Veröffentlicht: (2025)
von: Wang, George, et al.
Veröffentlicht: (2025)
Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
TranslationCorrect: A Unified Framework for Machine Translation Post-Editing with Predictive Error Assistance
von: Wasti, Syed Mekael, et al.
Veröffentlicht: (2025)
von: Wasti, Syed Mekael, et al.
Veröffentlicht: (2025)
Intention Analysis Makes LLMs A Good Jailbreak Defender
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
Can Automatic Metrics Assess High-Quality Translations?
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2024)
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
von: Tan, Shaomu, et al.
Veröffentlicht: (2025)
von: Tan, Shaomu, et al.
Veröffentlicht: (2025)
Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation
von: Cai, Shizhan, et al.
Veröffentlicht: (2025)
von: Cai, Shizhan, et al.
Veröffentlicht: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
APE-Bench: Evaluating Automated Proof Engineering for Formal Math Libraries
von: Xin, Huajian, et al.
Veröffentlicht: (2025)
von: Xin, Huajian, et al.
Veröffentlicht: (2025)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
von: Messing, Solomon
Veröffentlicht: (2026)
von: Messing, Solomon
Veröffentlicht: (2026)
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing
von: Deoghare, Sourabh, et al.
Veröffentlicht: (2025)
von: Deoghare, Sourabh, et al.
Veröffentlicht: (2025)
Towards Generating Automatic Anaphora Annotations
von: Taji, Dima, et al.
Veröffentlicht: (2025)
von: Taji, Dima, et al.
Veröffentlicht: (2025)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
von: Zan, Changtong, et al.
Veröffentlicht: (2024)
LangMark: A Multilingual Dataset for Automatic Post-Editing
von: Velazquez, Diego, et al.
Veröffentlicht: (2025)
von: Velazquez, Diego, et al.
Veröffentlicht: (2025)
Evaluating LLMs at Detecting Errors in LLM Responses
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
APE: Active Learning-based Tooling for Finding Informative Few-shot Examples for LLM-based Entity Matching
von: Qian, Kun, et al.
Veröffentlicht: (2024)
von: Qian, Kun, et al.
Veröffentlicht: (2024)
Enhancing LLM-Based Data Annotation with Error Decomposition
von: Xu, Zhen, et al.
Veröffentlicht: (2026)
von: Xu, Zhen, et al.
Veröffentlicht: (2026)
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
von: Liu, Hannah, et al.
Veröffentlicht: (2025)
von: Liu, Hannah, et al.
Veröffentlicht: (2025)
Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
von: Xiao, Yisong, et al.
Veröffentlicht: (2025)
OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
von: Alabdullah, Abdullah, et al.
Veröffentlicht: (2025)
von: Alabdullah, Abdullah, et al.
Veröffentlicht: (2025)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction
von: Ye, Jingheng, et al.
Veröffentlicht: (2024)
von: Ye, Jingheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
von: Lu, Qingyu, et al.
Veröffentlicht: (2026) -
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023) -
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
von: Lu, Qingyu, et al.
Veröffentlicht: (2025) -
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
von: Riley, Parker, et al.
Veröffentlicht: (2025) -
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
von: Li, Yunmeng, et al.
Veröffentlicht: (2024)