AskQE: Question Answering as Automatic Evaluation for Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ki, Dayeon, Duh, Kevin, Carpuat, Marine
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909762034073600
author Ki, Dayeon
Duh, Kevin
Carpuat, Marine
author_facet Ki, Dayeon
Duh, Kevin
Carpuat, Marine
contents How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection and quality estimation (QE) techniques do not address this practical scenario. We introduce AskQE, a question generation and answering framework designed to detect critical MT errors and provide actionable feedback, helping users decide whether to accept or reject MT outputs even without the knowledge of the target language. Using ContraTICO, a dataset of contrastive synthetic MT errors in the COVID-19 domain, we explore design choices for AskQE and develop an optimized version relying on LLaMA-3 70B and entailed facts to guide question generation. We evaluate the resulting system on the BioMQM dataset of naturally occurring MT errors, where AskQE has higher Kendall's Tau correlation and decision accuracy with human ratings compared to other QE metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AskQE: Question Answering as Automatic Evaluation for Machine Translation
Ki, Dayeon
Duh, Kevin
Carpuat, Marine
Computation and Language
How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection and quality estimation (QE) techniques do not address this practical scenario. We introduce AskQE, a question generation and answering framework designed to detect critical MT errors and provide actionable feedback, helping users decide whether to accept or reject MT outputs even without the knowledge of the target language. Using ContraTICO, a dataset of contrastive synthetic MT errors in the COVID-19 domain, we explore design choices for AskQE and develop an optimized version relying on LLaMA-3 70B and entailed facts to guide question generation. We evaluate the resulting system on the BioMQM dataset of naturally occurring MT errors, where AskQE has higher Kendall's Tau correlation and decision accuracy with human ratings compared to other QE metrics.
title AskQE: Question Answering as Automatic Evaluation for Machine Translation
topic Computation and Language
url https://arxiv.org/abs/2504.11582