AskQE: Question Answering as Automatic Evaluation for Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Ki, Dayeon, Duh, Kevin, Carpuat, Marine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
SpeechQE: Estimating the Quality of Direct Speech Translation
by: Han, HyoJung, et al.
Published: (2024)
by: Han, HyoJung, et al.
Published: (2024)
Automatic Input Rewriting Improves Translation with Large Language Models
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
Multiple LLM Agents Debate for Equitable Cultural Alignment
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations
by: Xiao, Yimin, et al.
Published: (2025)
by: Xiao, Yimin, et al.
Published: (2025)
Should We be Pedantic About Reasoning Errors in Machine Translation?
by: Bao, Calvin, et al.
Published: (2026)
by: Bao, Calvin, et al.
Published: (2026)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
How Multilingual Are Large Language Models Fine-Tuned for Translation?
by: Richburg, Aquia, et al.
Published: (2024)
by: Richburg, Aquia, et al.
Published: (2024)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
by: Agrawal, Sweta, et al.
Published: (2023)
by: Agrawal, Sweta, et al.
Published: (2023)
Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
by: Yang, Xinchen, et al.
Published: (2025)
by: Yang, Xinchen, et al.
Published: (2025)
Anti-LM Decoding for Zero-shot In-context Machine Translation
by: Sia, Suzanna, et al.
Published: (2023)
by: Sia, Suzanna, et al.
Published: (2023)
Keep It Private: Unsupervised Privatization of Online Text
by: Bao, Calvin, et al.
Published: (2024)
by: Bao, Calvin, et al.
Published: (2024)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
Exploring Representational Disparities Between Multilingual and Bilingual Translation Models
by: Verma, Neha, et al.
Published: (2023)
by: Verma, Neha, et al.
Published: (2023)
Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work
by: Bao, Calvin, et al.
Published: (2025)
by: Bao, Calvin, et al.
Published: (2025)
How often are errors in natural language reasoning due to paraphrastic variability?
by: Srikanth, Neha, et al.
Published: (2024)
by: Srikanth, Neha, et al.
Published: (2024)
QE-EBM: Using Quality Estimators as Energy Loss for Machine Translation
by: Yoo, Gahyun, et al.
Published: (2024)
by: Yoo, Gahyun, et al.
Published: (2024)
Where does In-context Translation Happen in Large Language Models
by: Sia, Suzanna, et al.
Published: (2024)
by: Sia, Suzanna, et al.
Published: (2024)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
by: Arias-Duart, Anna, et al.
Published: (2025)
by: Arias-Duart, Anna, et al.
Published: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
An Interdisciplinary Approach to Human-Centered Machine Translation
by: Carpuat, Marine, et al.
Published: (2025)
by: Carpuat, Marine, et al.
Published: (2025)
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
by: Li, Ruosen, et al.
Published: (2024)
by: Li, Ruosen, et al.
Published: (2024)
Low-Resource Named Entity Recognition with Cross-Lingual, Character-Level Neural Conditional Random Fields
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
AskSport: Web Application for Sports Question-Answering
by: Onofre, Enzo B, et al.
Published: (2025)
by: Onofre, Enzo B, et al.
Published: (2025)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
by: Ravisankar, Kartik, et al.
Published: (2025)
by: Ravisankar, Kartik, et al.
Published: (2025)
Long-Span Question-Answering: Automatic Question Generation and QA-System Ranking via Side-by-Side Evaluation
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Mitigating Semantic Leakage in Cross-lingual Embeddings via Orthogonality Constraint
by: Ki, Dayeon, et al.
Published: (2024)
by: Ki, Dayeon, et al.
Published: (2024)
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
Can QE-informed (Re)Translation lead to Error Correction?
by: Padmanabhan, Govardhan
Published: (2025)
by: Padmanabhan, Govardhan
Published: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
by: Wu, Guojun, et al.
Published: (2024)
by: Wu, Guojun, et al.
Published: (2024)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
by: Wang, Kuang-Da, et al.
Published: (2025)
by: Wang, Kuang-Da, et al.
Published: (2025)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
by: Proietti, Lorenzo, et al.
Published: (2026)
by: Proietti, Lorenzo, et al.
Published: (2026)
Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
Similar Items
-
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025) -
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
by: Ki, Dayeon, et al.
Published: (2026) -
SpeechQE: Estimating the Quality of Direct Speech Translation
by: Han, HyoJung, et al.
Published: (2024) -
Automatic Input Rewriting Improves Translation with Large Language Models
by: Ki, Dayeon, et al.
Published: (2025) -
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
by: Ki, Dayeon, et al.
Published: (2024)