Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Kocmi, Tom, Zouhar, Vilém, Avramidis, Eleftherios, Grundkiewicz, Roman, Karpinska, Marzena, Popović, Maja, Sachan, Mrinmaya, Shmatova, Mariya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026)
by: Zouhar, Vilém, et al.
Published: (2026)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Preliminary WMT24 Ranking of General MT Systems and LLMs
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)
by: Kocmi, Tom, et al.
Published: (2025)
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
by: Züfle, Maike, et al.
Published: (2025)
by: Züfle, Maike, et al.
Published: (2025)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Early-Exit and Instant Confidence Translation Quality Estimation
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Multilingual Performance Biases of Large Language Models in Education
by: Gupta, Vansh, et al.
Published: (2025)
by: Gupta, Vansh, et al.
Published: (2025)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
A Bayesian Optimization Approach to Machine Translation Reranking
by: Cheng, Julius, et al.
Published: (2024)
by: Cheng, Julius, et al.
Published: (2024)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
by: Proietti, Lorenzo, et al.
Published: (2026)
by: Proietti, Lorenzo, et al.
Published: (2026)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
Evaluating Optimal Reference Translations
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
by: Kasner, Zdeněk, et al.
Published: (2025)
by: Kasner, Zdeněk, et al.
Published: (2025)
A Formal Perspective on Byte-Pair Encoding
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
On Instruction-Finetuning Neural Machine Translation Models
by: Raunak, Vikas, et al.
Published: (2024)
by: Raunak, Vikas, et al.
Published: (2024)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
by: Lu, Qingyu, et al.
Published: (2023)
by: Lu, Qingyu, et al.
Published: (2023)
An Interdisciplinary Approach to Human-Centered Machine Translation
by: Carpuat, Marine, et al.
Published: (2025)
by: Carpuat, Marine, et al.
Published: (2025)
PyMarian: Fast Neural Machine Translation and Evaluation in Python
by: Gowda, Thamme, et al.
Published: (2024)
by: Gowda, Thamme, et al.
Published: (2024)
Probing for Arithmetic Errors in Language Models
by: Sun, Yucheng, et al.
Published: (2025)
by: Sun, Yucheng, et al.
Published: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
by: Ozyurt, Yilmazcan, et al.
Published: (2024)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
How Important is `Perfect' English for Machine Translation Prompts?
by: Schmidtová, Patrícia, et al.
Published: (2025)
by: Schmidtová, Patrícia, et al.
Published: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Multimodal Shannon Game with Images
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
by: Cui, Peng, et al.
Published: (2025)
by: Cui, Peng, et al.
Published: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
A Critical Study of Automatic Evaluation in Sign Language Translation
by: Yazdani, Shakib, et al.
Published: (2025)
by: Yazdani, Shakib, et al.
Published: (2025)
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
by: Nuutinen, Emil, et al.
Published: (2025)
by: Nuutinen, Emil, et al.
Published: (2025)
Similar Items
-
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024) -
Pearmut: Human Evaluation of Translation Made Trivial
by: Zouhar, Vilém, et al.
Published: (2026) -
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
by: Zouhar, Vilém, et al.
Published: (2025) -
Preliminary WMT24 Ranking of General MT Systems and LLMs
by: Kocmi, Tom, et al.
Published: (2024) -
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)