PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
Fuente:
arXiv
Guardado en:
| Autores principales: | Proietti, Lorenzo, Grundkiewicz, Roman, Post, Matt |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PyMarian: Fast Neural Machine Translation and Evaluation in Python
por: Gowda, Thamme, et al.
Publicado: (2024)
por: Gowda, Thamme, et al.
Publicado: (2024)
On Instruction-Finetuning Neural Machine Translation Models
por: Raunak, Vikas, et al.
Publicado: (2024)
por: Raunak, Vikas, et al.
Publicado: (2024)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
por: Raunak, Vikas, et al.
Publicado: (2023)
por: Raunak, Vikas, et al.
Publicado: (2023)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
por: Proietti, Lorenzo, et al.
Publicado: (2025)
por: Proietti, Lorenzo, et al.
Publicado: (2025)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
por: Kocmi, Tom, et al.
Publicado: (2024)
por: Kocmi, Tom, et al.
Publicado: (2024)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
por: Perrella, Stefano, et al.
Publicado: (2024)
por: Perrella, Stefano, et al.
Publicado: (2024)
Estimating Machine Translation Difficulty
por: Proietti, Lorenzo, et al.
Publicado: (2025)
por: Proietti, Lorenzo, et al.
Publicado: (2025)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
por: Perrella, Stefano, et al.
Publicado: (2024)
por: Perrella, Stefano, et al.
Publicado: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)
por: Kocmi, Tom, et al.
Publicado: (2024)
Preliminary Ranking of WMT25 General Machine Translation Systems
por: Kocmi, Tom, et al.
Publicado: (2025)
por: Kocmi, Tom, et al.
Publicado: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
por: Wu, Guojun, et al.
Publicado: (2024)
por: Wu, Guojun, et al.
Publicado: (2024)
AskQE: Question Answering as Automatic Evaluation for Machine Translation
por: Ki, Dayeon, et al.
Publicado: (2025)
por: Ki, Dayeon, et al.
Publicado: (2025)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
por: Wang, Kuang-Da, et al.
Publicado: (2025)
por: Wang, Kuang-Da, et al.
Publicado: (2025)
Escaping the sentence-level paradigm in machine translation
por: Post, Matt, et al.
Publicado: (2023)
por: Post, Matt, et al.
Publicado: (2023)
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
por: Chen, Danlu, et al.
Publicado: (2026)
por: Chen, Danlu, et al.
Publicado: (2026)
Recovering document annotations for sentence-level bitext
por: Wicks, Rachel, et al.
Publicado: (2024)
por: Wicks, Rachel, et al.
Publicado: (2024)
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation
por: Levtsov, Georgii, et al.
Publicado: (2025)
por: Levtsov, Georgii, et al.
Publicado: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
por: Zouhar, Vilém, et al.
Publicado: (2024)
por: Zouhar, Vilém, et al.
Publicado: (2024)
BiVert: Bidirectional Vocabulary Evaluation using Relations for Machine Translation
por: Cherf, Carinne, et al.
Publicado: (2024)
por: Cherf, Carinne, et al.
Publicado: (2024)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
por: Lawrence, Logan, et al.
Publicado: (2025)
por: Lawrence, Logan, et al.
Publicado: (2025)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
por: Thompson, Brian, et al.
Publicado: (2024)
por: Thompson, Brian, et al.
Publicado: (2024)
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
por: Lin, Leonard, et al.
Publicado: (2026)
por: Lin, Leonard, et al.
Publicado: (2026)
Non-Linear Scoring Model for Translation Quality Evaluation
por: Gladkoff, Serge, et al.
Publicado: (2025)
por: Gladkoff, Serge, et al.
Publicado: (2025)
The Comparison of Translationese in Machine Translation and Human Transation in terms of Translation Relations
por: Zhou, Fan
Publicado: (2024)
por: Zhou, Fan
Publicado: (2024)
Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
por: García-Romero, Cristian, et al.
Publicado: (2025)
por: García-Romero, Cristian, et al.
Publicado: (2025)
A Critical Study of Automatic Evaluation in Sign Language Translation
por: Yazdani, Shakib, et al.
Publicado: (2025)
por: Yazdani, Shakib, et al.
Publicado: (2025)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
por: Hallaç, İbrahim Rıza, et al.
Publicado: (2026)
por: Hallaç, İbrahim Rıza, et al.
Publicado: (2026)
Automatically Generating Chinese Homophone Words to Probe Machine Translation Estimation Systems
por: Qian, Shenbin, et al.
Publicado: (2025)
por: Qian, Shenbin, et al.
Publicado: (2025)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
por: Sutawika, Lintang, et al.
Publicado: (2026)
por: Sutawika, Lintang, et al.
Publicado: (2026)
GRRM: Group Relative Reward Modeling for Machine Translation
por: Yang, Sen, et al.
Publicado: (2026)
por: Yang, Sen, et al.
Publicado: (2026)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
por: Jiang, Zhaokun, et al.
Publicado: (2024)
por: Jiang, Zhaokun, et al.
Publicado: (2024)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
por: Moghe, Nikita, et al.
Publicado: (2024)
por: Moghe, Nikita, et al.
Publicado: (2024)
Beyond Scalar Scores: Reinforcement Learning for Error-Aware Quality Estimation of Machine Translation
por: Sindhujan, Archchana, et al.
Publicado: (2026)
por: Sindhujan, Archchana, et al.
Publicado: (2026)
The quasi-semantic competence of LLMs: a case study on the part-whole relation
por: Proietti, Mattia, et al.
Publicado: (2025)
por: Proietti, Mattia, et al.
Publicado: (2025)
Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related Language
por: Schoene, Annika Marie, et al.
Publicado: (2024)
por: Schoene, Annika Marie, et al.
Publicado: (2024)
Uncertainty Quantification for Evaluating Machine Translation Bias
por: Staliūnaitė, Ieva Raminta, et al.
Publicado: (2025)
por: Staliūnaitė, Ieva Raminta, et al.
Publicado: (2025)
Evaluating Structural Generalization in Neural Machine Translation
por: Kumon, Ryoma, et al.
Publicado: (2024)
por: Kumon, Ryoma, et al.
Publicado: (2024)
AI-Assisted Human Evaluation of Machine Translation
por: Zouhar, Vilém, et al.
Publicado: (2024)
por: Zouhar, Vilém, et al.
Publicado: (2024)
Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
por: Wu, Patrick Y., et al.
Publicado: (2023)
por: Wu, Patrick Y., et al.
Publicado: (2023)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
por: Kawabata, Akira, et al.
Publicado: (2024)
por: Kawabata, Akira, et al.
Publicado: (2024)
Ejemplares similares
-
PyMarian: Fast Neural Machine Translation and Evaluation in Python
por: Gowda, Thamme, et al.
Publicado: (2024) -
On Instruction-Finetuning Neural Machine Translation Models
por: Raunak, Vikas, et al.
Publicado: (2024) -
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
por: Raunak, Vikas, et al.
Publicado: (2023) -
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
por: Proietti, Lorenzo, et al.
Publicado: (2025) -
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
por: Kocmi, Tom, et al.
Publicado: (2024)