Pearmut: Human Evaluation of Translation Made Trivial

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zouhar, Vilém, Kocmi, Tom
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917419365171200
author Zouhar, Vilém
Kocmi, Tom
author_facet Zouhar, Vilém
Kocmi, Tom
contents Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is notoriously complex and slow to set up with existing tools with substantial engineering and operational overhead. We introduce Pearmut, a lightweight yet feature-rich platform that makes end-to-end human evaluation as easy to run as automatic evaluation. Pearmut removes common entry barriers and provides support for evaluating multilingual tasks, with a particular focus on machine translation. The platform implements standard evaluation protocols, including DA, ESA, and MQM, and is extensible to support new protocols. It features document-level context, absolute and contrastive evaluation, attention checks, ESAAI pre-annotations and both static and dynamic assignment strategies. Pearmut enables reliable human evaluation to become a practical, routine component of model development and diagnosis rather than an occasional effort.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02933
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Pearmut: Human Evaluation of Translation Made Trivial
Zouhar, Vilém
Kocmi, Tom
Computation and Language
Human-Computer Interaction
Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is notoriously complex and slow to set up with existing tools with substantial engineering and operational overhead. We introduce Pearmut, a lightweight yet feature-rich platform that makes end-to-end human evaluation as easy to run as automatic evaluation. Pearmut removes common entry barriers and provides support for evaluating multilingual tasks, with a particular focus on machine translation. The platform implements standard evaluation protocols, including DA, ESA, and MQM, and is extensible to support new protocols. It features document-level context, absolute and contrastive evaluation, attention checks, ESAAI pre-annotations and both static and dynamic assignment strategies. Pearmut enables reliable human evaluation to become a practical, routine component of model development and diagnosis rather than an occasional effort.
title Pearmut: Human Evaluation of Translation Made Trivial
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2601.02933