Pearmut: Human Evaluation of Translation Made Trivial
Fuente:
arXiv
Saved in:
| Main Authors: | Zouhar, Vilém, Kocmi, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
by: Chowdhury, Sankalan Pal, et al.
Published: (2024)
QE4PE: Word-level Quality Estimation for Human Post-Editing
by: Sarti, Gabriele, et al.
Published: (2025)
by: Sarti, Gabriele, et al.
Published: (2025)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
CafGa: Customizing Feature Attributions to Explain Language Models
by: Boyle, Alan, et al.
Published: (2025)
by: Boyle, Alan, et al.
Published: (2025)
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Context-Aware Monolingual Human Evaluation of Machine Translation
by: Picinini, Silvio, et al.
Published: (2025)
by: Picinini, Silvio, et al.
Published: (2025)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
by: Haq, Sami Ul, et al.
Published: (2025)
by: Haq, Sami Ul, et al.
Published: (2025)
Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation
by: Haavisto, Otso, et al.
Published: (2024)
by: Haavisto, Otso, et al.
Published: (2024)
Machine Translation in the Wild: User Reaction to Xiaohongshu's Built-In Translation Feature
by: He, Sui
Published: (2026)
by: He, Sui
Published: (2026)
Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows
by: Balashov, Yuri, et al.
Published: (2026)
by: Balashov, Yuri, et al.
Published: (2026)
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
by: Belz, Anya, et al.
Published: (2024)
by: Belz, Anya, et al.
Published: (2024)
Emojinize: Enriching Any Text with Emoji Translations
by: Klein, Lars Henning, et al.
Published: (2024)
by: Klein, Lars Henning, et al.
Published: (2024)
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
Media of Langue: The Interface for Exploring Word Translation Network/Space
by: Muramoto, Goki, et al.
Published: (2023)
by: Muramoto, Goki, et al.
Published: (2023)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication
by: Ghafouri, Bijean, et al.
Published: (2026)
by: Ghafouri, Bijean, et al.
Published: (2026)
Introducing Quality Estimation to Machine Translation Post-editing Workflow: An Empirical Study on Its Usefulness
by: Liu, Siqi, et al.
Published: (2025)
by: Liu, Siqi, et al.
Published: (2025)
Prompting ChatGPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
by: He, Sui
Published: (2024)
by: He, Sui
Published: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Efficient Machine Translation Corpus Generation: Integrating Human-in-the-Loop Post-Editing with Large Language Models
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
Understanding Large Language Model Behaviors through Interactive Counterfactual Generation and Analysis
by: Cheng, Furui, et al.
Published: (2024)
by: Cheng, Furui, et al.
Published: (2024)
Are Humans as Brittle as Large Language Models?
by: Li, Jiahui, et al.
Published: (2025)
by: Li, Jiahui, et al.
Published: (2025)
Evaluating LLMs as Human Surrogates in Controlled Experiments
by: Hoq, Adnan, et al.
Published: (2026)
by: Hoq, Adnan, et al.
Published: (2026)
Modeling Distinct Human Interaction in Web Agents
by: Huq, Faria, et al.
Published: (2026)
by: Huq, Faria, et al.
Published: (2026)
Lexical Indicators of Mind Perception in Human-AI Companionship
by: Banks, Jaime, et al.
Published: (2026)
by: Banks, Jaime, et al.
Published: (2026)
Practicing with Language Models Cultivates Human Empathic Communication
by: Kumar, Aakriti, et al.
Published: (2026)
by: Kumar, Aakriti, et al.
Published: (2026)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
Large Language Models for Virtual Human Gesture Selection
by: Torshizi, Parisa Ghanad, et al.
Published: (2025)
by: Torshizi, Parisa Ghanad, et al.
Published: (2025)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
Learning Next Action Predictors from Human-Computer Interaction
by: Shaikh, Omar, et al.
Published: (2026)
by: Shaikh, Omar, et al.
Published: (2026)
Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning
by: Churina, Svetlana, et al.
Published: (2025)
by: Churina, Svetlana, et al.
Published: (2025)
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
by: Zhang, Yusong, et al.
Published: (2024)
by: Zhang, Yusong, et al.
Published: (2024)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
Show or Tell? Modeling the evolution of request-making in Human-LLM conversations
by: Zhu, Shengqi, et al.
Published: (2025)
by: Zhu, Shengqi, et al.
Published: (2025)
Similar Items
-
AI-Assisted Human Evaluation of Machine Translation
by: Zouhar, Vilém, et al.
Published: (2024) -
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
by: Sarti, Gabriele, et al.
Published: (2025) -
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
by: Chowdhury, Sankalan Pal, et al.
Published: (2024) -
QE4PE: Word-level Quality Estimation for Human Post-Editing
by: Sarti, Gabriele, et al.
Published: (2025) -
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)