AI-Assisted Human Evaluation of Machine Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zouhar, Vilém, Kocmi, Tom, Sachan, Mrinmaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Estimating Machine Translation Difficulty
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2025)
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2025)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
von: Chowdhury, Sankalan Pal, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sankalan Pal, et al.
Veröffentlicht: (2024)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
von: Zouhar, Vilém
Veröffentlicht: (2024)
von: Zouhar, Vilém
Veröffentlicht: (2024)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
von: Cui, Peng, et al.
Veröffentlicht: (2024)
von: Cui, Peng, et al.
Veröffentlicht: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
Early-Exit and Instant Confidence Translation Quality Estimation
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
COMET-poly: Machine Translation Metric Grounded in Other Candidates
von: Züfle, Maike, et al.
Veröffentlicht: (2025)
von: Züfle, Maike, et al.
Veröffentlicht: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Multilingual Performance Biases of Large Language Models in Education
von: Gupta, Vansh, et al.
Veröffentlicht: (2025)
von: Gupta, Vansh, et al.
Veröffentlicht: (2025)
RELIC: Investigating Large Language Model Responses using Self-Consistency
von: Cheng, Furui, et al.
Veröffentlicht: (2023)
von: Cheng, Furui, et al.
Veröffentlicht: (2023)
Evaluating Optimal Reference Translations
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
A Bayesian Optimization Approach to Machine Translation Reranking
von: Cheng, Julius, et al.
Veröffentlicht: (2024)
von: Cheng, Julius, et al.
Veröffentlicht: (2024)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
von: Ni, Jingwei, et al.
Veröffentlicht: (2025)
von: Ni, Jingwei, et al.
Veröffentlicht: (2025)
A Formal Perspective on Byte-Pair Encoding
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
von: Raunak, Vikas, et al.
Veröffentlicht: (2023)
von: Raunak, Vikas, et al.
Veröffentlicht: (2023)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
How Important is `Perfect' English for Machine Translation Prompts?
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
von: Schmidtová, Patrícia, et al.
Veröffentlicht: (2025)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
Distributional Properties of Subword Regularization
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Multimodal Shannon Game with Images
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2023)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
von: Cui, Peng, et al.
Veröffentlicht: (2025)
von: Cui, Peng, et al.
Veröffentlicht: (2025)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
Generating Difficult-to-Translate Texts
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
Two Counterexamples to Tokenization and the Noiseless Channel
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Biased Tales: Cultural and Topic Bias in Generating Children's Stories
von: Rooein, Donya, et al.
Veröffentlicht: (2025)
von: Rooein, Donya, et al.
Veröffentlicht: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
von: Xu, Wenda, et al.
Veröffentlicht: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2024)
von: Ozyurt, Yilmazcan, et al.
Veröffentlicht: (2024)
Probing for Arithmetic Errors in Language Models
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
von: Sun, Yucheng, et al.
Veröffentlicht: (2025)
QE4PE: Word-level Quality Estimation for Human Post-Editing
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
von: Xiong, Chenfei, et al.
Veröffentlicht: (2025)
von: Xiong, Chenfei, et al.
Veröffentlicht: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026)
von: Do, Heejin, et al.
Veröffentlicht: (2026)
Preliminary Ranking of WMT25 General Machine Translation Systems
von: Kocmi, Tom, et al.
Veröffentlicht: (2025)
von: Kocmi, Tom, et al.
Veröffentlicht: (2025)
Efficiently Computing Susceptibility to Context in Language Models
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026) -
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
von: Kocmi, Tom, et al.
Veröffentlicht: (2024) -
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025) -
Estimating Machine Translation Difficulty
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2025) -
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
von: Chowdhury, Sankalan Pal, et al.
Veröffentlicht: (2024)