From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
Fuente:
arXiv
Saved in:
| Main Authors: | Finkelstein, Mara, Deutsch, Dan, Riley, Parker, Juraska, Juraj, Kovacs, Geza, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
Finding Replicable Human Evaluations via Stable Ranking Probability
by: Riley, Parker, et al.
Published: (2024)
by: Riley, Parker, et al.
Published: (2024)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods
by: Finkelstein, Mara, et al.
Published: (2023)
by: Finkelstein, Mara, et al.
Published: (2023)
Judging with Confidence: Calibrating Autoraters to Preference Distributions
by: Li, Zhuohang, et al.
Published: (2025)
by: Li, Zhuohang, et al.
Published: (2025)
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
by: Huynh, Jessica, et al.
Published: (2026)
by: Huynh, Jessica, et al.
Published: (2026)
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
by: Zhang, Lily H., et al.
Published: (2024)
by: Zhang, Lily H., et al.
Published: (2024)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Jack and Masters of all Trades: One-Pass Learning Sets of Model Sets From Large Pre-Trained Models
by: Choong, Han Xiang, et al.
Published: (2022)
by: Choong, Han Xiang, et al.
Published: (2022)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
by: Vu, Tu, et al.
Published: (2024)
by: Vu, Tu, et al.
Published: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
by: Luo, Zheheng, et al.
Published: (2024)
by: Luo, Zheheng, et al.
Published: (2024)
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Pulsation-driven helium transport as a potential source of the Blazhko effect
by: Kovacs, Geza
Published: (2026)
by: Kovacs, Geza
Published: (2026)
Digging Deeper for RR Lyrae Stars with Low Modulation Amplitudes
by: Kovacs, Geza
Published: (2025)
by: Kovacs, Geza
Published: (2025)
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
by: Gallouédec, Quentin, et al.
Published: (2024)
by: Gallouédec, Quentin, et al.
Published: (2024)
Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Goal Setting in Speech–Language Pathology: A Pilot Test of a ‘One‐Size‐Fits‐All’ Planning Framework
by: Justine Leigh Hamilton, et al.
Published: (2025)
by: Justine Leigh Hamilton, et al.
Published: (2025)
Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge
by: Vladika, Juraj, et al.
Published: (2025)
by: Vladika, Juraj, et al.
Published: (2025)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025)
by: Peter, Jan-Thorsten, et al.
Published: (2025)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
OneLLM: One Framework to Align All Modalities with Language
by: Han, Jiaming, et al.
Published: (2023)
by: Han, Jiaming, et al.
Published: (2023)
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
by: Yu, Jianxiang, et al.
Published: (2026)
by: Yu, Jianxiang, et al.
Published: (2026)
One LLM to Train Them All: Multi-Task Learning Framework for Fact-Checking
by: Larsson, Malin Astrid, et al.
Published: (2026)
by: Larsson, Malin Astrid, et al.
Published: (2026)
Efficient Terminology Integration for LLM-based Translation in Specialized Domains
by: Kim, Sejoon, et al.
Published: (2024)
by: Kim, Sejoon, et al.
Published: (2024)
Similar Items
-
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024) -
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025) -
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025) -
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025) -
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)