Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
Fuente:
arXiv
Salvato in:
| Autori principali: | Thompson, Brian, Mathur, Nitika, Deutsch, Daniel, Khayrallah, Huda |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating Metric Bias in Minimum Bayes Risk Decoding
di: Kovacs, Geza, et al.
Pubblicazione: (2024)
di: Kovacs, Geza, et al.
Pubblicazione: (2024)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
di: Soh, Yun Joon, et al.
Pubblicazione: (2024)
di: Soh, Yun Joon, et al.
Pubblicazione: (2024)
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
di: Luo, Shunyang, et al.
Pubblicazione: (2026)
di: Luo, Shunyang, et al.
Pubblicazione: (2026)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
di: Ramprasad, Sanjana, et al.
Pubblicazione: (2024)
di: Ramprasad, Sanjana, et al.
Pubblicazione: (2024)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
di: GUO, Jiaxin, et al.
Pubblicazione: (2025)
di: GUO, Jiaxin, et al.
Pubblicazione: (2025)
How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation
di: Khayrallah, Huda, et al.
Pubblicazione: (2023)
di: Khayrallah, Huda, et al.
Pubblicazione: (2023)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
di: Xu, Wenda, et al.
Pubblicazione: (2025)
di: Xu, Wenda, et al.
Pubblicazione: (2025)
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
di: Li, Senyu, et al.
Pubblicazione: (2025)
di: Li, Senyu, et al.
Pubblicazione: (2025)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
di: Zouhar, Vilém, et al.
Pubblicazione: (2024)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
di: Sun, Bian, et al.
Pubblicazione: (2026)
di: Sun, Bian, et al.
Pubblicazione: (2026)
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
di: Wang, Xinda, et al.
Pubblicazione: (2025)
di: Wang, Xinda, et al.
Pubblicazione: (2025)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
di: Kocyigit, Muhammed Yusuf, et al.
Pubblicazione: (2025)
di: Kocyigit, Muhammed Yusuf, et al.
Pubblicazione: (2025)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
di: Kirstein, Frederic, et al.
Pubblicazione: (2024)
di: Kirstein, Frederic, et al.
Pubblicazione: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
di: Cho, Yousang, et al.
Pubblicazione: (2025)
di: Cho, Yousang, et al.
Pubblicazione: (2025)
On-the-Fly Fusion of Large Language Models and Machine Translation
di: Hoang, Hieu, et al.
Pubblicazione: (2023)
di: Hoang, Hieu, et al.
Pubblicazione: (2023)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
di: Matsuda, Kazuki, et al.
Pubblicazione: (2024)
di: Matsuda, Kazuki, et al.
Pubblicazione: (2024)
JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
di: Lin, Leonard, et al.
Pubblicazione: (2026)
di: Lin, Leonard, et al.
Pubblicazione: (2026)
Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
di: Na, Injae, et al.
Pubblicazione: (2025)
di: Na, Injae, et al.
Pubblicazione: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
di: Wang, Tong, et al.
Pubblicazione: (2024)
di: Wang, Tong, et al.
Pubblicazione: (2024)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
di: Wang, Shouju, et al.
Pubblicazione: (2026)
di: Wang, Shouju, et al.
Pubblicazione: (2026)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
di: Collot, Stephane, et al.
Pubblicazione: (2025)
di: Collot, Stephane, et al.
Pubblicazione: (2025)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
di: Balamurali, Sai Shridhar, et al.
Pubblicazione: (2025)
di: Balamurali, Sai Shridhar, et al.
Pubblicazione: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
di: Schoenegger, Philipp, et al.
Pubblicazione: (2024)
di: Schoenegger, Philipp, et al.
Pubblicazione: (2024)
Human and Automatic Interpretation of Romanian Noun Compounds
di: Marinescu, Ioana, et al.
Pubblicazione: (2024)
di: Marinescu, Ioana, et al.
Pubblicazione: (2024)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
An Automatic Question Usability Evaluation Toolkit
di: Moore, Steven, et al.
Pubblicazione: (2024)
di: Moore, Steven, et al.
Pubblicazione: (2024)
Automatic Legal Writing Evaluation of LLMs
di: Pires, Ramon, et al.
Pubblicazione: (2025)
di: Pires, Ramon, et al.
Pubblicazione: (2025)
Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation
di: DiIanni, Colten, et al.
Pubblicazione: (2025)
di: DiIanni, Colten, et al.
Pubblicazione: (2025)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
di: Shibata, Takumi, et al.
Pubblicazione: (2025)
di: Shibata, Takumi, et al.
Pubblicazione: (2025)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
di: Anugraha, David, et al.
Pubblicazione: (2024)
di: Anugraha, David, et al.
Pubblicazione: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
di: Huang, Liangjie, et al.
Pubblicazione: (2025)
di: Huang, Liangjie, et al.
Pubblicazione: (2025)
Evaluation Metrics for Text Data Augmentation in NLP
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
di: Amadeus, Marcellus, et al.
Pubblicazione: (2024)
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
di: Dik, Selin, et al.
Pubblicazione: (2025)
di: Dik, Selin, et al.
Pubblicazione: (2025)
Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs
di: Tan, Nelvin, et al.
Pubblicazione: (2026)
di: Tan, Nelvin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Mitigating Metric Bias in Minimum Bayes Risk Decoding
di: Kovacs, Geza, et al.
Pubblicazione: (2024) -
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
di: Ryan, Michael J., et al.
Pubblicazione: (2025) -
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
di: Soh, Yun Joon, et al.
Pubblicazione: (2024) -
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
di: Luo, Shunyang, et al.
Pubblicazione: (2026) -
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
di: Ramprasad, Sanjana, et al.
Pubblicazione: (2024)