A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Soh, Yun Joon, Zhao, Jishen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Only Use Reactive Attention Slice For Long Context Retrieval
von: Soh, Yun Joon, et al.
Veröffentlicht: (2024)
von: Soh, Yun Joon, et al.
Veröffentlicht: (2024)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
von: Thompson, Brian, et al.
Veröffentlicht: (2024)
von: Thompson, Brian, et al.
Veröffentlicht: (2024)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
von: Do, Heejin, et al.
Veröffentlicht: (2025)
von: Do, Heejin, et al.
Veröffentlicht: (2025)
Language Models are Few-Shot Graders
von: Zhao, Chenyan, et al.
Veröffentlicht: (2025)
von: Zhao, Chenyan, et al.
Veröffentlicht: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
VietLyrics: A Large-Scale Dataset and Models for Vietnamese Automatic Lyrics Transcription
von: Nguyen, Quoc Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Quoc Anh, et al.
Veröffentlicht: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
von: Ryan, Michael J., et al.
Veröffentlicht: (2025)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
von: GUO, Jiaxin, et al.
Veröffentlicht: (2025)
Large Language Models As MOOCs Graders
von: Golchin, Shahriar, et al.
Veröffentlicht: (2024)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2024)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
von: Sun, Bian, et al.
Veröffentlicht: (2026)
von: Sun, Bian, et al.
Veröffentlicht: (2026)
Are Large Language Models Good Essay Graders?
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
Discovering Significant Topics from Legal Decisions with Selective Inference
von: Soh, Jerrold
Veröffentlicht: (2024)
von: Soh, Jerrold
Veröffentlicht: (2024)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
von: Liu, Yuliang, et al.
Veröffentlicht: (2025)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
von: Hao, Shibo, et al.
Veröffentlicht: (2024)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
von: Rusli, Andre, et al.
Veröffentlicht: (2024)
Learning to Maximize Mutual Information for Chain-of-Thought Distillation
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
von: Balamurali, Sai Shridhar, et al.
Veröffentlicht: (2025)
von: Balamurali, Sai Shridhar, et al.
Veröffentlicht: (2025)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
von: Barnes, Jeremy, et al.
Veröffentlicht: (2025)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
A Grey-box Text Attack Framework using Explainable AI
von: Chiramal, Esther, et al.
Veröffentlicht: (2025)
von: Chiramal, Esther, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
An Automatic Question Usability Evaluation Toolkit
von: Moore, Steven, et al.
Veröffentlicht: (2024)
von: Moore, Steven, et al.
Veröffentlicht: (2024)
Automatic Legal Writing Evaluation of LLMs
von: Pires, Ramon, et al.
Veröffentlicht: (2025)
von: Pires, Ramon, et al.
Veröffentlicht: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Emotion Identification for French in Written Texts: Considering their Modes of Expression as a Step Towards Text Complexity Analysis
von: Étienne, Aline, et al.
Veröffentlicht: (2024)
von: Étienne, Aline, et al.
Veröffentlicht: (2024)
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2026)
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2026)
Evaluation Metrics for Text Data Augmentation in NLP
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
von: Amadeus, Marcellus, et al.
Veröffentlicht: (2024)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
von: Qi, Siya, et al.
Veröffentlicht: (2024)
von: Qi, Siya, et al.
Veröffentlicht: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
von: Yang, Yongjin, et al.
Veröffentlicht: (2024)
Not All Layers Need Tuning: Selective Layer Restoration Recovers Diversity
von: Zhang, Bowen, et al.
Veröffentlicht: (2026)
von: Zhang, Bowen, et al.
Veröffentlicht: (2026)
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking
von: Brådland, Henrik, et al.
Veröffentlicht: (2025)
von: Brådland, Henrik, et al.
Veröffentlicht: (2025)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
von: Perrella, Stefano, et al.
Veröffentlicht: (2024)
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
von: Vu, Tu, et al.
Veröffentlicht: (2024)
von: Vu, Tu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
You Only Use Reactive Attention Slice For Long Context Retrieval
von: Soh, Yun Joon, et al.
Veröffentlicht: (2024) -
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
von: Thompson, Brian, et al.
Veröffentlicht: (2024) -
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
von: Do, Heejin, et al.
Veröffentlicht: (2025) -
Language Models are Few-Shot Graders
von: Zhao, Chenyan, et al.
Veröffentlicht: (2025) -
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)