Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
Fuente:
arXiv
Guardado en:
| Autores principales: | Deviyani, Athiya, Diaz, Fernando |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
por: Huynh, Jessica, et al.
Publicado: (2026)
por: Huynh, Jessica, et al.
Publicado: (2026)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
por: Gur-Arieh, Yoav, et al.
Publicado: (2026)
por: Gur-Arieh, Yoav, et al.
Publicado: (2026)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
por: Isaza-Giraldo, Andrés, et al.
Publicado: (2025)
por: Isaza-Giraldo, Andrés, et al.
Publicado: (2025)
A Critical Look at Meta-evaluating Summarisation Evaluation Metrics
por: Dai, Xiang, et al.
Publicado: (2024)
por: Dai, Xiang, et al.
Publicado: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)
por: Kocmi, Tom, et al.
Publicado: (2024)
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
por: Pauli, Amalie Brogaard, et al.
Publicado: (2025)
por: Pauli, Amalie Brogaard, et al.
Publicado: (2025)
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
por: Zhang, Luke, et al.
Publicado: (2026)
por: Zhang, Luke, et al.
Publicado: (2026)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
por: Perrella, Stefano, et al.
Publicado: (2024)
por: Perrella, Stefano, et al.
Publicado: (2024)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
por: Eigler, Lukáš, et al.
Publicado: (2026)
por: Eigler, Lukáš, et al.
Publicado: (2026)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
por: Thompson, Brian, et al.
Publicado: (2024)
por: Thompson, Brian, et al.
Publicado: (2024)
Nonparametric LLM Evaluation from Preference Data
por: Frauen, Dennis, et al.
Publicado: (2026)
por: Frauen, Dennis, et al.
Publicado: (2026)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
por: Anugraha, David, et al.
Publicado: (2024)
por: Anugraha, David, et al.
Publicado: (2024)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
por: Winata, Genta Indra, et al.
Publicado: (2024)
por: Winata, Genta Indra, et al.
Publicado: (2024)
KG-EDAS: A Meta-Metric Framework for Evaluating Knowledge Graph Completion Models
por: Gul, Haji, et al.
Publicado: (2025)
por: Gul, Haji, et al.
Publicado: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
por: Ramprasad, Sanjana, et al.
Publicado: (2024)
por: Ramprasad, Sanjana, et al.
Publicado: (2024)
Evaluating Metrics for Bias in Word Embeddings
por: Schröder, Sarah, et al.
Publicado: (2021)
por: Schröder, Sarah, et al.
Publicado: (2021)
Comparative Experimentation of Accuracy Metrics in Automated Medical Reporting: The Case of Otitis Consultations
por: Faber, Wouter, et al.
Publicado: (2023)
por: Faber, Wouter, et al.
Publicado: (2023)
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization
por: Guo, Yue, et al.
Publicado: (2023)
por: Guo, Yue, et al.
Publicado: (2023)
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
por: Aksu, Taha, et al.
Publicado: (2024)
por: Aksu, Taha, et al.
Publicado: (2024)
Faithful Model Evaluation for Model-Based Metrics
por: Goyal, Palash, et al.
Publicado: (2023)
por: Goyal, Palash, et al.
Publicado: (2023)
Towards Region-aware Bias Evaluation Metrics
por: Borah, Angana, et al.
Publicado: (2024)
por: Borah, Angana, et al.
Publicado: (2024)
CASPR: Automated Evaluation Metric for Contrastive Summarization
por: Ananthamurugan, Nirupan, et al.
Publicado: (2024)
por: Ananthamurugan, Nirupan, et al.
Publicado: (2024)
An Automatic Quality Metric for Evaluating Simultaneous Interpretation
por: Makinae, Mana, et al.
Publicado: (2024)
por: Makinae, Mana, et al.
Publicado: (2024)
Calibrating Model-Based Evaluation Metrics for Summarization
por: Liu, Hongye, et al.
Publicado: (2026)
por: Liu, Hongye, et al.
Publicado: (2026)
A Measure of the System Dependence of Automated Metrics
por: von Däniken, Pius, et al.
Publicado: (2024)
por: von Däniken, Pius, et al.
Publicado: (2024)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
por: Collot, Stephane, et al.
Publicado: (2025)
por: Collot, Stephane, et al.
Publicado: (2025)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
por: Mukherjee, Sourabrata, et al.
Publicado: (2025)
por: Mukherjee, Sourabrata, et al.
Publicado: (2025)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
por: Polák, Peter, et al.
Publicado: (2025)
por: Polák, Peter, et al.
Publicado: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
por: Clegg, Kester, et al.
Publicado: (2025)
por: Clegg, Kester, et al.
Publicado: (2025)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
KIEval: Evaluation Metric for Document Key Information Extraction
por: Khang, Minsoo, et al.
Publicado: (2025)
por: Khang, Minsoo, et al.
Publicado: (2025)
Identifying Reliable Evaluation Metrics for Scientific Text Revision
por: Jourdan, Léane, et al.
Publicado: (2025)
por: Jourdan, Léane, et al.
Publicado: (2025)
Hacking Neural Evaluation Metrics with Single Hub Text
por: Deguchi, Hiroyuki, et al.
Publicado: (2025)
por: Deguchi, Hiroyuki, et al.
Publicado: (2025)
Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
por: Hu, Taojun, et al.
Publicado: (2024)
por: Hu, Taojun, et al.
Publicado: (2024)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
por: Wu, Guojun, et al.
Publicado: (2024)
por: Wu, Guojun, et al.
Publicado: (2024)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Towards Explainable Evaluation Metrics for Machine Translation
por: Leiter, Christoph, et al.
Publicado: (2023)
por: Leiter, Christoph, et al.
Publicado: (2023)
Evaluation Metrics for Text Data Augmentation in NLP
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
por: Lee, Jisoo, et al.
Publicado: (2025)
por: Lee, Jisoo, et al.
Publicado: (2025)
Reference-free Evaluation Metrics for Text Generation: A Survey
por: Ito, Takumi, et al.
Publicado: (2025)
por: Ito, Takumi, et al.
Publicado: (2025)
Ejemplares similares
-
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
por: Huynh, Jessica, et al.
Publicado: (2026) -
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
por: Gur-Arieh, Yoav, et al.
Publicado: (2026) -
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
por: Isaza-Giraldo, Andrés, et al.
Publicado: (2025) -
A Critical Look at Meta-evaluating Summarisation Evaluation Metrics
por: Dai, Xiang, et al.
Publicado: (2024) -
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
por: Kocmi, Tom, et al.
Publicado: (2024)