A Measure of the System Dependence of Automated Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | von Däniken, Pius, Deriu, Jan, Cieliebak, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
by: von Däniken, Pius, et al.
Published: (2024)
by: von Däniken, Pius, et al.
Published: (2024)
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos
by: Giedemann, Patrick, et al.
Published: (2025)
by: Giedemann, Patrick, et al.
Published: (2025)
Error-preserving Automatic Speech Recognition of Young English Learners' Language
by: Michot, Janick, et al.
Published: (2024)
by: Michot, Janick, et al.
Published: (2024)
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
by: Stucki, Samuel, et al.
Published: (2025)
by: Stucki, Samuel, et al.
Published: (2025)
Voice Adaptation for Swiss German
by: Stucki, Samuel, et al.
Published: (2025)
by: Stucki, Samuel, et al.
Published: (2025)
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
by: Rusli, Andre, et al.
Published: (2024)
by: Rusli, Andre, et al.
Published: (2024)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
by: Yin, Joy Lim Jia, et al.
Published: (2025)
by: Yin, Joy Lim Jia, et al.
Published: (2025)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
by: Kirstein, Frederic, et al.
Published: (2024)
by: Kirstein, Frederic, et al.
Published: (2024)
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
by: Ormerod, Christopher
Published: (2025)
by: Ormerod, Christopher
Published: (2025)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
LABIIUM: AI-Enhanced Zero-configuration Measurement Automation System
by: Olowe, Emmanuel A., et al.
Published: (2024)
by: Olowe, Emmanuel A., et al.
Published: (2024)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026)
by: Chen, Yinzhu, et al.
Published: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
by: Joshi, Nisheeth, et al.
Published: (2025)
by: Joshi, Nisheeth, et al.
Published: (2025)
Measuring Retrieval Complexity in Question Answering Systems
by: Gabburo, Matteo, et al.
Published: (2024)
by: Gabburo, Matteo, et al.
Published: (2024)
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
by: DeLucia, Alexandra, et al.
Published: (2025)
by: DeLucia, Alexandra, et al.
Published: (2025)
Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
by: Lermen, Simon, et al.
Published: (2025)
by: Lermen, Simon, et al.
Published: (2025)
TreeMatch: A Fully Unsupervised WSD System Using Dependency Knowledge on a Specific Domain
by: Tran, Andrew, et al.
Published: (2025)
by: Tran, Andrew, et al.
Published: (2025)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
Remote Labor Index: Measuring AI Automation of Remote Work
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Automated Triaging and Transfer Learning of Incident Learning Safety Reports Using Large Language Representational Models
by: Beidler, Peter, et al.
Published: (2025)
by: Beidler, Peter, et al.
Published: (2025)
Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
by: Zhao, Yida, et al.
Published: (2024)
by: Zhao, Yida, et al.
Published: (2024)
What Are the Facts? Automated Extraction of Court-Established Facts from Criminal-Court Opinions
by: Bendová, Klára, et al.
Published: (2025)
by: Bendová, Klára, et al.
Published: (2025)
Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
by: Belem, Catarina G, et al.
Published: (2025)
by: Belem, Catarina G, et al.
Published: (2025)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
by: Park, Chanhee, et al.
Published: (2025)
by: Park, Chanhee, et al.
Published: (2025)
BBScore: A Brownian Bridge Based Metric for Assessing Text Coherence
by: Sheng, Zhecheng, et al.
Published: (2023)
by: Sheng, Zhecheng, et al.
Published: (2023)
Comparing Hallucination Detection Metrics for Multilingual Generation
by: Kang, Haoqiang, et al.
Published: (2024)
by: Kang, Haoqiang, et al.
Published: (2024)
Evaluation Metrics for Text Data Augmentation in NLP
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
Reward Models are Metrics in a Trench Coat
by: Gehrmann, Sebastian
Published: (2025)
by: Gehrmann, Sebastian
Published: (2025)
Automated Neural Patent Landscaping in the Small Data Regime
by: Erana, Tisa Islam, et al.
Published: (2024)
by: Erana, Tisa Islam, et al.
Published: (2024)
Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
by: West, Alva, et al.
Published: (2025)
by: West, Alva, et al.
Published: (2025)
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems
by: Chowdhury, Mohita, et al.
Published: (2025)
by: Chowdhury, Mohita, et al.
Published: (2025)
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Full-ECE: A Metric For Token-level Calibration on Large Language Models
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
Similar Items
-
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
by: von Däniken, Pius, et al.
Published: (2024) -
ViClaim: A Multilingual Multilabel Dataset for Automatic Claim Detection in Videos
by: Giedemann, Patrick, et al.
Published: (2025) -
Error-preserving Automatic Speech Recognition of Young English Learners' Language
by: Michot, Janick, et al.
Published: (2024) -
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
by: Stucki, Samuel, et al.
Published: (2025) -
Voice Adaptation for Swiss German
by: Stucki, Samuel, et al.
Published: (2025)