MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Veuthey, Jaime Raldua, Majid, Zainab Ali, Hariharan, Suhas, Haimes, Jacob
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916696628920320
author Veuthey, Jaime Raldua
Majid, Zainab Ali
Hariharan, Suhas
Haimes, Jacob
author_facet Veuthey, Jaime Raldua
Majid, Zainab Ali
Hariharan, Suhas
Haimes, Jacob
contents As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
Veuthey, Jaime Raldua
Majid, Zainab Ali
Hariharan, Suhas
Haimes, Jacob
Computation and Language
Artificial Intelligence
As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats.
title MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.14039