MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916696628920320 |
|---|---|
| author | Veuthey, Jaime Raldua Majid, Zainab Ali Hariharan, Suhas Haimes, Jacob |
| author_facet | Veuthey, Jaime Raldua Majid, Zainab Ali Hariharan, Suhas Haimes, Jacob |
| contents | As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_14039 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks Veuthey, Jaime Raldua Majid, Zainab Ali Hariharan, Suhas Haimes, Jacob Computation and Language Artificial Intelligence As Large Language Models (LLMs) advance, their potential for widespread societal impact grows simultaneously. Hence, rigorous LLM evaluations are both a technical necessity and social imperative. While numerous evaluation benchmarks have been developed, there remains a critical gap in meta-evaluation: effectively assessing benchmarks' quality. We propose MEQA, a framework for the meta-evaluation of question and answer (QA) benchmarks, to provide standardized assessments, quantifiable scores, and enable meaningful intra-benchmark comparisons. We demonstrate this approach on cybersecurity benchmarks, using human and LLM evaluators, highlighting the benchmarks' strengths and weaknesses. We motivate our choice of test domain by AI models' dual nature as powerful defensive tools and security threats. |
| title | MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2504.14039 |