MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Veuthey, Jaime Raldua, Majid, Zainab Ali, Hariharan, Suhas, Haimes, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
by: Hariharan, Suhas, et al.
Published: (2024)
by: Hariharan, Suhas, et al.
Published: (2024)
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
by: Chopra, Tanush, et al.
Published: (2024)
by: Chopra, Tanush, et al.
Published: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
by: Ji, Jiabao, et al.
Published: (2025)
by: Ji, Jiabao, et al.
Published: (2025)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
by: Reese, May Lynn, et al.
Published: (2026)
by: Reese, May Lynn, et al.
Published: (2026)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Integrated Framework for LLM Evaluation with Answer Generation
by: Lee, Sujeong, et al.
Published: (2025)
by: Lee, Sujeong, et al.
Published: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment
by: Yaacoub, Antoun, et al.
Published: (2025)
by: Yaacoub, Antoun, et al.
Published: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
by: Sturgeon, Benjamin, et al.
Published: (2025)
by: Sturgeon, Benjamin, et al.
Published: (2025)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
by: Moses, Movina, et al.
Published: (2025)
by: Moses, Movina, et al.
Published: (2025)
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
by: Kahana, Adar, et al.
Published: (2024)
by: Kahana, Adar, et al.
Published: (2024)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation
by: Sahu, Chandan Kumar, et al.
Published: (2026)
by: Sahu, Chandan Kumar, et al.
Published: (2026)
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
A Dataset of Open-Domain Question Answering with Multiple-Span Answers
by: Luo, Zhiyi, et al.
Published: (2024)
by: Luo, Zhiyi, et al.
Published: (2024)
Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation
by: Zhua, Fanwei, et al.
Published: (2025)
by: Zhua, Fanwei, et al.
Published: (2025)
Hallucination-Free Automatic Question & Answer Generation for Intuitive Learning
by: Wang, Nicholas X., et al.
Published: (2026)
by: Wang, Nicholas X., et al.
Published: (2026)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
by: Myrzakhan, Aidar, et al.
Published: (2024)
by: Myrzakhan, Aidar, et al.
Published: (2024)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
by: Cavalin, Paulo, et al.
Published: (2025)
by: Cavalin, Paulo, et al.
Published: (2025)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024)
by: Fu, Weiping, et al.
Published: (2024)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
by: Lai, Peichao, et al.
Published: (2025)
by: Lai, Peichao, et al.
Published: (2025)
Semantic Mastery: Enhancing LLMs with Advanced Natural Language Understanding
by: Hariharan, Mohanakrishnan
Published: (2025)
by: Hariharan, Mohanakrishnan
Published: (2025)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
by: Li, Jidong, et al.
Published: (2025)
by: Li, Jidong, et al.
Published: (2025)
General Table Question Answering via Answer-Formula Joint Generation
by: Wang, Zhongyuan, et al.
Published: (2025)
by: Wang, Zhongyuan, et al.
Published: (2025)
On Few-Shot Prompting for Controllable Question-Answer Generation in Narrative Comprehension
by: Leite, Bernardo, et al.
Published: (2024)
by: Leite, Bernardo, et al.
Published: (2024)
Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations
by: Agrawal, Garima, et al.
Published: (2024)
by: Agrawal, Garima, et al.
Published: (2024)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
by: Kostiuk, Yevhen, et al.
Published: (2025)
by: Kostiuk, Yevhen, et al.
Published: (2025)
A Benchmark for Long-Form Medical Question Answering
by: Hosseini, Pedram, et al.
Published: (2024)
by: Hosseini, Pedram, et al.
Published: (2024)
A Claim Decomposition Benchmark for Long-form Answer Verification
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
by: Liu, Yinhong, et al.
Published: (2025)
by: Liu, Yinhong, et al.
Published: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
by: Islamaj, Rezarta, et al.
Published: (2026)
by: Islamaj, Rezarta, et al.
Published: (2026)
Similar Items
-
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
by: Hariharan, Suhas, et al.
Published: (2024) -
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
by: Chopra, Tanush, et al.
Published: (2024) -
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
by: Ji, Jiabao, et al.
Published: (2025) -
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
by: Reese, May Lynn, et al.
Published: (2026) -
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)