DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khoramfar, Ali, Ramezani, Ali, Mohajeri, Mohammad Mahdi, Dousti, Mohammad Javad, Ahmadabadi, Majid Nili, Faili, Heshaam
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911473187422208
author Khoramfar, Ali
Ramezani, Ali
Mohajeri, Mohammad Mahdi
Dousti, Mohammad Javad
Ahmadabadi, Majid Nili
Faili, Heshaam
author_facet Khoramfar, Ali
Ramezani, Ali
Mohajeri, Mohammad Mahdi
Dousti, Mohammad Javad
Ahmadabadi, Majid Nili
Faili, Heshaam
contents While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets. Grounded in Bloom's taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models reveals a stark performance decline with accuracy dropping by up to 70% as tasks ascend the cognitive hierarchy. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
Khoramfar, Ali
Ramezani, Ali
Mohajeri, Mohammad Mahdi
Dousti, Mohammad Javad
Ahmadabadi, Majid Nili
Faili, Heshaam
Computation and Language
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets. Grounded in Bloom's taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models reveals a stark performance decline with accuracy dropping by up to 70% as tasks ascend the cognitive hierarchy. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.
title DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
topic Computation and Language
url https://arxiv.org/abs/2505.24532