Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Figueras, Banca Calvo, Agerri, Rodrigo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915507482918912
author Figueras, Banca Calvo
Agerri, Rodrigo
author_facet Figueras, Banca Calvo
Agerri, Rodrigo
contents The task of Critical Questions Generation (CQs-Gen) aims to foster critical thinking by enabling systems to generate questions that expose underlying assumptions and challenge the validity of argumentative reasoning structures. Despite growing interest in this area, progress has been hindered by the lack of suitable datasets and automatic evaluation standards. This paper presents a comprehensive approach to support the development and benchmarking of systems for this task. We construct the first large-scale dataset including ~5K manually annotated questions. We also investigate automatic evaluation methods and propose reference-based techniques as the strategy that best correlates with human judgments. Our zero-shot evaluation of 11 LLMs establishes a strong baseline while showcasing the difficulty of the task. Data and code plus a public leaderboard are provided to encourage further research, not only in terms of model performance, but also to explore the practical benefits of CQs-Gen for both automated reasoning and human critical thinking.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11341
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
Figueras, Banca Calvo
Agerri, Rodrigo
Computation and Language
The task of Critical Questions Generation (CQs-Gen) aims to foster critical thinking by enabling systems to generate questions that expose underlying assumptions and challenge the validity of argumentative reasoning structures. Despite growing interest in this area, progress has been hindered by the lack of suitable datasets and automatic evaluation standards. This paper presents a comprehensive approach to support the development and benchmarking of systems for this task. We construct the first large-scale dataset including ~5K manually annotated questions. We also investigate automatic evaluation methods and propose reference-based techniques as the strategy that best correlates with human judgments. Our zero-shot evaluation of 11 LLMs establishes a strong baseline while showcasing the difficulty of the task. Data and code plus a public leaderboard are provided to encourage further research, not only in terms of model performance, but also to explore the practical benefits of CQs-Gen for both automated reasoning and human critical thinking.
title Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.11341