Qworld: Question-Specific Evaluation Criteria for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Shanghua, Su, Yuchang, Sui, Pengwei, Ginder, Curtis, Zitnik, Marinka
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911542831742976
author Gao, Shanghua
Su, Yuchang
Sui, Pengwei
Ginder, Curtis
Zitnik, Marinka
author_facet Gao, Shanghua
Su, Yuchang
Sui, Pengwei
Ginder, Curtis
Zitnik, Marinka
contents Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements. Existing methods define criteria at the dataset level or generate them in a single pass, which limits their ability to explore the evaluation space implied by each question. We introduce One-Question-One-World (Qworld), a method that generates question-specific evaluation criteria using a recursive expansion tree. Given a question, Qworld decomposes it into scenarios, perspectives, and fine-grained binary criteria through structured hierarchical and horizontal expansion. The resulting criteria specify what a high-quality answer must address for that question. On HealthBench, Qworld covers 89% of expert-authored criteria and generates 79% novel criteria validated by human experts. Experts rate Qworld criteria higher in insight and granularity than those produced by prior methods. When applied to 11 frontier LLMs on HealthBench and Humanity's Last Exam, Qworld reveals capability differences in dimensions such as long-term impact, equity, error handling, and interdisciplinary reasoning that coarse rubrics do not distinguish. By formulating criteria generation as structured coverage of question-implied evaluation axes, Qworld enables evaluation that adapts to each question rather than relying on fixed task-level criteria.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23522
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Qworld: Question-Specific Evaluation Criteria for LLMs
Gao, Shanghua
Su, Yuchang
Sui, Pengwei
Ginder, Curtis
Zitnik, Marinka
Computation and Language
Artificial Intelligence
Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements. Existing methods define criteria at the dataset level or generate them in a single pass, which limits their ability to explore the evaluation space implied by each question. We introduce One-Question-One-World (Qworld), a method that generates question-specific evaluation criteria using a recursive expansion tree. Given a question, Qworld decomposes it into scenarios, perspectives, and fine-grained binary criteria through structured hierarchical and horizontal expansion. The resulting criteria specify what a high-quality answer must address for that question. On HealthBench, Qworld covers 89% of expert-authored criteria and generates 79% novel criteria validated by human experts. Experts rate Qworld criteria higher in insight and granularity than those produced by prior methods. When applied to 11 frontier LLMs on HealthBench and Humanity's Last Exam, Qworld reveals capability differences in dimensions such as long-term impact, equity, error handling, and interdisciplinary reasoning that coarse rubrics do not distinguish. By formulating criteria generation as structured coverage of question-implied evaluation axes, Qworld enables evaluation that adapts to each question rather than relying on fixed task-level criteria.
title Qworld: Question-Specific Evaluation Criteria for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.23522