IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sharma, Karun, Vats, Vidushee, Li, Shengzhi, Wang, Yuxiang, Sun, Zhongtian, Tiwari, Prayag
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910043257962496
author Sharma, Karun
Vats, Vidushee
Li, Shengzhi
Wang, Yuxiang
Sun, Zhongtian
Tiwari, Prayag
author_facet Sharma, Karun
Vats, Vidushee
Li, Shengzhi
Wang, Yuxiang
Sun, Zhongtian
Tiwari, Prayag
contents Peer review relies on substantive, evidence-based questions, yet current LLMs generate surface-level queries that perform worse than human reviewer questions in expert evaluation. To address this gap, we curate a high-quality dataset of reviewer questions from OpenReview and conduct a human preference study where expert annotators evaluate question-paper pairs across three dimensions: effort, evidence, and grounding. From these annotations, we train IntelliReward, a reward model built from a frozen autoregressive LLM with trainable multi-head transformers. Validated against expert judgments, IntelliReward predicts reviewer-question quality better than API-based SFT baselines and provides scalable evaluation. We apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) with IntelliReward to train IntelliAsk, a question-generation model aligned with human standards of effortful, evidence-based critique. Human evaluations show IntelliAsk generates more grounded, substantive and effortful questions than strong baselines and reduces reliance on first-page content. We also find improvements on reasoning and writing benchmarks, suggesting reviewer-question quality correlates with broader capabilities. Compared to Qwen3-32B, IntelliAsk improves MuSR (68.3 vs 64.7 Acc) and WritingBench (8.31 vs 8.07). We release our code, filtered review dataset, expert annotations, IntelliAsk and IntelliReward to support automatic evaluation of grounding, effort, and evidence in LLM-generated review questions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15849
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
Sharma, Karun
Vats, Vidushee
Li, Shengzhi
Wang, Yuxiang
Sun, Zhongtian
Tiwari, Prayag
Computation and Language
Artificial Intelligence
Peer review relies on substantive, evidence-based questions, yet current LLMs generate surface-level queries that perform worse than human reviewer questions in expert evaluation. To address this gap, we curate a high-quality dataset of reviewer questions from OpenReview and conduct a human preference study where expert annotators evaluate question-paper pairs across three dimensions: effort, evidence, and grounding. From these annotations, we train IntelliReward, a reward model built from a frozen autoregressive LLM with trainable multi-head transformers. Validated against expert judgments, IntelliReward predicts reviewer-question quality better than API-based SFT baselines and provides scalable evaluation. We apply Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) with IntelliReward to train IntelliAsk, a question-generation model aligned with human standards of effortful, evidence-based critique. Human evaluations show IntelliAsk generates more grounded, substantive and effortful questions than strong baselines and reduces reliance on first-page content. We also find improvements on reasoning and writing benchmarks, suggesting reviewer-question quality correlates with broader capabilities. Compared to Qwen3-32B, IntelliAsk improves MuSR (68.3 vs 64.7 Acc) and WritingBench (8.31 vs 8.07). We release our code, filtered review dataset, expert annotations, IntelliAsk and IntelliReward to support automatic evaluation of grounding, effort, and evidence in LLM-generated review questions.
title IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.15849