When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Jiale, Fang, Ke, Cheng, Lu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
von: Gulati, Anmol, et al.
Veröffentlicht: (2026)
von: Gulati, Anmol, et al.
Veröffentlicht: (2026)
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
von: Kumar, Abhijit, et al.
Veröffentlicht: (2026)
von: Kumar, Abhijit, et al.
Veröffentlicht: (2026)
IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
von: Sharma, Karun, et al.
Veröffentlicht: (2026)
von: Sharma, Karun, et al.
Veröffentlicht: (2026)
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
Generative Data Refinement: Just Ask for Better Data
von: Jiang, Minqi, et al.
Veröffentlicht: (2025)
von: Jiang, Minqi, et al.
Veröffentlicht: (2025)
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
von: Qin, Guanghui, et al.
Veröffentlicht: (2021)
von: Qin, Guanghui, et al.
Veröffentlicht: (2021)
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
von: Edwards, Nicholas, et al.
Veröffentlicht: (2026)
von: Edwards, Nicholas, et al.
Veröffentlicht: (2026)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
von: Xie, Qiming, et al.
Veröffentlicht: (2023)
von: Xie, Qiming, et al.
Veröffentlicht: (2023)
To Ask or Not to Ask: Learning to Require Human Feedback
von: Pugnana, Andrea, et al.
Veröffentlicht: (2025)
von: Pugnana, Andrea, et al.
Veröffentlicht: (2025)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
von: Rohera, Pritika, et al.
Veröffentlicht: (2025)
von: Rohera, Pritika, et al.
Veröffentlicht: (2025)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
von: Deng, Yihe, et al.
Veröffentlicht: (2023)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
CLARINET: Augmenting Language Models to Ask Clarification Questions for Retrieval
von: Chi, Yizhou, et al.
Veröffentlicht: (2024)
von: Chi, Yizhou, et al.
Veröffentlicht: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
von: Huang, Fanding, et al.
Veröffentlicht: (2025)
The Unlearnability Phenomenon in RLVR for Language Models
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
What Language is This? Ask Your Tokenizer
von: Meister, Clara, et al.
Veröffentlicht: (2026)
von: Meister, Clara, et al.
Veröffentlicht: (2026)
Rewards as Labels: Revisiting RLVR from a Classification Perspective
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
von: Zhai, Zepeng, et al.
Veröffentlicht: (2026)
DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2026)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
von: Yan, Xue, et al.
Veröffentlicht: (2023)
von: Yan, Xue, et al.
Veröffentlicht: (2023)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness
von: Deng, Haotian, et al.
Veröffentlicht: (2025)
von: Deng, Haotian, et al.
Veröffentlicht: (2025)
Self-Distilled RLVR
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
von: Yang, Chenxu, et al.
Veröffentlicht: (2026)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
von: Zhou, Zhanke, et al.
Veröffentlicht: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
von: Lee, Chanuk, et al.
Veröffentlicht: (2026)
Why Ask One When You Can Ask $k$? Learning-to-Defer to the Top-$k$ Experts
von: Montreuil, Yannis, et al.
Veröffentlicht: (2025)
von: Montreuil, Yannis, et al.
Veröffentlicht: (2025)
Learning to Ask: When LLM Agents Meet Unclear Instruction
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
Generalization of RLVR Using Causal Reasoning as a Testbed
von: Lu, Brian, et al.
Veröffentlicht: (2025)
von: Lu, Brian, et al.
Veröffentlicht: (2025)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
von: Xu, Ran, et al.
Veröffentlicht: (2026)
von: Xu, Ran, et al.
Veröffentlicht: (2026)
The Invisible Leash: Why RLVR May or May Not Escape Its Origin
von: Wu, Fang, et al.
Veröffentlicht: (2025)
von: Wu, Fang, et al.
Veröffentlicht: (2025)
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
When to Ask a Question: Understanding Communication Strategies in Generative AI Tools
von: Park, Charlotte, et al.
Veröffentlicht: (2026)
von: Park, Charlotte, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Rubric Anchors
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
Ask, and it shall be given: On the Turing completeness of prompting
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
Learning When to Ask: Simulation-Trained Humanoids for Mental-Health Diagnosis
von: Cenacchi, Filippo, et al.
Veröffentlicht: (2025)
von: Cenacchi, Filippo, et al.
Veröffentlicht: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Linear Dynamics in the RLVR Training of Large Language Models
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
von: Wang, Tianle, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
von: Gulati, Anmol, et al.
Veröffentlicht: (2026) -
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
von: Kumar, Abhijit, et al.
Veröffentlicht: (2026) -
IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
von: Sharma, Karun, et al.
Veröffentlicht: (2026) -
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024) -
Generative Data Refinement: Just Ask for Better Data
von: Jiang, Minqi, et al.
Veröffentlicht: (2025)