Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qingyang, Wu, Haitao, Zhang, Changqing, Zhao, Peilin, Bian, Yatao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
COME: Test-time adaption by Conservatively Minimizing Entropy
by: Zhang, Qingyang, et al.
Published: (2024)
by: Zhang, Qingyang, et al.
Published: (2024)
The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection
by: Zhang, Qingyang, et al.
Published: (2024)
by: Zhang, Qingyang, et al.
Published: (2024)
Energy-Guided Generative Modeling for Low-Energy Molecular Structure Discovery
by: Xu, Guikun, et al.
Published: (2025)
by: Xu, Guikun, et al.
Published: (2025)
ETDock: A Novel Equivariant Transformer for Protein-Ligand Docking
by: Yi, Yiqiang, et al.
Published: (2023)
by: Yi, Yiqiang, et al.
Published: (2023)
TEMPO: Scaling Test-time Training for Large Reasoning Models
by: Zhang, Qingyang, et al.
Published: (2026)
by: Zhang, Qingyang, et al.
Published: (2026)
Position: Graph Foundation Models are Already Here
by: Mao, Haitao, et al.
Published: (2024)
by: Mao, Haitao, et al.
Published: (2024)
Adapting Large Language Models for Content Moderation: Pitfalls in Data Engineering and Supervised Fine-tuning
by: Ma, Huan, et al.
Published: (2023)
by: Ma, Huan, et al.
Published: (2023)
Incentivizing LLMs to Self-Verify Their Answers
by: Zhang, Fuxiang, et al.
Published: (2025)
by: Zhang, Fuxiang, et al.
Published: (2025)
Enhancing Neural Subset Selection: Integrating Background Information into Set Representations
by: Xie, Binghui, et al.
Published: (2024)
by: Xie, Binghui, et al.
Published: (2024)
Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented Generation
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
Learning to Learn with Contrastive Meta-Objective
by: Wu, Shiguang, et al.
Published: (2024)
by: Wu, Shiguang, et al.
Published: (2024)
CrystalDiT: A Diffusion Transformer for Crystal Generation
by: Yi, Xiaohan, et al.
Published: (2025)
by: Yi, Xiaohan, et al.
Published: (2025)
Graph Unitary Message Passing
by: Qiu, Haiquan, et al.
Published: (2024)
by: Qiu, Haiquan, et al.
Published: (2024)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
by: Liu, Li, et al.
Published: (2024)
by: Liu, Li, et al.
Published: (2024)
Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
by: Chen, Zhikang, et al.
Published: (2025)
by: Chen, Zhikang, et al.
Published: (2025)
PACER: A Fully Push-forward-based Distributional Reinforcement Learning Algorithm
by: Bai, Wensong, et al.
Published: (2023)
by: Bai, Wensong, et al.
Published: (2023)
How Interpretable Are Interpretable Graph Neural Networks?
by: Chen, Yongqiang, et al.
Published: (2024)
by: Chen, Yongqiang, et al.
Published: (2024)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
by: Shen, Si, et al.
Published: (2025)
by: Shen, Si, et al.
Published: (2025)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Unsupervised Morphological Tree Tokenizer
by: Zhu, Qingyang, et al.
Published: (2024)
by: Zhu, Qingyang, et al.
Published: (2024)
CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answering
by: Liu, Yu, et al.
Published: (2026)
by: Liu, Yu, et al.
Published: (2026)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
by: Kwon, Omin, et al.
Published: (2026)
by: Kwon, Omin, et al.
Published: (2026)
If Concept Bottlenecks are the Question, are Foundation Models the Answer?
by: Debole, Nicola, et al.
Published: (2025)
by: Debole, Nicola, et al.
Published: (2025)
HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment
by: Chen, Yongqiang, et al.
Published: (2024)
by: Chen, Yongqiang, et al.
Published: (2024)
One Fits All: Learning Fair Graph Neural Networks for Various Sensitive Attributes
by: Zhu, Yuchang, et al.
Published: (2024)
by: Zhu, Yuchang, et al.
Published: (2024)
Communications-Incentivized Collaborative Reasoning in NetGPT through Agentic Reinforcement Learning
by: Yu, Xiaoxue, et al.
Published: (2026)
by: Yu, Xiaoxue, et al.
Published: (2026)
Out-of-Distribution Detection Methods Answer the Wrong Questions
by: Li, Yucen Lily, et al.
Published: (2025)
by: Li, Yucen Lily, et al.
Published: (2025)
W2T: LoRA Weights Already Know What They Can Do
by: Han, Xiaolong, et al.
Published: (2026)
by: Han, Xiaolong, et al.
Published: (2026)
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
by: Zhang, Weizhi, et al.
Published: (2025)
by: Zhang, Weizhi, et al.
Published: (2025)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
by: Xu, Ran, et al.
Published: (2025)
by: Xu, Ran, et al.
Published: (2025)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
by: Agrawal, Aakriti, et al.
Published: (2025)
by: Agrawal, Aakriti, et al.
Published: (2025)
Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoning
by: Wu, Yicong, et al.
Published: (2025)
by: Wu, Yicong, et al.
Published: (2025)
CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
by: Feng, Ruixiang, et al.
Published: (2025)
by: Feng, Ruixiang, et al.
Published: (2025)
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains
by: Chu, Xu, et al.
Published: (2025)
by: Chu, Xu, et al.
Published: (2025)
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
by: Fallahpour, Adibvafa, et al.
Published: (2025)
by: Fallahpour, Adibvafa, et al.
Published: (2025)
Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
by: Dong, Ziming, et al.
Published: (2026)
by: Dong, Ziming, et al.
Published: (2026)
Harnessing Reasoning Trajectories for Hallucination Detection via Answer-agreement Representation Shaping
by: Zhang, Jianxiong, et al.
Published: (2026)
by: Zhang, Jianxiong, et al.
Published: (2026)
SGFormer: Simplifying and Empowering Transformers for Large-Graph Representations
by: Wu, Qitian, et al.
Published: (2023)
by: Wu, Qitian, et al.
Published: (2023)
A Keyword-Based Technique to Evaluate Broad Question Answer Script
by: Mahmud, Tamim Al, et al.
Published: (2025)
by: Mahmud, Tamim Al, et al.
Published: (2025)
Similar Items
-
COME: Test-time adaption by Conservatively Minimizing Entropy
by: Zhang, Qingyang, et al.
Published: (2024) -
The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection
by: Zhang, Qingyang, et al.
Published: (2024) -
Energy-Guided Generative Modeling for Low-Energy Molecular Structure Discovery
by: Xu, Guikun, et al.
Published: (2025) -
ETDock: A Novel Equivariant Transformer for Protein-Ligand Docking
by: Yi, Yiqiang, et al.
Published: (2023) -
TEMPO: Scaling Test-time Training for Large Reasoning Models
by: Zhang, Qingyang, et al.
Published: (2026)