Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhicheng, Guo, Zhijiang, Huang, Yinya, Wang, Yongxin, Xie, Dongchun, Li, Hanhui, Wang, Yiwei, Liang, Xiaodan, Tang, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
AlignedCoT: Prompting Large Language Models via Native-Speaking Demonstrations
by: Yang, Zhicheng, et al.
Published: (2023)
by: Yang, Zhicheng, et al.
Published: (2023)
OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling
by: Yang, Zhicheng, et al.
Published: (2024)
by: Yang, Zhicheng, et al.
Published: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
by: Polonuer, Joaquín, et al.
Published: (2026)
by: Polonuer, Joaquín, et al.
Published: (2026)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
by: Hu, Hanxu, et al.
Published: (2026)
by: Hu, Hanxu, et al.
Published: (2026)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
by: Xiang, Kun, et al.
Published: (2026)
by: Xiang, Kun, et al.
Published: (2026)
FormalAlign: Automated Alignment Evaluation for Autoformalization
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
ORMind: A Cognitive-Inspired End-to-End Reasoning Framework for Operations Research
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
by: Dragoi, Marius, et al.
Published: (2025)
by: Dragoi, Marius, et al.
Published: (2025)
ATG: Benchmarking Automated Theorem Generation for Generative Language Models
by: Lin, Xiaohan, et al.
Published: (2024)
by: Lin, Xiaohan, et al.
Published: (2024)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
by: Huang, Zhuoxu, et al.
Published: (2026)
by: Huang, Zhuoxu, et al.
Published: (2026)
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
by: Ding, Yue, et al.
Published: (2025)
by: Ding, Yue, et al.
Published: (2025)
3D Visibility-aware Generalizable Neural Radiance Fields for Interacting Hands
by: Huang, Xuan, et al.
Published: (2024)
by: Huang, Xuan, et al.
Published: (2024)
Proving Theorems Recursively
by: Wang, Haiming, et al.
Published: (2024)
by: Wang, Haiming, et al.
Published: (2024)
Process-Driven Autoformalization in Lean 4
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
by: He, Paul, et al.
Published: (2026)
by: He, Paul, et al.
Published: (2026)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
by: Wang, Yongxin, et al.
Published: (2025)
by: Wang, Yongxin, et al.
Published: (2025)
Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
How Do Firms' Innovation Failures Affect the Quality of Subsequent Innovations? The Contingency Role of Knowledge Breadth and Depth
by: Guocai Chen, et al.
Published: (2024)
by: Guocai Chen, et al.
Published: (2024)
Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
by: Gao, Junqi, et al.
Published: (2025)
by: Gao, Junqi, et al.
Published: (2025)
Adaptive Negative Reinforcement for LLM Reasoning:Dynamically Balancing Correction and Diversity in RLVR
by: Ingle, Yash, et al.
Published: (2026)
by: Ingle, Yash, et al.
Published: (2026)
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
by: Robertson, Alex, et al.
Published: (2026)
by: Robertson, Alex, et al.
Published: (2026)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
by: Chen, Xuetian, et al.
Published: (2025)
by: Chen, Xuetian, et al.
Published: (2025)
DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
Breadth, Depth, and Flux of Course-Prerequisite Networks
by: Zuev, Konstantin, et al.
Published: (2025)
by: Zuev, Konstantin, et al.
Published: (2025)
The Depth and Breadth of Google Scholar: An Empirical Study
by: Neuhaus, Chris, et al.
Published: (2006)
by: Neuhaus, Chris, et al.
Published: (2006)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
by: Zhu, Yihua, et al.
Published: (2026)
by: Zhu, Yihua, et al.
Published: (2026)
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
by: Xiang, Kun, et al.
Published: (2025)
by: Xiang, Kun, et al.
Published: (2025)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
by: Xiang, Kun, et al.
Published: (2024)
by: Xiang, Kun, et al.
Published: (2024)
Similar Items
-
Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
by: Yang, Zhicheng, et al.
Published: (2026) -
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
by: Yang, Zhicheng, et al.
Published: (2025) -
TreeRPO: Tree Relative Policy Optimization
by: Yang, Zhicheng, et al.
Published: (2025) -
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026) -
AlignedCoT: Prompting Large Language Models via Native-Speaking Demonstrations
by: Yang, Zhicheng, et al.
Published: (2023)