HyPER: Bridging Exploration and Exploitation for Scalable LLM Reasoning with Hypothesis Path Expansion and Reduction
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Shengxuan, Huang, Haochen, Zhong, Shuzhang, Zuo, Pengfei, Li, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
by: Zhong, Shuzhang, et al.
Published: (2026)
by: Zhong, Shuzhang, et al.
Published: (2026)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement
by: Pasios, Stefanos, et al.
Published: (2026)
by: Pasios, Stefanos, et al.
Published: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Robust Answers, Fragile Logic: Probing the Decoupling Hypothesis in LLM Reasoning
by: Jiang, Enyi, et al.
Published: (2025)
by: Jiang, Enyi, et al.
Published: (2025)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
by: Tang, Shengji, et al.
Published: (2025)
by: Tang, Shengji, et al.
Published: (2025)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Query Decomposition for RAG: Balancing Exploration-Exploitation
by: Petcu, Roxana, et al.
Published: (2025)
by: Petcu, Roxana, et al.
Published: (2025)
HyKGE: A Hypothesis Knowledge Graph Enhanced Framework for Accurate and Reliable Medical LLMs Responses
by: Jiang, Xinke, et al.
Published: (2023)
by: Jiang, Xinke, et al.
Published: (2023)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
Scale-Adaptive Balancing of Exploration and Exploitation in Classical Planning
by: Wissow, Stephen, et al.
Published: (2023)
by: Wissow, Stephen, et al.
Published: (2023)
Exploration and Exploitation Errors Are Measurable for Language Model Agents
by: Park, Jaden, et al.
Published: (2026)
by: Park, Jaden, et al.
Published: (2026)
Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
Revisiting the UID Hypothesis in LLM Reasoning Traces
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
Structural Induced Exploration for Balanced and Scalable Multi-Robot Path Planning
by: Guo, Zikun, et al.
Published: (2025)
by: Guo, Zikun, et al.
Published: (2025)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
by: Liu, Jia, et al.
Published: (2025)
by: Liu, Jia, et al.
Published: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Scalable and Accurate Graph Reasoning with LLM-based Multi-Agents
by: Hu, Yuwei, et al.
Published: (2024)
by: Hu, Yuwei, et al.
Published: (2024)
Temporal Inductive Path Neural Network for Temporal Knowledge Graph Reasoning
by: Dong, Hao, et al.
Published: (2023)
by: Dong, Hao, et al.
Published: (2023)
HyEvo: Self-Evolving Hybrid Agentic Workflows for Efficient Reasoning
by: Xu, Beibei, et al.
Published: (2026)
by: Xu, Beibei, et al.
Published: (2026)
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
by: Gao, Xianjun, et al.
Published: (2025)
by: Gao, Xianjun, et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
by: Yang, Lu, et al.
Published: (2026)
by: Yang, Lu, et al.
Published: (2026)
An Improved FOX Optimization Algorithm Using Adaptive Exploration and Exploitation for Global Optimization
by: Jumaah, Mahmood A., et al.
Published: (2025)
by: Jumaah, Mahmood A., et al.
Published: (2025)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
by: Lam, Phong, et al.
Published: (2026)
by: Lam, Phong, et al.
Published: (2026)
KnowPath: Knowledge-enhanced Reasoning via LLM-generated Inference Paths over Knowledge Graphs
by: Zhao, Qi, et al.
Published: (2025)
by: Zhao, Qi, et al.
Published: (2025)
ReasonBridge: Efficient Reasoning Transfer from Closed to Open-Source Language Models
by: Zhong, Ziqi, et al.
Published: (2025)
by: Zhong, Ziqi, et al.
Published: (2025)
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
by: Hu, Wenbin, et al.
Published: (2025)
by: Hu, Wenbin, et al.
Published: (2025)
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
by: Chen, Yixiang, et al.
Published: (2025)
by: Chen, Yixiang, et al.
Published: (2025)
Similar Items
-
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
by: Zhong, Shuzhang, et al.
Published: (2026) -
WESE: Weak Exploration to Strong Exploitation for LLM Agents
by: Huang, Xu, et al.
Published: (2024) -
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement
by: Pasios, Stefanos, et al.
Published: (2026) -
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026) -
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)