HyPER: Bridging Exploration and Exploitation for Scalable LLM Reasoning with Hypothesis Path Expansion and Reduction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiu, Shengxuan, Huang, Haochen, Zhong, Shuzhang, Zuo, Pengfei, Li, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2026)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2026)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement
von: Pasios, Stefanos, et al.
Veröffentlicht: (2026)
von: Pasios, Stefanos, et al.
Veröffentlicht: (2026)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
von: Su, Xuerui, et al.
Veröffentlicht: (2025)
von: Su, Xuerui, et al.
Veröffentlicht: (2025)
Robust Answers, Fragile Logic: Probing the Decoupling Hypothesis in LLM Reasoning
von: Jiang, Enyi, et al.
Veröffentlicht: (2025)
von: Jiang, Enyi, et al.
Veröffentlicht: (2025)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
von: Liang, Kun, et al.
Veröffentlicht: (2026)
von: Liang, Kun, et al.
Veröffentlicht: (2026)
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
von: Tang, Shengji, et al.
Veröffentlicht: (2025)
von: Tang, Shengji, et al.
Veröffentlicht: (2025)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
Query Decomposition for RAG: Balancing Exploration-Exploitation
von: Petcu, Roxana, et al.
Veröffentlicht: (2025)
von: Petcu, Roxana, et al.
Veröffentlicht: (2025)
HyKGE: A Hypothesis Knowledge Graph Enhanced Framework for Accurate and Reliable Medical LLMs Responses
von: Jiang, Xinke, et al.
Veröffentlicht: (2023)
von: Jiang, Xinke, et al.
Veröffentlicht: (2023)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
von: Dai, Zhenwen, et al.
Veröffentlicht: (2024)
Exploitation Is All You Need... for Exploration
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
Scale-Adaptive Balancing of Exploration and Exploitation in Classical Planning
von: Wissow, Stephen, et al.
Veröffentlicht: (2023)
von: Wissow, Stephen, et al.
Veröffentlicht: (2023)
Exploration and Exploitation Errors Are Measurable for Language Model Agents
von: Park, Jaden, et al.
Veröffentlicht: (2026)
von: Park, Jaden, et al.
Veröffentlicht: (2026)
Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation
von: Zhang, Weiming, et al.
Veröffentlicht: (2026)
von: Zhang, Weiming, et al.
Veröffentlicht: (2026)
Revisiting the UID Hypothesis in LLM Reasoning Traces
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
Structural Induced Exploration for Balanced and Scalable Multi-Robot Path Planning
von: Guo, Zikun, et al.
Veröffentlicht: (2025)
von: Guo, Zikun, et al.
Veröffentlicht: (2025)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
von: Liu, Jia, et al.
Veröffentlicht: (2025)
von: Liu, Jia, et al.
Veröffentlicht: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
Scalable and Accurate Graph Reasoning with LLM-based Multi-Agents
von: Hu, Yuwei, et al.
Veröffentlicht: (2024)
von: Hu, Yuwei, et al.
Veröffentlicht: (2024)
Temporal Inductive Path Neural Network for Temporal Knowledge Graph Reasoning
von: Dong, Hao, et al.
Veröffentlicht: (2023)
von: Dong, Hao, et al.
Veröffentlicht: (2023)
HyEvo: Self-Evolving Hybrid Agentic Workflows for Efficient Reasoning
von: Xu, Beibei, et al.
Veröffentlicht: (2026)
von: Xu, Beibei, et al.
Veröffentlicht: (2026)
Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
von: Gao, Xianjun, et al.
Veröffentlicht: (2025)
von: Gao, Xianjun, et al.
Veröffentlicht: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
von: Su, Xun, et al.
Veröffentlicht: (2025)
von: Su, Xun, et al.
Veröffentlicht: (2025)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
von: Yang, Lu, et al.
Veröffentlicht: (2026)
von: Yang, Lu, et al.
Veröffentlicht: (2026)
An Improved FOX Optimization Algorithm Using Adaptive Exploration and Exploitation for Global Optimization
von: Jumaah, Mahmood A., et al.
Veröffentlicht: (2025)
von: Jumaah, Mahmood A., et al.
Veröffentlicht: (2025)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
von: Kim, JaeYoon, et al.
Veröffentlicht: (2024)
von: Kim, JaeYoon, et al.
Veröffentlicht: (2024)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
von: Lam, Phong, et al.
Veröffentlicht: (2026)
von: Lam, Phong, et al.
Veröffentlicht: (2026)
KnowPath: Knowledge-enhanced Reasoning via LLM-generated Inference Paths over Knowledge Graphs
von: Zhao, Qi, et al.
Veröffentlicht: (2025)
von: Zhao, Qi, et al.
Veröffentlicht: (2025)
ReasonBridge: Efficient Reasoning Transfer from Closed to Open-Source Language Models
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
von: Guo, Weiyang, et al.
Veröffentlicht: (2026)
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
von: Hu, Wenbin, et al.
Veröffentlicht: (2025)
von: Hu, Wenbin, et al.
Veröffentlicht: (2025)
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xuan, et al.
Veröffentlicht: (2025)
SELF-REDRAFT: Eliciting Intrinsic Exploration-Exploitation Balance in Test-Time Scaling for Code Generation
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
von: Chen, Yixiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2026) -
WESE: Weak Exploration to Strong Exploitation for LLM Agents
von: Huang, Xu, et al.
Veröffentlicht: (2024) -
HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement
von: Pasios, Stefanos, et al.
Veröffentlicht: (2026) -
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
von: Huang, Kexin, et al.
Veröffentlicht: (2026) -
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
von: Su, Xuerui, et al.
Veröffentlicht: (2025)