Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ziniu, Chen, Congliang, Yang, Tianyun, Ding, Tian, Sun, Ruoyu, Zhang, Ge, Huang, Wenhao, Luo, Zhi-Quan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Transformers Need Adam: A Hessian Perspective
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives
von: Liu, Yajiao, et al.
Veröffentlicht: (2025)
von: Liu, Yajiao, et al.
Veröffentlicht: (2025)
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
Off-Policy Value-Based Reinforcement Learning for Large Language Models
von: Wang, Peng-Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Peng-Yuan, et al.
Veröffentlicht: (2026)
Adam-mini: Use Fewer Learning Rates To Gain More
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
von: Zhang, Yushun, et al.
Veröffentlicht: (2024)
ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
von: Wang, Zhuohan, et al.
Veröffentlicht: (2025)
von: Wang, Zhuohan, et al.
Veröffentlicht: (2025)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Adam Converges Without Any Modification On Update Rules
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
von: Zhang, Yushun, et al.
Veröffentlicht: (2026)
Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
von: Yang, Tianyun, et al.
Veröffentlicht: (2025)
von: Yang, Tianyun, et al.
Veröffentlicht: (2025)
BEAR: Budgeted Evidence Allocation for Multi-Document Reasoning
von: Sun, Lin, et al.
Veröffentlicht: (2026)
von: Sun, Lin, et al.
Veröffentlicht: (2026)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
von: Liu, Mingjie, et al.
Veröffentlicht: (2025)
CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation
von: Petullo, James, et al.
Veröffentlicht: (2026)
von: Petullo, James, et al.
Veröffentlicht: (2026)
Knapsack Optimization-based Schema Linking for LLM-based Text-to-SQL Generation
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
von: Yuan, Zheng, et al.
Veröffentlicht: (2025)
SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
von: Sun, Huashan, et al.
Veröffentlicht: (2025)
von: Sun, Huashan, et al.
Veröffentlicht: (2025)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing LLMs
von: Kong, Chuyi, et al.
Veröffentlicht: (2024)
von: Kong, Chuyi, et al.
Veröffentlicht: (2024)
The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
von: Pouplin, Thomas, et al.
Veröffentlicht: (2024)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
von: Jang, Eyon, et al.
Veröffentlicht: (2026)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiangxiang, et al.
Veröffentlicht: (2026)
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
von: Li, Ziniu, et al.
Veröffentlicht: (2023)
von: Li, Ziniu, et al.
Veröffentlicht: (2023)
FuseRL: Dense Preference Optimization for Heterogeneous Model Fusion
von: Zhong, Longguang, et al.
Veröffentlicht: (2025)
von: Zhong, Longguang, et al.
Veröffentlicht: (2025)
Zero Token-Driven Deep Thinking in LLMs: Unlocking the Full Potential of Existing Parameters via Cyclic Refinement
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
von: Hu, Jian, et al.
Veröffentlicht: (2025)
von: Hu, Jian, et al.
Veröffentlicht: (2025)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
von: Yang, Hongming, et al.
Veröffentlicht: (2025)
von: Yang, Hongming, et al.
Veröffentlicht: (2025)
Self-Evolving Critique Abilities in Large Language Models
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
von: Hou, Zhenyu, et al.
Veröffentlicht: (2025)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
von: Wei, Zhepei, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Budget Allocation for Adaptive Test-Time Reasoning
von: Nguyen, Manh, et al.
Veröffentlicht: (2026)
von: Nguyen, Manh, et al.
Veröffentlicht: (2026)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
Teaching Language Models to Reason with Tools
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
von: Li, Chengpeng, et al.
Veröffentlicht: (2025)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
UIO-LLMs: Unbiased Incremental Optimization for Long-Context LLMs
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis
von: Xu, Tian, et al.
Veröffentlicht: (2022)
von: Xu, Tian, et al.
Veröffentlicht: (2022)
OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling
von: Chen, Yitian, et al.
Veröffentlicht: (2026)
von: Chen, Yitian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Why Transformers Need Adam: A Hessian Perspective
von: Zhang, Yushun, et al.
Veröffentlicht: (2024) -
Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives
von: Liu, Yajiao, et al.
Veröffentlicht: (2025) -
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
von: Li, Ziniu, et al.
Veröffentlicht: (2024) -
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
von: Deng, Wenhao, et al.
Veröffentlicht: (2025) -
Off-Policy Value-Based Reinforcement Learning for Large Language Models
von: Wang, Peng-Yuan, et al.
Veröffentlicht: (2026)