Gespeichert in:
| 1. Verfasser: | Li, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.06813 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
von: Wang, Junzhe, et al.
Veröffentlicht: (2026)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
von: Li, Yushu, et al.
Veröffentlicht: (2026)
von: Li, Yushu, et al.
Veröffentlicht: (2026)
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation
von: Tan, Heng, et al.
Veröffentlicht: (2025)
von: Tan, Heng, et al.
Veröffentlicht: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
Policy-Guided Search on Tree-of-Thoughts for Efficient Problem Solving with Bounded Language Model Queries
von: Pendurkar, Sumedh, et al.
Veröffentlicht: (2026)
von: Pendurkar, Sumedh, et al.
Veröffentlicht: (2026)
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
Offline Model-Based Optimization via Policy-Guided Gradient Search
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search
von: Li, Shuangtao, et al.
Veröffentlicht: (2025)
von: Li, Shuangtao, et al.
Veröffentlicht: (2025)
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
von: Wang, Yongqi, et al.
Veröffentlicht: (2026)
von: Wang, Yongqi, et al.
Veröffentlicht: (2026)
LLM Reasoning with Process Rewards for Outcome-Guided Steps
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
von: Rezaei, Mohammad, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
von: Mou, Zhiyu, et al.
Veröffentlicht: (2025)
von: Mou, Zhiyu, et al.
Veröffentlicht: (2025)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Continuous Optimization for Feature Selection with Permutation-Invariant Embedding and Policy-Guided Search
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
In Search of Trees: Decision-Tree Policy Synthesis for Black-Box Systems via Search
von: Demirović, Emir, et al.
Veröffentlicht: (2024)
von: Demirović, Emir, et al.
Veröffentlicht: (2024)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2025)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
von: Li, Jiawei, et al.
Veröffentlicht: (2024)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Tree Search for LLM Agent Reinforcement Learning
von: Ji, Yuxiang, et al.
Veröffentlicht: (2025)
von: Ji, Yuxiang, et al.
Veröffentlicht: (2025)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization
von: Liu, Shengchao, et al.
Veröffentlicht: (2025)
von: Liu, Shengchao, et al.
Veröffentlicht: (2025)
ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
von: Yin, Chenlong, et al.
Veröffentlicht: (2025)
von: Yin, Chenlong, et al.
Veröffentlicht: (2025)
LiteSearch: Efficacious Tree Search for LLM
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Variance-Aware Prior-Based Tree Policies for Monte Carlo Tree Search
von: Weichart, Maximilian
Veröffentlicht: (2025)
von: Weichart, Maximilian
Veröffentlicht: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
von: Mao, Yixiu, et al.
Veröffentlicht: (2026)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
von: Han, Dongge, et al.
Veröffentlicht: (2025)
von: Han, Dongge, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026) -
Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
von: Wang, Junzhe, et al.
Veröffentlicht: (2026) -
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025) -
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
von: Li, Yushu, et al.
Veröffentlicht: (2026) -
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)