Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Runze, Wang, Jiakang, Shi, Yuling, Xie, Zhihui, An, Chenxin, Zhang, Kaiyan, Zhao, Jian, Gu, Xiaodong, Lin, Lei, Hu, Wenping, Li, Xiu, Zhang, Fuzheng, Zhou, Guorui, Gai, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
von: Wang, Jiakang, et al.
Veröffentlicht: (2025)
von: Wang, Jiakang, et al.
Veröffentlicht: (2025)
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
von: Wang, Jiakang, et al.
Veröffentlicht: (2025)
von: Wang, Jiakang, et al.
Veröffentlicht: (2025)
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
Leanabell-Prover: Posttraining Scaling in Formal Reasoning
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
von: Ji, Xingguang, et al.
Veröffentlicht: (2025)
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
von: Zhao, Jian, et al.
Veröffentlicht: (2025)
von: Zhao, Jian, et al.
Veröffentlicht: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2025)
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
von: Fang, Yixiong, et al.
Veröffentlicht: (2025)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
von: Ou, Jiao, et al.
Veröffentlicht: (2024)
von: Ou, Jiao, et al.
Veröffentlicht: (2024)
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
von: Li, Huan, et al.
Veröffentlicht: (2025)
von: Li, Huan, et al.
Veröffentlicht: (2025)
MedForge: Interpretable Medical Deepfake Detection via Forgery-aware Reasoning
von: Chen, Zhihui, et al.
Veröffentlicht: (2026)
von: Chen, Zhihui, et al.
Veröffentlicht: (2026)
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
von: Lin, Lei, et al.
Veröffentlicht: (2023)
von: Lin, Lei, et al.
Veröffentlicht: (2023)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
von: Liu, Runze, et al.
Veröffentlicht: (2025)
von: Liu, Runze, et al.
Veröffentlicht: (2025)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
von: Xie, Can, et al.
Veröffentlicht: (2025)
von: Xie, Can, et al.
Veröffentlicht: (2025)
DialogBench: Evaluating LLMs as Human-like Dialogue Systems
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
von: Ou, Jiao, et al.
Veröffentlicht: (2023)
ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning
von: Nie, Shuaiyi, et al.
Veröffentlicht: (2026)
von: Nie, Shuaiyi, et al.
Veröffentlicht: (2026)
ReviewRL: Towards Automated Scientific Review with RL
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
Rethinking Code Complexity Through the Lens of Large Language Models
von: Xie, Chen, et al.
Veröffentlicht: (2026)
von: Xie, Chen, et al.
Veröffentlicht: (2026)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
Between Lines of Code: Unraveling the Distinct Patterns of Machine and Human Programmers
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning
von: Liu, Yihao, et al.
Veröffentlicht: (2025)
von: Liu, Yihao, et al.
Veröffentlicht: (2025)
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers
von: Huang, Chongxuan, et al.
Veröffentlicht: (2026)
von: Huang, Chongxuan, et al.
Veröffentlicht: (2026)
HoME: Hierarchy of Multi-Gate Experts for Multi-Task Learning at Kuaishou
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
von: Zeng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zeng, Wenhao, et al.
Veröffentlicht: (2025)
Improving RL Exploration for LLM Reasoning through Retrospective Replay
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models
von: Sun, Yuchong, et al.
Veröffentlicht: (2023)
von: Sun, Yuchong, et al.
Veröffentlicht: (2023)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Small Agent Can Also Rock! Empowering Small Language Models as Hallucination Detector
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Cheng, Xiaoxue, et al.
Veröffentlicht: (2024)
PROMISE: Process Reward Models Unlock Test-Time Scaling Laws in Generative Recommendations
von: Guo, Chengcheng, et al.
Veröffentlicht: (2026)
von: Guo, Chengcheng, et al.
Veröffentlicht: (2026)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
von: Fu, Jia, et al.
Veröffentlicht: (2025)
von: Fu, Jia, et al.
Veröffentlicht: (2025)
LongCodeZip: Compress Long Context for Code Language Models
von: Shi, Yuling, et al.
Veröffentlicht: (2025)
von: Shi, Yuling, et al.
Veröffentlicht: (2025)
ChorusCVR: Chorus Supervision for Entire Space Post-Click Conversion Rate Modeling
von: Cheng, Wei, et al.
Veröffentlicht: (2025)
von: Cheng, Wei, et al.
Veröffentlicht: (2025)
Exploration and Anti-Exploration with Distributional Random Network Distillation
von: Yang, Kai, et al.
Veröffentlicht: (2024)
von: Yang, Kai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
von: Wang, Jiakang, et al.
Veröffentlicht: (2025) -
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
von: Wang, Jiakang, et al.
Veröffentlicht: (2025) -
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025) -
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025) -
Leanabell-Prover: Posttraining Scaling in Formal Reasoning
von: Zhang, Jingyuan, et al.
Veröffentlicht: (2025)