Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Tianyi, Li, Long, Guo, Hongcan, Chen, Yibiao, Li, Yixia, Wang, Yong, Chen, Yun, Chen, Guanhua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
di: Wang, Tianyi, et al.
Pubblicazione: (2026)
di: Wang, Tianyi, et al.
Pubblicazione: (2026)
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
di: Li, Yixia, et al.
Pubblicazione: (2024)
di: Li, Yixia, et al.
Pubblicazione: (2024)
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
di: Wang, Wenkai, et al.
Pubblicazione: (2025)
di: Wang, Wenkai, et al.
Pubblicazione: (2025)
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
di: Li, Zeping, et al.
Pubblicazione: (2026)
di: Li, Zeping, et al.
Pubblicazione: (2026)
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning
di: Li, Zhicong, et al.
Pubblicazione: (2026)
di: Li, Zhicong, et al.
Pubblicazione: (2026)
JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification
di: Wang, Xi, et al.
Pubblicazione: (2026)
di: Wang, Xi, et al.
Pubblicazione: (2026)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
di: Chen, Zhifei, et al.
Pubblicazione: (2025)
di: Chen, Zhifei, et al.
Pubblicazione: (2025)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
di: Sun, Han, et al.
Pubblicazione: (2026)
di: Sun, Han, et al.
Pubblicazione: (2026)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
di: Wu, Yuning, et al.
Pubblicazione: (2026)
di: Wu, Yuning, et al.
Pubblicazione: (2026)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization
di: Zhang, Siwei, et al.
Pubblicazione: (2026)
di: Zhang, Siwei, et al.
Pubblicazione: (2026)
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
di: Wang, Boxiao, et al.
Pubblicazione: (2026)
di: Wang, Boxiao, et al.
Pubblicazione: (2026)
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
di: Liu, Jinmei, et al.
Pubblicazione: (2026)
di: Liu, Jinmei, et al.
Pubblicazione: (2026)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
di: Meng, Xiangtao, et al.
Pubblicazione: (2025)
di: Meng, Xiangtao, et al.
Pubblicazione: (2025)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
di: Wu, Wanxing, et al.
Pubblicazione: (2026)
di: Wu, Wanxing, et al.
Pubblicazione: (2026)
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
di: Li, Yu, et al.
Pubblicazione: (2026)
di: Li, Yu, et al.
Pubblicazione: (2026)
DREAM: Diffusion Rectification and Estimation-Adaptive Models
di: Zhou, Jinxin, et al.
Pubblicazione: (2023)
di: Zhou, Jinxin, et al.
Pubblicazione: (2023)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
di: Wang, Youkang, et al.
Pubblicazione: (2025)
di: Wang, Youkang, et al.
Pubblicazione: (2025)
SPAR: Support-Preserving Action Rectification
di: Zhao, Jiaxin, et al.
Pubblicazione: (2026)
di: Zhao, Jiaxin, et al.
Pubblicazione: (2026)
Geometric Manifold Rectification for Imbalanced Learning
di: Wang, Xubin, et al.
Pubblicazione: (2026)
di: Wang, Xubin, et al.
Pubblicazione: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026)
di: Lai, Peng, et al.
Pubblicazione: (2026)
Annealing Self-Distillation Rectification Improves Adversarial Training
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification
di: Luo, Yaoyang, et al.
Pubblicazione: (2026)
di: Luo, Yaoyang, et al.
Pubblicazione: (2026)
Efficient Exploration for Iterative Nash Preference Optimization
di: Nan, Tianlong, et al.
Pubblicazione: (2026)
di: Nan, Tianlong, et al.
Pubblicazione: (2026)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
di: Xiong, Boya, et al.
Pubblicazione: (2025)
di: Xiong, Boya, et al.
Pubblicazione: (2025)
Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
di: Wu, Wenxun, et al.
Pubblicazione: (2025)
di: Wu, Wenxun, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping
di: Li, Hanze, et al.
Pubblicazione: (2026)
di: Li, Hanze, et al.
Pubblicazione: (2026)
Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification
di: Fan, Zexia, et al.
Pubblicazione: (2026)
di: Fan, Zexia, et al.
Pubblicazione: (2026)
AgenticRec: End-to-End Tool-Integrated Policy Optimization for Ranking-Oriented Recommender Agents
di: Li, Tianyi, et al.
Pubblicazione: (2026)
di: Li, Tianyi, et al.
Pubblicazione: (2026)
Distract Large Language Models for Automatic Jailbreak Attack
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
di: Xiao, Zeguan, et al.
Pubblicazione: (2024)
VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
di: Cao, Xinye, et al.
Pubblicazione: (2025)
di: Cao, Xinye, et al.
Pubblicazione: (2025)
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics
di: Cao, Bowen, et al.
Pubblicazione: (2026)
di: Cao, Bowen, et al.
Pubblicazione: (2026)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
di: You, Yuyang, et al.
Pubblicazione: (2026)
di: You, Yuyang, et al.
Pubblicazione: (2026)
Unbiased Rectification for Sequential Recommender Systems Under Fake Orders
di: Qin, Qiyu, et al.
Pubblicazione: (2026)
di: Qin, Qiyu, et al.
Pubblicazione: (2026)
Offline Critic-Guided Diffusion Policy for Multi-User Delay-Constrained Scheduling
di: Li, Zhuoran, et al.
Pubblicazione: (2025)
di: Li, Zhuoran, et al.
Pubblicazione: (2025)
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
di: Wang, Hanqing, et al.
Pubblicazione: (2024)
di: Wang, Hanqing, et al.
Pubblicazione: (2024)
M-GRPO: Stabilizing Self-Supervised Reinforcement Learning for Large Language Models with Momentum-Anchored Policy Optimization
di: Bai, Bizhe, et al.
Pubblicazione: (2025)
di: Bai, Bizhe, et al.
Pubblicazione: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
di: Chen, Peter, et al.
Pubblicazione: (2025)
di: Chen, Peter, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
di: Wang, Tianyi, et al.
Pubblicazione: (2026) -
SeTAR: Out-of-Distribution Detection with Selective Low-Rank Approximation
di: Li, Yixia, et al.
Pubblicazione: (2024) -
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
di: Wang, Wenkai, et al.
Pubblicazione: (2025) -
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
di: Li, Zeping, et al.
Pubblicazione: (2026) -
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning
di: Li, Zhicong, et al.
Pubblicazione: (2026)