Proactive Constrained Policy Optimization with Preemptive Penalty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Ning, Wang, Pengyu, Liu, Guoqing, Zhang, Haifeng, Lv, Pin, Wang, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exterior Penalty Policy Optimization with Penalty Metric Network under Constraints
von: Gao, Shiqing, et al.
Veröffentlicht: (2024)
von: Gao, Shiqing, et al.
Veröffentlicht: (2024)
Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
Self-Supervised Penalty-Based Learning for Robust Constrained Optimization
von: Benslimane, Wyame, et al.
Veröffentlicht: (2025)
von: Benslimane, Wyame, et al.
Veröffentlicht: (2025)
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
Stop Guessing: Optimizing Goalkeeper Policies for Soccer Penalty Kicks
von: Bransen, Lotte, et al.
Veröffentlicht: (2025)
von: Bransen, Lotte, et al.
Veröffentlicht: (2025)
State-wise Constrained Policy Optimization
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Yongcheng, et al.
Veröffentlicht: (2025)
Stochastic Penalty-Barrier Methods for Constrained Machine Learning
von: Bosák, Adam, et al.
Veröffentlicht: (2026)
von: Bosák, Adam, et al.
Veröffentlicht: (2026)
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization
von: He, Longxiang, et al.
Veröffentlicht: (2024)
von: He, Longxiang, et al.
Veröffentlicht: (2024)
Solving Richly Constrained Reinforcement Learning through State Augmentation and Reward Penalties
von: Jiang, Hao, et al.
Veröffentlicht: (2023)
von: Jiang, Hao, et al.
Veröffentlicht: (2023)
e-COP : Episodic Constrained Optimization of Policies
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2024)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
von: Wang, Mingyi, et al.
Veröffentlicht: (2026)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
von: Niu, Yifan, et al.
Veröffentlicht: (2025)
von: Niu, Yifan, et al.
Veröffentlicht: (2025)
Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance
von: Huang, RenMing, et al.
Veröffentlicht: (2024)
von: Huang, RenMing, et al.
Veröffentlicht: (2024)
Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk
von: Tangri, Rohan, et al.
Veröffentlicht: (2026)
von: Tangri, Rohan, et al.
Veröffentlicht: (2026)
Scaling Multi-Hop Training Data via Graph-Constrained Path Selection
von: Chen, Pengyu, et al.
Veröffentlicht: (2026)
von: Chen, Pengyu, et al.
Veröffentlicht: (2026)
Constrained Group Relative Policy Optimization
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
Autoregressive Policy Optimization for Constrained Allocation Tasks
von: Winkel, David, et al.
Veröffentlicht: (2024)
von: Winkel, David, et al.
Veröffentlicht: (2024)
Variational Delayed Policy Optimization
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2024)
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
von: Nathani, Deepak, et al.
Veröffentlicht: (2026)
von: Nathani, Deepak, et al.
Veröffentlicht: (2026)
Large Language Models as Generalist Policies for Network Optimization
von: Wu, Duo, et al.
Veröffentlicht: (2025)
von: Wu, Duo, et al.
Veröffentlicht: (2025)
skscope: Fast Sparsity-Constrained Optimization in Python
von: Wang, Zezhi, et al.
Veröffentlicht: (2024)
von: Wang, Zezhi, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
von: Wang, Wenkai, et al.
Veröffentlicht: (2025)
von: Wang, Wenkai, et al.
Veröffentlicht: (2025)
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
GIPO: Gaussian Importance Sampling Policy Optimization
von: Lu, Chengxuan, et al.
Veröffentlicht: (2026)
von: Lu, Chengxuan, et al.
Veröffentlicht: (2026)
Diffusion Actor-Critic: Formulating Constrained Policy Iteration as Diffusion Noise Regression for Offline Reinforcement Learning
von: Fang, Linjiajie, et al.
Veröffentlicht: (2024)
von: Fang, Linjiajie, et al.
Veröffentlicht: (2024)
Score Regularized Policy Optimization through Diffusion Behavior
von: Chen, Huayu, et al.
Veröffentlicht: (2023)
von: Chen, Huayu, et al.
Veröffentlicht: (2023)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
Mildly Constrained Evaluation Policy for Offline Reinforcement Learning
von: Xu, Linjie, et al.
Veröffentlicht: (2023)
von: Xu, Linjie, et al.
Veröffentlicht: (2023)
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
von: Yao, Yihang, et al.
Veröffentlicht: (2026)
von: Yao, Yihang, et al.
Veröffentlicht: (2026)
Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization
von: Huang, Yancheng, et al.
Veröffentlicht: (2026)
von: Huang, Yancheng, et al.
Veröffentlicht: (2026)
Trajectory-Oriented Policy Optimization with Sparse Rewards
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exterior Penalty Policy Optimization with Penalty Metric Network under Constraints
von: Gao, Shiqing, et al.
Veröffentlicht: (2024) -
Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction
von: Wang, Haoyu, et al.
Veröffentlicht: (2026) -
Self-Supervised Penalty-Based Learning for Robust Constrained Optimization
von: Benslimane, Wyame, et al.
Veröffentlicht: (2025) -
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
von: Shen, Yiyang, et al.
Veröffentlicht: (2026) -
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
von: Li, Zihao, et al.
Veröffentlicht: (2024)