e-COP : Episodic Constrained Optimization of Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Agnihotri, Akhil, Jain, Rahul, Ramachandran, Deepak, Singla, Sahil |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Best Policy Learning from Trajectory Preference Feedback
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
by: Agnihotri, Akhil
Published: (2025)
by: Agnihotri, Akhil
Published: (2025)
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
Bayesian Learning in Episodic Zero-Sum Games
by: Yueh, Chang-Wei, et al.
Published: (2026)
by: Yueh, Chang-Wei, et al.
Published: (2026)
Bandit Sequential Posted Pricing via Half-Concavity
by: Singla, Sahil, et al.
Published: (2023)
by: Singla, Sahil, et al.
Published: (2023)
SAPG: Split and Aggregate Policy Gradients
by: Singla, Jayesh, et al.
Published: (2024)
by: Singla, Jayesh, et al.
Published: (2024)
Improved and Oracle-Efficient Online $\ell_1$-Multicalibration
by: Ghuge, Rohan, et al.
Published: (2025)
by: Ghuge, Rohan, et al.
Published: (2025)
Memoryless Policy Iteration for Episodic POMDPs
by: van Zuijlen, Roy, et al.
Published: (2025)
by: van Zuijlen, Roy, et al.
Published: (2025)
VIP-COP: Context Optimization for Tabular Foundation Models
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
Posterior Sampling-based Online Learning for Episodic POMDPs
by: Tang, Dengwang, et al.
Published: (2023)
by: Tang, Dengwang, et al.
Published: (2023)
Pure Exploration for Constrained Best Mixed Arm Identification with a Fixed Budget
by: Tang, Dengwang, et al.
Published: (2024)
by: Tang, Dengwang, et al.
Published: (2024)
RANDPOL: Parameter-Efficient End-to-End Quadruped Locomotion via Randomized Policy Learning
by: Liu, Zhuochen, et al.
Published: (2025)
by: Liu, Zhuochen, et al.
Published: (2025)
Single-Sample and Robust Online Resource Allocation
by: Ghuge, Rohan, et al.
Published: (2025)
by: Ghuge, Rohan, et al.
Published: (2025)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
by: Tzannetos, Georgios, et al.
Published: (2025)
by: Tzannetos, Georgios, et al.
Published: (2025)
State-wise Constrained Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Proactive Constrained Policy Optimization with Preemptive Penalty
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Evolutionary Policy Optimization
by: Wang, Jianren, et al.
Published: (2025)
by: Wang, Jianren, et al.
Published: (2025)
Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk
by: Tangri, Rohan, et al.
Published: (2026)
by: Tangri, Rohan, et al.
Published: (2026)
Architecting Digital Twins for Intelligent Transportation Systems
by: Bhatt, Hiya, et al.
Published: (2025)
by: Bhatt, Hiya, et al.
Published: (2025)
Constrained Best Arm Identification in Grouped Bandits
by: Dharod, Sahil, et al.
Published: (2024)
by: Dharod, Sahil, et al.
Published: (2024)
Constrained Group Relative Policy Optimization
by: Girgis, Roger, et al.
Published: (2026)
by: Girgis, Roger, et al.
Published: (2026)
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2024)
by: Nika, Andi, et al.
Published: (2024)
Autoregressive Policy Optimization for Constrained Allocation Tasks
by: Winkel, David, et al.
Published: (2024)
by: Winkel, David, et al.
Published: (2024)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
AscendOptimizer: Episodic Agent for Ascend NPU Operator Optimization
by: Wu, Jiehao, et al.
Published: (2026)
by: Wu, Jiehao, et al.
Published: (2026)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
by: Niu, Yifan, et al.
Published: (2025)
by: Niu, Yifan, et al.
Published: (2025)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
by: Zhang, Jing, et al.
Published: (2023)
by: Zhang, Jing, et al.
Published: (2023)
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
by: Agnihotri, Rudransh, et al.
Published: (2025)
by: Agnihotri, Rudransh, et al.
Published: (2025)
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
by: Lei, Yuheng, et al.
Published: (2026)
by: Lei, Yuheng, et al.
Published: (2026)
$O(\sqrt{T})$ Static Regret and Instance Dependent Constraint Violation for Constrained Online Convex Optimization
by: Vaze, Rahul, et al.
Published: (2025)
by: Vaze, Rahul, et al.
Published: (2025)
AlignIQL: Policy Alignment in Implicit Q-Learning through Constrained Optimization
by: He, Longxiang, et al.
Published: (2024)
by: He, Longxiang, et al.
Published: (2024)
Co2PO: Coordinated Constrained Policy Optimization for Multi-Agent RL
by: Patel, Shrenik, et al.
Published: (2026)
by: Patel, Shrenik, et al.
Published: (2026)
Constrained Policy Optimization via Sampling-Based Weight-Space Projection
by: Cao, Shengfan, et al.
Published: (2025)
by: Cao, Shengfan, et al.
Published: (2025)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
by: Si, Wenwen, et al.
Published: (2025)
by: Si, Wenwen, et al.
Published: (2025)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
by: Wang, Mingyi, et al.
Published: (2026)
by: Wang, Mingyi, et al.
Published: (2026)
Similar Items
-
Best Policy Learning from Trajectory Preference Feedback
by: Agnihotri, Akhil, et al.
Published: (2025) -
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
by: Agnihotri, Akhil, et al.
Published: (2025) -
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024) -
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023) -
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
by: Agnihotri, Akhil
Published: (2025)