Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Hazra, Somnath, Dasgupta, Pallab, Dey, Soumyajit |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tackling Uncertainties in Multi-Agent Reinforcement Learning through Integration of Agent Termination Dynamics
por: Hazra, Somnath, et al.
Publicado: (2025)
por: Hazra, Somnath, et al.
Publicado: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
por: K, Swaminathan S, et al.
Publicado: (2026)
por: K, Swaminathan S, et al.
Publicado: (2026)
Towards Adaptive IMFs -- Generalization of utility functions in Multi-Agent Frameworks
por: Dey, Kaushik, et al.
Publicado: (2024)
por: Dey, Kaushik, et al.
Publicado: (2024)
Constrained Latent Action Policies for Model-Based Offline Reinforcement Learning
por: Alles, Marvin, et al.
Publicado: (2024)
por: Alles, Marvin, et al.
Publicado: (2024)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
por: Zhang, Tianle, et al.
Publicado: (2024)
por: Zhang, Tianle, et al.
Publicado: (2024)
On Safer Reinforcement Learning for Sedation and Analgesia in Intensive Care
por: Romero-Hernandez, Joel, et al.
Publicado: (2026)
por: Romero-Hernandez, Joel, et al.
Publicado: (2026)
Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
por: Gao, Shujian, et al.
Publicado: (2026)
por: Gao, Shujian, et al.
Publicado: (2026)
Leveraging Constraint Violation Signals For Action-Constrained Reinforcement Learning
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2025)
por: Brahmanage, Janaka Chathuranga, et al.
Publicado: (2025)
Reinforcing Language Agents via Policy Optimization with Action Decomposition
por: Wen, Muning, et al.
Publicado: (2024)
por: Wen, Muning, et al.
Publicado: (2024)
AD$^2$: Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving Systems
por: Sahu, Ishan, et al.
Publicado: (2026)
por: Sahu, Ishan, et al.
Publicado: (2026)
Two-Step Offline Preference-Based Reinforcement Learning with Constrained Actions
por: Xu, Yinglun, et al.
Publicado: (2023)
por: Xu, Yinglun, et al.
Publicado: (2023)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
por: Rietz, Finn, et al.
Publicado: (2024)
por: Rietz, Finn, et al.
Publicado: (2024)
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
por: Tiofack, Franki Nguimatsia, et al.
Publicado: (2025)
por: Tiofack, Franki Nguimatsia, et al.
Publicado: (2025)
Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
por: Terence, Ng Wen Zheng, et al.
Publicado: (2024)
por: Terence, Ng Wen Zheng, et al.
Publicado: (2024)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
por: Zhu, Yuanyang, et al.
Publicado: (2024)
por: Zhu, Yuanyang, et al.
Publicado: (2024)
Autoregressive Policy Optimization for Constrained Allocation Tasks
por: Winkel, David, et al.
Publicado: (2024)
por: Winkel, David, et al.
Publicado: (2024)
An Advantage-based Optimization Method for Reinforcement Learning in Large Action Space
por: Lin, Hai, et al.
Publicado: (2024)
por: Lin, Hai, et al.
Publicado: (2024)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
por: Thoma, Vinzenz, et al.
Publicado: (2024)
por: Thoma, Vinzenz, et al.
Publicado: (2024)
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
por: Hu, Jifeng, et al.
Publicado: (2025)
por: Hu, Jifeng, et al.
Publicado: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
por: Gao, Chen-Xiao, et al.
Publicado: (2025)
por: Gao, Chen-Xiao, et al.
Publicado: (2025)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
por: Kang, Hyungkyu, et al.
Publicado: (2025)
por: Kang, Hyungkyu, et al.
Publicado: (2025)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
por: Yao, Yihang, et al.
Publicado: (2023)
por: Yao, Yihang, et al.
Publicado: (2023)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
por: Hao, Ruijie, et al.
Publicado: (2026)
por: Hao, Ruijie, et al.
Publicado: (2026)
VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
por: Wang, Haozhe, et al.
Publicado: (2025)
por: Wang, Haozhe, et al.
Publicado: (2025)
Continual Learning as Computationally Constrained Reinforcement Learning
por: Kumar, Saurabh, et al.
Publicado: (2023)
por: Kumar, Saurabh, et al.
Publicado: (2023)
State-Constrained Offline Reinforcement Learning
por: Hepburn, Charles A., et al.
Publicado: (2024)
por: Hepburn, Charles A., et al.
Publicado: (2024)
Locally Constrained Representations in Reinforcement Learning
por: Nath, Somjit, et al.
Publicado: (2022)
por: Nath, Somjit, et al.
Publicado: (2022)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
por: Goodall, Alexander W., et al.
Publicado: (2025)
por: Goodall, Alexander W., et al.
Publicado: (2025)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
por: Zhan, Simon Sinong, et al.
Publicado: (2025)
por: Zhan, Simon Sinong, et al.
Publicado: (2025)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
por: Zhang, Han, et al.
Publicado: (2025)
por: Zhang, Han, et al.
Publicado: (2025)
Efficient Deep Reinforcement Learning with Predictive Processing Proximal Policy Optimization
por: Küçükoğlu, Burcu, et al.
Publicado: (2022)
por: Küçükoğlu, Burcu, et al.
Publicado: (2022)
Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization
por: Alvo, Matias, et al.
Publicado: (2023)
por: Alvo, Matias, et al.
Publicado: (2023)
Valid Property-Enhanced Contrastive Learning for Targeted Optimization & Resampling for Novel Drug Design
por: Banerjee, Amartya, et al.
Publicado: (2025)
por: Banerjee, Amartya, et al.
Publicado: (2025)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
por: Si, Wenwen, et al.
Publicado: (2025)
por: Si, Wenwen, et al.
Publicado: (2025)
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
por: Wang, Mingyi, et al.
Publicado: (2026)
por: Wang, Mingyi, et al.
Publicado: (2026)
Quantum-Enhanced Vision Transformer for Flood Detection using Remote Sensing Imagery
por: Maity, Soumyajit, et al.
Publicado: (2026)
por: Maity, Soumyajit, et al.
Publicado: (2026)
Consensus Sampling for Safer Generative AI
por: Kalai, Adam Tauman, et al.
Publicado: (2025)
por: Kalai, Adam Tauman, et al.
Publicado: (2025)
Memory Allocation in Resource-Constrained Reinforcement Learning
por: Tamborski, Massimiliano, et al.
Publicado: (2025)
por: Tamborski, Massimiliano, et al.
Publicado: (2025)
Large Scale Constrained Clustering With Reinforcement Learning
por: Schesch, Benedikt, et al.
Publicado: (2024)
por: Schesch, Benedikt, et al.
Publicado: (2024)
Imitating Cost-Constrained Behaviors in Reinforcement Learning
por: Shao, Qian, et al.
Publicado: (2024)
por: Shao, Qian, et al.
Publicado: (2024)
Ejemplares similares
-
Tackling Uncertainties in Multi-Agent Reinforcement Learning through Integration of Agent Termination Dynamics
por: Hazra, Somnath, et al.
Publicado: (2025) -
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
por: K, Swaminathan S, et al.
Publicado: (2026) -
Towards Adaptive IMFs -- Generalization of utility functions in Multi-Agent Frameworks
por: Dey, Kaushik, et al.
Publicado: (2024) -
Constrained Latent Action Policies for Model-Based Offline Reinforcement Learning
por: Alles, Marvin, et al.
Publicado: (2024) -
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
por: Zhang, Tianle, et al.
Publicado: (2024)