SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | K, Swaminathan S, Hazra, Aritra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
von: Hazra, Somnath, et al.
Veröffentlicht: (2025)
von: Hazra, Somnath, et al.
Veröffentlicht: (2025)
DiffClone: Enhanced Behaviour Cloning in Robotics with Diffusion-Driven Policy Learning
von: Mani, Sabariswaran, et al.
Veröffentlicht: (2024)
von: Mani, Sabariswaran, et al.
Veröffentlicht: (2024)
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
von: Hu, Jiaheng, et al.
Veröffentlicht: (2025)
von: Hu, Jiaheng, et al.
Veröffentlicht: (2025)
Sample-efficient and Scalable Exploration in Continuous-Time RL
von: Iten, Klemens, et al.
Veröffentlicht: (2025)
von: Iten, Klemens, et al.
Veröffentlicht: (2025)
Enhanced Importance Sampling through Latent Space Exploration in Normalizing Flows
von: Kruse, Liam A., et al.
Veröffentlicht: (2025)
von: Kruse, Liam A., et al.
Veröffentlicht: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
von: Park, Seohong, et al.
Veröffentlicht: (2023)
von: Park, Seohong, et al.
Veröffentlicht: (2023)
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
von: Yuan, Xiu, et al.
Veröffentlicht: (2024)
von: Yuan, Xiu, et al.
Veröffentlicht: (2024)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
Safe Exploration via Policy Priors
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
von: Wendl, Manuel, et al.
Veröffentlicht: (2026)
CaRL: Learning Scalable Planning Policies with Simple Rewards
von: Jaeger, Bernhard, et al.
Veröffentlicht: (2025)
von: Jaeger, Bernhard, et al.
Veröffentlicht: (2025)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
von: Guo, Jian-Ting, et al.
Veröffentlicht: (2025)
von: Guo, Jian-Ting, et al.
Veröffentlicht: (2025)
Harnessing Bounded-Support Evolution Strategies for Policy Refinement
von: Hirschowitz, Ethan, et al.
Veröffentlicht: (2025)
von: Hirschowitz, Ethan, et al.
Veröffentlicht: (2025)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
von: Wagenmaker, Andrew, et al.
Veröffentlicht: (2025)
von: Wagenmaker, Andrew, et al.
Veröffentlicht: (2025)
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
von: Choi, Wonhyeok, et al.
Veröffentlicht: (2026)
von: Choi, Wonhyeok, et al.
Veröffentlicht: (2026)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
First Order Model-Based RL through Decoupled Backpropagation
von: Amigo, Joseph, et al.
Veröffentlicht: (2025)
von: Amigo, Joseph, et al.
Veröffentlicht: (2025)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
von: Jin, Yang, et al.
Veröffentlicht: (2025)
von: Jin, Yang, et al.
Veröffentlicht: (2025)
Extremum-Seeking Action Selection for Accelerating Policy Optimization
von: Chang, Ya-Chien, et al.
Veröffentlicht: (2024)
von: Chang, Ya-Chien, et al.
Veröffentlicht: (2024)
Real-Time Execution of Action Chunking Flow Policies
von: Black, Kevin, et al.
Veröffentlicht: (2025)
von: Black, Kevin, et al.
Veröffentlicht: (2025)
Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters
von: Kong, Lingxiao, et al.
Veröffentlicht: (2026)
von: Kong, Lingxiao, et al.
Veröffentlicht: (2026)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
von: Wu, Yilin, et al.
Veröffentlicht: (2025)
von: Wu, Yilin, et al.
Veröffentlicht: (2025)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
von: Cai, Shizhe, et al.
Veröffentlicht: (2025)
von: Cai, Shizhe, et al.
Veröffentlicht: (2025)
Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
von: Terence, Ng Wen Zheng, et al.
Veröffentlicht: (2024)
von: Terence, Ng Wen Zheng, et al.
Veröffentlicht: (2024)
Don't Start from Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion
von: Chen, Kaiqi, et al.
Veröffentlicht: (2024)
von: Chen, Kaiqi, et al.
Veröffentlicht: (2024)
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
von: Hong, Matthew M., et al.
Veröffentlicht: (2026)
von: Hong, Matthew M., et al.
Veröffentlicht: (2026)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
von: Jin, Yang, et al.
Veröffentlicht: (2025)
von: Jin, Yang, et al.
Veröffentlicht: (2025)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
von: Sukhija, Bhavya, et al.
Veröffentlicht: (2024)
von: Sukhija, Bhavya, et al.
Veröffentlicht: (2024)
GRAIL: Goal Recognition Alignment through Imitation Learning
von: Elhadad, Osher, et al.
Veröffentlicht: (2026)
von: Elhadad, Osher, et al.
Veröffentlicht: (2026)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
von: Honari, Homayoun, et al.
Veröffentlicht: (2026)
von: Honari, Homayoun, et al.
Veröffentlicht: (2026)
AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
von: Pappalardo, Octavio
Veröffentlicht: (2026)
von: Pappalardo, Octavio
Veröffentlicht: (2026)
Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
von: Yokozawa, Riko, et al.
Veröffentlicht: (2025)
von: Yokozawa, Riko, et al.
Veröffentlicht: (2025)
ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
von: Cherepanov, Egor, et al.
Veröffentlicht: (2025)
von: Cherepanov, Egor, et al.
Veröffentlicht: (2025)
Stochastic Q-learning for Large Discrete Action Spaces
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
von: Fourati, Fares, et al.
Veröffentlicht: (2024)
Automatic Environment Shaping is the Next Frontier in RL
von: Park, Younghyo, et al.
Veröffentlicht: (2024)
von: Park, Younghyo, et al.
Veröffentlicht: (2024)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
von: Park, Seohong, et al.
Veröffentlicht: (2023)
von: Park, Seohong, et al.
Veröffentlicht: (2023)
Language-Conditioned Offline RL for Multi-Robot Navigation
von: Morad, Steven, et al.
Veröffentlicht: (2024)
von: Morad, Steven, et al.
Veröffentlicht: (2024)
Investigating Memory in Model-Free RL with POPGym Arcade
von: Wang, Zekang, et al.
Veröffentlicht: (2025)
von: Wang, Zekang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
von: Hazra, Somnath, et al.
Veröffentlicht: (2025) -
DiffClone: Enhanced Behaviour Cloning in Robotics with Diffusion-Driven Policy Learning
von: Mani, Sabariswaran, et al.
Veröffentlicht: (2024) -
SLAC: Simulation-Pretrained Latent Action Space for Whole-Body Real-World RL
von: Hu, Jiaheng, et al.
Veröffentlicht: (2025) -
Sample-efficient and Scalable Exploration in Continuous-Time RL
von: Iten, Klemens, et al.
Veröffentlicht: (2025) -
Enhanced Importance Sampling through Latent Space Exploration in Normalizing Flows
von: Kruse, Liam A., et al.
Veröffentlicht: (2025)