OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Burnwal, Returaj, Bhatt, Nirav Pravinbhai, Ravindran, Balaraman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
Learning from Observation: A Survey of Recent Advances
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025)
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
von: Acharjee, Jashaswimalya, et al.
Veröffentlicht: (2026)
von: Acharjee, Jashaswimalya, et al.
Veröffentlicht: (2026)
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
von: Verma, Abhishek, et al.
Veröffentlicht: (2025)
von: Verma, Abhishek, et al.
Veröffentlicht: (2025)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
von: Verma, Richa, et al.
Veröffentlicht: (2026)
von: Verma, Richa, et al.
Veröffentlicht: (2026)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
von: Verma, Abhishek, et al.
Veröffentlicht: (2025)
von: Verma, Abhishek, et al.
Veröffentlicht: (2025)
Offline Safe Reinforcement Learning Using Trajectory Classification
von: Gong, Ze, et al.
Veröffentlicht: (2024)
von: Gong, Ze, et al.
Veröffentlicht: (2024)
Offline Reinforcement Learning with Generative Trajectory Policies
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
Functional Groups are All you Need for Chemically Interpretable Molecular Property Prediction
von: Balaji, Roshan, et al.
Veröffentlicht: (2025)
von: Balaji, Roshan, et al.
Veröffentlicht: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
von: S, Santhosh G, et al.
Veröffentlicht: (2025)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
von: Hu, Hao, et al.
Veröffentlicht: (2025)
von: Hu, Hao, et al.
Veröffentlicht: (2025)
Constraint-Adaptive Policy Switching for Offline Safe Reinforcement Learning
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
von: Chemingui, Yassine, et al.
Veröffentlicht: (2024)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
DITTO: Offline Imitation Learning with World Models
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
von: DeMoss, Branton, et al.
Veröffentlicht: (2023)
Offline Trajectory Optimization for Offline Reinforcement Learning
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2024)
OLLIE: Imitation Learning from Offline Pretraining to Online Finetuning
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
Offline Imitation Learning with Model-based Reverse Augmentation
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
von: Shao, Jie-Jing, et al.
Veröffentlicht: (2024)
How to Leverage Diverse Demonstrations in Offline Imitation Learning
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
von: Yue, Sheng, et al.
Veröffentlicht: (2024)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
von: Low, Siow Meng, et al.
Veröffentlicht: (2024)
von: Low, Siow Meng, et al.
Veröffentlicht: (2024)
Balance Equation-based Distributionally Robust Offline Imitation Learning
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2025)
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2025)
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
von: Hoang, Huy, et al.
Veröffentlicht: (2024)
von: Hoang, Huy, et al.
Veröffentlicht: (2024)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
von: Xue, Ruiqi, et al.
Veröffentlicht: (2026)
von: Xue, Ruiqi, et al.
Veröffentlicht: (2026)
Online Optimization for Offline Safe Reinforcement Learning
von: Chemingui, Yassine, et al.
Veröffentlicht: (2025)
von: Chemingui, Yassine, et al.
Veröffentlicht: (2025)
Offline Imitation Learning Through Graph Search and Retrieval
von: Yin, Zhao-Heng, et al.
Veröffentlicht: (2024)
von: Yin, Zhao-Heng, et al.
Veröffentlicht: (2024)
SEABO: A Simple Search-Based Method for Offline Imitation Learning
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
von: Lyu, Jiafei, et al.
Veröffentlicht: (2024)
Align Your Intents: Offline Imitation Learning via Optimal Transport
von: Bobrin, Maksim, et al.
Veröffentlicht: (2024)
von: Bobrin, Maksim, et al.
Veröffentlicht: (2024)
When should we prefer Decision Transformers for Offline Reinforcement Learning?
von: Bhargava, Prajjwal, et al.
Veröffentlicht: (2023)
von: Bhargava, Prajjwal, et al.
Veröffentlicht: (2023)
A Dual Approach to Imitation Learning from Observations with Offline Datasets
von: Sikchi, Harshit, et al.
Veröffentlicht: (2024)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2024)
Robust Probabilistic Shielding for Safe Offline Reinforcement Learning
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2026)
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2026)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
von: Hoang, Huy, et al.
Veröffentlicht: (2025)
von: Hoang, Huy, et al.
Veröffentlicht: (2025)
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression
von: Ge, Zichang, et al.
Veröffentlicht: (2025)
von: Ge, Zichang, et al.
Veröffentlicht: (2025)
Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
Markov Balance Satisfaction Improves Performance in Strictly Batch Offline Imitation Learning
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2024)
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2024)
Decoupled Guidance Diffusion for Adaptive Offline Safe Reinforcement Learning
von: Chen, Rufeng, et al.
Veröffentlicht: (2026)
von: Chen, Rufeng, et al.
Veröffentlicht: (2026)
TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning
von: Sestini, Alessandro, et al.
Veröffentlicht: (2025)
von: Sestini, Alessandro, et al.
Veröffentlicht: (2025)
Offline Diversity Maximization Under Imitation Constraints
von: Vlastelica, Marin, et al.
Veröffentlicht: (2023)
von: Vlastelica, Marin, et al.
Veröffentlicht: (2023)
Imitate the Good and Avoid the Bad: An Incremental Approach to Safe Reinforcement Learning
von: Hoang, Huy, et al.
Veröffentlicht: (2023)
von: Hoang, Huy, et al.
Veröffentlicht: (2023)
Online Adaptation for Enhancing Imitation Learning Policies
von: Malato, Federico, et al.
Veröffentlicht: (2024)
von: Malato, Federico, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025) -
Learning from Observation: A Survey of Recent Advances
von: Burnwal, Returaj, et al.
Veröffentlicht: (2025) -
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
von: Acharjee, Jashaswimalya, et al.
Veröffentlicht: (2026) -
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
von: Verma, Abhishek, et al.
Veröffentlicht: (2025) -
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
von: Verma, Richa, et al.
Veröffentlicht: (2026)