Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Junyu, Zheng, Zhi, Ying, Donghao, Jin, Ming, Gu, Shangding, Spanos, Costas, Lavaei, Javad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
LLMs Should Express Uncertainty Explicitly
by: Guo, Junyu, et al.
Published: (2026)
by: Guo, Junyu, et al.
Published: (2026)
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
by: Gu, Shangding, et al.
Published: (2025)
by: Gu, Shangding, et al.
Published: (2025)
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
by: Zheng, Zhi, et al.
Published: (2024)
by: Zheng, Zhi, et al.
Published: (2024)
Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization
by: Gu, Shangding
Published: (2026)
by: Gu, Shangding
Published: (2026)
Pausing Policy Learning in Non-stationary Reinforcement Learning
by: Lee, Hyunin, et al.
Published: (2024)
by: Lee, Hyunin, et al.
Published: (2024)
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
by: Ying, Donghao, et al.
Published: (2022)
by: Ying, Donghao, et al.
Published: (2022)
TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
by: Ding, Yuhao, et al.
Published: (2021)
by: Ding, Yuhao, et al.
Published: (2021)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
by: Bonnaire, Tony, et al.
Published: (2025)
by: Bonnaire, Tony, et al.
Published: (2025)
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
RLBenchNet: The Right Network for the Right Reinforcement Learning Task
by: Smirnov, Ivan, et al.
Published: (2025)
by: Smirnov, Ivan, et al.
Published: (2025)
A CMDP-within-online framework for Meta-Safe Reinforcement Learning
by: Khattar, Vanshaj, et al.
Published: (2024)
by: Khattar, Vanshaj, et al.
Published: (2024)
Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL
by: Luo, Qin-Wen, et al.
Published: (2025)
by: Luo, Qin-Wen, et al.
Published: (2025)
Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation
by: Gu, Shangding, et al.
Published: (2024)
by: Gu, Shangding, et al.
Published: (2024)
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
by: Luo, Qin-Wen, et al.
Published: (2024)
by: Luo, Qin-Wen, et al.
Published: (2024)
Don't Waste Mistakes: Leveraging Negative RL-Groups via Confidence Reweighting
by: Feng, Yunzhen, et al.
Published: (2025)
by: Feng, Yunzhen, et al.
Published: (2025)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
by: Tarasov, Denis, et al.
Published: (2024)
by: Tarasov, Denis, et al.
Published: (2024)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
by: Ma, Ziye, et al.
Published: (2024)
by: Ma, Ziye, et al.
Published: (2024)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
by: Kim, Taesan, et al.
Published: (2026)
by: Kim, Taesan, et al.
Published: (2026)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
by: Sokar, Ghada, et al.
Published: (2024)
by: Sokar, Ghada, et al.
Published: (2024)
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming
by: Kavumba, Pride, et al.
Published: (2026)
by: Kavumba, Pride, et al.
Published: (2026)
Structural Correspondence and Universal Approximation in Diagonal plus Low-Rank Neural Networks
by: Chen, Ying, et al.
Published: (2026)
by: Chen, Ying, et al.
Published: (2026)
Why is Normalization Preferred? A Worst-Case Complexity Theory for Stochastically Preconditioned SGD under Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Don't Play Favorites: Minority Guidance for Diffusion Models
by: Um, Soobin, et al.
Published: (2023)
by: Um, Soobin, et al.
Published: (2023)
High Probability Complexity Bounds of Trust-Region Stochastic Sequential Quadratic Programming with Heavy-Tailed Noise
by: Fang, Yuchen, et al.
Published: (2025)
by: Fang, Yuchen, et al.
Published: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
by: Liu, Xianyang, et al.
Published: (2026)
by: Liu, Xianyang, et al.
Published: (2026)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
by: Jeha, Paul, et al.
Published: (2026)
by: Jeha, Paul, et al.
Published: (2026)
Agentic Web: Weaving the Next Web with AI Agents
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
by: Chen, Zhaoyang, et al.
Published: (2025)
by: Chen, Zhaoyang, et al.
Published: (2025)
Similar Items
-
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025) -
LLMs Should Express Uncertainty Explicitly
by: Guo, Junyu, et al.
Published: (2026) -
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
by: Gu, Shangding, et al.
Published: (2025) -
Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation
by: Gu, Shangding, et al.
Published: (2024) -
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
by: Gu, Shangding, et al.
Published: (2025)