Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pandit, Kartik, Ganguly, Sourav, Banerjee, Arnesh, Angizi, Shaahin, Ghosh, Arnob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026)
by: Ganguly, Sourav, et al.
Published: (2026)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
SPICEPilot: Navigating SPICE Code Generation and Simulation with AI Guidance
by: Vungarala, Deepak, et al.
Published: (2024)
by: Vungarala, Deepak, et al.
Published: (2024)
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
by: Vungarala, Deepak, et al.
Published: (2025)
by: Vungarala, Deepak, et al.
Published: (2025)
A Framework for Inherently Safer AGI through Language-Mediated Active Inference
by: Wen, Bo
Published: (2025)
by: Wen, Bo
Published: (2025)
Dependability in Embedded Systems: A Survey of Fault Tolerance Methods and Software-Based Mitigation Techniques
by: Solouki, Mohammadreza Amel, et al.
Published: (2024)
by: Solouki, Mohammadreza Amel, et al.
Published: (2024)
Runtime-Certified Bounded-Error Quantized Attention
by: Calver, Dean
Published: (2026)
by: Calver, Dean
Published: (2026)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
Lyapunov-Certified Direct Switching Theory for Q-Learning
by: Lee, Donghwan
Published: (2026)
by: Lee, Donghwan
Published: (2026)
DeepSafeMPC: Deep Learning-Based Model Predictive Control for Safe Multi-Agent Reinforcement Learning
by: Wang, Xuefeng, et al.
Published: (2024)
by: Wang, Xuefeng, et al.
Published: (2024)
Approximate Model-Based Shielding for Safe Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2023)
by: Goodall, Alexander W., et al.
Published: (2023)
CPS-LLM: Large Language Model based Safe Usage Plan Generator for Human-in-the-Loop Human-in-the-Plant Cyber-Physical System
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Safe Deep Model-Based Reinforcement Learning with Lyapunov Functions
by: Zhang, Harry
Published: (2024)
by: Zhang, Harry
Published: (2024)
Reinforcement Learning Constrained Beam Search for Parameter Optimization of Paper Drying Under Flexible Constraints
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
Learning to Drive Safely with Hybrid Options
by: De Cooman, Bram, et al.
Published: (2025)
by: De Cooman, Bram, et al.
Published: (2025)
Large Language Model Powered Automated Modeling and Optimization of Active Distribution Network Dispatch Problems
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Parallel Differentiable Reachability for Learning and Planning with Certified Neural Dynamics and Controllers
by: Shen, Keyi, et al.
Published: (2026)
by: Shen, Keyi, et al.
Published: (2026)
Certified Training with Branch-and-Bound for Lyapunov-stable Neural Control
by: Shi, Zhouxing, et al.
Published: (2024)
by: Shi, Zhouxing, et al.
Published: (2024)
Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates
by: Wang, Chengxiao, et al.
Published: (2026)
by: Wang, Chengxiao, et al.
Published: (2026)
Intersection of Reinforcement Learning and Bayesian Optimization for Intelligent Control of Industrial Processes: A Safe MPC-based DPG using Multi-Objective BO
by: Esfahani, Hossein Nejatbakhsh, et al.
Published: (2025)
by: Esfahani, Hossein Nejatbakhsh, et al.
Published: (2025)
Provably Safe Generative Sampling with Constricting Barrier Functions
by: Gadginmath, Darshan, et al.
Published: (2026)
by: Gadginmath, Darshan, et al.
Published: (2026)
Constrained Reinforcement Learning for Safe Heat Pump Control
by: Zhang, Baohe, et al.
Published: (2024)
by: Zhang, Baohe, et al.
Published: (2024)
GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
by: Zhou, Zhehua, et al.
Published: (2024)
by: Zhou, Zhehua, et al.
Published: (2024)
Learning Contextual Runtime Monitors for Safe AI-Based Autonomy
by: Luque-Cerpa, Alejandro, et al.
Published: (2026)
by: Luque-Cerpa, Alejandro, et al.
Published: (2026)
Balancing SoC in Battery Cells using Safe Action Perturbations
by: Yadav, E Harshith Kumar, et al.
Published: (2025)
by: Yadav, E Harshith Kumar, et al.
Published: (2025)
Guided Safe Shooting: model based reinforcement learning with safety constraints
by: Paolo, Giuseppe, et al.
Published: (2022)
by: Paolo, Giuseppe, et al.
Published: (2022)
Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction
by: Durkin, Alex, et al.
Published: (2025)
by: Durkin, Alex, et al.
Published: (2025)
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
by: Banerjee, Chayan, et al.
Published: (2025)
by: Banerjee, Chayan, et al.
Published: (2025)
Action Mapping for Reinforcement Learning in Continuous Environments with Constraints
by: Theile, Mirco, et al.
Published: (2024)
by: Theile, Mirco, et al.
Published: (2024)
A Safe Reinforcement Learning driven Weights-varying Model Predictive Control for Autonomous Vehicle Motion Control
by: Zarrouki, Baha, et al.
Published: (2024)
by: Zarrouki, Baha, et al.
Published: (2024)
DCcluster-Opt: Benchmarking Dynamic Multi-Objective Optimization for Geo-Distributed Data Center Workloads
by: Guillen-Perez, Antonio, et al.
Published: (2025)
by: Guillen-Perez, Antonio, et al.
Published: (2025)
Learning Hidden Subgoals under Temporal Ordering Constraints in Reinforcement Learning
by: Xu, Duo, et al.
Published: (2024)
by: Xu, Duo, et al.
Published: (2024)
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
by: Fard, Amir, et al.
Published: (2025)
by: Fard, Amir, et al.
Published: (2025)
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
by: Spoor, Lindsay, et al.
Published: (2025)
by: Spoor, Lindsay, et al.
Published: (2025)
Benchmarking Model Predictive Control Algorithms in Building Optimization Testing Framework (BOPTEST)
by: Mostafavi, Saman, et al.
Published: (2023)
by: Mostafavi, Saman, et al.
Published: (2023)
Trustworthy and Explainable Deep Reinforcement Learning for Safe and Energy-Efficient Process Control: A Use Case in Industrial Compressed Air Systems
by: Bezold, Vincent, et al.
Published: (2025)
by: Bezold, Vincent, et al.
Published: (2025)
Deep Model Predictive Optimization
by: Sacks, Jacob, et al.
Published: (2023)
by: Sacks, Jacob, et al.
Published: (2023)
Test Time Training for AC Power Flow Surrogates via Physics and Operational Constraint Refinement
by: Dogoulis, Panteleimon, et al.
Published: (2025)
by: Dogoulis, Panteleimon, et al.
Published: (2025)
Similar Items
-
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
by: Ganguly, Sourav, et al.
Published: (2026) -
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025) -
Provably Efficient Sample Complexity for Robust CMDP
by: Ganguly, Sourav, et al.
Published: (2025) -
SPICEPilot: Navigating SPICE Code Generation and Simulation with AI Guidance
by: Vungarala, Deepak, et al.
Published: (2024) -
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
by: Vungarala, Deepak, et al.
Published: (2025)