Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pandit, Kartik, Ganguly, Sourav, Banerjee, Arnesh, Angizi, Shaahin, Ghosh, Arnob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
Provably Efficient Sample Complexity for Robust CMDP
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025)
SPICEPilot: Navigating SPICE Code Generation and Simulation with AI Guidance
von: Vungarala, Deepak, et al.
Veröffentlicht: (2024)
von: Vungarala, Deepak, et al.
Veröffentlicht: (2024)
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)
A Framework for Inherently Safer AGI through Language-Mediated Active Inference
von: Wen, Bo
Veröffentlicht: (2025)
von: Wen, Bo
Veröffentlicht: (2025)
Dependability in Embedded Systems: A Survey of Fault Tolerance Methods and Software-Based Mitigation Techniques
von: Solouki, Mohammadreza Amel, et al.
Veröffentlicht: (2024)
von: Solouki, Mohammadreza Amel, et al.
Veröffentlicht: (2024)
Runtime-Certified Bounded-Error Quantized Attention
von: Calver, Dean
Veröffentlicht: (2026)
von: Calver, Dean
Veröffentlicht: (2026)
Certifiably Robust Policies for Uncertain Parametric Environments
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2024)
Lyapunov-Certified Direct Switching Theory for Q-Learning
von: Lee, Donghwan
Veröffentlicht: (2026)
von: Lee, Donghwan
Veröffentlicht: (2026)
DeepSafeMPC: Deep Learning-Based Model Predictive Control for Safe Multi-Agent Reinforcement Learning
von: Wang, Xuefeng, et al.
Veröffentlicht: (2024)
von: Wang, Xuefeng, et al.
Veröffentlicht: (2024)
Approximate Model-Based Shielding for Safe Reinforcement Learning
von: Goodall, Alexander W., et al.
Veröffentlicht: (2023)
von: Goodall, Alexander W., et al.
Veröffentlicht: (2023)
CPS-LLM: Large Language Model based Safe Usage Plan Generator for Human-in-the-Loop Human-in-the-Plant Cyber-Physical System
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
Safe Deep Model-Based Reinforcement Learning with Lyapunov Functions
von: Zhang, Harry
Veröffentlicht: (2024)
von: Zhang, Harry
Veröffentlicht: (2024)
Reinforcement Learning Constrained Beam Search for Parameter Optimization of Paper Drying Under Flexible Constraints
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
von: Chen, Siyuan, et al.
Veröffentlicht: (2025)
Learning to Drive Safely with Hybrid Options
von: De Cooman, Bram, et al.
Veröffentlicht: (2025)
von: De Cooman, Bram, et al.
Veröffentlicht: (2025)
Large Language Model Powered Automated Modeling and Optimization of Active Distribution Network Dispatch Problems
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
Parallel Differentiable Reachability for Learning and Planning with Certified Neural Dynamics and Controllers
von: Shen, Keyi, et al.
Veröffentlicht: (2026)
von: Shen, Keyi, et al.
Veröffentlicht: (2026)
Certified Training with Branch-and-Bound for Lyapunov-stable Neural Control
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
von: Shi, Zhouxing, et al.
Veröffentlicht: (2024)
Formal Synthesis of Certifiably Robust Neural Lyapunov-Barrier Certificates
von: Wang, Chengxiao, et al.
Veröffentlicht: (2026)
von: Wang, Chengxiao, et al.
Veröffentlicht: (2026)
Intersection of Reinforcement Learning and Bayesian Optimization for Intelligent Control of Industrial Processes: A Safe MPC-based DPG using Multi-Objective BO
von: Esfahani, Hossein Nejatbakhsh, et al.
Veröffentlicht: (2025)
von: Esfahani, Hossein Nejatbakhsh, et al.
Veröffentlicht: (2025)
Provably Safe Generative Sampling with Constricting Barrier Functions
von: Gadginmath, Darshan, et al.
Veröffentlicht: (2026)
von: Gadginmath, Darshan, et al.
Veröffentlicht: (2026)
Constrained Reinforcement Learning for Safe Heat Pump Control
von: Zhang, Baohe, et al.
Veröffentlicht: (2024)
von: Zhang, Baohe, et al.
Veröffentlicht: (2024)
GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
von: Zhou, Zhehua, et al.
Veröffentlicht: (2024)
von: Zhou, Zhehua, et al.
Veröffentlicht: (2024)
Learning Contextual Runtime Monitors for Safe AI-Based Autonomy
von: Luque-Cerpa, Alejandro, et al.
Veröffentlicht: (2026)
von: Luque-Cerpa, Alejandro, et al.
Veröffentlicht: (2026)
Balancing SoC in Battery Cells using Safe Action Perturbations
von: Yadav, E Harshith Kumar, et al.
Veröffentlicht: (2025)
von: Yadav, E Harshith Kumar, et al.
Veröffentlicht: (2025)
Guided Safe Shooting: model based reinforcement learning with safety constraints
von: Paolo, Giuseppe, et al.
Veröffentlicht: (2022)
von: Paolo, Giuseppe, et al.
Veröffentlicht: (2022)
Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction
von: Durkin, Alex, et al.
Veröffentlicht: (2025)
von: Durkin, Alex, et al.
Veröffentlicht: (2025)
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
von: Banerjee, Chayan, et al.
Veröffentlicht: (2025)
von: Banerjee, Chayan, et al.
Veröffentlicht: (2025)
Action Mapping for Reinforcement Learning in Continuous Environments with Constraints
von: Theile, Mirco, et al.
Veröffentlicht: (2024)
von: Theile, Mirco, et al.
Veröffentlicht: (2024)
A Safe Reinforcement Learning driven Weights-varying Model Predictive Control for Autonomous Vehicle Motion Control
von: Zarrouki, Baha, et al.
Veröffentlicht: (2024)
von: Zarrouki, Baha, et al.
Veröffentlicht: (2024)
DCcluster-Opt: Benchmarking Dynamic Multi-Objective Optimization for Geo-Distributed Data Center Workloads
von: Guillen-Perez, Antonio, et al.
Veröffentlicht: (2025)
von: Guillen-Perez, Antonio, et al.
Veröffentlicht: (2025)
Learning Hidden Subgoals under Temporal Ordering Constraints in Reinforcement Learning
von: Xu, Duo, et al.
Veröffentlicht: (2024)
von: Xu, Duo, et al.
Veröffentlicht: (2024)
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
von: Fard, Amir, et al.
Veröffentlicht: (2025)
von: Fard, Amir, et al.
Veröffentlicht: (2025)
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
von: Spoor, Lindsay, et al.
Veröffentlicht: (2025)
von: Spoor, Lindsay, et al.
Veröffentlicht: (2025)
Benchmarking Model Predictive Control Algorithms in Building Optimization Testing Framework (BOPTEST)
von: Mostafavi, Saman, et al.
Veröffentlicht: (2023)
von: Mostafavi, Saman, et al.
Veröffentlicht: (2023)
Trustworthy and Explainable Deep Reinforcement Learning for Safe and Energy-Efficient Process Control: A Use Case in Industrial Compressed Air Systems
von: Bezold, Vincent, et al.
Veröffentlicht: (2025)
von: Bezold, Vincent, et al.
Veröffentlicht: (2025)
Deep Model Predictive Optimization
von: Sacks, Jacob, et al.
Veröffentlicht: (2023)
von: Sacks, Jacob, et al.
Veröffentlicht: (2023)
Test Time Training for AC Power Flow Surrogates via Physics and Operational Constraint Refinement
von: Dogoulis, Panteleimon, et al.
Veröffentlicht: (2025)
von: Dogoulis, Panteleimon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026) -
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025) -
Provably Efficient Sample Complexity for Robust CMDP
von: Ganguly, Sourav, et al.
Veröffentlicht: (2025) -
SPICEPilot: Navigating SPICE Code Generation and Simulation with AI Guidance
von: Vungarala, Deepak, et al.
Veröffentlicht: (2024) -
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
von: Vungarala, Deepak, et al.
Veröffentlicht: (2025)