Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Dongkyu Derek, Song, Huan, Chowdhury, Arijit Ghosh, An, Haotian, Wang, Yawei, Thekkanal, Rohit, Sokhandan, Negin, Keshava, Sharlina, Marlowe, Hannah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
by: Song, Huan, et al.
Published: (2025)
by: Song, Huan, et al.
Published: (2025)
Improved Few-Shot Image Classification Through Multiple-Choice Questions
by: Khullar, Dipika, et al.
Published: (2024)
by: Khullar, Dipika, et al.
Published: (2024)
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Automated Virtual Product Placement and Assessment in Images using Diffusion Models
by: Alam, Mohammad Mahmudul, et al.
Published: (2024)
by: Alam, Mohammad Mahmudul, et al.
Published: (2024)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Breaking Free: How to Hack Safety Guardrails in Black-Box Diffusion Models!
by: Kotyan, Shashank, et al.
Published: (2024)
by: Kotyan, Shashank, et al.
Published: (2024)
Building Effective Safety Guardrails in AI Education Tools
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Safety Guardrails for LLM-Enabled Robots
by: Ravichandran, Zachary, et al.
Published: (2025)
by: Ravichandran, Zachary, et al.
Published: (2025)
AGrail: A Lifelong Agent Guardrail with Effective and Adaptive Safety Detection
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents
by: Sah, Tanmay, et al.
Published: (2026)
by: Sah, Tanmay, et al.
Published: (2026)
Test-Time Training Undermines Safety Guardrails
by: Antonelli, Simone, et al.
Published: (2026)
by: Antonelli, Simone, et al.
Published: (2026)
Bypassing Safety Guardrails in LLMs Using Humor
by: Cisneros-Velarde, Pedro
Published: (2025)
by: Cisneros-Velarde, Pedro
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
Backtracking for Safety
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
The Safety-Privacy Tradeoff in Linear Bandits
by: Zibaie, Arghavan, et al.
Published: (2025)
by: Zibaie, Arghavan, et al.
Published: (2025)
Fast Computer Model Calibration using Annealed and Transformed Variational Inference
by: Cho, Dongkyu Derek, et al.
Published: (2022)
by: Cho, Dongkyu Derek, et al.
Published: (2022)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
by: Nagireddy, Manish, et al.
Published: (2024)
by: Nagireddy, Manish, et al.
Published: (2024)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
by: Zheng, Aaron, et al.
Published: (2024)
by: Zheng, Aaron, et al.
Published: (2024)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
by: Bach, Thong, et al.
Published: (2026)
by: Bach, Thong, et al.
Published: (2026)
A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
by: Zhang, Bingjie, et al.
Published: (2025)
by: Zhang, Bingjie, et al.
Published: (2025)
RLNVR: Reinforcement Learning from Non-Verified Real-World Rewards
by: Krishnan, Rohit, et al.
Published: (2025)
by: Krishnan, Rohit, et al.
Published: (2025)
Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
Belonging and Transnational Refugee Settlement
by: Marlowe, Jay
Published: (2021)
by: Marlowe, Jay
Published: (2021)
Belonging and Transnational Refugee Settlement
by: Marlowe, Jay
Published: (2021)
by: Marlowe, Jay
Published: (2021)
Cutting Loose.
by: Marlowe, John
Published: (1994)
by: Marlowe, John
Published: (1994)
Dealing with the Evil Twins: Improving Random Augmentation by Addressing Catastrophic Forgetting of Diverse Augmentations
by: Cho, Dongkyu, et al.
Published: (2025)
by: Cho, Dongkyu, et al.
Published: (2025)
The Supportiveness-Safety Tradeoff in LLM Well-Being Agents
by: Lalwani, Himanshi, et al.
Published: (2026)
by: Lalwani, Himanshi, et al.
Published: (2026)
Elucidating Optimal Reward-Diversity Tradeoffs in Text-to-Image Diffusion Models
by: Jena, Rohit, et al.
Published: (2024)
by: Jena, Rohit, et al.
Published: (2024)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
by: Wu, Yaozu, et al.
Published: (2025)
by: Wu, Yaozu, et al.
Published: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
by: Liu, Zhe, et al.
Published: (2026)
by: Liu, Zhe, et al.
Published: (2026)
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
by: Han, Shanshan, et al.
Published: (2025)
by: Han, Shanshan, et al.
Published: (2025)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
by: Yang, Yahan, et al.
Published: (2025)
by: Yang, Yahan, et al.
Published: (2025)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
by: Abdulkadir, Perry
Published: (2025)
by: Abdulkadir, Perry
Published: (2025)
Similar Items
-
Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
by: Song, Huan, et al.
Published: (2025) -
Improved Few-Shot Image Classification Through Multiple-Choice Questions
by: Khullar, Dipika, et al.
Published: (2024) -
Implicit Updates for Average-Reward Temporal Difference Learning
by: Kim, Hwanwoo, et al.
Published: (2025) -
Automated Virtual Product Placement and Assessment in Images using Diffusion Models
by: Alam, Mohammad Mahmudul, et al.
Published: (2024) -
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)