Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
Fuente:
arXiv
Saved in:
| Main Authors: | Ilin, Aleksei, Matevosyan, Gor, Ma, Xueying, Eremin, Vladimir, Dada, Suhaa, Li, Muqun, Shaik, Riyaaz, Tokgozoglu, Haluk Noyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025)
by: Guo, Yihao, et al.
Published: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026)
by: Islam, Md Asiful, et al.
Published: (2026)
FUELVISION: A Multimodal Data Fusion and Multimodel Ensemble Algorithm for Wildfire Fuels Mapping
by: Shaik, Riyaaz Uddien, et al.
Published: (2024)
by: Shaik, Riyaaz Uddien, et al.
Published: (2024)
Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
by: Zheng, Aaron, et al.
Published: (2024)
by: Zheng, Aaron, et al.
Published: (2024)
Eigencone Constellations on Ranked Spheres
by: Matevosyan, Norayr
Published: (2026)
by: Matevosyan, Norayr
Published: (2026)
Weak (non)conservation and stochastic dynamics of angular momentum
by: Matevosyan, Ashot
Published: (2024)
by: Matevosyan, Ashot
Published: (2024)
Test-Time Training Undermines Safety Guardrails
by: Antonelli, Simone, et al.
Published: (2026)
by: Antonelli, Simone, et al.
Published: (2026)
Bethe subspaces and wonderful models for toric arrangements
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
Bethe subalgebras in Yangians and the wonderful compactification
by: Ilin, Aleksei, et al.
Published: (2018)
by: Ilin, Aleksei, et al.
Published: (2018)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
by: Niu, Muqun, et al.
Published: (2024)
by: Niu, Muqun, et al.
Published: (2024)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
by: Mazza, Arnon, et al.
Published: (2026)
by: Mazza, Arnon, et al.
Published: (2026)
Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
by: Cheng, Kellen Tan, et al.
Published: (2025)
by: Cheng, Kellen Tan, et al.
Published: (2025)
Mit Russlandhintergrund in Deutschland: Ansichten zu Politik, Gesellschaft und Geschichte
by: Krawatzek, Félix, et al.
Published: (2024)
by: Krawatzek, Félix, et al.
Published: (2024)
Gaudin models and moduli space of flower curves
by: Ilin, Aleksei, et al.
Published: (2024)
by: Ilin, Aleksei, et al.
Published: (2024)
Beable-guided measurement theory
by: Aleshin, Aleksei M., et al.
Published: (2024)
by: Aleshin, Aleksei M., et al.
Published: (2024)
Safety Guardrails for LLM-Enabled Robots
by: Ravichandran, Zachary, et al.
Published: (2025)
by: Ravichandran, Zachary, et al.
Published: (2025)
System-Bath Approach to Rotating Brownian Motion
by: Matevosyan, Ashot, et al.
Published: (2025)
by: Matevosyan, Ashot, et al.
Published: (2025)
Three-spheres theorem for harmonic functions (non-concentric case)
by: Arakelian, Norair U., et al.
Published: (2026)
by: Arakelian, Norair U., et al.
Published: (2026)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Bypassing Safety Guardrails in LLMs Using Humor
by: Cisneros-Velarde, Pedro
Published: (2025)
by: Cisneros-Velarde, Pedro
Published: (2025)
Partitioning the set of natural numbers into Mersenne trees and into arithmetic progressions; Natural Matrix and Linnik's constant
by: Eremin, Gennady
Published: (2024)
by: Eremin, Gennady
Published: (2024)
Unsupervised anomaly detection on cybersecurity data streams: a case with BETH dataset
by: Eremin, Evgeniy
Published: (2025)
by: Eremin, Evgeniy
Published: (2025)
Infinite matrix of odd natural numbers. A bit about Sophie Germain prime numbers
by: Eremin, Gennady
Published: (2025)
by: Eremin, Gennady
Published: (2025)
Sequential Synthetic Difference in Differences
by: Arkhangelsky, Dmitry, et al.
Published: (2024)
by: Arkhangelsky, Dmitry, et al.
Published: (2024)
Nonequilibrium noise emerging from broken detailed balance in active gels
by: Matevosyan, Ashot, et al.
Published: (2026)
by: Matevosyan, Ashot, et al.
Published: (2026)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Building Effective Safety Guardrails in AI Education Tools
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
PoseGuard: Pose-Guided Generation with Safety Guardrails
by: Wang, Kongxin, et al.
Published: (2025)
by: Wang, Kongxin, et al.
Published: (2025)
GADT: Enhancing Transferable Adversarial Attacks through Gradient-guided Adversarial Data Transformation
by: Ma, Yating, et al.
Published: (2024)
by: Ma, Yating, et al.
Published: (2024)
Evaluation of Library Utilization by Students Enrolled in External Degree Programme in University of Nairobi, Kenya
by: Gor, Peter Ochieng
Published: (2012)
by: Gor, Peter Ochieng
Published: (2012)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
by: Sreedhar, Makesh Narsimhan, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
by: Bach, Thong, et al.
Published: (2026)
by: Bach, Thong, et al.
Published: (2026)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
DGPO: RL-Steered Graph Diffusion for Neural Architecture Generation
by: Liuliakov, Aleksei, et al.
Published: (2026)
by: Liuliakov, Aleksei, et al.
Published: (2026)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
by: Bertollo, Giacomo, et al.
Published: (2025)
by: Bertollo, Giacomo, et al.
Published: (2025)
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
by: Zhao, Yunhan, et al.
Published: (2026)
by: Zhao, Yunhan, et al.
Published: (2026)
Similar Items
-
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025) -
A Lightweight Explainable Guardrail for Prompt Safety
by: Islam, Md Asiful, et al.
Published: (2026) -
FUELVISION: A Multimodal Data Fusion and Multimodel Ensemble Algorithm for Wildfire Fuels Mapping
by: Shaik, Riyaaz Uddien, et al.
Published: (2024) -
Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
by: Zheng, Aaron, et al.
Published: (2024) -
Eigencone Constellations on Ranked Spheres
by: Matevosyan, Norayr
Published: (2026)