From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
Fuente:
arXiv
Guardado en:
| Autores principales: | Pandya, Ravi, Bland, Madison, Nguyen, Duy P., Liu, Changliu, Fisac, Jaime Fernández, Bajcsy, Andrea |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
por: Bajcsy, Andrea, et al.
Publicado: (2024)
por: Bajcsy, Andrea, et al.
Publicado: (2024)
Robots that Learn to Safely Influence via Prediction-Informed Reach-Avoid Dynamic Games
por: Pandya, Ravi, et al.
Publicado: (2024)
por: Pandya, Ravi, et al.
Publicado: (2024)
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
por: Wang, Justin, et al.
Publicado: (2024)
por: Wang, Justin, et al.
Publicado: (2024)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
por: Alagharu, Rishab, et al.
Publicado: (2026)
por: Alagharu, Rishab, et al.
Publicado: (2026)
Toward Reusability of AI Models Using Dynamic Updates of AI Documentation
por: Bajcsy, Peter, et al.
Publicado: (2026)
por: Bajcsy, Peter, et al.
Publicado: (2026)
Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
por: Liang, Kaiqu, et al.
Publicado: (2024)
por: Liang, Kaiqu, et al.
Publicado: (2024)
Lattice: Generative Guardrails for Conversational Agents
por: Broadhurst, Emily, et al.
Publicado: (2026)
por: Broadhurst, Emily, et al.
Publicado: (2026)
Multimodal Safe Control for Human-Robot Interaction
por: Pandya, Ravi, et al.
Publicado: (2023)
por: Pandya, Ravi, et al.
Publicado: (2023)
AssemblyComplete: 3D Combinatorial Construction with Deep Reinforcement Learning
por: Chen, Alan, et al.
Publicado: (2024)
por: Chen, Alan, et al.
Publicado: (2024)
Enhancing Guardrails for Safe and Secure Healthcare AI
por: Gangavarapu, Ananya
Publicado: (2024)
por: Gangavarapu, Ananya
Publicado: (2024)
Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
por: García-Ferrero, Iker, et al.
Publicado: (2025)
por: García-Ferrero, Iker, et al.
Publicado: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
por: Shu, Huizhen, et al.
Publicado: (2025)
por: Shu, Huizhen, et al.
Publicado: (2025)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
por: Liang, Kaiqu, et al.
Publicado: (2025)
por: Liang, Kaiqu, et al.
Publicado: (2025)
Building Effective Safety Guardrails in AI Education Tools
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
por: Clark, Hannah-Beth, et al.
Publicado: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
por: Li, Qi, et al.
Publicado: (2026)
por: Li, Qi, et al.
Publicado: (2026)
Estimating Neural Network Robustness via Lipschitz Constant and Architecture Sensitivity
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
por: Kholkar, Gauri, et al.
Publicado: (2025)
por: Kholkar, Gauri, et al.
Publicado: (2025)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
por: Ren, Xuancheng, et al.
Publicado: (2026)
por: Ren, Xuancheng, et al.
Publicado: (2026)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Control Invariant Sets for Neural Network Dynamical Systems and Recursive Feasibility in Model Predictive Control
por: Li, Xiao, et al.
Publicado: (2025)
por: Li, Xiao, et al.
Publicado: (2025)
To Use or to Refuse? Re-Centering Student Agency with Generative AI in Engineering Design Education
por: Willems, Thijs, et al.
Publicado: (2025)
por: Willems, Thijs, et al.
Publicado: (2025)
From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI
por: Koch, Christopher
Publicado: (2026)
por: Koch, Christopher
Publicado: (2026)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
por: She, Yining, et al.
Publicado: (2025)
por: She, Yining, et al.
Publicado: (2025)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
por: Giarrusso, Francesco, et al.
Publicado: (2025)
por: Giarrusso, Francesco, et al.
Publicado: (2025)
Meta-Control: Automatic Model-based Control Synthesis for Heterogeneous Robot Skills
por: Wei, Tianhao, et al.
Publicado: (2024)
por: Wei, Tianhao, et al.
Publicado: (2024)
Current state of LLM Risks and AI Guardrails
por: Ayyamperumal, Suriya Ganesh, et al.
Publicado: (2024)
por: Ayyamperumal, Suriya Ganesh, et al.
Publicado: (2024)
Learn to Refuse: Making Large Language Models More Controllable and Reliable through Knowledge Scope Limitation and Refusal Mechanism
por: Cao, Lang
Publicado: (2023)
por: Cao, Lang
Publicado: (2023)
No Free Lunch with Guardrails
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
Simultaneous Task Allocation and Planning for Multi-Robots under Hierarchical Temporal Logic Specifications
por: Luo, Xusheng, et al.
Publicado: (2024)
por: Luo, Xusheng, et al.
Publicado: (2024)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
por: Jin, Xisen, et al.
Publicado: (2026)
por: Jin, Xisen, et al.
Publicado: (2026)
Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries
por: Noever, David, et al.
Publicado: (2025)
por: Noever, David, et al.
Publicado: (2025)
Initial Risk Probing and Feasibility Testing of Glow: a Generative AI-Powered Dialectical Behavior Therapy Skills Coach for Substance Use Recovery and HIV Prevention
por: Wang, Liying, et al.
Publicado: (2026)
por: Wang, Liying, et al.
Publicado: (2026)
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules
por: Pattison, Cameron, et al.
Publicado: (2026)
por: Pattison, Cameron, et al.
Publicado: (2026)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
por: Bertollo, Giacomo, et al.
Publicado: (2025)
por: Bertollo, Giacomo, et al.
Publicado: (2025)
HanoiWorld : A Joint Embedding Predictive Architecture BasedWorld Model for Autonomous Vehicle Controller
por: Dat, Tran Tien, et al.
Publicado: (2026)
por: Dat, Tran Tien, et al.
Publicado: (2026)
Guardrails for avoiding harmful medical product recommendations and off-label promotion in generative AI models
por: Lopez-Martinez, Daniel
Publicado: (2024)
por: Lopez-Martinez, Daniel
Publicado: (2024)
OneShield -- the Next Generation of LLM Guardrails
por: DeLuca, Chad, et al.
Publicado: (2025)
por: DeLuca, Chad, et al.
Publicado: (2025)
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
por: Mavi, John, et al.
Publicado: (2025)
por: Mavi, John, et al.
Publicado: (2025)
Ejemplares similares
-
Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
por: Bajcsy, Andrea, et al.
Publicado: (2024) -
Robots that Learn to Safely Influence via Prediction-Informed Reach-Avoid Dynamic Games
por: Pandya, Ravi, et al.
Publicado: (2024) -
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
por: Wang, Justin, et al.
Publicado: (2024) -
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
por: Alagharu, Rishab, et al.
Publicado: (2026) -
Toward Reusability of AI Models Using Dynamic Updates of AI Documentation
por: Bajcsy, Peter, et al.
Publicado: (2026)