No Free Lunch with Guardrails
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Divyanshu, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quantifying CBRN Risk in Frontier Models
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
por: Kumar, Anurakt, et al.
Publicado: (2024)
por: Kumar, Anurakt, et al.
Publicado: (2024)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
por: Kumar, Divyanshu, et al.
Publicado: (2024)
por: Kumar, Divyanshu, et al.
Publicado: (2024)
VERA: Validation and Enhancement for Retrieval Augmented systems
por: Birur, Nitin Aravind, et al.
Publicado: (2024)
por: Birur, Nitin Aravind, et al.
Publicado: (2024)
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
por: Kumar, Divyanshu, et al.
Publicado: (2026)
por: Kumar, Divyanshu, et al.
Publicado: (2026)
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
por: Kumar, Divyanshu, et al.
Publicado: (2026)
por: Kumar, Divyanshu, et al.
Publicado: (2026)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
por: Kumar, Divyanshu, et al.
Publicado: (2024)
por: Kumar, Divyanshu, et al.
Publicado: (2024)
No Free Lunch Theorem for Privacy-Preserving LLM Inference
por: Zhang, Xiaojin, et al.
Publicado: (2024)
por: Zhang, Xiaojin, et al.
Publicado: (2024)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
por: Xue, Zhiyu, et al.
Publicado: (2024)
por: Xue, Zhiyu, et al.
Publicado: (2024)
Provably Secure Agent Guardrail
por: Wu, Benlong, et al.
Publicado: (2026)
por: Wu, Benlong, et al.
Publicado: (2026)
Enhancing Guardrails for Safe and Secure Healthcare AI
por: Gangavarapu, Ananya
Publicado: (2024)
por: Gangavarapu, Ananya
Publicado: (2024)
AgentWall: A Runtime Safety Layer for Local AI Agents
por: Aravind, Ashwin
Publicado: (2026)
por: Aravind, Ashwin
Publicado: (2026)
A Comparative Evaluation of AI Agent Security Guardrails
por: Li, Qi, et al.
Publicado: (2026)
por: Li, Qi, et al.
Publicado: (2026)
Cognitive Cybersecurity for Artificial Intelligence: Guardrail Engineering with CCS-7
por: Aydin, Yuksel
Publicado: (2025)
por: Aydin, Yuksel
Publicado: (2025)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
por: Wang, Xunguang, et al.
Publicado: (2025)
por: Wang, Xunguang, et al.
Publicado: (2025)
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails
por: Li, Nanxi, et al.
Publicado: (2026)
por: Li, Nanxi, et al.
Publicado: (2026)
Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
por: Bertollo, Giacomo, et al.
Publicado: (2025)
por: Bertollo, Giacomo, et al.
Publicado: (2025)
Precision Guided Approach to Mitigate Data Poisoning Attacks in Federated Learning
por: Kumar, K Naveen, et al.
Publicado: (2024)
por: Kumar, K Naveen, et al.
Publicado: (2024)
Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks
por: Wu, ChenYu, et al.
Publicado: (2025)
por: Wu, ChenYu, et al.
Publicado: (2025)
Free Lunch for Federated Remote Sensing Target Fine-Grained Classification: A Parameter-Efficient Framework
por: Chen, Shengchao, et al.
Publicado: (2024)
por: Chen, Shengchao, et al.
Publicado: (2024)
Quantum Gatekeeper: Multi-Factor Context-Bound Image Steganography with VQC Based Key Derivation on Quantum Hardware
por: Tomar, Sahil, et al.
Publicado: (2026)
por: Tomar, Sahil, et al.
Publicado: (2026)
Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection
por: Hasan, Najmul, et al.
Publicado: (2026)
por: Hasan, Najmul, et al.
Publicado: (2026)
Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts
por: Rani, Nanda, et al.
Publicado: (2024)
por: Rani, Nanda, et al.
Publicado: (2024)
OneShield -- the Next Generation of LLM Guardrails
por: DeLuca, Chad, et al.
Publicado: (2025)
por: DeLuca, Chad, et al.
Publicado: (2025)
A Survey on Offensive AI Within Cybersecurity
por: Girhepuje, Sahil, et al.
Publicado: (2024)
por: Girhepuje, Sahil, et al.
Publicado: (2024)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
por: Jin, Xisen, et al.
Publicado: (2026)
por: Jin, Xisen, et al.
Publicado: (2026)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
por: Das, Saswat, et al.
Publicado: (2026)
por: Das, Saswat, et al.
Publicado: (2026)
SGuard-v1: Safety Guardrail for Large Language Models
por: Lee, JoonHo, et al.
Publicado: (2025)
por: Lee, JoonHo, et al.
Publicado: (2025)
SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
por: Liu, Zhe, et al.
Publicado: (2026)
por: Liu, Zhe, et al.
Publicado: (2026)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
por: Zafar, Osama, et al.
Publicado: (2026)
por: Zafar, Osama, et al.
Publicado: (2026)
Current state of LLM Risks and AI Guardrails
por: Ayyamperumal, Suriya Ganesh, et al.
Publicado: (2024)
por: Ayyamperumal, Suriya Ganesh, et al.
Publicado: (2024)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
por: Biswas, Anjanava, et al.
Publicado: (2026)
por: Biswas, Anjanava, et al.
Publicado: (2026)
ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
por: Owiredu-Ashley, Harry
Publicado: (2026)
por: Owiredu-Ashley, Harry
Publicado: (2026)
SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection
por: Hasan, Touseef, et al.
Publicado: (2026)
por: Hasan, Touseef, et al.
Publicado: (2026)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
por: Hong, Yining, et al.
Publicado: (2026)
por: Hong, Yining, et al.
Publicado: (2026)
In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
por: Durner, Nils
Publicado: (2025)
por: Durner, Nils
Publicado: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
por: Wang, Libo
Publicado: (2024)
por: Wang, Libo
Publicado: (2024)
Ejemplares similares
-
Quantifying CBRN Risk in Frontier Models
por: Kumar, Divyanshu, et al.
Publicado: (2025) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
por: Kumar, Anurakt, et al.
Publicado: (2024) -
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
por: Kumar, Divyanshu, et al.
Publicado: (2025) -
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
por: Kumar, Divyanshu, et al.
Publicado: (2024) -
VERA: Validation and Enhancement for Retrieval Augmented systems
por: Birur, Nitin Aravind, et al.
Publicado: (2024)