Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
Fuente:
arXiv
Guardado en:
| Autores principales: | Kezins, Nikita, Ekka, Urbas, Berrang, Pascal, Arnaboldi, Luca |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Automatic LLM Red Teaming
por: Belaire, Roman, et al.
Publicado: (2025)
por: Belaire, Roman, et al.
Publicado: (2025)
PAC-Bayesian Generalization Guarantees for Fairness on Stochastic and Deterministic Classifiers
por: Bastian, Julien, et al.
Publicado: (2026)
por: Bastian, Julien, et al.
Publicado: (2026)
3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models
por: Yin, Jianyao, et al.
Publicado: (2025)
por: Yin, Jianyao, et al.
Publicado: (2025)
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
por: Arnaboldi, Luca, et al.
Publicado: (2023)
por: Arnaboldi, Luca, et al.
Publicado: (2023)
Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis
por: Hafez, Ahmad, et al.
Publicado: (2025)
por: Hafez, Ahmad, et al.
Publicado: (2025)
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
por: Arnaboldi, Luca, et al.
Publicado: (2024)
por: Arnaboldi, Luca, et al.
Publicado: (2024)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
por: Bach, Thong, et al.
Publicado: (2026)
por: Bach, Thong, et al.
Publicado: (2026)
Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
por: Pan, Jiazhen, et al.
Publicado: (2025)
por: Pan, Jiazhen, et al.
Publicado: (2025)
LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
por: Sagar, Som, et al.
Publicado: (2024)
por: Sagar, Som, et al.
Publicado: (2024)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
por: He, Pengfei, et al.
Publicado: (2026)
por: He, Pengfei, et al.
Publicado: (2026)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
por: Sharma, Mrinank, et al.
Publicado: (2025)
por: Sharma, Mrinank, et al.
Publicado: (2025)
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
por: Li, Xinyu, et al.
Publicado: (2026)
por: Li, Xinyu, et al.
Publicado: (2026)
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
por: Schoepf, Stefan, et al.
Publicado: (2025)
por: Schoepf, Stefan, et al.
Publicado: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
Link Stealing Attacks Against Inductive Graph Neural Networks
por: Wu, Yixin, et al.
Publicado: (2024)
por: Wu, Yixin, et al.
Publicado: (2024)
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
por: Arnaboldi, Luca, et al.
Publicado: (2024)
por: Arnaboldi, Luca, et al.
Publicado: (2024)
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
por: Dandi, Yatin, et al.
Publicado: (2024)
por: Dandi, Yatin, et al.
Publicado: (2024)
Abstractive Red-Teaming of Language Model Character
por: Rahn, Nate, et al.
Publicado: (2026)
por: Rahn, Nate, et al.
Publicado: (2026)
Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance
por: Kwon, Minchan, et al.
Publicado: (2026)
por: Kwon, Minchan, et al.
Publicado: (2026)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Adversarial Robustness Guarantees for Quantum Classifiers
por: Dowling, Neil, et al.
Publicado: (2024)
por: Dowling, Neil, et al.
Publicado: (2024)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
por: Taheri, Hossein, et al.
Publicado: (2024)
por: Taheri, Hossein, et al.
Publicado: (2024)
Explainable Clustering Beyond Worst-Case Guarantees
por: Fleissner, Maximilian, et al.
Publicado: (2024)
por: Fleissner, Maximilian, et al.
Publicado: (2024)
Geometric Red-Teaming for Robotic Manipulation
por: Goel, Divyam, et al.
Publicado: (2025)
por: Goel, Divyam, et al.
Publicado: (2025)
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
por: Zheng, Xiang, et al.
Publicado: (2025)
por: Zheng, Xiang, et al.
Publicado: (2025)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
por: Sinha, Anusha, et al.
Publicado: (2025)
por: Sinha, Anusha, et al.
Publicado: (2025)
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
por: Hadad, Itamar, et al.
Publicado: (2026)
por: Hadad, Itamar, et al.
Publicado: (2026)
Quantifying Multimodal Capabilities: Formal Generalization Guarantees in Pairwise Metric Learning
por: Zhou, Richeng, et al.
Publicado: (2026)
por: Zhou, Richeng, et al.
Publicado: (2026)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
por: Deng, Yihe, et al.
Publicado: (2025)
por: Deng, Yihe, et al.
Publicado: (2025)
Efficient Evaluation of LLM Performance with Statistical Guarantees
por: Wu, Skyler, et al.
Publicado: (2026)
por: Wu, Skyler, et al.
Publicado: (2026)
ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
por: Guo, Taicheng, et al.
Publicado: (2023)
por: Guo, Taicheng, et al.
Publicado: (2023)
The 4/$δ$ Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee
por: Dantas, PIerre, et al.
Publicado: (2025)
por: Dantas, PIerre, et al.
Publicado: (2025)
Interpretability Guarantees with Merlin-Arthur Classifiers
por: Wäldchen, Stephan, et al.
Publicado: (2022)
por: Wäldchen, Stephan, et al.
Publicado: (2022)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
por: Dandi, Yatin, et al.
Publicado: (2026)
por: Dandi, Yatin, et al.
Publicado: (2026)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
por: Arnaboldi, Luca, et al.
Publicado: (2025)
por: Arnaboldi, Luca, et al.
Publicado: (2025)
Embodied Red Teaming for Auditing Robotic Foundation Models
por: Karnik, Sathwik, et al.
Publicado: (2024)
por: Karnik, Sathwik, et al.
Publicado: (2024)
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
por: Pavlova, Maya, et al.
Publicado: (2024)
por: Pavlova, Maya, et al.
Publicado: (2024)
Red-Teaming for Inducing Societal Bias in Large Language Models
por: Luo, Chu Fei, et al.
Publicado: (2024)
por: Luo, Chu Fei, et al.
Publicado: (2024)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
por: Zhang, Jiawei, et al.
Publicado: (2025)
por: Zhang, Jiawei, et al.
Publicado: (2025)
Red-Teaming Segment Anything Model
por: Jankowski, Krzysztof, et al.
Publicado: (2024)
por: Jankowski, Krzysztof, et al.
Publicado: (2024)
Ejemplares similares
-
Automatic LLM Red Teaming
por: Belaire, Roman, et al.
Publicado: (2025) -
PAC-Bayesian Generalization Guarantees for Fairness on Stochastic and Deterministic Classifiers
por: Bastian, Julien, et al.
Publicado: (2026) -
3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models
por: Yin, Jianyao, et al.
Publicado: (2025) -
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
por: Arnaboldi, Luca, et al.
Publicado: (2023) -
Safe LLM-Controlled Robots with Formal Guarantees via Reachability Analysis
por: Hafez, Ahmad, et al.
Publicado: (2025)