LLM-Assisted Red Teaming of Diffusion Models through "Failures Are Fated, But Can Be Faded"
Fuente:
arXiv
Saved in:
| Main Authors: | Sagar, Som, Taparia, Aditya, Senanayake, Ransalu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations
by: Taparia, Aditya, et al.
Published: (2024)
by: Taparia, Aditya, et al.
Published: (2024)
BaTCAVe: Trustworthy Explanations for Robot Behaviors
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
ExpressivityBench: Can LLMs Communicate Implicitly?
by: Tint, Joshua, et al.
Published: (2024)
by: Tint, Joshua, et al.
Published: (2024)
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM
by: Nguyen, Son, et al.
Published: (2026)
by: Nguyen, Son, et al.
Published: (2026)
PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?
by: Gundawar, Atharva, et al.
Published: (2025)
by: Gundawar, Atharva, et al.
Published: (2025)
Viewpoint-Agnostic Manipulation Policies with Strategic Vantage Selection
by: Vasudevan, Sreevishakh, et al.
Published: (2025)
by: Vasudevan, Sreevishakh, et al.
Published: (2025)
Multiple Distribution Shift -- Aerial (MDS-A): A Dataset for Test-Time Error Detection and Model Adaptation
by: Ngu, Noel, et al.
Published: (2025)
by: Ngu, Noel, et al.
Published: (2025)
Towards Adapting Reinforcement Learning Agents to New Tasks: Insights from Q-Values
by: Ramaswamy, Ashwin, et al.
Published: (2024)
by: Ramaswamy, Ashwin, et al.
Published: (2024)
The Anatomy of Uncertainty in LLMs
by: Taparia, Aditya, et al.
Published: (2026)
by: Taparia, Aditya, et al.
Published: (2026)
Consistency-based Abductive Reasoning over Perceptual Errors of Multiple Pre-trained Models in Novel Environments
by: Leiva, Mario, et al.
Published: (2025)
by: Leiva, Mario, et al.
Published: (2025)
The Role of Predictive Uncertainty and Diversity in Embodied AI and Robot Learning
by: Senanayake, Ransalu
Published: (2024)
by: Senanayake, Ransalu
Published: (2024)
Automatic LLM Red Teaming
by: Belaire, Roman, et al.
Published: (2025)
by: Belaire, Roman, et al.
Published: (2025)
Modeling Multi-Objective Tradeoffs with Monotonic Utility Functions
by: Chen, Edward, et al.
Published: (2024)
by: Chen, Edward, et al.
Published: (2024)
Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
by: Kezins, Nikita, et al.
Published: (2026)
by: Kezins, Nikita, et al.
Published: (2026)
Abstractive Red-Teaming of Language Model Character
by: Rahn, Nate, et al.
Published: (2026)
by: Rahn, Nate, et al.
Published: (2026)
LLM Routing as Reasoning: A MaxSAT View
by: Nguyen, Son, et al.
Published: (2026)
by: Nguyen, Son, et al.
Published: (2026)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
by: Nöther, Jonathan, et al.
Published: (2025)
by: Nöther, Jonathan, et al.
Published: (2025)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
Consistent Diffusion Language Models
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
by: Nunez, Elvis, et al.
Published: (2024)
by: Nunez, Elvis, et al.
Published: (2024)
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
by: Schoepf, Stefan, et al.
Published: (2025)
by: Schoepf, Stefan, et al.
Published: (2025)
OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode
by: He, Zeyuan, et al.
Published: (2026)
by: He, Zeyuan, et al.
Published: (2026)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
Embodied Red Teaming for Auditing Robotic Foundation Models
by: Karnik, Sathwik, et al.
Published: (2024)
by: Karnik, Sathwik, et al.
Published: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Red-Teaming for Inducing Societal Bias in Large Language Models
by: Luo, Chu Fei, et al.
Published: (2024)
by: Luo, Chu Fei, et al.
Published: (2024)
Quantile Activation: Correcting a Failure Mode of ML Models
by: Challa, Aditya, et al.
Published: (2024)
by: Challa, Aditya, et al.
Published: (2024)
Geometric Red-Teaming for Robotic Manipulation
by: Goel, Divyam, et al.
Published: (2025)
by: Goel, Divyam, et al.
Published: (2025)
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
by: Zheng, Xiang, et al.
Published: (2025)
by: Zheng, Xiang, et al.
Published: (2025)
Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
by: Patel, Nyal, et al.
Published: (2025)
by: Patel, Nyal, et al.
Published: (2025)
The Feedback Hamiltonian is the Score Function: A Diffusion-Model Framework for Quantum Trajectory Reversal
by: Dubey, Sagar, et al.
Published: (2026)
by: Dubey, Sagar, et al.
Published: (2026)
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
by: Li, Chunxiao, et al.
Published: (2026)
by: Li, Chunxiao, et al.
Published: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025)
by: Sinha, Anusha, et al.
Published: (2025)
Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
by: Pan, Jiazhen, et al.
Published: (2025)
by: Pan, Jiazhen, et al.
Published: (2025)
Quantifying Out-of-Training Uncertainty of Neural-Network based Turbulence Closures
by: Grogan, Cody, et al.
Published: (2025)
by: Grogan, Cody, et al.
Published: (2025)
Similar Items
-
Failures Are Fated, But Can Be Faded: Characterizing and Mitigating Unwanted Behaviors in Large-Scale Vision and Language Models
by: Sagar, Som, et al.
Published: (2024) -
Explainable Concept Generation through Vision-Language Preference Learning for Understanding Neural Networks' Internal Representations
by: Taparia, Aditya, et al.
Published: (2024) -
BaTCAVe: Trustworthy Explanations for Robot Behaviors
by: Sagar, Som, et al.
Published: (2024) -
ExpressivityBench: Can LLMs Communicate Implicitly?
by: Tint, Joshua, et al.
Published: (2024) -
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
by: Sagar, Som, et al.
Published: (2024)