Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Krishiv, Kaur, Ramneet, Samplawski, Colin, Acharya, Manoj, Roy, Anirban, Elenius, Daniel, Matejek, Brian, Cobb, Adam D., Jha, Susmit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Diffusion Models Dream of Electric Planes? Discrete and Continuous Simulation-Based Inference for Aircraft Design
by: Ghiglino, Aurelien, et al.
Published: (2026)
by: Ghiglino, Aurelien, et al.
Published: (2026)
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
by: Padhi, Trilok, et al.
Published: (2026)
by: Padhi, Trilok, et al.
Published: (2026)
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
by: Samplawski, Colin, et al.
Published: (2025)
by: Samplawski, Colin, et al.
Published: (2025)
Addressing Uncertainty in LLMs to Enhance Reliability in Generative AI
by: Kaur, Ramneet, et al.
Published: (2024)
by: Kaur, Ramneet, et al.
Published: (2024)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)
by: Padhi, Trilok, et al.
Published: (2025)
Privacy Preserving In-Context-Learning Framework for Large Language Models
by: Bhusal, Bishnu, et al.
Published: (2025)
by: Bhusal, Bishnu, et al.
Published: (2025)
AGENT: An Aerial Vehicle Generation and Design Tool Using Large Language Models
by: Samplawski, Colin, et al.
Published: (2025)
by: Samplawski, Colin, et al.
Published: (2025)
Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
by: Gupta, Ayush, et al.
Published: (2025)
by: Gupta, Ayush, et al.
Published: (2025)
TeleLoRA: Teleporting Model-Specific Alignment Across LLMs
by: Lin, Xiao, et al.
Published: (2025)
by: Lin, Xiao, et al.
Published: (2025)
Resource-Constrained Heuristic for Max-SAT
by: Matejek, Brian, et al.
Published: (2024)
by: Matejek, Brian, et al.
Published: (2024)
Spatio-Temporal Pruning for Compressed Spiking Large Language Models
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
Backpropagation-Free Metropolis-Adjusted Langevin Algorithm
by: Cobb, Adam D., et al.
Published: (2025)
by: Cobb, Adam D., et al.
Published: (2025)
Safety Monitoring for Learning-Enabled Cyber-Physical Systems in Out-of-Distribution Scenarios
by: Lin, Vivian, et al.
Published: (2025)
by: Lin, Vivian, et al.
Published: (2025)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026)
by: Babu, Abhijith, et al.
Published: (2026)
Taylor-Model Physics-Informed Neural Networks (PINNs) for Ordinary Differential Equations
by: Nagesh, Chandra Kanth, et al.
Published: (2025)
by: Nagesh, Chandra Kanth, et al.
Published: (2025)
SAFE-NID: Self-Attention with Normalizing-Flow Encodings for Network Intrusion Detection Dataset
by: Matejek, Brian, et al.
Published: (2025)
by: Matejek, Brian, et al.
Published: (2025)
Optimal Abstractions for Verifying Properties of Kolmogorov-Arnold Networks (KANs)
by: Schwartz, Noah, et al.
Published: (2026)
by: Schwartz, Noah, et al.
Published: (2026)
Second-Order Forward-Mode Automatic Differentiation for Optimization
by: Cobb, Adam D., et al.
Published: (2024)
by: Cobb, Adam D., et al.
Published: (2024)
Shrinking POMCP: A Framework for Real-Time UAV Search and Rescue
by: Zhang, Yunuo, et al.
Published: (2024)
by: Zhang, Yunuo, et al.
Published: (2024)
TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision
by: Gupta, Ayush, et al.
Published: (2025)
by: Gupta, Ayush, et al.
Published: (2025)
Concept-based Analysis of Neural Networks via Vision-Language Models
by: Mangal, Ravi, et al.
Published: (2024)
by: Mangal, Ravi, et al.
Published: (2024)
Characterizations and properties of solutions to parabolic problems of linear growth
by: Elenius, Theo
Published: (2025)
by: Elenius, Theo
Published: (2025)
Non-Markovian Quantum Control via Model Maximum Likelihood Estimation and Reinforcement Learning
by: Neema, Tanmay, et al.
Published: (2024)
by: Neema, Tanmay, et al.
Published: (2024)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
by: Goyal, Agam, et al.
Published: (2025)
by: Goyal, Agam, et al.
Published: (2025)
Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
by: Pramanik, Vishal, et al.
Published: (2026)
by: Pramanik, Vishal, et al.
Published: (2026)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
by: Aravindan, Ashwath Vaithinathan, et al.
Published: (2025)
On the Dataless Training of Neural Networks
by: Velasquez, Alvaro, et al.
Published: (2025)
by: Velasquez, Alvaro, et al.
Published: (2025)
The Chemistry of Character in Breaking Bad
by: Mittell, Jason
Published: (2025)
by: Mittell, Jason
Published: (2025)
Analyzing Memorization in Large Language Models through the Lens of Model Attribution
by: Menta, Tarun Ram, et al.
Published: (2025)
by: Menta, Tarun Ram, et al.
Published: (2025)
INTERLEAVE: A Faster Symbolic Algorithm for Maximal End Component Decomposition
by: Bansal, Suguman, et al.
Published: (2025)
by: Bansal, Suguman, et al.
Published: (2025)
Melanoma Detection with Uncertainty Quantification
by: Kim, SangHyuk, et al.
Published: (2024)
by: Kim, SangHyuk, et al.
Published: (2024)
Understanding Interpretability by generalized distillation in Supervised Classification
by: Agarwal, Adit, et al.
Published: (2020)
by: Agarwal, Adit, et al.
Published: (2020)
Doubling measures, Poincaré inequalities and parabolic Harnack inequalities for a doubly nonlinear equation
by: Elenius, Theo, et al.
Published: (2026)
by: Elenius, Theo, et al.
Published: (2026)
Breaking Bad: How Compilers Break Constant-Time Implementations
by: Schneider, Moritz, et al.
Published: (2024)
by: Schneider, Moritz, et al.
Published: (2024)
From Breaking Bad to Breaking Bonds—Mass Spectrometry in the Classroom
by: Russell J. Mortishire‐Smith, et al.
Published: (2025)
by: Russell J. Mortishire‐Smith, et al.
Published: (2025)
Abstract Art Interpretation Using ControlNet
by: Srivastava, Rishabh, et al.
Published: (2024)
by: Srivastava, Rishabh, et al.
Published: (2024)
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning
by: Das, Susmit
Published: (2026)
by: Das, Susmit
Published: (2026)
TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing
by: Neogi, Susmit
Published: (2025)
by: Neogi, Susmit
Published: (2025)
LUMOS: Large User MOdels for User Behavior Prediction
by: Nigam, Dhruv, et al.
Published: (2025)
by: Nigam, Dhruv, et al.
Published: (2025)
Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies
by: Mittal, Avni
Published: (2026)
by: Mittal, Avni
Published: (2026)
Similar Items
-
Do Diffusion Models Dream of Electric Planes? Discrete and Continuous Simulation-Based Inference for Aircraft Design
by: Ghiglino, Aurelien, et al.
Published: (2026) -
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
by: Padhi, Trilok, et al.
Published: (2026) -
Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace Inference
by: Samplawski, Colin, et al.
Published: (2025) -
Addressing Uncertainty in LLMs to Enhance Reliability in Generative AI
by: Kaur, Ramneet, et al.
Published: (2024) -
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
by: Padhi, Trilok, et al.
Published: (2025)