Improving Alignment and Robustness with Circuit Breakers
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Andy, Phan, Long, Wang, Justin, Duenas, Derek, Lin, Maxwell, Andriushchenko, Maksym, Wang, Rowan, Kolter, Zico, Fredrikson, Matt, Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025)
by: Krishna, Satyapriya, et al.
Published: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025)
by: Zou, Andy, et al.
Published: (2025)
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)
by: Zou, Andy, et al.
Published: (2023)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
Evaluating Language Model Reasoning about Confidential Information
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
Consistency Models Made Easy
by: Geng, Zhengyang, et al.
Published: (2024)
by: Geng, Zhengyang, et al.
Published: (2024)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Reducing Political Manipulation with Consistency Training
by: Phan, Long, et al.
Published: (2026)
by: Phan, Long, et al.
Published: (2026)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
Superintelligence Strategy: Expert Version
by: Hendrycks, Dan, et al.
Published: (2025)
by: Hendrycks, Dan, et al.
Published: (2025)
Compressed Sensing for Capability Localization in Large Language Models
by: Bair, Anna, et al.
Published: (2026)
by: Bair, Anna, et al.
Published: (2026)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Massive Activations in Large Language Models
by: Sun, Mingjie, et al.
Published: (2024)
by: Sun, Mingjie, et al.
Published: (2024)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
Improved Mean Flows: On the Challenges of Fastforward Generative Models
by: Geng, Zhengyang, et al.
Published: (2025)
by: Geng, Zhengyang, et al.
Published: (2025)
Idiosyncrasies in Large Language Models
by: Sun, Mingjie, et al.
Published: (2025)
by: Sun, Mingjie, et al.
Published: (2025)
Forcing Diffuse Distributions out of Language Models
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Is Your Text-to-Image Model Robust to Caption Noise?
by: Yu, Weichen, et al.
Published: (2024)
by: Yu, Weichen, et al.
Published: (2024)
A Simple and Effective Pruning Approach for Large Language Models
by: Sun, Mingjie, et al.
Published: (2023)
by: Sun, Mingjie, et al.
Published: (2023)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
From Variance to Veracity: Unbundling and Mitigating Gradient Variance in Differentiable Bundle Adjustment Layers
by: Gurumurthy, Swaminathan, et al.
Published: (2024)
by: Gurumurthy, Swaminathan, et al.
Published: (2024)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Similar Items
-
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024) -
D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
by: Krishna, Satyapriya, et al.
Published: (2025) -
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025) -
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
by: Zou, Andy, et al.
Published: (2025) -
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)