Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sheshadri, Abhay, Ewart, Aidan, Guo, Phillip, Lynch, Aengus, Wu, Cindy, Hebbar, Vivek, Sleight, Henry, Stickland, Asa Cooper, Perez, Ethan, Hadfield-Menell, Dylan, Casper, Stephen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Eight Methods to Evaluate Robust Unlearning in LLMs
von: Lynch, Aengus, et al.
Veröffentlicht: (2024)
von: Lynch, Aengus, et al.
Veröffentlicht: (2024)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
von: Khan, Ariba, et al.
Veröffentlicht: (2025)
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
von: Casper, Stephen, et al.
Veröffentlicht: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
von: Guo, Phillip, et al.
Veröffentlicht: (2024)
The Persistent Vulnerability of Aligned AI Systems
von: Lynch, Aengus
Veröffentlicht: (2026)
von: Lynch, Aengus
Veröffentlicht: (2026)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
von: Doshi, Jai, et al.
Veröffentlicht: (2024)
Layered Unlearning for Adversarial Relearning
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2026)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
von: Soni, Prajna, et al.
Veröffentlicht: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Diverse Preference Learning for Capabilities and Alignment
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
von: Slocum, Stewart, et al.
Veröffentlicht: (2025)
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
von: Haupt, Andreas A., et al.
Veröffentlicht: (2022)
Cooperative Inverse Reinforcement Learning
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
von: Hadfield-Menell, Dylan, et al.
Veröffentlicht: (2016)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Why Do Language Model Agents Whistleblow?
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
von: Agrawal, Kushal, et al.
Veröffentlicht: (2025)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
von: Price, Sara, et al.
Veröffentlicht: (2024)
von: Price, Sara, et al.
Veröffentlicht: (2024)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
von: Khan, Emaan Bilal, et al.
Veröffentlicht: (2026)
Obfuscated Activations Bypass LLM Latent-Space Defenses
von: Bailey, Luke, et al.
Veröffentlicht: (2024)
von: Bailey, Luke, et al.
Veröffentlicht: (2024)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
von: Stickland, Asa Cooper, et al.
Veröffentlicht: (2024)
von: Stickland, Asa Cooper, et al.
Veröffentlicht: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
von: Che, Zora, et al.
Veröffentlicht: (2025)
von: Che, Zora, et al.
Veröffentlicht: (2025)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wen, Jiaxin, et al.
Veröffentlicht: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
von: Shenoy, Keshav, et al.
Veröffentlicht: (2026)
von: Shenoy, Keshav, et al.
Veröffentlicht: (2026)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
von: Peng, Alwin, et al.
Veröffentlicht: (2024)
von: Peng, Alwin, et al.
Veröffentlicht: (2024)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
von: Brinkmann, Jannik, et al.
Veröffentlicht: (2024)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
von: Ensign, Danielle, et al.
Veröffentlicht: (2025)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
von: Youstra, Jack, et al.
Veröffentlicht: (2025)
von: Youstra, Jack, et al.
Veröffentlicht: (2025)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
von: Hägele, Alexander, et al.
Veröffentlicht: (2026)
Causal Machine Learning: A Survey and Open Problems
von: Kaddour, Jean, et al.
Veröffentlicht: (2022)
von: Kaddour, Jean, et al.
Veröffentlicht: (2022)
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
von: Stickland, Asa Cooper, et al.
Veröffentlicht: (2025)
von: Stickland, Asa Cooper, et al.
Veröffentlicht: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
von: Lynch, Aengus, et al.
Veröffentlicht: (2025)
von: Lynch, Aengus, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Eight Methods to Evaluate Robust Unlearning in LLMs
von: Lynch, Aengus, et al.
Veröffentlicht: (2024) -
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024) -
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
von: Khan, Ariba, et al.
Veröffentlicht: (2025) -
Pitfalls of Evidence-Based AI Policy
von: Casper, Stephen, et al.
Veröffentlicht: (2025) -
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
von: Guo, Phillip, et al.
Veröffentlicht: (2024)