Latent Adversarial Training Improves the Representation of Refusal
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abbas, Alexandra, Petrova, Nora, Lyons, Helios Ael, Perez-Campanero, Natalia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
von: Marshall, Thomas, et al.
Veröffentlicht: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
von: Lan, Wenhao, et al.
Veröffentlicht: (2026)
von: Lan, Wenhao, et al.
Veröffentlicht: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
von: Formento, Brian, et al.
Veröffentlicht: (2024)
von: Formento, Brian, et al.
Veröffentlicht: (2024)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
von: Muhamed, Aashiq, et al.
Veröffentlicht: (2025)
Silenced Biases: The Dark Side LLMs Learned to Refuse
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
von: Himelstein, Rom, et al.
Veröffentlicht: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Thinking into the Future: Latent Lookahead Training for Transformers
von: Noci, Lorenzo, et al.
Veröffentlicht: (2026)
von: Noci, Lorenzo, et al.
Veröffentlicht: (2026)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Interpreting Latent Student Knowledge Representations in Programming Assignments
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
von: Wu, Yuan
Veröffentlicht: (2024)
von: Wu, Yuan
Veröffentlicht: (2024)
Vocabulary-Defined Semantics: Latent Space Clustering for Improving In-Context Learning
von: Gu, Jian, et al.
Veröffentlicht: (2024)
von: Gu, Jian, et al.
Veröffentlicht: (2024)
Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
von: Maiya, Sharan, et al.
Veröffentlicht: (2025)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025)
von: Li, Ran, et al.
Veröffentlicht: (2025)
Improving LLM Final Representations with Inter-Layer Geometry
von: Ulanovski, Tom, et al.
Veröffentlicht: (2026)
von: Ulanovski, Tom, et al.
Veröffentlicht: (2026)
Self-Improving World Modelling with Latent Actions
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
von: Aviss, Thea
Veröffentlicht: (2026)
von: Aviss, Thea
Veröffentlicht: (2026)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
$\textbf{AGT$^{AO}$}$: Robust and Stabilized LLM Unlearning via Adversarial Gating Training with Adaptive Orthogonality
von: Li, Pengyu, et al.
Veröffentlicht: (2026)
von: Li, Pengyu, et al.
Veröffentlicht: (2026)
Enhancing Latent Computation in Transformers with Latent Tokens
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
von: Piterbarg, Ulyana, et al.
Veröffentlicht: (2024)
von: Piterbarg, Ulyana, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Refusal in LLMs is an Affine Function
von: Marshall, Thomas, et al.
Veröffentlicht: (2024) -
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024) -
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
von: Lan, Wenhao, et al.
Veröffentlicht: (2026) -
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026) -
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)