Jailbreak Transferability Emerges from Shared Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Angell, Rico, Brinkmann, Jannik, He, He |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026)
by: Shah, Avidan, et al.
Published: (2026)
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
by: Angell, Rico
Published: (2024)
by: Angell, Rico
Published: (2024)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023)
by: Angell, Rico, et al.
Published: (2023)
Estimating Tail Risks in Language Model Output Distributions
by: Angell, Rico, et al.
Published: (2026)
by: Angell, Rico, et al.
Published: (2026)
A Mechanistic Analysis of a Transformer Trained on a Symbolic Multi-Step Reasoning Task
by: Brinkmann, Jannik, et al.
Published: (2024)
by: Brinkmann, Jannik, et al.
Published: (2024)
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025)
by: Todd, Eric, et al.
Published: (2025)
Representational Transfer Learning for Matrix Completion
by: He, Yong, et al.
Published: (2024)
by: He, Yong, et al.
Published: (2024)
Disentangling Shared and Task-Specific Representations from Multi-Modal Clinical Data
by: Lyu, He, et al.
Published: (2026)
by: Lyu, He, et al.
Published: (2026)
Understanding and Enhancing the Transferability of Jailbreaking Attacks
by: Lin, Runqi, et al.
Published: (2025)
by: Lin, Runqi, et al.
Published: (2025)
Mechanisms of AI Protein Folding in ESMFold
by: Lu, Kevin, et al.
Published: (2026)
by: Lu, Kevin, et al.
Published: (2026)
Transfer Learning of Multiobjective Indirect Low-Thrust Trajectories Using Diffusion Models and Markov Chain Monte Carlo
by: Graebner, Jannik, et al.
Published: (2026)
by: Graebner, Jannik, et al.
Published: (2026)
On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation
by: Han, Andy, et al.
Published: (2026)
by: Han, Andy, et al.
Published: (2026)
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
by: Li, Songze, et al.
Published: (2026)
by: Li, Songze, et al.
Published: (2026)
Trustworthy Transfer Learning: A Survey
by: Wu, Jun, et al.
Published: (2024)
by: Wu, Jun, et al.
Published: (2024)
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
by: Lin, Runqi, et al.
Published: (2025)
by: Lin, Runqi, et al.
Published: (2025)
The Environmental Impact of Ensemble Techniques in Recommender Systems
by: Nitschke, Jannik
Published: (2025)
by: Nitschke, Jannik
Published: (2025)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
by: Mukherjee, Arpan, et al.
Published: (2025)
by: Mukherjee, Arpan, et al.
Published: (2025)
Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study
by: Li, Danya, et al.
Published: (2026)
by: Li, Danya, et al.
Published: (2026)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
by: Nurlanov, Zhakshylyk, et al.
Published: (2026)
by: Nurlanov, Zhakshylyk, et al.
Published: (2026)
Learning Shared Representations for Multi-Task Linear Bandits
by: Lin, Jiabin, et al.
Published: (2026)
by: Lin, Jiabin, et al.
Published: (2026)
Learning with Shared Representations: Statistical Rates and Efficient Algorithms
by: Niu, Xiaochun, et al.
Published: (2024)
by: Niu, Xiaochun, et al.
Published: (2024)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
Hierarchical Successor Representation for Robust Transfer
by: Yu, Changmin, et al.
Published: (2026)
by: Yu, Changmin, et al.
Published: (2026)
Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
by: Li, Kevin Y., et al.
Published: (2026)
by: Li, Kevin Y., et al.
Published: (2026)
Robust Knowledge Transfer in Tiered Reinforcement Learning
by: Huang, Jiawei, et al.
Published: (2023)
by: Huang, Jiawei, et al.
Published: (2023)
TRYLOCK: Defense-in-Depth Against LLM Jailbreaks via Layered Preference and Representation Engineering
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Learning Shared Representations from Unpaired Data
by: Yacobi, Amitai, et al.
Published: (2025)
by: Yacobi, Amitai, et al.
Published: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
Learning Relational Tabular Data without Shared Features
by: Wu, Zhaomin, et al.
Published: (2025)
by: Wu, Zhaomin, et al.
Published: (2025)
Compound Fault Diagnosis for Train Transmission Systems Using Deep Learning with Fourier-enhanced Representation
by: Rico, Jonathan Adam, et al.
Published: (2025)
by: Rico, Jonathan Adam, et al.
Published: (2025)
Towards Predicting the Success of Transfer-based Attacks by Quantifying Shared Feature Representations
by: Dale, Ashley S., et al.
Published: (2024)
by: Dale, Ashley S., et al.
Published: (2024)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
by: Wurgaft, Daniel, et al.
Published: (2026)
by: Wurgaft, Daniel, et al.
Published: (2026)
Meta-RL with Shared Representations Enables Fast Adaptation in Energy Systems
by: Zangato, Théo, et al.
Published: (2026)
by: Zangato, Théo, et al.
Published: (2026)
Identifiable Multimodal Causal Representation Learning under Partial Latent Sharing
by: Benhamza, Manal, et al.
Published: (2026)
by: Benhamza, Manal, et al.
Published: (2026)
Predictive Representations for Skill Transfer in Reinforcement Learning
by: Vereecken, Ruben, et al.
Published: (2026)
by: Vereecken, Ruben, et al.
Published: (2026)
Understanding the Transferability of Representations via Task-Relatedness
by: Mehra, Akshay, et al.
Published: (2023)
by: Mehra, Akshay, et al.
Published: (2023)
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
by: Miao, Albert, et al.
Published: (2025)
by: Miao, Albert, et al.
Published: (2025)
How Data Augmentation Shapes Neural Representations
by: He, Tianxiao, et al.
Published: (2026)
by: He, Tianxiao, et al.
Published: (2026)
Similar Items
-
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
by: Shah, Avidan, et al.
Published: (2026) -
Polynomial Precision Dependence Solutions to Alignment Research Center Matrix Completion Problems
by: Angell, Rico
Published: (2024) -
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024) -
Fast, Scalable, Warm-Start Semidefinite Programming with Spectral Bundling and Sketching
by: Angell, Rico, et al.
Published: (2023) -
Estimating Tail Risks in Language Model Output Distributions
by: Angell, Rico, et al.
Published: (2026)