Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ball, Sarah, Kreuter, Frauke, Panickssery, Nina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
von: Mura, Raffaele, et al.
Veröffentlicht: (2025)
von: Mura, Raffaele, et al.
Veröffentlicht: (2025)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
von: Ackerman, Christopher, et al.
Veröffentlicht: (2024)
von: Ackerman, Christopher, et al.
Veröffentlicht: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
von: Arditi, Andy, et al.
Veröffentlicht: (2024)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
von: Chew, Robert, et al.
Veröffentlicht: (2026)
von: Chew, Robert, et al.
Veröffentlicht: (2026)
Mitigating Many-Shot Jailbreaking
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
von: Ackerman, Christopher M., et al.
Veröffentlicht: (2025)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
von: Lee, Isack, et al.
Veröffentlicht: (2024)
von: Lee, Isack, et al.
Veröffentlicht: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
von: Candogan, Leyla Naz, et al.
Veröffentlicht: (2025)
von: Candogan, Leyla Naz, et al.
Veröffentlicht: (2025)
Steering Llama 2 via Contrastive Activation Addition
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
von: Panickssery, Nina, et al.
Veröffentlicht: (2023)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
von: Zhou, Yefan, et al.
Veröffentlicht: (2025)
von: Zhou, Yefan, et al.
Veröffentlicht: (2025)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Large Language Models Explore by Latent Distilling
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
von: Zeng, Yuanhao, et al.
Veröffentlicht: (2026)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
von: Leyton-Brown, Kevin, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
von: Hu, Xulin, et al.
Veröffentlicht: (2026)
Understanding Subword Compositionality of Large Language Models
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
von: Xu, Yi, et al.
Veröffentlicht: (2026)
von: Xu, Yi, et al.
Veröffentlicht: (2026)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
GUNDAM: Aligning Large Language Models with Graph Understanding
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
von: Ouyang, Sheng, et al.
Veröffentlicht: (2024)
Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction
von: Ruppik, Benjamin Matthias, et al.
Veröffentlicht: (2024)
von: Ruppik, Benjamin Matthias, et al.
Veröffentlicht: (2024)
Demystifying Embedding Spaces using Large Language Models
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2023)
von: Tennenholtz, Guy, et al.
Veröffentlicht: (2023)
Meanings and Feelings of Large Language Models: Observability of Latent States in Generative AI
von: Liu, Tian Yu, et al.
Veröffentlicht: (2024)
von: Liu, Tian Yu, et al.
Veröffentlicht: (2024)
Counterfactual Token Generation in Large Language Models
von: Chatzi, Ivi, et al.
Veröffentlicht: (2024)
von: Chatzi, Ivi, et al.
Veröffentlicht: (2024)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
von: Hua, Peichun, et al.
Veröffentlicht: (2025)
LSEBMCL: A Latent Space Energy-Based Model for Continual Learning
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
von: Li, Xiaodi, et al.
Veröffentlicht: (2025)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
Can Large Language Models Understand Intermediate Representations in Compilers?
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
von: Cho, Seungho, et al.
Veröffentlicht: (2025)
Model-diff: A Tool for Comparative Study of Language Models in the Input Space
von: Liu, Weitang, et al.
Veröffentlicht: (2024)
von: Liu, Weitang, et al.
Veröffentlicht: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Xue, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand Molecules?
von: Sadeghi, Shaghayegh, et al.
Veröffentlicht: (2024)
von: Sadeghi, Shaghayegh, et al.
Veröffentlicht: (2024)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
von: Lin, Shi, et al.
Veröffentlicht: (2024)
von: Lin, Shi, et al.
Veröffentlicht: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
von: Mahran, Mariam, et al.
Veröffentlicht: (2025)
von: Mahran, Mariam, et al.
Veröffentlicht: (2025)
Rethinking How to Evaluate Language Model Jailbreak
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
von: Cai, Hongyu, et al.
Veröffentlicht: (2024)
SpaceByte: Towards Deleting Tokenization from Large Language Modeling
von: Slagle, Kevin
Veröffentlicht: (2024)
von: Slagle, Kevin
Veröffentlicht: (2024)
Ähnliche Einträge
-
Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction
von: Ball, Sarah, et al.
Veröffentlicht: (2025) -
LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
von: Mura, Raffaele, et al.
Veröffentlicht: (2025) -
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
von: Ackerman, Christopher, et al.
Veröffentlicht: (2024) -
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
von: Ball, Sarah, et al.
Veröffentlicht: (2025) -
Refusal in Language Models Is Mediated by a Single Direction
von: Arditi, Andy, et al.
Veröffentlicht: (2024)