LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Mura, Raffaele, Piras, Giorgio, Lukošiūtė, Kamilė, Pintor, Maura, Karbasi, Amin, Biggio, Battista |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent-space Attacks for Refusal Evasion in Language Models
by: Piras, Giorgio, et al.
Published: (2026)
by: Piras, Giorgio, et al.
Published: (2026)
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
by: Piras, Giorgio, et al.
Published: (2025)
by: Piras, Giorgio, et al.
Published: (2025)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
S2AP: Score-space Sharpness Minimization for Adversarial Pruning
by: Piras, Giorgio, et al.
Published: (2025)
by: Piras, Giorgio, et al.
Published: (2025)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Evaluating Line-level Localization Ability of Learning-based Code Vulnerability Detection Models
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
(Im)possibility of Automated Hallucination Detection in Large Language Models
by: Karbasi, Amin, et al.
Published: (2025)
by: Karbasi, Amin, et al.
Published: (2025)
HO-FMN: Hyperparameter Optimization for Fast Minimum-Norm Attacks
by: Mura, Raffaele, et al.
Published: (2024)
by: Mura, Raffaele, et al.
Published: (2024)
Adversarial Pruning: A Survey and Benchmark of Pruning Methods for Adversarial Robustness
by: Piras, Giorgio, et al.
Published: (2024)
by: Piras, Giorgio, et al.
Published: (2024)
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
by: Xu, Zihao, et al.
Published: (2026)
by: Xu, Zihao, et al.
Published: (2026)
Emotions Where Art Thou: Understanding and Characterizing the Emotional Latent Space of Large Language Models
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
Contextual Categorization Enhancement through LLMs Latent-Space
by: Bettouche, Zineddine, et al.
Published: (2024)
by: Bettouche, Zineddine, et al.
Published: (2024)
On the Failure of Latent State Persistence in Large Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
SAGE-5GC: Security-Aware Guidelines for Evaluating Anomaly Detection in the 5G Core Network
by: Manca, Cristian, et al.
Published: (2026)
by: Manca, Cristian, et al.
Published: (2026)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
A Language Model's Guide Through Latent Space
by: von Rütte, Dimitri, et al.
Published: (2024)
by: von Rütte, Dimitri, et al.
Published: (2024)
Unlocking the Working Memory of Large Language Models for Latent Reasoning
by: Aichberger, Lukas, et al.
Published: (2026)
by: Aichberger, Lukas, et al.
Published: (2026)
On Effects of Steering Latent Representation for Large Language Model Unlearning
by: Huu-Tien, Dang, et al.
Published: (2024)
by: Huu-Tien, Dang, et al.
Published: (2024)
Large Language Models Explore by Latent Distilling
by: Zeng, Yuanhao, et al.
Published: (2026)
by: Zeng, Yuanhao, et al.
Published: (2026)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
by: Liu, Yuliang, et al.
Published: (2026)
by: Liu, Yuliang, et al.
Published: (2026)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
by: Zhou, Weikang, et al.
Published: (2024)
by: Zhou, Weikang, et al.
Published: (2024)
SeLaR: Selective Latent Reasoning in Large Language Models
by: Fu, Renyu, et al.
Published: (2026)
by: Fu, Renyu, et al.
Published: (2026)
Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Models
by: Silvestri, Matteo, et al.
Published: (2025)
by: Silvestri, Matteo, et al.
Published: (2025)
Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
Self-Specialization: Uncovering Latent Expertise within Large Language Models
by: Kang, Junmo, et al.
Published: (2023)
by: Kang, Junmo, et al.
Published: (2023)
Gödel Test: Can Large Language Models Solve Easy Conjectures?
by: Feldman, Moran, et al.
Published: (2025)
by: Feldman, Moran, et al.
Published: (2025)
Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
by: Vijay, Supriti, et al.
Published: (2025)
by: Vijay, Supriti, et al.
Published: (2025)
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
iCLP: Large Language Model Reasoning with Implicit Cognition Latent Planning
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
Active Use of Latent Constituency Representation in both Humans and Large Language Models
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
by: Pandey, Sanskar, et al.
Published: (2025)
by: Pandey, Sanskar, et al.
Published: (2025)
Evaluating the Evaluators: Trust in Adversarial Robustness Tests
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
by: Cinà, Antonio Emanuele, et al.
Published: (2025)
Buffer-free Class-Incremental Learning with Out-of-Distribution Detection
by: Gupta, Srishti, et al.
Published: (2025)
by: Gupta, Srishti, et al.
Published: (2025)
Learning to Ponder: Adaptive Reasoning in Latent Space
by: He, Yixin, et al.
Published: (2025)
by: He, Yixin, et al.
Published: (2025)
Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
Layer-Order Inversion: Rethinking Latent Multi-Hop Reasoning in Large Language Models
by: Liu, Xukai, et al.
Published: (2026)
by: Liu, Xukai, et al.
Published: (2026)
Uncovering Latent Chain of Thought Vectors in Language Models
by: Zhang, Jason, et al.
Published: (2024)
by: Zhang, Jason, et al.
Published: (2024)
Latent Multi-Head Attention for Small Language Models
by: Mehta, Sushant, et al.
Published: (2025)
by: Mehta, Sushant, et al.
Published: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
by: Xu, Yi, et al.
Published: (2026)
by: Xu, Yi, et al.
Published: (2026)
Similar Items
-
Latent-space Attacks for Refusal Evasion in Language Models
by: Piras, Giorgio, et al.
Published: (2026) -
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
by: Piras, Giorgio, et al.
Published: (2025) -
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025) -
S2AP: Score-space Sharpness Minimization for Adversarial Pruning
by: Piras, Giorgio, et al.
Published: (2025) -
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)