AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zijun, Tu, Haoqin, Mei, Jieru, Zhao, Bingchen, Wang, Yisen, Xie, Cihang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
por: Wang, Zijun, et al.
Publicado: (2025)
por: Wang, Zijun, et al.
Publicado: (2025)
What If We Recaption Billions of Web Images with LLaMA-3?
por: Li, Xianhang, et al.
Publicado: (2024)
por: Li, Xianhang, et al.
Publicado: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
por: Yang, Siwei, et al.
Publicado: (2024)
por: Yang, Siwei, et al.
Publicado: (2024)
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
por: Wu, Juncheng, et al.
Publicado: (2026)
por: Wu, Juncheng, et al.
Publicado: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)
por: Mu, Junjie, et al.
Publicado: (2025)
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?
por: Xie, Yunfei, et al.
Publicado: (2024)
por: Xie, Yunfei, et al.
Publicado: (2024)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
por: Batra, Hunar, et al.
Publicado: (2025)
por: Batra, Hunar, et al.
Publicado: (2025)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
por: Wang, Zijun, et al.
Publicado: (2026)
por: Wang, Zijun, et al.
Publicado: (2026)
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
por: Wu, Juncheng, et al.
Publicado: (2025)
por: Wu, Juncheng, et al.
Publicado: (2025)
Attn-GS: Attention-Guided Context Compression for Efficient Personalized LLMs
por: Zeng, Shenglai, et al.
Publicado: (2026)
por: Zeng, Shenglai, et al.
Publicado: (2026)
AHELM: A Holistic Evaluation of Audio-Language Models
por: Lee, Tony, et al.
Publicado: (2025)
por: Lee, Tony, et al.
Publicado: (2025)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
por: Meng, Wenlong, et al.
Publicado: (2025)
por: Meng, Wenlong, et al.
Publicado: (2025)
GCG Attack On A Diffusion LLM
por: Neyroud, Ruben, et al.
Publicado: (2025)
por: Neyroud, Ruben, et al.
Publicado: (2025)
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
por: Liao, Zeyi, et al.
Publicado: (2024)
por: Liao, Zeyi, et al.
Publicado: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
por: Li, Xiao, et al.
Publicado: (2024)
por: Li, Xiao, et al.
Publicado: (2024)
ViLBench: A Suite for Vision-Language Process Reward Modeling
por: Tu, Haoqin, et al.
Publicado: (2025)
por: Tu, Haoqin, et al.
Publicado: (2025)
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
por: Kumar, Vishal, et al.
Publicado: (2024)
por: Kumar, Vishal, et al.
Publicado: (2024)
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
por: Chen, Hardy, et al.
Publicado: (2025)
por: Chen, Hardy, et al.
Publicado: (2025)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
por: Chen, Taiye, et al.
Publicado: (2025)
por: Chen, Taiye, et al.
Publicado: (2025)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
por: Chen, Hardy, et al.
Publicado: (2026)
por: Chen, Hardy, et al.
Publicado: (2026)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
por: Li, Yu, et al.
Publicado: (2025)
por: Li, Yu, et al.
Publicado: (2025)
ProxyAttn: Guided Sparse Attention via Representative Heads
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
por: Pu, Rui, et al.
Publicado: (2024)
por: Pu, Rui, et al.
Publicado: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
por: Tan, Yuting, et al.
Publicado: (2025)
por: Tan, Yuting, et al.
Publicado: (2025)
SpecAttn: Speculating Sparse Attention
por: Shah, Harsh
Publicado: (2025)
por: Shah, Harsh
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
Defending LLMs against Jailbreaking Attacks via Backtranslation
por: Wang, Yihan, et al.
Publicado: (2024)
por: Wang, Yihan, et al.
Publicado: (2024)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
por: Xu, Zhao, et al.
Publicado: (2024)
por: Xu, Zhao, et al.
Publicado: (2024)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
por: Chen, Kexin, et al.
Publicado: (2024)
por: Chen, Kexin, et al.
Publicado: (2024)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
por: Mei, Jieru, et al.
Publicado: (2024)
por: Mei, Jieru, et al.
Publicado: (2024)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
Enhancing Jailbreak Attacks with Diversity Guidance
por: Zhang, Xu, et al.
Publicado: (2024)
por: Zhang, Xu, et al.
Publicado: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
LongAttn: Selecting Long-context Training Data via Token-level Attention
por: Wu, Longyun, et al.
Publicado: (2025)
por: Wu, Longyun, et al.
Publicado: (2025)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
por: Xie, Yunfei, et al.
Publicado: (2024)
por: Xie, Yunfei, et al.
Publicado: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
por: Tu, Shangqing, et al.
Publicado: (2024)
por: Tu, Shangqing, et al.
Publicado: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
por: Mo, Yichuan, et al.
Publicado: (2024)
por: Mo, Yichuan, et al.
Publicado: (2024)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
por: Ni, Ziyi, et al.
Publicado: (2025)
por: Ni, Ziyi, et al.
Publicado: (2025)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
por: Wang, Hao, et al.
Publicado: (2024)
por: Wang, Hao, et al.
Publicado: (2024)
Ejemplares similares
-
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
por: Wang, Zijun, et al.
Publicado: (2025) -
What If We Recaption Billions of Web Images with LLaMA-3?
por: Li, Xianhang, et al.
Publicado: (2024) -
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
por: Yang, Siwei, et al.
Publicado: (2024) -
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
por: Wu, Juncheng, et al.
Publicado: (2026) -
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)