Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zaree, Pedram, Mamun, Md Abdullah Al, Alam, Quazi Mishkatul, Dong, Yue, Alouani, Ihsen, Abu-Ghazaleh, Nael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AttenMIA: LLM Membership Inference Attack through Attention Signals
by: Zaree, Pedram, et al.
Published: (2026)
by: Zaree, Pedram, et al.
Published: (2026)
Co(ve)rtex: ML Models as storage channels and their (mis-)applications
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
Bypassing Prompt Injection Detectors through Evasive Injections
by: Rahman, Md Jahedur, et al.
Published: (2026)
by: Rahman, Md Jahedur, et al.
Published: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023)
by: Slocum, Carter, et al.
Published: (2023)
I Know What You Sync: Covert and Side Channel Attacks on File Systems via syncfs
by: Gu, Cheng, et al.
Published: (2024)
by: Gu, Cheng, et al.
Published: (2024)
Siren Song: Manipulating Pose Estimation in XR Headsets Using Acoustic Attacks
by: Huang, Zijian, et al.
Published: (2025)
by: Huang, Zijian, et al.
Published: (2025)
Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis
by: Park, Jeonghwan, et al.
Published: (2025)
by: Park, Jeonghwan, et al.
Published: (2025)
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
Stealth by Conformity: Evading Robust Aggregation through Adaptive Poisoning
by: McGaughey, Ryan, et al.
Published: (2025)
by: McGaughey, Ryan, et al.
Published: (2025)
Watermarking Neuromorphic Brains: Intellectual Property Protection in Spiking Neural Networks
by: Poursiami, Hamed, et al.
Published: (2024)
by: Poursiami, Hamed, et al.
Published: (2024)
Are Neuromorphic Architectures Inherently Privacy-preserving? An Exploratory Study
by: Moshruba, Ayana, et al.
Published: (2024)
by: Moshruba, Ayana, et al.
Published: (2024)
BrainLeaks: On the Privacy-Preserving Properties of Neuromorphic Architectures against Model Inversion Attacks
by: Poursiami, Hamed, et al.
Published: (2024)
by: Poursiami, Hamed, et al.
Published: (2024)
Token-based Vehicular Security System (TVSS): Scalable, Secure, Low-latency Public Key Infrastructure for Connected Vehicles
by: Rabiah, Abdulrahman Bin, et al.
Published: (2024)
by: Rabiah, Abdulrahman Bin, et al.
Published: (2024)
Evasive Hardware Trojan through Adversarial Power Trace
by: Omidi, Behnam, et al.
Published: (2024)
by: Omidi, Behnam, et al.
Published: (2024)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
NVBleed: Covert and Side-Channel Attacks on NVIDIA Multi-GPU Interconnect
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
Beyond the Bridge: Contention-Based Covert and Side Channel Attacks on Multi-GPU Interconnect
by: Zhang, Yicheng, et al.
Published: (2024)
by: Zhang, Yicheng, et al.
Published: (2024)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
by: Shi, Yunfan
Published: (2024)
by: Shi, Yunfan
Published: (2024)
AdvART: Adversarial Art for Camouflaged Object Detection Attacks
by: Guesmi, Amira, et al.
Published: (2023)
by: Guesmi, Amira, et al.
Published: (2023)
On Jailbreaking Quantized Language Models Through Fault Injection Attacks
by: Zahran, Noureldin, et al.
Published: (2025)
by: Zahran, Noureldin, et al.
Published: (2025)
ShadowScope: GPU Monitoring and Validation via Composable Side Channel Signals
by: Almusaddar, Ghadeer, et al.
Published: (2025)
by: Almusaddar, Ghadeer, et al.
Published: (2025)
On the Feasibility of Hybrid Homomorphic Encryption for Intelligent Transportation Systems
by: Yates, Kyle, et al.
Published: (2026)
by: Yates, Kyle, et al.
Published: (2026)
IoT-Enabled Smart Car Parking System through Integrated Sensors and Mobile Applications
by: Mamun, Abdullah Al, et al.
Published: (2024)
by: Mamun, Abdullah Al, et al.
Published: (2024)
SnatchML: Hijacking ML models without Training Access
by: Ghorbel, Mahmoud, et al.
Published: (2024)
by: Ghorbel, Mahmoud, et al.
Published: (2024)
Hardware Design and Security Needs Attention: From Survey to Path Forward
by: Ghimire, Sujan, et al.
Published: (2025)
by: Ghimire, Sujan, et al.
Published: (2025)
CANGuard: A Spatio-Temporal CNN-GRU-Attention Hybrid Architecture for Intrusion Detection in In-Vehicle CAN Networks
by: Sajib, Rakib Hossain, et al.
Published: (2026)
by: Sajib, Rakib Hossain, et al.
Published: (2026)
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
by: Yuan, Shuai, et al.
Published: (2025)
by: Yuan, Shuai, et al.
Published: (2025)
Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4
by: Polyakov, Alex, et al.
Published: (2026)
by: Polyakov, Alex, et al.
Published: (2026)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Post-Quantum Cryptography for Intelligent Transportation Systems: An Implementation-Focused Review
by: Mamun, Abdullah Al, et al.
Published: (2026)
by: Mamun, Abdullah Al, et al.
Published: (2026)
Enhancing Transportation Cyber-Physical Systems Security: A Shift to Post-Quantum Cryptography
by: Mamun, Abdullah Al, et al.
Published: (2024)
by: Mamun, Abdullah Al, et al.
Published: (2024)
Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
by: Alam, Md Takrim Ul, et al.
Published: (2026)
by: Alam, Md Takrim Ul, et al.
Published: (2026)
To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Experimental Evaluation of Post-Quantum Homomorphic Encryption for Privacy-Preserving I2I Communication in ITS
by: Mamun, Abdullah Al, et al.
Published: (2025)
by: Mamun, Abdullah Al, et al.
Published: (2025)
Assessing Cyber Risks in Hydropower Systems Through HAZOP and Bow-Tie Analysis
by: Frempong-Kore, Kwabena Opoku, et al.
Published: (2026)
by: Frempong-Kore, Kwabena Opoku, et al.
Published: (2026)
Powering the Future of IoT: Federated Learning for Optimized Power Consumption and Enhanced Privacy
by: Shirvani, Ghazaleh, et al.
Published: (2024)
by: Shirvani, Ghazaleh, et al.
Published: (2024)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
Similar Items
-
AttenMIA: LLM Membership Inference Attack through Attention Signals
by: Zaree, Pedram, et al.
Published: (2026) -
Co(ve)rtex: ML Models as storage channels and their (mis-)applications
by: Mamun, Md Abdullah Al, et al.
Published: (2023) -
Bypassing Prompt Injection Detectors through Evasive Injections
by: Rahman, Md Jahedur, et al.
Published: (2026) -
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025) -
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023)