Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Shuai, Zhang, Zhibo, Li, Yuxi, Bai, Guangdong, Kailong, Wang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Beyond Fidelity: Explaining Vulnerability Localization of Learning-based Detectors
by: Cheng, Baijun, et al.
Published: (2024)
by: Cheng, Baijun, et al.
Published: (2024)
Poincaré Differential Privacy for Hierarchy-Aware Graph Embedding
by: Wei, Yuecen, et al.
Published: (2023)
by: Wei, Yuecen, et al.
Published: (2023)
Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
by: Churina, Svetlana, et al.
Published: (2025)
by: Churina, Svetlana, et al.
Published: (2025)
Practical Poisoning Attacks against Retrieval-Augmented Generation
by: Zhang, Baolei, et al.
Published: (2025)
by: Zhang, Baolei, et al.
Published: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
LMAE4Eth: Generalizable and Robust Ethereum Fraud Detection by Exploring Transaction Semantics and Masked Graph Embedding
by: Jia, Yifan, et al.
Published: (2025)
by: Jia, Yifan, et al.
Published: (2025)
Supervised Robustness-preserving Data-free Neural Network Pruning
by: Meng, Mark Huasong, et al.
Published: (2022)
by: Meng, Mark Huasong, et al.
Published: (2022)
Information Leakage from Embedding in Large Language Models
by: Wan, Zhipeng, et al.
Published: (2024)
by: Wan, Zhipeng, et al.
Published: (2024)
FedRecAttack: Model Poisoning Attack to Federated Recommendation
by: Rong, Dazhong, et al.
Published: (2022)
by: Rong, Dazhong, et al.
Published: (2022)
FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
by: Long, Yuzhen, et al.
Published: (2025)
by: Long, Yuzhen, et al.
Published: (2025)
EASE: Practical and Efficient Safety Alignment for Small Language Models
by: Shi, Haonan, et al.
Published: (2025)
by: Shi, Haonan, et al.
Published: (2025)
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Embedding Attack Project (Work Report)
by: Pu, Jiameng, et al.
Published: (2024)
by: Pu, Jiameng, et al.
Published: (2024)
BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding
by: Benoit, Tristan, et al.
Published: (2024)
by: Benoit, Tristan, et al.
Published: (2024)
What Really is a Member? Discrediting Membership Inference via Poisoning
by: Mangaokar, Neal, et al.
Published: (2025)
by: Mangaokar, Neal, et al.
Published: (2025)
SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking
by: Gao, Zheng, et al.
Published: (2026)
by: Gao, Zheng, et al.
Published: (2026)
When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems
by: Zhang, Lingxi, et al.
Published: (2026)
by: Zhang, Lingxi, et al.
Published: (2026)
Private Training & Data Generation by Clustering Embeddings
by: Zhou, Felix, et al.
Published: (2025)
by: Zhou, Felix, et al.
Published: (2025)
RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement
by: Wang, Ziye, et al.
Published: (2026)
by: Wang, Ziye, et al.
Published: (2026)
Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
by: Zhang, Baolei, et al.
Published: (2025)
by: Zhang, Baolei, et al.
Published: (2025)
Debiased Graph Poisoning Attack via Contrastive Surrogate Objective
by: Yoon, Kanghoon, et al.
Published: (2024)
by: Yoon, Kanghoon, et al.
Published: (2024)
Practicable Black-box Evasion Attacks on Link Prediction in Dynamic Graphs -- A Graph Sequential Embedding Method
by: Li, Jiate, et al.
Published: (2024)
by: Li, Jiate, et al.
Published: (2024)
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
by: Li, Jianhui, et al.
Published: (2024)
by: Li, Jianhui, et al.
Published: (2024)
Bypassing Prompt Guards in Production with Controlled-Release Prompting
by: Fairoze, Jaiden, et al.
Published: (2025)
by: Fairoze, Jaiden, et al.
Published: (2025)
Preventing the Popular Item Embedding Based Attack in Federated Recommendations
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Hiding Backdoors within Event Sequence Data via Poisoning Attacks
by: Ermilova, Alina, et al.
Published: (2023)
by: Ermilova, Alina, et al.
Published: (2023)
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Data Poisoning Attacks in Intelligent Transportation Systems: A Survey
by: Wang, Feilong, et al.
Published: (2024)
by: Wang, Feilong, et al.
Published: (2024)
Salsa Fresca: Angular Embeddings and Pre-Training for ML Attacks on Learning With Errors
by: Stevens, Samuel, et al.
Published: (2024)
by: Stevens, Samuel, et al.
Published: (2024)
Timber! Poisoning Decision Trees
by: Calzavara, Stefano, et al.
Published: (2024)
by: Calzavara, Stefano, et al.
Published: (2024)
How to Defend Against Large-scale Model Poisoning Attacks in Federated Learning: A Vertical Solution
by: Wang, Jinbo, et al.
Published: (2024)
by: Wang, Jinbo, et al.
Published: (2024)
Embedding-based classifiers can detect prompt injection attacks
by: Ayub, Md. Ahsan, et al.
Published: (2024)
by: Ayub, Md. Ahsan, et al.
Published: (2024)
Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
by: Wang, Yujing, et al.
Published: (2024)
by: Wang, Yujing, et al.
Published: (2024)
Similar Items
-
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024) -
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024) -
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025) -
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026) -
Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
by: Wang, Hao, et al.
Published: (2024)