Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Jiawei, Fang, Hao, Yang, Xiaochen, Gao, Kuofeng, Chen, Bin, Xia, Shu-Tao, Xu, Ke, Qiu, Han |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
by: Sun, Shuoyang, et al.
Published: (2026)
by: Sun, Shuoyang, et al.
Published: (2026)
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
VAGUEGAN: Stealthy Poisoning and Backdoor Attacks on Image Generative Pipelines
by: Faisal, Mostafa Mohaimen Akand, et al.
Published: (2025)
by: Faisal, Mostafa Mohaimen Akand, et al.
Published: (2025)
A Proxy Attack-Free Strategy for Practically Improving the Poisoning Efficiency in Backdoor Attacks
by: Li, Ziqiang, et al.
Published: (2023)
by: Li, Ziqiang, et al.
Published: (2023)
GaussTrap: Stealthy Poisoning Attacks on 3D Gaussian Splatting for Targeted Scene Confusion
by: Hong, Jiaxin, et al.
Published: (2025)
by: Hong, Jiaxin, et al.
Published: (2025)
StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided Backdoors
by: Yu, Qinkai, et al.
Published: (2026)
by: Yu, Qinkai, et al.
Published: (2026)
Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
by: Qiu, Yixiang, et al.
Published: (2025)
by: Qiu, Yixiang, et al.
Published: (2025)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
Stealthy Backdoor Attack via Confidence-driven Sampling
by: He, Pengfei, et al.
Published: (2023)
by: He, Pengfei, et al.
Published: (2023)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
by: Zhang, Bokang, et al.
Published: (2025)
by: Zhang, Bokang, et al.
Published: (2025)
Poisoning the Pixels: Revisiting Backdoor Attacks on Semantic Segmentation
by: Zhang, Guangsheng, et al.
Published: (2026)
by: Zhang, Guangsheng, et al.
Published: (2026)
Bloodroot: When Watermarking Turns Poisonous For Stealthy Backdoor
by: Chen, Kuan-Yu, et al.
Published: (2025)
by: Chen, Kuan-Yu, et al.
Published: (2025)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
by: Yang, Xiaochen, et al.
Published: (2026)
by: Yang, Xiaochen, et al.
Published: (2026)
JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
by: Liao, Yifan, et al.
Published: (2025)
by: Liao, Yifan, et al.
Published: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
by: Li, Ziqiang, et al.
Published: (2024)
by: Li, Ziqiang, et al.
Published: (2024)
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
by: Guo, Ji, et al.
Published: (2026)
by: Guo, Ji, et al.
Published: (2026)
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Stealthy Targeted Backdoor Attacks against Image Captioning
by: Fan, Wenshu, et al.
Published: (2024)
by: Fan, Wenshu, et al.
Published: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
by: Hsu, Chia-Yi, et al.
Published: (2026)
by: Hsu, Chia-Yi, et al.
Published: (2026)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
by: Zhong, Zhiyuan, et al.
Published: (2025)
by: Zhong, Zhiyuan, et al.
Published: (2025)
Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs
by: Luo, Qi, et al.
Published: (2026)
by: Luo, Qi, et al.
Published: (2026)
FlowMur: A Stealthy and Practical Audio Backdoor Attack with Limited Knowledge
by: Lan, Jiahe, et al.
Published: (2023)
by: Lan, Jiahe, et al.
Published: (2023)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Privacy Leakage on DNNs: A Survey of Model Inversion Attacks and Defenses
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
Planning Stealthy Backdoor Attacks in MDPs with Observation-Based Triggers
by: Wei, Xinyi, et al.
Published: (2025)
by: Wei, Xinyi, et al.
Published: (2025)
HoneyImage: Verifiable, Harmless, and Stealthy Dataset Ownership Verification for Image Models
by: Zhu, Zhihao, et al.
Published: (2025)
by: Zhu, Zhihao, et al.
Published: (2025)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness
by: Feng, Yu, et al.
Published: (2025)
by: Feng, Yu, et al.
Published: (2025)
Similar Items
-
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025) -
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
by: Sun, Shuoyang, et al.
Published: (2026) -
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025) -
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025) -
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)