Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Gejian, Wu, Hanzhou, Zhang, Xinpeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
by: Zhao, Gejian, et al.
Published: (2025)
by: Zhao, Gejian, et al.
Published: (2025)
An Information Asymmetry Game for Trigger-based DNN Model Watermarking
by: Huang, Chaoyue, et al.
Published: (2025)
by: Huang, Chaoyue, et al.
Published: (2025)
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
by: Zhao, Xiangyu, et al.
Published: (2024)
by: Zhao, Xiangyu, et al.
Published: (2024)
Robust and Imperceptible Black-box DNN Watermarking Based on Fourier Perturbation Analysis and Frequency Sensitivity Clustering
by: Liu, Yong, et al.
Published: (2022)
by: Liu, Yong, et al.
Published: (2022)
Generative Model Watermarking Suppressing High-Frequency Artifacts
by: Zhang, Li, et al.
Published: (2023)
by: Zhang, Li, et al.
Published: (2023)
A Game Between the Defender and the Attacker for Trigger-based Black-box Model Watermarking
by: Huang, Chaoyue, et al.
Published: (2025)
by: Huang, Chaoyue, et al.
Published: (2025)
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
by: Zhang, Ziming, et al.
Published: (2026)
by: Zhang, Ziming, et al.
Published: (2026)
Yet Another Watermark for Large Language Models
by: Bao, Siyuan, et al.
Published: (2025)
by: Bao, Siyuan, et al.
Published: (2025)
Defining Cost Function of Steganography with Large Language Models
by: Wu, Hanzhou, et al.
Published: (2025)
by: Wu, Hanzhou, et al.
Published: (2025)
Watermarking Quantum Neural Networks Based on Sample Grouped and Paired Training
by: Zhou, Limengnan, et al.
Published: (2025)
by: Zhou, Limengnan, et al.
Published: (2025)
A Fingerprint for Large Language Models
by: Yang, Zhiguang, et al.
Published: (2024)
by: Yang, Zhiguang, et al.
Published: (2024)
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
by: Sun, Chengrui, et al.
Published: (2025)
by: Sun, Chengrui, et al.
Published: (2025)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
A Practical Trigger-Free Backdoor Attack on Neural Networks
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers
by: Zhao, Xin, et al.
Published: (2025)
by: Zhao, Xin, et al.
Published: (2025)
Hardware-Triggered Backdoors
by: Möller, Jonas, et al.
Published: (2026)
by: Möller, Jonas, et al.
Published: (2026)
Backdoor Contrastive Learning via Bi-level Trigger Optimization
by: Sun, Weiyu, et al.
Published: (2024)
by: Sun, Weiyu, et al.
Published: (2024)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
by: Lin, Chenhao, et al.
Published: (2025)
by: Lin, Chenhao, et al.
Published: (2025)
Rotation, Scale, and Translation Resilient Black-box Fingerprinting for Intellectual Property Protection of EaaS Models
by: Zhang, Hongjie, et al.
Published: (2025)
by: Zhang, Hongjie, et al.
Published: (2025)
Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors
by: Abad, Gorka, et al.
Published: (2026)
by: Abad, Gorka, et al.
Published: (2026)
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
by: Bullwinkel, Blake, et al.
Published: (2026)
by: Bullwinkel, Blake, et al.
Published: (2026)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
by: Yao, Duanyi, et al.
Published: (2026)
by: Yao, Duanyi, et al.
Published: (2026)
Revocable Backdoor for Deep Model Trading
by: Xu, Yiran, et al.
Published: (2024)
by: Xu, Yiran, et al.
Published: (2024)
Mitigating Backdoor Triggered and Targeted Data Poisoning Attacks in Voice Authentication Systems
by: Mohammadi, Alireza, et al.
Published: (2025)
by: Mohammadi, Alireza, et al.
Published: (2025)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
M-to-N Backdoor Paradigm: A Multi-Trigger and Multi-Target Attack to Deep Learning Models
by: Hou, Linshan, et al.
Published: (2022)
by: Hou, Linshan, et al.
Published: (2022)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
by: Zhang, Bokang, et al.
Published: (2025)
by: Zhang, Bokang, et al.
Published: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
by: Ashcraft, Chace, et al.
Published: (2025)
by: Ashcraft, Chace, et al.
Published: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
by: Hou, Yang, et al.
Published: (2024)
by: Hou, Yang, et al.
Published: (2024)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
by: Wang, Ruofei, et al.
Published: (2024)
by: Wang, Ruofei, et al.
Published: (2024)
Attack by Yourself: Effective and Unnoticeable Multi-Category Graph Backdoor Attacks with Subgraph Triggers Pool
by: Li, Jiangtong, et al.
Published: (2024)
by: Li, Jiangtong, et al.
Published: (2024)
ME: Trigger Element Combination Backdoor Attack on Copyright Infringement
by: Yang, Feiyu, et al.
Published: (2025)
by: Yang, Feiyu, et al.
Published: (2025)
Cross-Paradigm Graph Backdoor Attacks with Promptable Subgraph Triggers
by: Liu, Dongyi, et al.
Published: (2025)
by: Liu, Dongyi, et al.
Published: (2025)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
by: Ma, Yuan, et al.
Published: (2024)
by: Ma, Yuan, et al.
Published: (2024)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
Exposing Hidden Backdoors in NFT Smart Contracts: A Static Security Analysis of Rug Pull Patterns
by: Pathade, Chetan, et al.
Published: (2025)
by: Pathade, Chetan, et al.
Published: (2025)
BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors
by: Hsu, Chia-Yi, et al.
Published: (2025)
by: Hsu, Chia-Yi, et al.
Published: (2025)
Similar Items
-
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
by: Zhao, Gejian, et al.
Published: (2025) -
An Information Asymmetry Game for Trigger-based DNN Model Watermarking
by: Huang, Chaoyue, et al.
Published: (2025) -
Transferable Watermarking to Self-supervised Pre-trained Graph Encoders by Trigger Embeddings
by: Zhao, Xiangyu, et al.
Published: (2024) -
Robust and Imperceptible Black-box DNN Watermarking Based on Fourier Perturbation Analysis and Frequency Sensitivity Clustering
by: Liu, Yong, et al.
Published: (2022) -
Generative Model Watermarking Suppressing High-Frequency Artifacts
by: Zhang, Li, et al.
Published: (2023)