Gespeichert in:
| Hauptverfasser: | Fu, Hang, Peng, Wanli, Zhou, Yinghan, Wu, Jiaxuan, Wen, Juan, Xue, Yiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.04261 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EditMF: Drawing an Invisible Fingerprint for Your Large Language Models
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
von: Zhou, Yinghan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinghan, et al.
Veröffentlicht: (2025)
Retrieval-Confused Generation is a Good Defender for Privacy Violation Attack of Large Language Models
von: Peng, Wanli, et al.
Veröffentlicht: (2025)
von: Peng, Wanli, et al.
Veröffentlicht: (2025)
Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing
von: Zhang, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhang, Ziwei, et al.
Veröffentlicht: (2025)
GTSD: Generative Text Steganography Based on Diffusion Model
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
ImF: Implicit Fingerprint for Large Language Models
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025)
A Fingerprint for Large Language Models
von: Yang, Zhiguang, et al.
Veröffentlicht: (2024)
von: Yang, Zhiguang, et al.
Veröffentlicht: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
von: Yuan, Zenghui, et al.
Veröffentlicht: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
von: Qu, Wenjie, et al.
Veröffentlicht: (2025)
von: Qu, Wenjie, et al.
Veröffentlicht: (2025)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
von: Ni, Zhenyang, et al.
Veröffentlicht: (2024)
von: Ni, Zhenyang, et al.
Veröffentlicht: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
von: Hossen, Md Imran, et al.
Veröffentlicht: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Collusion-Driven Impersonation Attack on Channel-Resistant RF Fingerprinting
von: Xu, Zhou, et al.
Veröffentlicht: (2025)
von: Xu, Zhou, et al.
Veröffentlicht: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
von: Gao, Jianxin, et al.
Veröffentlicht: (2026)
von: Gao, Jianxin, et al.
Veröffentlicht: (2026)
Clean-image Backdoor Attacks
von: Rong, Dazhong, et al.
Veröffentlicht: (2024)
von: Rong, Dazhong, et al.
Veröffentlicht: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
von: Yu, Miao, et al.
Veröffentlicht: (2025)
von: Yu, Miao, et al.
Veröffentlicht: (2025)
Transferring Backdoors between Large Language Models by Knowledge Distillation
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Unlearning-Enhanced Website Fingerprinting Attack: Against Backdoor Poisoning in Anonymous Networks
von: Yuan, Yali, et al.
Veröffentlicht: (2025)
von: Yuan, Yali, et al.
Veröffentlicht: (2025)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
von: Yuan, Danni, et al.
Veröffentlicht: (2023)
von: Yuan, Danni, et al.
Veröffentlicht: (2023)
FFCBA: Feature-based Full-target Clean-label Backdoor Attacks
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
von: Yin, Yangxu, et al.
Veröffentlicht: (2025)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
Backdooring Masked Diffusion Language Models
von: Cao, Daniel Yiming, et al.
Veröffentlicht: (2026)
von: Cao, Daniel Yiming, et al.
Veröffentlicht: (2026)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
von: Chen, Yulin, et al.
Veröffentlicht: (2024)
Concept-Guided Backdoor Attack on Vision Language Models
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EditMF: Drawing an Invisible Fingerprint for Your Large Language Models
von: Wu, Jiaxuan, et al.
Veröffentlicht: (2025) -
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025) -
SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025) -
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025) -
Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
von: Zhou, Yinghan, et al.
Veröffentlicht: (2025)