Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Chen, Sun, Yuchen, Gao, Jiaxin, Jia, Yanwen, Gong, Xueluan, Wang, Qian, Lam, Kwok-Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025)
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
Hidden Data Privacy Breaches in Federated Learning
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
TrojanDam: Detection-Free Backdoor Defense in Federated Learning through Proactive Model Robustification utilizing OOD Data
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
An Effective and Resilient Backdoor Attack Framework against Deep Neural Networks and Vision Transformers
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning
von: Shui, Yuhan, et al.
Veröffentlicht: (2026)
von: Shui, Yuhan, et al.
Veröffentlicht: (2026)
Backdoor Attack with Sparse and Invisible Trigger
von: Gao, Yinghua, et al.
Veröffentlicht: (2023)
von: Gao, Yinghua, et al.
Veröffentlicht: (2023)
Threats, Attacks, and Defenses in Machine Unlearning: A Survey
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
von: Liu, Ziyao, et al.
Veröffentlicht: (2024)
Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
von: Zhou, Kaiyu, et al.
Veröffentlicht: (2026)
Proactive Detection of Physical Inter-rule Vulnerabilities in IoT Services Using a Deep Learning Approach
von: Huang, Bing, et al.
Veröffentlicht: (2024)
von: Huang, Bing, et al.
Veröffentlicht: (2024)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
von: Yang, Yuxin, et al.
Veröffentlicht: (2024)
TED-LaST: Towards Robust Backdoor Defense Against Adaptive Attacks
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Mo, Xiaoxing, et al.
Veröffentlicht: (2025)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
von: Liu, Mingrui, et al.
Veröffentlicht: (2025)
Towards Physical World Backdoor Attacks against Skeleton Action Recognition
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
Neutralizing Backdoors through Information Conflicts for Large Language Models
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
BDFirewall: Towards Effective and Expeditiously Black-Box Backdoor Defense in MLaaS
von: Li, Ye, et al.
Veröffentlicht: (2025)
von: Li, Ye, et al.
Veröffentlicht: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Backdoor Attacks and Defenses in Computer Vision Domain: A Survey
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
von: Abbasi, Bilal Hussain, et al.
Veröffentlicht: (2025)
Universal Graph Backdoor Defense: A Feature-based Homophily Perspective
von: Pan, Mengting, et al.
Veröffentlicht: (2026)
von: Pan, Mengting, et al.
Veröffentlicht: (2026)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
von: Zhai, Shengfang, et al.
Veröffentlicht: (2025)
von: Zhai, Shengfang, et al.
Veröffentlicht: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
von: Li, Xiaohu, et al.
Veröffentlicht: (2025)
Evolutionary Trigger Detection and Lightweight Model Repair Based Backdoor Defense
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
von: Yan, Lu, et al.
Veröffentlicht: (2024)
von: Yan, Lu, et al.
Veröffentlicht: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
von: Xiong, Chen, et al.
Veröffentlicht: (2024)
Backdoor Threats in Variational Quantum Circuits: Taxonomy, Attacks, and Defenses
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
Prototype-Guided Robust Learning against Backdoor Attacks
von: Guo, Wei, et al.
Veröffentlicht: (2025)
von: Guo, Wei, et al.
Veröffentlicht: (2025)
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing
von: Fan, Kaisheng, et al.
Veröffentlicht: (2026)
von: Fan, Kaisheng, et al.
Veröffentlicht: (2026)
Robustness Inspired Graph Backdoor Defense
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiwei, et al.
Veröffentlicht: (2024)
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
von: Yu, Haiyang, et al.
Veröffentlicht: (2024)
Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
von: Wang, He, et al.
Veröffentlicht: (2026)
von: Wang, He, et al.
Veröffentlicht: (2026)
ARMOR: Shielding Unlearnable Examples against Data Augmentation
von: Gong, Xueluan, et al.
Veröffentlicht: (2025)
von: Gong, Xueluan, et al.
Veröffentlicht: (2025)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
von: Liu, Mingrui, et al.
Veröffentlicht: (2026)
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024)
von: Yi, Biao, et al.
Veröffentlicht: (2024)
SoK: The Last Line of Defense: On Backdoor Defense Evaluation
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
von: Abad, Gorka, et al.
Veröffentlicht: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
von: Zhou, Yihe, et al.
Veröffentlicht: (2025)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
Seal Your Backdoor with Variational Defense
von: Sabolić, Ivan, et al.
Veröffentlicht: (2025)
von: Sabolić, Ivan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems
von: Gao, Jiaxin, et al.
Veröffentlicht: (2025) -
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
von: Gong, Xueluan, et al.
Veröffentlicht: (2024) -
LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
von: Li, Songze, et al.
Veröffentlicht: (2025) -
Hidden Data Privacy Breaches in Federated Learning
von: Gong, Xueluan, et al.
Veröffentlicht: (2024) -
TrojanDam: Detection-Free Backdoor Defense in Federated Learning through Proactive Model Robustification utilizing OOD Data
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)