Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qin, Mo, Wenjie, Tong, Terry, Xu, Jiashu, Wang, Fei, Xiao, Chaowei, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
Instructional Fingerprinting of Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
von: Xu, Jiashu, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Detecting and Mitigating Backdoor Attacks in OTA-FL Systems: A Two-Stage Robust Aggregation Scheme
von: Ma, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoyan, et al.
Veröffentlicht: (2026)
Rethinking Backdoor Detection Evaluation for Language Models
von: Yan, Jun, et al.
Veröffentlicht: (2024)
von: Yan, Jun, et al.
Veröffentlicht: (2024)
Triaging Threats to Specialized Guardrails
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2026)
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2026)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
Threat Modeling and Attack Surface Analysis of IoT-Enabled Controlled Environment Agriculture Systems
von: Vakhnovskyi, Andrii
Veröffentlicht: (2026)
von: Vakhnovskyi, Andrii
Veröffentlicht: (2026)
Cyberscurity Threats and Defense Mechanisms in IoT network
von: Dao, Trung, et al.
Veröffentlicht: (2026)
von: Dao, Trung, et al.
Veröffentlicht: (2026)
Graph Representation-based Model Poisoning on Federated Large Language Models
von: Cai, Hanlin, et al.
Veröffentlicht: (2025)
von: Cai, Hanlin, et al.
Veröffentlicht: (2025)
Test-time Backdoor Mitigation for Black-Box Large Language Models with Defensive Demonstrations
von: Mo, Wenjie, et al.
Veröffentlicht: (2023)
von: Mo, Wenjie, et al.
Veröffentlicht: (2023)
Threat-based Security Controls to Protect Industrial Control Systems
von: Srinivasan, Haritha, et al.
Veröffentlicht: (2025)
von: Srinivasan, Haritha, et al.
Veröffentlicht: (2025)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications
von: Zaboli, Aydin, et al.
Veröffentlicht: (2023)
von: Zaboli, Aydin, et al.
Veröffentlicht: (2023)
Open Sky, Open Threats: Replay Attacks in Space Launch and Re-entry Phases
von: Benchoubane, Nesrine, et al.
Veröffentlicht: (2025)
von: Benchoubane, Nesrine, et al.
Veröffentlicht: (2025)
Evaluation of Real-Time Mitigation Techniques for Cyber Security in IEC 61850 / IEC 62351 Substations
von: Herath, Akila, et al.
Veröffentlicht: (2025)
von: Herath, Akila, et al.
Veröffentlicht: (2025)
FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation
von: Mitra, Shaswata, et al.
Veröffentlicht: (2025)
von: Mitra, Shaswata, et al.
Veröffentlicht: (2025)
SG-ML: Smart Grid Cyber Range Modelling Language
von: Roomi, Muhammad M., et al.
Veröffentlicht: (2025)
von: Roomi, Muhammad M., et al.
Veröffentlicht: (2025)
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
von: Chen, Xiangsen, et al.
Veröffentlicht: (2026)
von: Chen, Xiangsen, et al.
Veröffentlicht: (2026)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yu, Zhiyuan, et al.
Veröffentlicht: (2024)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
von: Yi, Biao, et al.
Veröffentlicht: (2025)
von: Yi, Biao, et al.
Veröffentlicht: (2025)
Enhancing Cyber-Resiliency of DER-based SmartGrid: A Survey
von: Liu, Mengxiang, et al.
Veröffentlicht: (2023)
von: Liu, Mengxiang, et al.
Veröffentlicht: (2023)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Limits of Residual-Based Detection for Physically Consistent False Data Injection
von: Xiao, Chenhan, et al.
Veröffentlicht: (2026)
von: Xiao, Chenhan, et al.
Veröffentlicht: (2026)
Cybersecurity Threats to Power Grid Operations from the Demand-Side Response Ecosystem
von: Lakshminarayana, Subhash, et al.
Veröffentlicht: (2023)
von: Lakshminarayana, Subhash, et al.
Veröffentlicht: (2023)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
Cooking Up Risks: Benchmarking and Reducing Food Safety Risks in Large Language Models
von: Luo, Weidi, et al.
Veröffentlicht: (2026)
von: Luo, Weidi, et al.
Veröffentlicht: (2026)
Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems
von: Sharshar, Muhammad, et al.
Veröffentlicht: (2025)
von: Sharshar, Muhammad, et al.
Veröffentlicht: (2025)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
von: Liang, Zi, et al.
Veröffentlicht: (2024)
von: Liang, Zi, et al.
Veröffentlicht: (2024)
Privacy-Aware Smart Cameras: View Coverage via Socially Responsible Coordination
von: Qin, Chuhao, et al.
Veröffentlicht: (2026)
von: Qin, Chuhao, et al.
Veröffentlicht: (2026)
A Survey of Machine Learning-based Physical-Layer Authentication in Wireless Communications
von: Meng, Rui, et al.
Veröffentlicht: (2024)
von: Meng, Rui, et al.
Veröffentlicht: (2024)
Leveraging Functional Encryption and Deep Learning for Privacy-Preserving Traffic Forecasting
von: Adom, Isaac, et al.
Veröffentlicht: (2025)
von: Adom, Isaac, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024) -
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023) -
Instructional Fingerprinting of Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2024) -
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)