Backdoor Collapse: Eliminating Unknown Threats via Known Backdoor Aggregation in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Liang, Yu, Miao, Aloqaily, Moayad, Zhou, Zhenhong, Wang, Kun, Pang, Linsey, Mehrotra, Prakhar, Wen, Qingsong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
by: Zhou, Zhenhong, et al.
Published: (2026)
by: Zhou, Zhenhong, et al.
Published: (2026)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
SafeSeek: Universal Attribution of Safety Circuits in Language Models
by: Yu, Miao, et al.
Published: (2026)
by: Yu, Miao, et al.
Published: (2026)
Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
by: Luo, Kaiwen, et al.
Published: (2026)
by: Luo, Kaiwen, et al.
Published: (2026)
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
by: Min, Nay Myat, et al.
Published: (2024)
by: Min, Nay Myat, et al.
Published: (2024)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
by: Yang, Wenkai, et al.
Published: (2024)
by: Yang, Wenkai, et al.
Published: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
by: Wen, Rui, et al.
Published: (2026)
by: Wen, Rui, et al.
Published: (2026)
Editorial for the IJNM Special Issue From the Best Papers of IEEE ICBC 2023 “Advancing Blockchain and Cryptocurrency”
by: Laura Ricci, et al.
Published: (2024)
by: Laura Ricci, et al.
Published: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Test-Time Backdoor Attacks on Multimodal Large Language Models
by: Lu, Dong, et al.
Published: (2024)
by: Lu, Dong, et al.
Published: (2024)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
by: Fortier, Alexandrine, et al.
Published: (2025)
by: Fortier, Alexandrine, et al.
Published: (2025)
Rethinking Backdoor Detection Evaluation for Language Models
by: Yan, Jun, et al.
Published: (2024)
by: Yan, Jun, et al.
Published: (2024)
MEGen: Generative Backdoor into Large Language Models via Model Editing
by: Qiu, Jiyang, et al.
Published: (2024)
by: Qiu, Jiyang, et al.
Published: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)
by: Hao, Yunzhuo, et al.
Published: (2024)
Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
by: Amayuelas, Alfonso, et al.
Published: (2023)
by: Amayuelas, Alfonso, et al.
Published: (2023)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
A Survey on Trustworthy LLM Agents: Threats and Countermeasures
by: Yu, Miao, et al.
Published: (2025)
by: Yu, Miao, et al.
Published: (2025)
Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models
by: Chen, Jinwen, et al.
Published: (2025)
by: Chen, Jinwen, et al.
Published: (2025)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Neutralizing Backdoors through Information Conflicts for Large Language Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)
by: Cao, Yuanpu, et al.
Published: (2023)
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
by: Li, Zherui, et al.
Published: (2025)
by: Li, Zherui, et al.
Published: (2025)
Resource Consumption Threats in Large Language Models
by: Zhang, Yuanhe, et al.
Published: (2026)
by: Zhang, Yuanhe, et al.
Published: (2026)
Backdoor Attacks on Dense Retrieval via Public and Unintentional Triggers
by: Long, Quanyu, et al.
Published: (2024)
by: Long, Quanyu, et al.
Published: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
Backdoor Attack on Multilingual Machine Translation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Blockchain-enabled Energy Trading and Battery-based Sharing in Microgrids
by: Zekiye, Abdulrezzak, et al.
Published: (2024)
by: Zekiye, Abdulrezzak, et al.
Published: (2024)
ED-DAO: Energy Donation Algorithms based on Decentralized Autonomous Organization
by: Zekiye, Abdulrezzak, et al.
Published: (2025)
by: Zekiye, Abdulrezzak, et al.
Published: (2025)
Similar Items
-
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
by: Zhou, Zhenhong, et al.
Published: (2026) -
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
by: Yu, Miao, et al.
Published: (2025) -
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
by: Wang, Jin, et al.
Published: (2026) -
SafeSeek: Universal Attribution of Safety Circuits in Language Models
by: Yu, Miao, et al.
Published: (2026) -
Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment through Latent Acoustic Pattern Triggers
by: Lin, Liang, et al.
Published: (2025)