Rethinking Backdoor Detection Evaluation for Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Jun, Mo, Wenjie Jacky, Ren, Xiang, Jia, Robin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023)
por: Yan, Jun, et al.
Publicado: (2023)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
por: Hu, Man, et al.
Publicado: (2025)
por: Hu, Man, et al.
Publicado: (2025)
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)
por: Wen, Rui, et al.
Publicado: (2026)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
por: Li, Haoran, et al.
Publicado: (2024)
por: Li, Haoran, et al.
Publicado: (2024)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
por: Wen, Xiaofei, et al.
Publicado: (2025)
por: Wen, Xiaofei, et al.
Publicado: (2025)
Triaging Threats to Specialized Guardrails
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
por: Li, Ziqiang, et al.
Publicado: (2024)
por: Li, Ziqiang, et al.
Publicado: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
por: Ding, Yidong, et al.
Publicado: (2025)
por: Ding, Yidong, et al.
Publicado: (2025)
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
por: Fortier, Alexandrine, et al.
Publicado: (2025)
por: Fortier, Alexandrine, et al.
Publicado: (2025)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
por: Xue, Eric, et al.
Publicado: (2025)
por: Xue, Eric, et al.
Publicado: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
por: Jiang, Peihai, et al.
Publicado: (2025)
por: Jiang, Peihai, et al.
Publicado: (2025)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
por: Cheng, Pengzhou, et al.
Publicado: (2024)
por: Cheng, Pengzhou, et al.
Publicado: (2024)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
por: Wang, Zhuoshang, et al.
Publicado: (2026)
por: Wang, Zhuoshang, et al.
Publicado: (2026)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
por: Wu, Zhengxian, et al.
Publicado: (2025)
por: Wu, Zhengxian, et al.
Publicado: (2025)
Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
por: Yi, Biao, et al.
Publicado: (2025)
por: Yi, Biao, et al.
Publicado: (2025)
Composite Backdoor Attacks Against Large Language Models
por: Huang, Hai, et al.
Publicado: (2023)
por: Huang, Hai, et al.
Publicado: (2023)
CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
por: Zeng, Rui, et al.
Publicado: (2024)
por: Zeng, Rui, et al.
Publicado: (2024)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
por: He, Xuanli, et al.
Publicado: (2024)
por: He, Xuanli, et al.
Publicado: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
por: He, Xuanli, et al.
Publicado: (2024)
por: He, Xuanli, et al.
Publicado: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
por: Du, Wei, et al.
Publicado: (2023)
por: Du, Wei, et al.
Publicado: (2023)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
por: Yuan, Xiaohan, et al.
Publicado: (2024)
por: Yuan, Xiaohan, et al.
Publicado: (2024)
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
por: Yin, Rui, et al.
Publicado: (2026)
por: Yin, Rui, et al.
Publicado: (2026)
Rethinking How to Evaluate Language Model Jailbreak
por: Cai, Hongyu, et al.
Publicado: (2024)
por: Cai, Hongyu, et al.
Publicado: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
por: Cao, Yuanpu, et al.
Publicado: (2023)
por: Cao, Yuanpu, et al.
Publicado: (2023)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2023)
por: Xu, Jiashu, et al.
Publicado: (2023)
Task-Agnostic Detector for Insertion-Based Backdoor Attacks
por: Lyu, Weimin, et al.
Publicado: (2024)
por: Lyu, Weimin, et al.
Publicado: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
por: Wu, Zongru, et al.
Publicado: (2024)
por: Wu, Zongru, et al.
Publicado: (2024)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
por: Cui, Xinyue, et al.
Publicado: (2025)
por: Cui, Xinyue, et al.
Publicado: (2025)
Safely Learning with Private Data: A Federated Learning Framework for Large Language Model
por: Zheng, JiaYing, et al.
Publicado: (2024)
por: Zheng, JiaYing, et al.
Publicado: (2024)
BadActs: A Universal Backdoor Defense in the Activation Space
por: Yi, Biao, et al.
Publicado: (2024)
por: Yi, Biao, et al.
Publicado: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
por: Peng, Yuefeng, et al.
Publicado: (2024)
por: Peng, Yuefeng, et al.
Publicado: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
por: Hao, Yunzhuo, et al.
Publicado: (2024)
por: Hao, Yunzhuo, et al.
Publicado: (2024)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
por: Wu, Zongru, et al.
Publicado: (2024)
por: Wu, Zongru, et al.
Publicado: (2024)
Ejemplares similares
-
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
por: Yan, Jun, et al.
Publicado: (2023) -
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
por: Hu, Man, et al.
Publicado: (2025) -
ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
por: Wang, Zihan, et al.
Publicado: (2025) -
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026) -
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
por: Li, Haoran, et al.
Publicado: (2024)