ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zihan, Zhang, Rui, Li, Hongwei, Fan, Wenshu, Jiang, Wenbo, Zhao, Qingchuan, Xu, Guowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908642868985856
author Wang, Zihan
Zhang, Rui
Li, Hongwei
Fan, Wenshu
Jiang, Wenbo
Zhao, Qingchuan
Xu, Guowen
author_facet Wang, Zihan
Zhang, Rui
Li, Hongwei
Fan, Wenshu
Jiang, Wenbo
Zhao, Qingchuan
Xu, Guowen
contents Backdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100\% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
Wang, Zihan
Zhang, Rui
Li, Hongwei
Fan, Wenshu
Jiang, Wenbo
Zhao, Qingchuan
Xu, Guowen
Cryptography and Security
Computation and Language
Backdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100\% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments.
title ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2508.01365