STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Park, Seong-Gyu, Park, Sohee, Lee, Jisu, Na, Hyunsik, Choi, Daeseon
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908761578274816
author Park, Seong-Gyu
Park, Sohee
Lee, Jisu
Na, Hyunsik
Choi, Daeseon
author_facet Park, Seong-Gyu
Park, Sohee
Lee, Jisu
Na, Hyunsik
Choi, Daeseon
contents Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model's general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\approx$ 1.0) with approximately $42\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
Park, Seong-Gyu
Park, Sohee
Lee, Jisu
Na, Hyunsik
Choi, Daeseon
Computation and Language
Cryptography and Security
Machine Learning
Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model's general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\approx$ 1.0) with approximately $42\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection.
title STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2601.08511