STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908761578274816 |
|---|---|
| author | Park, Seong-Gyu Park, Sohee Lee, Jisu Na, Hyunsik Choi, Daeseon |
| author_facet | Park, Seong-Gyu Park, Sohee Lee, Jisu Na, Hyunsik Choi, Daeseon |
| contents | Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model's general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\approx$ 1.0) with approximately $42\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_08511 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio Park, Seong-Gyu Park, Sohee Lee, Jisu Na, Hyunsik Choi, Daeseon Computation and Language Cryptography and Security Machine Learning Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model's general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\approx$ 1.0) with approximately $42\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection. |
| title | STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio |
| topic | Computation and Language Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2601.08511 |