Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917143498457088 |
|---|---|
| author | Liu, Lijia Kondo, Takumi Atarashi, Kyohei Takeuchi, Koh Li, Jiyi Saito, Shigeru Kashima, Hisashi |
| author_facet | Liu, Lijia Kondo, Takumi Atarashi, Kyohei Takeuchi, Koh Li, Jiyi Saito, Shigeru Kashima, Hisashi |
| contents | This paper investigates defenses for LLM-based evaluation systems against prompt injection. We formalize a class of threats called blind attacks, where a candidate answer is crafted independently of the true answer to deceive the evaluator. To counter such attacks, we propose a framework that augments Standard Evaluation (SE) with Counterfactual Evaluation (CFE), which re-evaluates the submission against a deliberately false ground-truth answer. An attack is detected if the system validates an answer under both standard and counterfactual conditions. Experiments show that while standard evaluation is highly vulnerable, our SE+CFE framework significantly improves security by boosting attack detection with minimal performance trade-offs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_23453 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems Liu, Lijia Kondo, Takumi Atarashi, Kyohei Takeuchi, Koh Li, Jiyi Saito, Shigeru Kashima, Hisashi Cryptography and Security Computation and Language This paper investigates defenses for LLM-based evaluation systems against prompt injection. We formalize a class of threats called blind attacks, where a candidate answer is crafted independently of the true answer to deceive the evaluator. To counter such attacks, we propose a framework that augments Standard Evaluation (SE) with Counterfactual Evaluation (CFE), which re-evaluates the submission against a deliberately false ground-truth answer. An attack is detected if the system validates an answer under both standard and counterfactual conditions. Experiments show that while standard evaluation is highly vulnerable, our SE+CFE framework significantly improves security by boosting attack detection with minimal performance trade-offs. |
| title | Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems |
| topic | Cryptography and Security Computation and Language |
| url | https://arxiv.org/abs/2507.23453 |