Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Lijia, Kondo, Takumi, Atarashi, Kyohei, Takeuchi, Koh, Li, Jiyi, Saito, Shigeru, Kashima, Hisashi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917143498457088
author Liu, Lijia
Kondo, Takumi
Atarashi, Kyohei
Takeuchi, Koh
Li, Jiyi
Saito, Shigeru
Kashima, Hisashi
author_facet Liu, Lijia
Kondo, Takumi
Atarashi, Kyohei
Takeuchi, Koh
Li, Jiyi
Saito, Shigeru
Kashima, Hisashi
contents This paper investigates defenses for LLM-based evaluation systems against prompt injection. We formalize a class of threats called blind attacks, where a candidate answer is crafted independently of the true answer to deceive the evaluator. To counter such attacks, we propose a framework that augments Standard Evaluation (SE) with Counterfactual Evaluation (CFE), which re-evaluates the submission against a deliberately false ground-truth answer. An attack is detected if the system validates an answer under both standard and counterfactual conditions. Experiments show that while standard evaluation is highly vulnerable, our SE+CFE framework significantly improves security by boosting attack detection with minimal performance trade-offs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
Liu, Lijia
Kondo, Takumi
Atarashi, Kyohei
Takeuchi, Koh
Li, Jiyi
Saito, Shigeru
Kashima, Hisashi
Cryptography and Security
Computation and Language
This paper investigates defenses for LLM-based evaluation systems against prompt injection. We formalize a class of threats called blind attacks, where a candidate answer is crafted independently of the true answer to deceive the evaluator. To counter such attacks, we propose a framework that augments Standard Evaluation (SE) with Counterfactual Evaluation (CFE), which re-evaluates the submission against a deliberately false ground-truth answer. An attack is detected if the system validates an answer under both standard and counterfactual conditions. Experiments show that while standard evaluation is highly vulnerable, our SE+CFE framework significantly improves security by boosting attack detection with minimal performance trade-offs.
title Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2507.23453