Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maloyan, Narek, Namiot, Dmitry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912346544275456
author Maloyan, Narek
Namiot, Dmitry
author_facet Maloyan, Narek
Namiot, Dmitry
contents LLM as judge systems used to assess text quality code correctness and argument strength are vulnerable to prompt injection attacks. We introduce a framework that separates content author attacks from system prompt attacks and evaluate five models Gemma 3.27B Gemma 3.4B Llama 3.2 3B GPT 4 and Claude 3 Opus on four tasks with various defenses using fifty prompts per condition. Attacks achieved up to seventy three point eight percent success smaller models proved more vulnerable and transferability ranged from fifty point five to sixty two point six percent. Our results contrast with Universal Prompt Injection and AdvPrompter We recommend multi model committees and comparative scoring and release all code and datasets
format Preprint
id arxiv_https___arxiv_org_abs_2504_18333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
Maloyan, Narek
Namiot, Dmitry
Cryptography and Security
Computation and Language
LLM as judge systems used to assess text quality code correctness and argument strength are vulnerable to prompt injection attacks. We introduce a framework that separates content author attacks from system prompt attacks and evaluate five models Gemma 3.27B Gemma 3.4B Llama 3.2 3B GPT 4 and Claude 3 Opus on four tasks with various defenses using fifty prompts per condition. Attacks achieved up to seventy three point eight percent success smaller models proved more vulnerable and transferability ranged from fifty point five to sixty two point six percent. Our results contrast with Universal Prompt Injection and AdvPrompter We recommend multi model committees and comparative scoring and release all code and datasets
title Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2504.18333