Selective Adversarial Attacks on LLM Benchmarks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dubrovsky, Ivan, Orlova, Anastasia, Iov, Illarion, Gubina, Nina, Gureeva, Irena, Zaytsev, Alexey
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914095111864320
author Dubrovsky, Ivan
Orlova, Anastasia
Iov, Illarion
Gubina, Nina
Gureeva, Irena
Zaytsev, Alexey
author_facet Dubrovsky, Ivan
Orlova, Anastasia
Iov, Illarion
Gubina, Nina
Gureeva, Irena
Zaytsev, Alexey
contents Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Prior work on adversarial robustness in NLP has emphasized text attacks that affect many models equally, leaving open the question of whether it is possible to selectively degrade or enhance performance while minimally affecting other models. We formalize this problem and study selective adversarial attacks on MMLU - a widely used benchmark designed to measure a language model's broad general knowledge and reasoning ability across different subjects. Using canonical attacks integrated into TextAttack framework, we introduce a protocol for selectivity assessment, develop a custom constraint to increase selectivity of attacks and propose a surrogate-LLM pipeline that generates selective perturbations. Empirically, we find that selective adversarial attacks exist and can materially alter relative rankings, challenging the fairness, reproducibility, and transparency of leaderboard-driven evaluation. Our results motivate perturbation-aware reporting and robustness diagnostics for LLM evaluation and demonstrate that even subtle edits can shift comparative judgments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13570
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Selective Adversarial Attacks on LLM Benchmarks
Dubrovsky, Ivan
Orlova, Anastasia
Iov, Illarion
Gubina, Nina
Gureeva, Irena
Zaytsev, Alexey
Machine Learning
Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Prior work on adversarial robustness in NLP has emphasized text attacks that affect many models equally, leaving open the question of whether it is possible to selectively degrade or enhance performance while minimally affecting other models. We formalize this problem and study selective adversarial attacks on MMLU - a widely used benchmark designed to measure a language model's broad general knowledge and reasoning ability across different subjects. Using canonical attacks integrated into TextAttack framework, we introduce a protocol for selectivity assessment, develop a custom constraint to increase selectivity of attacks and propose a surrogate-LLM pipeline that generates selective perturbations. Empirically, we find that selective adversarial attacks exist and can materially alter relative rankings, challenging the fairness, reproducibility, and transparency of leaderboard-driven evaluation. Our results motivate perturbation-aware reporting and robustness diagnostics for LLM evaluation and demonstrate that even subtle edits can shift comparative judgments.
title Selective Adversarial Attacks on LLM Benchmarks
topic Machine Learning
url https://arxiv.org/abs/2510.13570