Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fazla, Arnisa, Krauter, Lucas, Piedrahita, David Guzman, Michail, Andrianos
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909674063790080
author Fazla, Arnisa
Krauter, Lucas
Piedrahita, David Guzman
Michail, Andrianos
author_facet Fazla, Arnisa
Krauter, Lucas
Piedrahita, David Guzman
Michail, Andrianos
contents We extend BeamAttack, an adversarial attack algorithm designed to evaluate the robustness of text classification systems through word-level modifications guided by beam search. Our extensions include support for word deletions and the option to skip substitutions, enabling the discovery of minimal modifications that alter model predictions. We also integrate LIME to better prioritize word replacements. Evaluated across multiple datasets and victim models (BiLSTM, BERT, and adversarially trained RoBERTa) within the BODEGA framework, our approach achieves over a 99\% attack success rate while preserving the semantic and lexical similarity of the original texts. Through both quantitative and qualitative analysis, we highlight BeamAttack's effectiveness and its limitations. Our implementation is available at https://github.com/LucK1Y/BeamAttack
format Preprint
id arxiv_https___arxiv_org_abs_2506_23661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
Fazla, Arnisa
Krauter, Lucas
Piedrahita, David Guzman
Michail, Andrianos
Computation and Language
We extend BeamAttack, an adversarial attack algorithm designed to evaluate the robustness of text classification systems through word-level modifications guided by beam search. Our extensions include support for word deletions and the option to skip substitutions, enabling the discovery of minimal modifications that alter model predictions. We also integrate LIME to better prioritize word replacements. Evaluated across multiple datasets and victim models (BiLSTM, BERT, and adversarially trained RoBERTa) within the BODEGA framework, our approach achieves over a 99\% attack success rate while preserving the semantic and lexical similarity of the original texts. Through both quantitative and qualitative analysis, we highlight BeamAttack's effectiveness and its limitations. Our implementation is available at https://github.com/LucK1Y/BeamAttack
title Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
topic Computation and Language
url https://arxiv.org/abs/2506.23661