Differential Robustness in Transformer Language Models: Empirical Evaluation Under Adversarial Text Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gidatkar, Taniya, Ajao, Oluwaseun, Shardlow, Matthew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912584492384256
author Gidatkar, Taniya
Ajao, Oluwaseun
Shardlow, Matthew
author_facet Gidatkar, Taniya
Ajao, Oluwaseun
Shardlow, Matthew
contents This study evaluates the resilience of large language models (LLMs) against adversarial attacks, specifically focusing on Flan-T5, BERT, and RoBERTa-Base. Using systematically designed adversarial tests through TextFooler and BERTAttack, we found significant variations in model robustness. RoBERTa-Base and FlanT5 demonstrated remarkable resilience, maintaining accuracy even when subjected to sophisticated attacks, with attack success rates of 0%. In contrast. BERT-Base showed considerable vulnerability, with TextFooler achieving a 93.75% success rate in reducing model accuracy from 48% to just 3%. Our research reveals that while certain LLMs have developed effective defensive mechanisms, these safeguards often require substantial computational resources. This study contributes to the understanding of LLM security by identifying existing strengths and weaknesses in current safeguarding approaches and proposes practical recommendations for developing more efficient and effective defensive strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Differential Robustness in Transformer Language Models: Empirical Evaluation Under Adversarial Text Attacks
Gidatkar, Taniya
Ajao, Oluwaseun
Shardlow, Matthew
Cryptography and Security
Artificial Intelligence
Computation and Language
I.2; H.3.3
This study evaluates the resilience of large language models (LLMs) against adversarial attacks, specifically focusing on Flan-T5, BERT, and RoBERTa-Base. Using systematically designed adversarial tests through TextFooler and BERTAttack, we found significant variations in model robustness. RoBERTa-Base and FlanT5 demonstrated remarkable resilience, maintaining accuracy even when subjected to sophisticated attacks, with attack success rates of 0%. In contrast. BERT-Base showed considerable vulnerability, with TextFooler achieving a 93.75% success rate in reducing model accuracy from 48% to just 3%. Our research reveals that while certain LLMs have developed effective defensive mechanisms, these safeguards often require substantial computational resources. This study contributes to the understanding of LLM security by identifying existing strengths and weaknesses in current safeguarding approaches and proposes practical recommendations for developing more efficient and effective defensive strategies.
title Differential Robustness in Transformer Language Models: Empirical Evaluation Under Adversarial Text Attacks
topic Cryptography and Security
Artificial Intelligence
Computation and Language
I.2; H.3.3
url https://arxiv.org/abs/2509.09706