On Adversarial Examples for Text Classification by Perturbing Latent Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sooksatra, Korn, Khanal, Bikram, Rivas, Pablo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913343970738176
author Sooksatra, Korn
Khanal, Bikram
Rivas, Pablo
author_facet Sooksatra, Korn
Khanal, Bikram
Rivas, Pablo
contents Recently, with the advancement of deep learning, several applications in text classification have advanced significantly. However, this improvement comes with a cost because deep learning is vulnerable to adversarial examples. This weakness indicates that deep learning is not very robust. Fortunately, the input of a text classifier is discrete. Hence, it can prevent the classifier from state-of-the-art attacks. Nonetheless, previous works have generated black-box attacks that successfully manipulate the discrete values of the input to find adversarial examples. Therefore, instead of changing the discrete values, we transform the input into its embedding vector containing real values to perform the state-of-the-art white-box attacks. Then, we convert the perturbed embedding vector back into a text and name it an adversarial example. In summary, we create a framework that measures the robustness of a text classifier by using the gradients of the classifier.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03789
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Adversarial Examples for Text Classification by Perturbing Latent Representations
Sooksatra, Korn
Khanal, Bikram
Rivas, Pablo
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
68T01, 68T50
I.2.7
Recently, with the advancement of deep learning, several applications in text classification have advanced significantly. However, this improvement comes with a cost because deep learning is vulnerable to adversarial examples. This weakness indicates that deep learning is not very robust. Fortunately, the input of a text classifier is discrete. Hence, it can prevent the classifier from state-of-the-art attacks. Nonetheless, previous works have generated black-box attacks that successfully manipulate the discrete values of the input to find adversarial examples. Therefore, instead of changing the discrete values, we transform the input into its embedding vector containing real values to perform the state-of-the-art white-box attacks. Then, we convert the perturbed embedding vector back into a text and name it an adversarial example. In summary, we create a framework that measures the robustness of a text classifier by using the gradients of the classifier.
title On Adversarial Examples for Text Classification by Perturbing Latent Representations
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
68T01, 68T50
I.2.7
url https://arxiv.org/abs/2405.03789