Claim-Guided Textual Backdoor Attack for Practical Applications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Minkyoo, Kim, Hanna, Kim, Jaehan, Jin, Youngjin, Shin, Seungwon
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916410733625344
author Song, Minkyoo
Kim, Hanna
Kim, Jaehan
Jin, Youngjin
Shin, Seungwon
author_facet Song, Minkyoo
Kim, Hanna
Kim, Jaehan
Jin, Youngjin
Shin, Seungwon
contents Recent advances in natural language processing and the increased use of large language models have exposed new security vulnerabilities, such as backdoor attacks. Previous backdoor attacks require input manipulation after model distribution to activate the backdoor, posing limitations in real-world applicability. Addressing this gap, we introduce a novel Claim-Guided Backdoor Attack (CGBA), which eliminates the need for such manipulations by utilizing inherent textual claims as triggers. CGBA leverages claim extraction, clustering, and targeted training to trick models to misbehave on targeted claims without affecting their performance on clean data. CGBA demonstrates its effectiveness and stealthiness across various datasets and models, significantly enhancing the feasibility of practical backdoor attacks. Our code and data will be available at https://github.com/PaperCGBA/CGBA.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16618
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Claim-Guided Textual Backdoor Attack for Practical Applications
Song, Minkyoo
Kim, Hanna
Kim, Jaehan
Jin, Youngjin
Shin, Seungwon
Computation and Language
Artificial Intelligence
Cryptography and Security
Recent advances in natural language processing and the increased use of large language models have exposed new security vulnerabilities, such as backdoor attacks. Previous backdoor attacks require input manipulation after model distribution to activate the backdoor, posing limitations in real-world applicability. Addressing this gap, we introduce a novel Claim-Guided Backdoor Attack (CGBA), which eliminates the need for such manipulations by utilizing inherent textual claims as triggers. CGBA leverages claim extraction, clustering, and targeted training to trick models to misbehave on targeted claims without affecting their performance on clean data. CGBA demonstrates its effectiveness and stealthiness across various datasets and models, significantly enhancing the feasibility of practical backdoor attacks. Our code and data will be available at https://github.com/PaperCGBA/CGBA.
title Claim-Guided Textual Backdoor Attack for Practical Applications
topic Computation and Language
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2409.16618