On the Security Vulnerabilities of Text-to-SQL Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Xutan, Zhang, Yipeng, Yang, Jingfeng, Stevenson, Mark
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913347083960320
author Peng, Xutan
Zhang, Yipeng
Yang, Jingfeng
Stevenson, Mark
author_facet Peng, Xutan
Zhang, Yipeng
Yang, Jingfeng
Stevenson, Mark
contents Although it has been demonstrated that Natural Language Processing (NLP) algorithms are vulnerable to deliberate attacks, the question of whether such weaknesses can lead to software security threats is under-explored. To bridge this gap, we conducted vulnerability tests on Text-to-SQL systems that are commonly used to create natural language interfaces to databases. We showed that the Text-to-SQL modules within six commercial applications can be manipulated to produce malicious code, potentially leading to data breaches and Denial of Service attacks. This is the first demonstration that NLP models can be exploited as attack vectors in the wild. In addition, experiments using four open-source language models verified that straightforward backdoor attacks on Text-to-SQL systems achieve a 100% success rate without affecting their performance. The aim of this work is to draw the community's attention to potential software security issues associated with NLP algorithms and encourage exploration of methods to mitigate against them.
format Preprint
id arxiv_https___arxiv_org_abs_2211_15363
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle On the Security Vulnerabilities of Text-to-SQL Models
Peng, Xutan
Zhang, Yipeng
Yang, Jingfeng
Stevenson, Mark
Computation and Language
Cryptography and Security
Databases
Machine Learning
Software Engineering
Although it has been demonstrated that Natural Language Processing (NLP) algorithms are vulnerable to deliberate attacks, the question of whether such weaknesses can lead to software security threats is under-explored. To bridge this gap, we conducted vulnerability tests on Text-to-SQL systems that are commonly used to create natural language interfaces to databases. We showed that the Text-to-SQL modules within six commercial applications can be manipulated to produce malicious code, potentially leading to data breaches and Denial of Service attacks. This is the first demonstration that NLP models can be exploited as attack vectors in the wild. In addition, experiments using four open-source language models verified that straightforward backdoor attacks on Text-to-SQL systems achieve a 100% success rate without affecting their performance. The aim of this work is to draw the community's attention to potential software security issues associated with NLP algorithms and encourage exploration of methods to mitigate against them.
title On the Security Vulnerabilities of Text-to-SQL Models
topic Computation and Language
Cryptography and Security
Databases
Machine Learning
Software Engineering
url https://arxiv.org/abs/2211.15363