Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Poulami, Dabre, Raj, Bhattacharyya, Pushpak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929630297980928
author Ghosh, Poulami
Dabre, Raj
Bhattacharyya, Pushpak
author_facet Ghosh, Poulami
Dabre, Raj
Bhattacharyya, Pushpak
contents Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, which are subtle and more prevalent in nature. In this paper, we study whether PLMs are agnostic to linguistically grounded attacks or not. To this end, we offer the first study addressing this, investigating different Indic languages and various downstream tasks. Our findings reveal that although PLMs are susceptible to linguistic perturbations, when compared to non-linguistic attacks, PLMs exhibit a slightly lower susceptibility to linguistic attacks. This highlights that even constrained attacks are effective. Moreover, we investigate the implications of these outcomes across a range of languages, encompassing diverse language families and different scripts.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10805
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Ghosh, Poulami
Dabre, Raj
Bhattacharyya, Pushpak
Computation and Language
Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, which are subtle and more prevalent in nature. In this paper, we study whether PLMs are agnostic to linguistically grounded attacks or not. To this end, we offer the first study addressing this, investigating different Indic languages and various downstream tasks. Our findings reveal that although PLMs are susceptible to linguistic perturbations, when compared to non-linguistic attacks, PLMs exhibit a slightly lower susceptibility to linguistic attacks. This highlights that even constrained attacks are effective. Moreover, we investigate the implications of these outcomes across a range of languages, encompassing diverse language families and different scripts.
title Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
topic Computation and Language
url https://arxiv.org/abs/2412.10805