Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chrabąszcz, Maciej, Lorenc, Katarzyna, Seweryn, Karolina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912419581788160
author Chrabąszcz, Maciej
Lorenc, Katarzyna
Seweryn, Karolina
author_facet Chrabąszcz, Maciej
Lorenc, Katarzyna
Seweryn, Karolina
contents Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jailbreaks and perturbations necessitates additional evaluations. Many LLMs are multilingual, but safety-related training data contains mainly high-resource languages like English. This can leave them vulnerable to perturbations in low-resource languages such as Polish. We show how surprisingly strong attacks can be cheaply created by altering just a few characters and using a small proxy model for word importance calculation. We find that these character and word-level attacks drastically alter the predictions of different LLMs, suggesting a potential vulnerability that can be used to circumvent their internal safety mechanisms. We validate our attack construction methodology on Polish, a low-resource language, and find potential vulnerabilities of LLMs in this language. Additionally, we show how it can be extended to other languages. We release the created datasets and code for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
Chrabąszcz, Maciej
Lorenc, Katarzyna
Seweryn, Karolina
Computation and Language
Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jailbreaks and perturbations necessitates additional evaluations. Many LLMs are multilingual, but safety-related training data contains mainly high-resource languages like English. This can leave them vulnerable to perturbations in low-resource languages such as Polish. We show how surprisingly strong attacks can be cheaply created by altering just a few characters and using a small proxy model for word importance calculation. We find that these character and word-level attacks drastically alter the predictions of different LLMs, suggesting a potential vulnerability that can be used to circumvent their internal safety mechanisms. We validate our attack construction methodology on Polish, a low-resource language, and find potential vulnerabilities of LLMs in this language. Additionally, we show how it can be extended to other languages. We release the created datasets and code for further research.
title Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
topic Computation and Language
url https://arxiv.org/abs/2506.07645