Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yiyi, Li, Qiongxiu, Biswas, Russa, Bjerva, Johannes
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917917182918656
author Chen, Yiyi
Li, Qiongxiu
Biswas, Russa
Bjerva, Johannes
author_facet Chen, Yiyi
Li, Qiongxiu
Biswas, Russa
Bjerva, Johannes
contents Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate language. This phenomenon presents a critical challenge in text generation by LLMs, often appearing as erratic and unpredictable behavior. We hypothesize that there are linguistic regularities to this inherent vulnerability in LLMs and shed light on patterns of language confusion across LLMs. We introduce a novel metric, Language Confusion Entropy, designed to directly measure and quantify this confusion, based on language distributions informed by linguistic typology and lexical variation. Comprehensive comparisons with the Language Confusion Benchmark (Marchisio et al., 2024) confirm the effectiveness of our metric, revealing patterns of language confusion across LLMs. We further link language confusion to LLM security, and find patterns in the case of multilingual embedding inversion attacks. Our analysis demonstrates that linguistic typology offers theoretically grounded interpretation, and valuable insights into leveraging language similarities as a prior for LLM alignment and security.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13237
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
Chen, Yiyi
Li, Qiongxiu
Biswas, Russa
Bjerva, Johannes
Computation and Language
Artificial Intelligence
Cryptography and Security
I.1.2; I.1.5
Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate language. This phenomenon presents a critical challenge in text generation by LLMs, often appearing as erratic and unpredictable behavior. We hypothesize that there are linguistic regularities to this inherent vulnerability in LLMs and shed light on patterns of language confusion across LLMs. We introduce a novel metric, Language Confusion Entropy, designed to directly measure and quantify this confusion, based on language distributions informed by linguistic typology and lexical variation. Comprehensive comparisons with the Language Confusion Benchmark (Marchisio et al., 2024) confirm the effectiveness of our metric, revealing patterns of language confusion across LLMs. We further link language confusion to LLM security, and find patterns in the case of multilingual embedding inversion attacks. Our analysis demonstrates that linguistic typology offers theoretically grounded interpretation, and valuable insights into leveraging language similarities as a prior for LLM alignment and security.
title Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
topic Computation and Language
Artificial Intelligence
Cryptography and Security
I.1.2; I.1.5
url https://arxiv.org/abs/2410.13237