REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Buyun, Luo, Jinqi, Peng, Liangzu, Chan, Kwan Ho Ryan, Thaker, Darshan, Kinfu, Kaleab A., Tian, Fengrui, Hassani, Hamed, Vidal, René
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917552001646592
author Liang, Buyun
Luo, Jinqi
Peng, Liangzu
Chan, Kwan Ho Ryan
Thaker, Darshan
Kinfu, Kaleab A.
Tian, Fengrui
Hassani, Hamed
Vidal, René
author_facet Liang, Buyun
Luo, Jinqi
Peng, Liangzu
Chan, Kwan Ho Ryan
Thaker, Darshan
Kinfu, Kaleab A.
Tian, Fengrui
Hassani, Hamed
Vidal, René
contents Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, making it important to systematically evaluate their reliability under realistic adversarial inputs. We formulate hallucination elicitation as a constrained optimization problem, where the goal is to find semantically coherent adversarial prompts that are equivalent to benign user prompts. Existing attack methods remain limited: discrete prompt-based attacks preserve semantic equivalence and coherence but search only over a limited set of prompt variations, while continuous latent-space attacks explore a richer space but often decode into prompts that are no longer valid rephrasings. To address these limitations, we propose REALISTA, a realistic latent-space attack framework. REALISTA constructs an input-dependent dictionary of valid editing directions, each corresponding to a semantically equivalent and coherent rephrasing, and optimizes continuous combinations of these directions in latent space. This design combines the optimization flexibility of continuous attacks with the semantic realism of discrete rephrasing-based attacks. Experiments demonstrate that REALISTA achieves superior or comparable performance to state-of-the-art realistic attacks on open-source LLMs and, crucially, succeeds in attacking large reasoning models under free-form response settings, where prior realistic attacks fail. Code is available at https://github.com/Buyun-Liang/REALISTA.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
Liang, Buyun
Luo, Jinqi
Peng, Liangzu
Chan, Kwan Ho Ryan
Thaker, Darshan
Kinfu, Kaleab A.
Tian, Fengrui
Hassani, Hamed
Vidal, René
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, making it important to systematically evaluate their reliability under realistic adversarial inputs. We formulate hallucination elicitation as a constrained optimization problem, where the goal is to find semantically coherent adversarial prompts that are equivalent to benign user prompts. Existing attack methods remain limited: discrete prompt-based attacks preserve semantic equivalence and coherence but search only over a limited set of prompt variations, while continuous latent-space attacks explore a richer space but often decode into prompts that are no longer valid rephrasings. To address these limitations, we propose REALISTA, a realistic latent-space attack framework. REALISTA constructs an input-dependent dictionary of valid editing directions, each corresponding to a semantically equivalent and coherent rephrasing, and optimizes continuous combinations of these directions in latent space. This design combines the optimization flexibility of continuous attacks with the semantic realism of discrete rephrasing-based attacks. Experiments demonstrate that REALISTA achieves superior or comparable performance to state-of-the-art realistic attacks on open-source LLMs and, crucially, succeeds in attacking large reasoning models under free-form response settings, where prior realistic attacks fail. Code is available at https://github.com/Buyun-Liang/REALISTA.
title REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.12813