REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Buyun, Luo, Jinqi, Peng, Liangzu, Chan, Kwan Ho Ryan, Thaker, Darshan, Kinfu, Kaleab A., Tian, Fengrui, Hassani, Hamed, Vidal, René |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
by: Kinfu, Kaleab A., et al.
Published: (2025)
by: Kinfu, Kaleab A., et al.
Published: (2025)
PaCE: Parsimonious Concept Engineering for Large Language Models
by: Luo, Jinqi, et al.
Published: (2024)
by: Luo, Jinqi, et al.
Published: (2024)
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
Contextual Knowledge Pursuit for Faithful Visual Synthesis
by: Luo, Jinqi, et al.
Published: (2023)
by: Luo, Jinqi, et al.
Published: (2023)
Frequency-Guided Posterior Sampling for Diffusion-Based Image Restoration
by: Thaker, Darshan, et al.
Published: (2024)
by: Thaker, Darshan, et al.
Published: (2024)
Mathematics of Continual Learning
by: Peng, Liangzu, et al.
Published: (2025)
by: Peng, Liangzu, et al.
Published: (2025)
Voyaging into Perpetual Dynamic Scenes from a Single View
by: Tian, Fengrui, et al.
Published: (2025)
by: Tian, Fengrui, et al.
Published: (2025)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
by: Delgosha, Payam, et al.
Published: (2024)
by: Delgosha, Payam, et al.
Published: (2024)
Computerized Assessment of Motor Imitation for Distinguishing Autism in Video (CAMI-2DNet)
by: Kinfu, Kaleab A., et al.
Published: (2025)
by: Kinfu, Kaleab A., et al.
Published: (2025)
IP-CRR: Information Pursuit for Interpretable Classification of Chest Radiology Reports
by: Ge, Yuyan, et al.
Published: (2025)
by: Ge, Yuyan, et al.
Published: (2025)
Scalable 3D Registration via Truncated Entry-wise Absolute Residuals
by: Huang, Tianyu, et al.
Published: (2024)
by: Huang, Tianyu, et al.
Published: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
LoRanPAC: Low-rank Random Features and Pre-trained Models for Bridging Theory and Practice in Continual Learning
by: Peng, Liangzu, et al.
Published: (2024)
by: Peng, Liangzu, et al.
Published: (2024)
Watermark Smoothing Attacks against Language Models
by: Chang, Hongyan, et al.
Published: (2024)
by: Chang, Hongyan, et al.
Published: (2024)
Learning Interpretable Queries for Explainable Image Classification with Information Pursuit
by: Kolek, Stefan, et al.
Published: (2023)
by: Kolek, Stefan, et al.
Published: (2023)
Adversarial Attacks on Robotic Vision Language Action Models
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
by: Jones, Eliot Krzysztof, et al.
Published: (2025)
Recovery Guarantees for Continual Learning of Dependent Tasks: Memory, Data-Dependent Regularization, and Data-Dependent Weights
by: Peng, Liangzu, et al.
Published: (2026)
by: Peng, Liangzu, et al.
Published: (2026)
Scale-Cascaded Diffusion Models for Super-Resolution in Medical Imaging
by: Thaker, Darshan, et al.
Published: (2026)
by: Thaker, Darshan, et al.
Published: (2026)
Adversarial Elicitation
by: Iakovlev, Andrei
Published: (2026)
by: Iakovlev, Andrei
Published: (2026)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
by: More, Yash, et al.
Published: (2024)
by: More, Yash, et al.
Published: (2024)
Block Acceleration Without Momentum: On Optimal Stepsizes of Block Gradient Descent for Least-Squares
by: Peng, Liangzu, et al.
Published: (2024)
by: Peng, Liangzu, et al.
Published: (2024)
Steer LLM Latents for Hallucination Detection
by: Park, Seongheon, et al.
Published: (2025)
by: Park, Seongheon, et al.
Published: (2025)
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026)
by: Moniri, Behrad, et al.
Published: (2026)
The curse of overparametrization in adversarial training: Precise analysis of robust generalization for random features regression
by: Hassani, Hamed, et al.
Published: (2022)
by: Hassani, Hamed, et al.
Published: (2022)
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
by: Moniri, Behrad, et al.
Published: (2025)
by: Moniri, Behrad, et al.
Published: (2025)
Adversarial Reasoning at Jailbreaking Time
by: Sabbaghi, Mahdi, et al.
Published: (2025)
by: Sabbaghi, Mahdi, et al.
Published: (2025)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Can Implicit Bias Imply Adversarial Robustness?
by: Min, Hancheng, et al.
Published: (2024)
by: Min, Hancheng, et al.
Published: (2024)
A Survey on Large Language Model Hallucination via a Creativity Perspective
by: Jiang, Xuhui, et al.
Published: (2024)
by: Jiang, Xuhui, et al.
Published: (2024)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
by: Kulkarni, Prashant
Published: (2026)
by: Kulkarni, Prashant
Published: (2026)
Adversarial Examples Might be Avoidable: The Role of Data Concentration in Adversarial Robustness
by: Pal, Ambar, et al.
Published: (2023)
by: Pal, Ambar, et al.
Published: (2023)
Fool the Stoplight: Realistic Adversarial Patch Attacks on Traffic Light Detectors
by: Pavlitska, Svetlana, et al.
Published: (2025)
by: Pavlitska, Svetlana, et al.
Published: (2025)
Elicitive Curricular Development
by: Echavarría Alvarez, Josefina, et al.
Published: (2020)
by: Echavarría Alvarez, Josefina, et al.
Published: (2020)
Membership Inference Attacks for Unseen Classes
by: Thaker, Pratiksha, et al.
Published: (2025)
by: Thaker, Pratiksha, et al.
Published: (2025)
Similar Items
-
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2025) -
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025) -
Transformers with Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
by: Kinfu, Kaleab A., et al.
Published: (2025) -
PaCE: Parsimonious Concept Engineering for Large Language Models
by: Luo, Jinqi, et al.
Published: (2024) -
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)