Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Qiming, Tang, Jinwen, Huang, Xingran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909776499179520
author Guo, Qiming
Tang, Jinwen
Huang, Xingran
author_facet Guo, Qiming
Tang, Jinwen
Huang, Xingran
contents We introduce Advertisement Embedding Attacks (AEA), a new class of LLM security threats that stealthily inject promotional or malicious content into model outputs and AI agents. AEA operate through two low-cost vectors: (1) hijacking third-party service-distribution platforms to prepend adversarial prompts, and (2) publishing back-doored open-source checkpoints fine-tuned with attacker data. Unlike conventional attacks that degrade accuracy, AEA subvert information integrity, causing models to return covert ads, propaganda, or hate speech while appearing normal. We detail the attack pipeline, map five stakeholder victim groups, and present an initial prompt-based self-inspection defense that mitigates these injections without additional model retraining. Our findings reveal an urgent, under-addressed gap in LLM security and call for coordinated detection, auditing, and policy responses from the AI-safety community.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
Guo, Qiming
Tang, Jinwen
Huang, Xingran
Cryptography and Security
Artificial Intelligence
Machine Learning
We introduce Advertisement Embedding Attacks (AEA), a new class of LLM security threats that stealthily inject promotional or malicious content into model outputs and AI agents. AEA operate through two low-cost vectors: (1) hijacking third-party service-distribution platforms to prepend adversarial prompts, and (2) publishing back-doored open-source checkpoints fine-tuned with attacker data. Unlike conventional attacks that degrade accuracy, AEA subvert information integrity, causing models to return covert ads, propaganda, or hate speech while appearing normal. We detail the attack pipeline, map five stakeholder victim groups, and present an initial prompt-based self-inspection defense that mitigates these injections without additional model retraining. Our findings reveal an urgent, under-addressed gap in LLM security and call for coordinated detection, auditing, and policy responses from the AI-safety community.
title Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.17674