LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xingyu, Liu, Xiaolei, Liu, Cheng, Xu, Yixiao, Ding, Kangyi, Xin, Bangzhou, Yin, Jia-Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917072957603840
author Li, Xingyu
Liu, Xiaolei
Liu, Cheng
Xu, Yixiao
Ding, Kangyi
Xin, Bangzhou
Yin, Jia-Li
author_facet Li, Xingyu
Liu, Xiaolei
Liu, Cheng
Xu, Yixiao
Ding, Kangyi
Xin, Bangzhou
Yin, Jia-Li
contents As large language models (LLMs) scale, their inference incurs substantial computational resources, exposing them to energy-latency attacks, where crafted prompts induce high energy and latency cost. Existing attack methods aim to prolong output by delaying the generation of termination symbols. However, as the output grows longer, controlling the termination symbols through input becomes difficult, making these methods less effective. Therefore, we propose LoopLLM, an energy-latency attack framework based on the observation that repetitive generation can trigger low-entropy decoding loops, reliably compelling LLMs to generate until their output limits. LoopLLM introduces (1) a repetition-inducing prompt optimization that exploits autoregressive vulnerabilities to induce repetitive generation, and (2) a token-aligned ensemble optimization that aggregates gradients to improve cross-model transferability. Extensive experiments on 12 open-source and 2 commercial LLMs show that LoopLLM significantly outperforms existing methods, achieving over 90% of the maximum output length, compared to 20% for baselines, and improving transferability by around 40% to DeepSeek-V3 and Gemini 2.5 Flash.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
Li, Xingyu
Liu, Xiaolei
Liu, Cheng
Xu, Yixiao
Ding, Kangyi
Xin, Bangzhou
Yin, Jia-Li
Cryptography and Security
Artificial Intelligence
Computation and Language
As large language models (LLMs) scale, their inference incurs substantial computational resources, exposing them to energy-latency attacks, where crafted prompts induce high energy and latency cost. Existing attack methods aim to prolong output by delaying the generation of termination symbols. However, as the output grows longer, controlling the termination symbols through input becomes difficult, making these methods less effective. Therefore, we propose LoopLLM, an energy-latency attack framework based on the observation that repetitive generation can trigger low-entropy decoding loops, reliably compelling LLMs to generate until their output limits. LoopLLM introduces (1) a repetition-inducing prompt optimization that exploits autoregressive vulnerabilities to induce repetitive generation, and (2) a token-aligned ensemble optimization that aggregates gradients to improve cross-model transferability. Extensive experiments on 12 open-source and 2 commercial LLMs show that LoopLLM significantly outperforms existing methods, achieving over 90% of the maximum output length, compared to 20% for baselines, and improving transferability by around 40% to DeepSeek-V3 and Gemini 2.5 Flash.
title LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.07876