Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Haochun, Yan, Yuliang, Lu, Jiahua, Liu, Huaxiao, Dai, Enyan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913038323417088
author Tang, Haochun
Yan, Yuliang
Lu, Jiahua
Liu, Huaxiao
Dai, Enyan
author_facet Tang, Haochun
Yan, Yuliang
Lu, Jiahua
Liu, Huaxiao
Dai, Enyan
contents Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing strategy introduces a new security concern that adversaries may manipulate the router to consistently select expensive high-capability models. Existing routing attacks depend on either white-box access or heuristic prompts, rendering them ineffective in real-world black-box scenarios. In this work, we propose R$^2$A, which aims to mislead black-box LLM routers to expensive models via adversarial suffix optimization. Specifically, R$^2$A deploys a hybrid ensemble surrogate router to mimic the black-box router. A suffix optimization algorithm is further adapted for the ensemble-based surrogate. Extensive experiments on multiple open-source and commercial routing systems demonstrate that {R$^2$A} significantly increases the routing rate to expensive models on queries of different distributions. Code and examples: https://github.com/thcxiker/R2A-Attack.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15022
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
Tang, Haochun
Yan, Yuliang
Lu, Jiahua
Liu, Huaxiao
Dai, Enyan
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing strategy introduces a new security concern that adversaries may manipulate the router to consistently select expensive high-capability models. Existing routing attacks depend on either white-box access or heuristic prompts, rendering them ineffective in real-world black-box scenarios. In this work, we propose R$^2$A, which aims to mislead black-box LLM routers to expensive models via adversarial suffix optimization. Specifically, R$^2$A deploys a hybrid ensemble surrogate router to mimic the black-box router. A suffix optimization algorithm is further adapted for the ensemble-based surrogate. Extensive experiments on multiple open-source and commercial routing systems demonstrate that {R$^2$A} significantly increases the routing rate to expensive models on queries of different distributions. Code and examples: https://github.com/thcxiker/R2A-Attack.
title Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.15022