IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Aosong, Srinivasan, Balasubramaniam, Zhou, Yun, Xu, Zhichao, Zhou, Kang, Guan, Sheng, Chen, Yueyan, Wu, Xian, Kulkarni, Ninad, Zhang, Yi, Shen, Zhengyuan, Bespalov, Dmitriy, Mishra, Soumya Smruti, Teng, Yifei, Wang, Darren Yow-Bang, Ding, Haibo, Cheong, Lin Lee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915540667203584
author Feng, Aosong
Srinivasan, Balasubramaniam
Zhou, Yun
Xu, Zhichao
Zhou, Kang
Guan, Sheng
Chen, Yueyan
Wu, Xian
Kulkarni, Ninad
Zhang, Yi
Shen, Zhengyuan
Bespalov, Dmitriy
Mishra, Soumya Smruti
Teng, Yifei
Wang, Darren Yow-Bang
Ding, Haibo
Cheong, Lin Lee
author_facet Feng, Aosong
Srinivasan, Balasubramaniam
Zhou, Yun
Xu, Zhichao
Zhou, Kang
Guan, Sheng
Chen, Yueyan
Wu, Xian
Kulkarni, Ninad
Zhang, Yi
Shen, Zhengyuan
Bespalov, Dmitriy
Mishra, Soumya Smruti
Teng, Yifei
Wang, Darren Yow-Bang
Ding, Haibo
Cheong, Lin Lee
contents Routing incoming queries to the most cost-effective LLM while maintaining response quality poses a fundamental challenge in optimizing performance-cost trade-offs for large-scale commercial systems. We present IPR\, -- \,a quality-constrained \textbf{I}ntelligent \textbf{P}rompt \textbf{R}outing framework that dynamically selects optimal models based on predicted response quality and user-specified tolerance levels. IPR introduces three key innovations: (1) a modular architecture with lightweight quality estimators trained on 1.5M prompts annotated with calibrated quality scores, enabling fine-grained quality prediction across model families; (2) a user-controlled routing mechanism with tolerance parameter $τ\in [0,1]$ that provides explicit control over quality-cost trade-offs; and (3) an extensible design using frozen encoders with model-specific adapters, reducing new model integration from days to hours. To rigorously train and evaluate IPR, we curate an industrial-level dataset IPRBench\footnote{IPRBench will be released upon legal approval.}, a comprehensive benchmark containing 1.5 million examples with response quality annotations across 11 LLM candidates. Deployed on a major cloud platform, IPR achieves 43.9\% cost reduction while maintaining quality parity with the strongest model in the Claude family and processes requests with sub-150ms latency. The deployed system and additional product details are publicly available at https://aws.amazon.com/bedrock/intelligent-prompt-routing/
format Preprint
id arxiv_https___arxiv_org_abs_2509_06274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs
Feng, Aosong
Srinivasan, Balasubramaniam
Zhou, Yun
Xu, Zhichao
Zhou, Kang
Guan, Sheng
Chen, Yueyan
Wu, Xian
Kulkarni, Ninad
Zhang, Yi
Shen, Zhengyuan
Bespalov, Dmitriy
Mishra, Soumya Smruti
Teng, Yifei
Wang, Darren Yow-Bang
Ding, Haibo
Cheong, Lin Lee
Machine Learning
Routing incoming queries to the most cost-effective LLM while maintaining response quality poses a fundamental challenge in optimizing performance-cost trade-offs for large-scale commercial systems. We present IPR\, -- \,a quality-constrained \textbf{I}ntelligent \textbf{P}rompt \textbf{R}outing framework that dynamically selects optimal models based on predicted response quality and user-specified tolerance levels. IPR introduces three key innovations: (1) a modular architecture with lightweight quality estimators trained on 1.5M prompts annotated with calibrated quality scores, enabling fine-grained quality prediction across model families; (2) a user-controlled routing mechanism with tolerance parameter $τ\in [0,1]$ that provides explicit control over quality-cost trade-offs; and (3) an extensible design using frozen encoders with model-specific adapters, reducing new model integration from days to hours. To rigorously train and evaluate IPR, we curate an industrial-level dataset IPRBench\footnote{IPRBench will be released upon legal approval.}, a comprehensive benchmark containing 1.5 million examples with response quality annotations across 11 LLM candidates. Deployed on a major cloud platform, IPR achieves 43.9\% cost reduction while maintaining quality parity with the strongest model in the Claude family and processes requests with sub-150ms latency. The deployed system and additional product details are publicly available at https://aws.amazon.com/bedrock/intelligent-prompt-routing/
title IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs
topic Machine Learning
url https://arxiv.org/abs/2509.06274