TensorOpera Router: A Multi-Model Router for Efficient LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stripelis, Dimitris, Hu, Zijian, Zhang, Jipeng, Xu, Zhaozhuo, Shah, Alay Dilipbhai, Jin, Han, Yao, Yuhang, Avestimehr, Salman, He, Chaoyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914986608033792
author Stripelis, Dimitris
Hu, Zijian
Zhang, Jipeng
Xu, Zhaozhuo
Shah, Alay Dilipbhai
Jin, Han
Yao, Yuhang
Avestimehr, Salman
He, Chaoyang
author_facet Stripelis, Dimitris
Hu, Zijian
Zhang, Jipeng
Xu, Zhaozhuo
Shah, Alay Dilipbhai
Jin, Han
Yao, Yuhang
Avestimehr, Salman
He, Chaoyang
contents With the rapid growth of Large Language Models (LLMs) across various domains, numerous new LLMs have emerged, each possessing domain-specific expertise. This proliferation has highlighted the need for quick, high-quality, and cost-effective LLM query response methods. Yet, no single LLM exists to efficiently balance this trilemma. Some models are powerful but extremely costly, while others are fast and inexpensive but qualitatively inferior. To address this challenge, we present TO-Router, a non-monolithic LLM querying system that seamlessly integrates various LLM experts into a single query interface and dynamically routes incoming queries to the most high-performant expert based on query's requirements. Through extensive experiments, we demonstrate that when compared to standalone expert models, TO-Router improves query efficiency by up to 40\%, and leads to significant cost reductions of up to 30%, while maintaining or enhancing model performance by up to 10%.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12320
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
Stripelis, Dimitris
Hu, Zijian
Zhang, Jipeng
Xu, Zhaozhuo
Shah, Alay Dilipbhai
Jin, Han
Yao, Yuhang
Avestimehr, Salman
He, Chaoyang
Artificial Intelligence
Machine Learning
I.2; I.5
With the rapid growth of Large Language Models (LLMs) across various domains, numerous new LLMs have emerged, each possessing domain-specific expertise. This proliferation has highlighted the need for quick, high-quality, and cost-effective LLM query response methods. Yet, no single LLM exists to efficiently balance this trilemma. Some models are powerful but extremely costly, while others are fast and inexpensive but qualitatively inferior. To address this challenge, we present TO-Router, a non-monolithic LLM querying system that seamlessly integrates various LLM experts into a single query interface and dynamically routes incoming queries to the most high-performant expert based on query's requirements. Through extensive experiments, we demonstrate that when compared to standalone expert models, TO-Router improves query efficiency by up to 40\%, and leads to significant cost reductions of up to 30%, while maintaining or enhancing model performance by up to 10%.
title TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
topic Artificial Intelligence
Machine Learning
I.2; I.5
url https://arxiv.org/abs/2408.12320