Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Na, Injae, Noh, Keonwoong, Jung, Woohwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910973154033664
author Na, Injae
Noh, Keonwoong
Jung, Woohwan
author_facet Na, Injae
Noh, Keonwoong
Jung, Woohwan
contents LLM providers typically offer multiple LLM tiers, varying in performance and price. As NLP tasks become more complex and modularized, selecting the suitable LLM tier for each subtask is a key challenge to balance between cost and performance. To address the problem, we introduce LLM Automatic Transmission (LLM-AT) framework that automatically selects LLM tiers without training. LLM-AT consists of Starter, Generator, and Judge. The starter selects the initial LLM tier expected to solve the given question, the generator produces a response using the LLM of the selected tier, and the judge evaluates the validity of the response. If the response is invalid, LLM-AT iteratively upgrades to a higher-tier model, generates a new response, and re-evaluates until a valid response is obtained. Additionally, we propose accuracy estimator, which enables the suitable initial LLM tier selection without training. Given an input question, accuracy estimator estimates the expected accuracy of each LLM tier by computing the valid response rate across top-k similar queries from past inference records. Experiments demonstrate that LLM-AT achieves superior performance while reducing costs, making it a practical solution for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
Na, Injae
Noh, Keonwoong
Jung, Woohwan
Computation and Language
Artificial Intelligence
LLM providers typically offer multiple LLM tiers, varying in performance and price. As NLP tasks become more complex and modularized, selecting the suitable LLM tier for each subtask is a key challenge to balance between cost and performance. To address the problem, we introduce LLM Automatic Transmission (LLM-AT) framework that automatically selects LLM tiers without training. LLM-AT consists of Starter, Generator, and Judge. The starter selects the initial LLM tier expected to solve the given question, the generator produces a response using the LLM of the selected tier, and the judge evaluates the validity of the response. If the response is invalid, LLM-AT iteratively upgrades to a higher-tier model, generates a new response, and re-evaluates until a valid response is obtained. Additionally, we propose accuracy estimator, which enables the suitable initial LLM tier selection without training. Given an input question, accuracy estimator estimates the expected accuracy of each LLM tier by computing the valid response rate across top-k similar queries from past inference records. Experiments demonstrate that LLM-AT achieves superior performance while reducing costs, making it a practical solution for real-world applications.
title Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.20921