SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Yuanzhe, Liu, Yide, Huang, Zisu, Yin, Ruicheng, Zheng, Xiaoqing, Huang, Xuanjing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909828512743424
author Shen, Yuanzhe
Liu, Yide
Huang, Zisu
Yin, Ruicheng
Zheng, Xiaoqing
Huang, Xuanjing
author_facet Shen, Yuanzhe
Liu, Yide
Huang, Zisu
Yin, Ruicheng
Zheng, Xiaoqing
Huang, Xuanjing
contents Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness frequently depends on costly commercial APIs or cloud services. Model selection thus entails a critical trade-off between performance and cost: high-performing LLMs typically incur substantial expenses, whereas budget-friendly small language models (SLMs) are constrained by limited capabilities. Current research primarily proposes two routing strategies: pre-generation routing and cascade routing. Both approaches have distinct characteristics, with cascade routing typically offering superior cost-effectiveness and accuracy despite its higher latency. To further address the limitations of both approaches, we introduce SATER, a dual-mode compatible approach that fine-tunes models through shortest-response preference optimization and a confidence-aware rejection mechanism. SATER significantly reduces redundant outputs and response times, while improving both the performance of pre-generation routing and the efficiency of cascade routing. Experiments across three SLMs and six datasets, varying in type and complexity, demonstrate that SATER achieves comparable performance while consistently reducing computational costs by over 50\% and cascade latency by over 80\%.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
Shen, Yuanzhe
Liu, Yide
Huang, Zisu
Yin, Ruicheng
Zheng, Xiaoqing
Huang, Xuanjing
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness frequently depends on costly commercial APIs or cloud services. Model selection thus entails a critical trade-off between performance and cost: high-performing LLMs typically incur substantial expenses, whereas budget-friendly small language models (SLMs) are constrained by limited capabilities. Current research primarily proposes two routing strategies: pre-generation routing and cascade routing. Both approaches have distinct characteristics, with cascade routing typically offering superior cost-effectiveness and accuracy despite its higher latency. To further address the limitations of both approaches, we introduce SATER, a dual-mode compatible approach that fine-tunes models through shortest-response preference optimization and a confidence-aware rejection mechanism. SATER significantly reduces redundant outputs and response times, while improving both the performance of pre-generation routing and the efficiency of cascade routing. Experiments across three SLMs and six datasets, varying in type and complexity, demonstrate that SATER achieves comparable performance while consistently reducing computational costs by over 50\% and cascade latency by over 80\%.
title SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.05164