MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yiqun, Li, Hao, Wang, Zihan, Feng, Shi, Yang, Xiaocui, Wang, Daling, Zhang, Bo, Bai, Lei, Hu, Shuyue
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910166649143296
author Zhang, Yiqun
Li, Hao
Wang, Zihan
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Bo
Bai, Lei
Hu, Shuyue
author_facet Zhang, Yiqun
Li, Hao
Wang, Zihan
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Bo
Bai, Lei
Hu, Shuyue
contents Multi-turn, long-horizon tasks are increasingly common for large language models (LLMs), but solving them typically requires many sequential model invocations, accumulating substantial inference costs. Here, we study cost-aware multi-turn LLM routing: selecting which model to invoke at each turn from a model pool, given a fixed cost budget. We propose MTRouter, which encodes the interaction history and candidate models into joint history-model embeddings, and learns an outcome estimator from logged trajectories to predict turn-level model utility. Experiments show that MTRouter improves the performance-cost trade-off: on ScienceWorld, it surpasses GPT-5 while reducing total cost by 58.7%; on Humanity's Last Exam (HLE), it achieves competitive accuracy while reducing total cost by 43.4% relative to GPT-5, and these gains even carry over to held-out tasks. Further analyses reveal several mechanisms underlying its effectiveness: relative to prior multi-turn routers, MTRouter makes fewer model switches, is more tolerant to transient errors, and exhibits emergent specialization across models. Code: https://github.com/ZhangYiqun018/MTRouter
format Preprint
id arxiv_https___arxiv_org_abs_2604_23530
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
Zhang, Yiqun
Li, Hao
Wang, Zihan
Feng, Shi
Yang, Xiaocui
Wang, Daling
Zhang, Bo
Bai, Lei
Hu, Shuyue
Computation and Language
Artificial Intelligence
Multi-turn, long-horizon tasks are increasingly common for large language models (LLMs), but solving them typically requires many sequential model invocations, accumulating substantial inference costs. Here, we study cost-aware multi-turn LLM routing: selecting which model to invoke at each turn from a model pool, given a fixed cost budget. We propose MTRouter, which encodes the interaction history and candidate models into joint history-model embeddings, and learns an outcome estimator from logged trajectories to predict turn-level model utility. Experiments show that MTRouter improves the performance-cost trade-off: on ScienceWorld, it surpasses GPT-5 while reducing total cost by 58.7%; on Humanity's Last Exam (HLE), it achieves competitive accuracy while reducing total cost by 43.4% relative to GPT-5, and these gains even carry over to held-out tasks. Further analyses reveal several mechanisms underlying its effectiveness: relative to prior multi-turn routers, MTRouter makes fewer model switches, is more tolerant to transient errors, and exhibits emergent specialization across models. Code: https://github.com/ZhangYiqun018/MTRouter
title MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.23530