Token-Level LLM Collaboration via FusionRoute

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiong, Nuoya, Zhou, Yuhang, Zeng, Hanqing, Chen, Zhaorun, Huang, Furong, Bi, Shuchao, Zhang, Lizhu, Zhao, Zhuokai
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918515302203392
author Xiong, Nuoya
Zhou, Yuhang
Zeng, Hanqing
Chen, Zhaorun
Huang, Furong
Bi, Shuchao
Zhang, Lizhu
Zhao, Zhuokai
author_facet Xiong, Nuoya
Zhou, Yuhang
Zeng, Hanqing
Chen, Zhaorun
Huang, Furong
Bi, Shuchao
Zhang, Lizhu
Zhao, Zhuokai
contents Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a single general-purpose model typically requires scaling to sizes that are prohibitively expensive to train and deploy. On the other hand, while smaller domain-specialized models are much more efficient, they struggle to generalize beyond their training distributions. To address this dilemma, we propose FusionRoute, a robust and effective token-level multi-LLM collaboration framework in which a lightweight router simultaneously (i) selects the most suitable expert at each decoding step and (ii) contributes a complementary logit that refines or corrects the selected expert's next-token distribution via logit addition. Unlike existing token-level collaboration methods that rely solely on fixed expert outputs, we provide a theoretical analysis showing that pure expert-only routing is fundamentally limited: unless strong global coverage assumptions hold, it cannot in general realize the optimal decoding policy. By augmenting expert selection with a trainable complementary generator, FusionRoute expands the effective policy class and enables recovery of optimal value functions under mild conditions. Empirically, across both Llama-3 and Gemma-2 families and diverse benchmarks spanning mathematical reasoning, code generation, and instruction following, FusionRoute outperforms both sequence- and token-level collaboration, model merging, and direct fine-tuning, while remaining competitive with domain experts on their respective tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05106
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Token-Level LLM Collaboration via FusionRoute
Xiong, Nuoya
Zhou, Yuhang
Zeng, Hanqing
Chen, Zhaorun
Huang, Furong
Bi, Shuchao
Zhang, Lizhu
Zhao, Zhuokai
Artificial Intelligence
Computation and Language
Machine Learning
Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a single general-purpose model typically requires scaling to sizes that are prohibitively expensive to train and deploy. On the other hand, while smaller domain-specialized models are much more efficient, they struggle to generalize beyond their training distributions. To address this dilemma, we propose FusionRoute, a robust and effective token-level multi-LLM collaboration framework in which a lightweight router simultaneously (i) selects the most suitable expert at each decoding step and (ii) contributes a complementary logit that refines or corrects the selected expert's next-token distribution via logit addition. Unlike existing token-level collaboration methods that rely solely on fixed expert outputs, we provide a theoretical analysis showing that pure expert-only routing is fundamentally limited: unless strong global coverage assumptions hold, it cannot in general realize the optimal decoding policy. By augmenting expert selection with a trainable complementary generator, FusionRoute expands the effective policy class and enables recovery of optimal value functions under mild conditions. Empirically, across both Llama-3 and Gemma-2 families and diverse benchmarks spanning mathematical reasoning, code generation, and instruction following, FusionRoute outperforms both sequence- and token-level collaboration, model merging, and direct fine-tuning, while remaining competitive with domain experts on their respective tasks.
title Token-Level LLM Collaboration via FusionRoute
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.05106