RelayLLM: Efficient Reasoning via Collaborative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chengsong, Zheng, Tong, Huang, Langlin, Li, Jinyuan, Liu, Haolin, Huang, Jiaxin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911361541341184
author Huang, Chengsong
Zheng, Tong
Huang, Langlin
Li, Jinyuan
Liu, Haolin
Huang, Jiaxin
author_facet Huang, Chengsong
Zheng, Tong
Huang, Langlin
Li, Jinyuan
Liu, Haolin
Huang, Jiaxin
contents Large Language Models (LLMs) for complex reasoning is often hindered by high computational costs and latency, while resource-efficient Small Language Models (SLMs) typically lack the necessary reasoning capacity. Existing collaborative approaches, such as cascading or routing, operate at a coarse granularity by offloading entire queries to LLMs, resulting in significant computational waste when the SLM is capable of handling the majority of reasoning steps. To address this, we propose RelayLLM, a novel framework for efficient reasoning via token-level collaborative decoding. Unlike routers, RelayLLM empowers the SLM to act as an active controller that dynamically invokes the LLM only for critical tokens via a special command, effectively "relaying" the generation process. We introduce a two-stage training framework, including warm-up and Group Relative Policy Optimization (GRPO) to teach the model to balance independence with strategic help-seeking. Empirical results across six benchmarks demonstrate that RelayLLM achieves an average accuracy of 49.52%, effectively bridging the performance gap between the two models. Notably, this is achieved by invoking the LLM for only 1.07% of the total generated tokens, offering a 98.2% cost reduction compared to performance-matched random routers.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05167
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RelayLLM: Efficient Reasoning via Collaborative Decoding
Huang, Chengsong
Zheng, Tong
Huang, Langlin
Li, Jinyuan
Liu, Haolin
Huang, Jiaxin
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) for complex reasoning is often hindered by high computational costs and latency, while resource-efficient Small Language Models (SLMs) typically lack the necessary reasoning capacity. Existing collaborative approaches, such as cascading or routing, operate at a coarse granularity by offloading entire queries to LLMs, resulting in significant computational waste when the SLM is capable of handling the majority of reasoning steps. To address this, we propose RelayLLM, a novel framework for efficient reasoning via token-level collaborative decoding. Unlike routers, RelayLLM empowers the SLM to act as an active controller that dynamically invokes the LLM only for critical tokens via a special command, effectively "relaying" the generation process. We introduce a two-stage training framework, including warm-up and Group Relative Policy Optimization (GRPO) to teach the model to balance independence with strategic help-seeking. Empirical results across six benchmarks demonstrate that RelayLLM achieves an average accuracy of 49.52%, effectively bridging the performance gap between the two models. Notably, this is achieved by invoking the LLM for only 1.07% of the total generated tokens, offering a 98.2% cost reduction compared to performance-matched random routers.
title RelayLLM: Efficient Reasoning via Collaborative Decoding
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.05167