TARo: Token-level Adaptive Routing for LLM Test-time Alignment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rai, Arushi, Zhang, Qiang, Zeng, Hanqing, Zhang, Yunkai, Tamboli, Dipesh, Fan, Xiangjun, Zhao, Zhuokai, Zhang, Lizhu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908901242306560
author Rai, Arushi
Zhang, Qiang
Zeng, Hanqing
Zhang, Yunkai
Tamboli, Dipesh
Fan, Xiangjun
Zhao, Zhuokai
Zhang, Lizhu
author_facet Rai, Arushi
Zhang, Qiang
Zeng, Hanqing
Zhang, Yunkai
Tamboli, Dipesh
Fan, Xiangjun
Zhao, Zhuokai
Zhang, Lizhu
contents Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. Recent test-time alignment methods offer a lightweight alternative, but have been explored mainly for preference alignment rather than reasoning. To bridge this gap, we propose, Token-level Adaptive Routing (TARo), which steers frozen LLMs toward structured reasoning entirely at inference time. Specifically, we first train reward models on step-wise mathematical traces to capture fine-grained logical consistency signals, then introduce a learnable token-level router that automatically controls the guidance of the reward model to the base model. Extensive experiments show that TARo significantly improves reasoning performance by up to +22.4% over base model and +8.4% over existing token-level test-time alignment methods, while also boosting out-of-distribution clinical reasoning (MedXpertQA) and instruction following (AlpacaEval). Furthermore, TARo also generalizes from small to large backbones without retraining, extending test-time alignment from preference optimization to robust, cross-domain reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18411
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TARo: Token-level Adaptive Routing for LLM Test-time Alignment
Rai, Arushi
Zhang, Qiang
Zeng, Hanqing
Zhang, Yunkai
Tamboli, Dipesh
Fan, Xiangjun
Zhao, Zhuokai
Zhang, Lizhu
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. Recent test-time alignment methods offer a lightweight alternative, but have been explored mainly for preference alignment rather than reasoning. To bridge this gap, we propose, Token-level Adaptive Routing (TARo), which steers frozen LLMs toward structured reasoning entirely at inference time. Specifically, we first train reward models on step-wise mathematical traces to capture fine-grained logical consistency signals, then introduce a learnable token-level router that automatically controls the guidance of the reward model to the base model. Extensive experiments show that TARo significantly improves reasoning performance by up to +22.4% over base model and +8.4% over existing token-level test-time alignment methods, while also boosting out-of-distribution clinical reasoning (MedXpertQA) and instruction following (AlpacaEval). Furthermore, TARo also generalizes from small to large backbones without retraining, extending test-time alignment from preference optimization to robust, cross-domain reasoning.
title TARo: Token-level Adaptive Routing for LLM Test-time Alignment
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.18411