Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jaemin, Chang, Hangeol, Hwang, Hyunmin, Kim, Choonghan, Ye, Jong Chul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914580909785088
author Kim, Jaemin
Chang, Hangeol
Hwang, Hyunmin
Kim, Choonghan
Ye, Jong Chul
author_facet Kim, Jaemin
Chang, Hangeol
Hwang, Hyunmin
Kim, Choonghan
Ye, Jong Chul
contents Large Language Models (LLMs) have demonstrated remarkable general capabilities, but enhancing skills such as reasoning often demands substantial computational resources and may compromise generalization. While Parameter-Efficient Fine-Tuning (PEFT) methods offer a more resource-conscious alternative, they typically require retraining for each LLM backbone due to architectural dependencies. To address these challenges, we propose Universal Reasoner (UniR)-a modular, composable, and plug-and-play reasoning module that can be used with larger frozen LLMs to provide specialized reasoning capabilities with a shared or aligned token space. Specifically, UniR decomposes the reward into a standalone reasoning module trained in a decoupled manner using verifiable rewards, effectively translating trajectory-level signals into token-level guidance. Once trained, UniR is combined with frozen LLMs at inference time by simply adding its output logits to those of the backbone. This additive structure enables modular composition: multiple UniR modules trained for different tasks can be jointly applied by summing their logits, enabling complex reasoning via composition. Furthermore, UniR demonstrates weak-to-strong generalization, where reasoning modules trained on smaller models effectively guide much larger LLMs in the same model family, and generalize across domains such as in vision language models and medical reasoning. Experiments on mathematical reasoning and machine translation show that UniR surpasses existing fine-tuning methods. Code is open-sourced at https://github.com/hangeol/UniR.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
Kim, Jaemin
Chang, Hangeol
Hwang, Hyunmin
Kim, Choonghan
Ye, Jong Chul
Artificial Intelligence
Computation and Language
Machine Learning
Large Language Models (LLMs) have demonstrated remarkable general capabilities, but enhancing skills such as reasoning often demands substantial computational resources and may compromise generalization. While Parameter-Efficient Fine-Tuning (PEFT) methods offer a more resource-conscious alternative, they typically require retraining for each LLM backbone due to architectural dependencies. To address these challenges, we propose Universal Reasoner (UniR)-a modular, composable, and plug-and-play reasoning module that can be used with larger frozen LLMs to provide specialized reasoning capabilities with a shared or aligned token space. Specifically, UniR decomposes the reward into a standalone reasoning module trained in a decoupled manner using verifiable rewards, effectively translating trajectory-level signals into token-level guidance. Once trained, UniR is combined with frozen LLMs at inference time by simply adding its output logits to those of the backbone. This additive structure enables modular composition: multiple UniR modules trained for different tasks can be jointly applied by summing their logits, enabling complex reasoning via composition. Furthermore, UniR demonstrates weak-to-strong generalization, where reasoning modules trained on smaller models effectively guide much larger LLMs in the same model family, and generalize across domains such as in vision language models and medical reasoning. Experiments on mathematical reasoning and machine translation show that UniR surpasses existing fine-tuning methods. Code is open-sourced at https://github.com/hangeol/UniR.
title Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2505.19075