Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Jaeseong, Kwon, Dayoung, hwang, seung-won
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909831059734528
author Lee, Jaeseong
Kwon, Dayoung
hwang, seung-won
author_facet Lee, Jaeseong
Kwon, Dayoung
hwang, seung-won
contents Large Reasoning Models (LRMs) excel in structured tasks by emulating deliberate human reasoning but often suffer from overthinking, degrading performance and wasting resources. One possible baseline is to deploy both LLM and LRM, then route input by predicting whether it requires reasoning and may cause overthinking. However, deploying multiple models can be costly or impractical. We propose a superposed deployment strategy with a lightweight, training-free regulation to optimize inference by switching one model on and off. Instead of routing, we selectively unlearn from LRM at inference, scaling down computation while preserving reasoning. By analyzing the cumulative energy of singular values, we identify optimal low-rank projections to adjust reasoning just right.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06750
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs
Lee, Jaeseong
Kwon, Dayoung
hwang, seung-won
Computation and Language
Large Reasoning Models (LRMs) excel in structured tasks by emulating deliberate human reasoning but often suffer from overthinking, degrading performance and wasting resources. One possible baseline is to deploy both LLM and LRM, then route input by predicting whether it requires reasoning and may cause overthinking. However, deploying multiple models can be costly or impractical. We propose a superposed deployment strategy with a lightweight, training-free regulation to optimize inference by switching one model on and off. Instead of routing, we selectively unlearn from LRM at inference, scaling down computation while preserving reasoning. By analyzing the cumulative energy of singular values, we identify optimal low-rank projections to adjust reasoning just right.
title Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs
topic Computation and Language
url https://arxiv.org/abs/2510.06750