Optimizing Reasoning Efficiency through Prompt Difficulty Prediction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Bo, Kapusuzoglu, Berkcan, Balasubramaniam, Kartik, Sahu, Sambit, Chakraborty, Supriyo, Winata, Genta Indra
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918188038488064
author Zhao, Bo
Kapusuzoglu, Berkcan
Balasubramaniam, Kartik
Sahu, Sambit
Chakraborty, Supriyo
Winata, Genta Indra
author_facet Zhao, Bo
Kapusuzoglu, Berkcan
Balasubramaniam, Kartik
Sahu, Sambit
Chakraborty, Supriyo
Winata, Genta Indra
contents Reasoning language models perform well on complex tasks but are costly to deploy due to their size and long reasoning traces. We propose a routing approach that assigns each problem to the smallest model likely to solve it, reducing compute without sacrificing accuracy. Using intermediate representations from s1.1-32B, we train lightweight predictors of problem difficulty or model correctness to guide routing across a pool of reasoning models. On diverse math benchmarks, routing improves efficiency over random assignment and matches s1.1-32B's performance while using significantly less compute. Our results demonstrate that difficulty-aware routing is effective for cost-efficient deployment of reasoning models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_03808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
Zhao, Bo
Kapusuzoglu, Berkcan
Balasubramaniam, Kartik
Sahu, Sambit
Chakraborty, Supriyo
Winata, Genta Indra
Machine Learning
Artificial Intelligence
Reasoning language models perform well on complex tasks but are costly to deploy due to their size and long reasoning traces. We propose a routing approach that assigns each problem to the smallest model likely to solve it, reducing compute without sacrificing accuracy. Using intermediate representations from s1.1-32B, we train lightweight predictors of problem difficulty or model correctness to guide routing across a pool of reasoning models. On diverse math benchmarks, routing improves efficiency over random assignment and matches s1.1-32B's performance while using significantly less compute. Our results demonstrate that difficulty-aware routing is effective for cost-efficient deployment of reasoning models.
title Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.03808