Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Troshin, Sergey, Mohammed, Wafaa, Meng, Yan, Monz, Christof, Fokkens, Antske, Niculae, Vlad
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912622225391616
author Troshin, Sergey
Mohammed, Wafaa
Meng, Yan
Monz, Christof
Fokkens, Antske
Niculae, Vlad
author_facet Troshin, Sergey
Mohammed, Wafaa
Meng, Yan
Monz, Christof
Fokkens, Antske
Niculae, Vlad
contents Diversity is an essential metric for evaluating the creativity of outputs generated by language models. Temperature-based sampling is a common strategy to increase diversity. However, for tasks that require high precision, e.g., mathematical reasoning, uncontrolled high temperature sampling, e.g., min-$p$ or top-$p$, degrades reasoning quality. We demonstrate that the loss of accuracy is caused by sampling incorrect continuations in sensitive decoding positions. To address this, in this paper, we propose \textbf{selective sampling}, a method that dynamically switches between greedy and high-temperature sampling based on a sampling risk metric. This risk metric estimates the likelihood of output errors when applying high-temperature sampling on the current token position. To predict sampling risk, we train a lightweight classifier on a small subset of verifiable problems. The trained classifier can be integrated with the base language model with minimal latency overhead. Experiments on mathematical reasoning tasks demonstrate that selective sampling enhances the quality-diversity trade-off, even in high-temperature settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
Troshin, Sergey
Mohammed, Wafaa
Meng, Yan
Monz, Christof
Fokkens, Antske
Niculae, Vlad
Machine Learning
Artificial Intelligence
Computation and Language
Diversity is an essential metric for evaluating the creativity of outputs generated by language models. Temperature-based sampling is a common strategy to increase diversity. However, for tasks that require high precision, e.g., mathematical reasoning, uncontrolled high temperature sampling, e.g., min-$p$ or top-$p$, degrades reasoning quality. We demonstrate that the loss of accuracy is caused by sampling incorrect continuations in sensitive decoding positions. To address this, in this paper, we propose \textbf{selective sampling}, a method that dynamically switches between greedy and high-temperature sampling based on a sampling risk metric. This risk metric estimates the likelihood of output errors when applying high-temperature sampling on the current token position. To predict sampling risk, we train a lightweight classifier on a small subset of verifiable problems. The trained classifier can be integrated with the base language model with minimal latency overhead. Experiments on mathematical reasoning tasks demonstrate that selective sampling enhances the quality-diversity trade-off, even in high-temperature settings.
title Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.01218