The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pinhanez, Claudio, Cavalin, Paulo, Sanctos, Cassia, Grave, Marcelo, Primerano, Yago
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915491261448192
author Pinhanez, Claudio
Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Primerano, Yago
author_facet Pinhanez, Claudio
Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Primerano, Yago
contents This work explores the consistency of small LLMs (2B-8B parameters) in answering multiple times the same question. We present a study on known, open-source LLMs responding to 10 repetitions of questions from the multiple-choice benchmarks MMLU-Redux and MedQA, considering different inference temperatures, small vs. medium models (50B-80B), finetuned vs. base models, and other parameters. We also look into the effects of requiring multi-trial answer consistency on accuracy and the trade-offs involved in deciding which model best provides both of them. To support those studies, we propose some new analytical and graphical tools. Results show that the number of questions which can be answered consistently vary considerably among models but are typically in the 50%-80% range for small models at low inference temperatures. Also, accuracy among consistent answers seems to reasonably correlate with overall accuracy. Results for medium-sized models seem to indicate much higher levels of answer consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09705
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
Pinhanez, Claudio
Cavalin, Paulo
Sanctos, Cassia
Grave, Marcelo
Primerano, Yago
Computation and Language
Artificial Intelligence
This work explores the consistency of small LLMs (2B-8B parameters) in answering multiple times the same question. We present a study on known, open-source LLMs responding to 10 repetitions of questions from the multiple-choice benchmarks MMLU-Redux and MedQA, considering different inference temperatures, small vs. medium models (50B-80B), finetuned vs. base models, and other parameters. We also look into the effects of requiring multi-trial answer consistency on accuracy and the trade-offs involved in deciding which model best provides both of them. To support those studies, we propose some new analytical and graphical tools. Results show that the number of questions which can be answered consistently vary considerably among models but are typically in the 50%-80% range for small models at low inference temperatures. Also, accuracy among consistent answers seems to reasonably correlate with overall accuracy. Results for medium-sized models seem to indicate much higher levels of answer consistency.
title The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.09705