Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Balepur, Nishant, Desai, Atrey, Rudinger, Rachel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917421546209280
author Balepur, Nishant
Desai, Atrey
Rudinger, Rachel
author_facet Balepur, Nishant
Desai, Atrey
Rudinger, Rachel
contents Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern is that LLMs do not solve MCQs as intended, as work finds LLMs sans reasoning succeed in MCQA without using the question, i.e., choices-only. Such partial-input success is often linked to trivial shortcuts, but reasoning traces could reveal if choices-only strategies are truly shallow. To examine these strategies, we have reasoning LLMs solve MCQs in full and choices-only inputs; test-time reasoning often boosts accuracy in full and in choices-only, half the time. While possibly due to shallow shortcuts, choices-only success is barely affected by the length of reasoning traces, and after finding traces pass faithfulness tests, we show they use less problematic strategies like inferring missing questions. In all, we challenge claims that partial-input success is always a flaw, so we propose how reasoning traces could separate problematic data from less problematic reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
Balepur, Nishant
Desai, Atrey
Rudinger, Rachel
Computation and Language
Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern is that LLMs do not solve MCQs as intended, as work finds LLMs sans reasoning succeed in MCQA without using the question, i.e., choices-only. Such partial-input success is often linked to trivial shortcuts, but reasoning traces could reveal if choices-only strategies are truly shallow. To examine these strategies, we have reasoning LLMs solve MCQs in full and choices-only inputs; test-time reasoning often boosts accuracy in full and in choices-only, half the time. While possibly due to shallow shortcuts, choices-only success is barely affected by the length of reasoning traces, and after finding traces pass faithfulness tests, we show they use less problematic strategies like inferring missing questions. In all, we challenge claims that partial-input success is always a flaw, so we propose how reasoning traces could separate problematic data from less problematic reasoning.
title Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
topic Computation and Language
url https://arxiv.org/abs/2510.07761