Saved in:
Bibliographic Details
Main Authors: Wu, Tsung-Han, Miroyan, Mihran, Chan, David M., Darrell, Trevor, Norouzi, Narges, Gonzalez, Joseph E.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.11713
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911740249243648
author Wu, Tsung-Han
Miroyan, Mihran
Chan, David M.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
author_facet Wu, Tsung-Han
Miroyan, Mihran
Chan, David M.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
contents Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge the frozen world assumption and evaluate LRM robustness under two realistic dynamic scenarios: interruptions, which test the accuracy of model responses under budget-constrained outputs, and dynamic context, which tests model adaptation to in-flight changes. Across mathematics and programming benchmarks that require long-form reasoning, static evaluations consistently overestimate robustness: even state-of-the-art LRMs, which achieve high accuracy in static settings, can fail unpredictably when interrupted or exposed to changing context, with performance dropping by up to 60% when updates are introduced late in the reasoning process. Our analysis further reveals several novel failure modes, including reasoning leakage, where models fold the reasoning into their final answer when interrupted; panic, where under time pressure models abandon reasoning entirely and return incorrect answers; and self-doubt, where performance degrades when trying to incorporate updated information. Project Page: http://dynamic-lm.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2510_11713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are Large Reasoning Models Interruptible?
Wu, Tsung-Han
Miroyan, Mihran
Chan, David M.
Darrell, Trevor
Norouzi, Narges
Gonzalez, Joseph E.
Computation and Language
Machine Learning
Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge the frozen world assumption and evaluate LRM robustness under two realistic dynamic scenarios: interruptions, which test the accuracy of model responses under budget-constrained outputs, and dynamic context, which tests model adaptation to in-flight changes. Across mathematics and programming benchmarks that require long-form reasoning, static evaluations consistently overestimate robustness: even state-of-the-art LRMs, which achieve high accuracy in static settings, can fail unpredictably when interrupted or exposed to changing context, with performance dropping by up to 60% when updates are introduced late in the reasoning process. Our analysis further reveals several novel failure modes, including reasoning leakage, where models fold the reasoning into their final answer when interrupted; panic, where under time pressure models abandon reasoning entirely and return incorrect answers; and self-doubt, where performance degrades when trying to incorporate updated information. Project Page: http://dynamic-lm.github.io/
title Are Large Reasoning Models Interruptible?
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.11713