Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Haotian, Xu, Bo, Shu, Yao, Yang, Menglin, Qin, Chengwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918159055847424
author Wu, Haotian
Xu, Bo
Shu, Yao
Yang, Menglin
Qin, Chengwei
author_facet Wu, Haotian
Xu, Bo
Shu, Yao
Yang, Menglin
Qin, Chengwei
contents Reasoning large language models (RLLMs) have recently demonstrated remarkable capabilities through structured and multi-step reasoning. While prior research has primarily focused on improving their training and inference strategies, their potential for in-context learning (ICL) remains largely underexplored. To fill this gap, we propose Thinking with Nothinking Calibration (JointThinking), a new ICL paradigm that prompts the model to generate two answers in parallel: one in Thinking mode and the other in Nothinking mode. A second round of Thinking is triggered only when the two initial responses are inconsistent, using a single prompt with two different answers. Extensive experiments across multiple reasoning benchmarks demonstrate that JointThinking significantly outperforms few-shot chain-of-thought (CoT), thinking twice and majority voting. Moreover, it achieves comparable in-distribution performance to training-based SOTA reasoning method, while substantially outperforming on out-of-distribution tasks. We further conduct a systematic analysis of the calibration mechanism, showing the importance of structural thinking diversity and the benefits of consistency check. Additionally, we observe that the performance gap between actual and ideal reasoning narrows as model size increases in the second thinking, indicating the strong scalability of our approach. Finally, we discuss current limitations and outline promising directions for future ICL research in RLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
Wu, Haotian
Xu, Bo
Shu, Yao
Yang, Menglin
Qin, Chengwei
Computation and Language
Reasoning large language models (RLLMs) have recently demonstrated remarkable capabilities through structured and multi-step reasoning. While prior research has primarily focused on improving their training and inference strategies, their potential for in-context learning (ICL) remains largely underexplored. To fill this gap, we propose Thinking with Nothinking Calibration (JointThinking), a new ICL paradigm that prompts the model to generate two answers in parallel: one in Thinking mode and the other in Nothinking mode. A second round of Thinking is triggered only when the two initial responses are inconsistent, using a single prompt with two different answers. Extensive experiments across multiple reasoning benchmarks demonstrate that JointThinking significantly outperforms few-shot chain-of-thought (CoT), thinking twice and majority voting. Moreover, it achieves comparable in-distribution performance to training-based SOTA reasoning method, while substantially outperforming on out-of-distribution tasks. We further conduct a systematic analysis of the calibration mechanism, showing the importance of structural thinking diversity and the benefits of consistency check. Additionally, we observe that the performance gap between actual and ideal reasoning narrows as model size increases in the second thinking, indicating the strong scalability of our approach. Finally, we discuss current limitations and outline promising directions for future ICL research in RLLMs.
title Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2508.03363