Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yue, Liu, Qiuzhi, Xu, Jiahao, Liang, Tian, Chen, Xingyu, He, Zhiwei, Song, Linfeng, Yu, Dian, Li, Juntao, Zhang, Zhuosheng, Wang, Rui, Tu, Zhaopeng, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915156777238528
author Wang, Yue
Liu, Qiuzhi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Song, Linfeng
Yu, Dian
Li, Juntao
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
author_facet Wang, Yue
Liu, Qiuzhi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Song, Linfeng
Yu, Dian
Li, Juntao
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
contents Large language models (LLMs) such as OpenAI's o1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where o1-like LLMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source o1-like models, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty TIP that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in o1-like LLMs and offer a practical solution to enhance their problem-solving capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Wang, Yue
Liu, Qiuzhi
Xu, Jiahao
Liang, Tian
Chen, Xingyu
He, Zhiwei
Song, Linfeng
Yu, Dian
Li, Juntao
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
Computation and Language
Large language models (LLMs) such as OpenAI's o1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where o1-like LLMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source o1-like models, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty TIP that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in o1-like LLMs and offer a practical solution to enhance their problem-solving capabilities.
title Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
topic Computation and Language
url https://arxiv.org/abs/2501.18585