Your thoughts tell who you are: Characterize the reasoning patterns of LRMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yida, Mao, Yuning, Yang, Xianjun, Ge, Suyu, Bi, Shengjie, Liu, Lijuan, Hosseini, Saghar, Tan, Liang, Nie, Yixin, Nie, Shaoliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914062639562752
author Chen, Yida
Mao, Yuning
Yang, Xianjun
Ge, Suyu
Bi, Shengjie
Liu, Lijuan
Hosseini, Saghar
Tan, Liang
Nie, Yixin
Nie, Shaoliang
author_facet Chen, Yida
Mao, Yuning
Yang, Xianjun
Ge, Suyu
Bi, Shengjie
Liu, Lijuan
Hosseini, Saghar
Tan, Liang
Nie, Yixin
Nie, Shaoliang
contents Current comparisons of large reasoning models (LRMs) focus on macro-level statistics such as task accuracy or reasoning length. Whether different LRMs reason differently remains an open question. To address this gap, we introduce the LLM-proposed Open Taxonomy (LOT), a classification method that uses a generative language model to compare reasoning traces from two LRMs and articulate their distinctive features in words. LOT then models how these features predict the source LRM of a reasoning trace based on their empirical distributions across LRM outputs. Iterating this process over a dataset of reasoning traces yields a human-readable taxonomy that characterizes how models think. We apply LOT to compare the reasoning of 12 open-source LRMs on tasks in math, science, and coding. LOT identifies systematic differences in their thoughts, achieving 80-100% accuracy in distinguishing reasoning traces from LRMs that differ in scale, base model family, or objective domain. Beyond classification, LOT's natural-language taxonomy provides qualitative explanations of how LRMs think differently. Finally, in a case study, we link the reasoning differences to performance: aligning the reasoning style of smaller Qwen3 models with that of the largest Qwen3 during test time improves their accuracy on GPQA by 3.3-5.7%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24147
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
Chen, Yida
Mao, Yuning
Yang, Xianjun
Ge, Suyu
Bi, Shengjie
Liu, Lijuan
Hosseini, Saghar
Tan, Liang
Nie, Yixin
Nie, Shaoliang
Computation and Language
Artificial Intelligence
Machine Learning
Current comparisons of large reasoning models (LRMs) focus on macro-level statistics such as task accuracy or reasoning length. Whether different LRMs reason differently remains an open question. To address this gap, we introduce the LLM-proposed Open Taxonomy (LOT), a classification method that uses a generative language model to compare reasoning traces from two LRMs and articulate their distinctive features in words. LOT then models how these features predict the source LRM of a reasoning trace based on their empirical distributions across LRM outputs. Iterating this process over a dataset of reasoning traces yields a human-readable taxonomy that characterizes how models think. We apply LOT to compare the reasoning of 12 open-source LRMs on tasks in math, science, and coding. LOT identifies systematic differences in their thoughts, achieving 80-100% accuracy in distinguishing reasoning traces from LRMs that differ in scale, base model family, or objective domain. Beyond classification, LOT's natural-language taxonomy provides qualitative explanations of how LRMs think differently. Finally, in a case study, we link the reasoning differences to performance: aligning the reasoning style of smaller Qwen3 models with that of the largest Qwen3 during test time improves their accuracy on GPQA by 3.3-5.7%.
title Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.24147