DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Marjanović, Sara Vera, Patel, Arkil, Adlakha, Vaibhav, Aghajohari, Milad, BehnamGhader, Parishad, Bhatia, Mehar, Khandelwal, Aditi, Kraft, Austin, Krojer, Benno, Lù, Xing Han, Meade, Nicholas, Shin, Dongchan, Kazemnejad, Amirhossein, Kamath, Gaurav, Mosbach, Marius, Stańczak, Karolina, Reddy, Siva
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908769178353664
author Marjanović, Sara Vera
Patel, Arkil
Adlakha, Vaibhav
Aghajohari, Milad
BehnamGhader, Parishad
Bhatia, Mehar
Khandelwal, Aditi
Kraft, Austin
Krojer, Benno
Lù, Xing Han
Meade, Nicholas
Shin, Dongchan
Kazemnejad, Amirhossein
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Reddy, Siva
author_facet Marjanović, Sara Vera
Patel, Arkil
Adlakha, Vaibhav
Aghajohari, Milad
BehnamGhader, Parishad
Bhatia, Mehar
Khandelwal, Aditi
Kraft, Austin
Krojer, Benno
Lù, Xing Han
Meade, Nicholas
Shin, Dongchan
Kazemnejad, Amirhossein
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Reddy, Siva
contents Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains, seemingly "thinking" about a problem before providing an answer. This reasoning process is publicly available to the user, creating endless opportunities for studying the reasoning behaviour of the model and opening up the field of Thoughtology. Starting from a taxonomy of DeepSeek-R1's basic building blocks of reasoning, our analyses on DeepSeek-R1 investigate the impact and controllability of thought length, management of long or confusing contexts, cultural and safety concerns, and the status of DeepSeek-R1 vis-à-vis cognitive phenomena, such as human-like language processing and world modelling. Our findings paint a nuanced picture. Notably, we show DeepSeek-R1 has a 'sweet spot' of reasoning, where extra inference time can impair model performance. Furthermore, we find a tendency for DeepSeek-R1 to persistently ruminate on previously explored problem formulations, obstructing further exploration. We also note strong safety vulnerabilities of DeepSeek-R1 compared to its non-reasoning counterpart, which can also compromise safety-aligned LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07128
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
Marjanović, Sara Vera
Patel, Arkil
Adlakha, Vaibhav
Aghajohari, Milad
BehnamGhader, Parishad
Bhatia, Mehar
Khandelwal, Aditi
Kraft, Austin
Krojer, Benno
Lù, Xing Han
Meade, Nicholas
Shin, Dongchan
Kazemnejad, Amirhossein
Kamath, Gaurav
Mosbach, Marius
Stańczak, Karolina
Reddy, Siva
Computation and Language
Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains, seemingly "thinking" about a problem before providing an answer. This reasoning process is publicly available to the user, creating endless opportunities for studying the reasoning behaviour of the model and opening up the field of Thoughtology. Starting from a taxonomy of DeepSeek-R1's basic building blocks of reasoning, our analyses on DeepSeek-R1 investigate the impact and controllability of thought length, management of long or confusing contexts, cultural and safety concerns, and the status of DeepSeek-R1 vis-à-vis cognitive phenomena, such as human-like language processing and world modelling. Our findings paint a nuanced picture. Notably, we show DeepSeek-R1 has a 'sweet spot' of reasoning, where extra inference time can impair model performance. Furthermore, we find a tendency for DeepSeek-R1 to persistently ruminate on previously explored problem formulations, obstructing further exploration. We also note strong safety vulnerabilities of DeepSeek-R1 compared to its non-reasoning counterpart, which can also compromise safety-aligned LLMs.
title DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
topic Computation and Language
url https://arxiv.org/abs/2504.07128