A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912233492054016 |
|---|---|
| author | Wang, Jun |
| author_facet | Wang, Jun |
| contents | OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_10867 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1 Wang, Jun Artificial Intelligence Computation and Language OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework. |
| title | A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1 |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2502.10867 |