A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
Fuente:
arXiv
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912233492054016 |
|---|---|
| author | Wang, Jun |
| author_facet | Wang, Jun |
| contents | OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_10867 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1 Wang, Jun Artificial Intelligence Computation and Language OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework. |
| title | A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1 |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2502.10867 |