A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Wang, Jun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912233492054016
author Wang, Jun
author_facet Wang, Jun
contents OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
Wang, Jun
Artificial Intelligence
Computation and Language
OpenAI o1 has shown that applying reinforcement learning to integrate reasoning steps directly during inference can significantly improve a model's reasoning capabilities. This result is exciting as the field transitions from the conventional autoregressive method of generating answers to a more deliberate approach that models the slow-thinking process through step-by-step reasoning training. Reinforcement learning plays a key role in both the model's training and decoding processes. In this article, we present a comprehensive formulation of reasoning problems and investigate the use of both model-based and model-free approaches to better support this slow-thinking framework.
title A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.10867