LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Komal, Ashraf, Tajamul, Thawakar, Omkar, Anwer, Rao Muhammad, Cholakkal, Hisham, Shah, Mubarak, Yang, Ming-Hsuan, Torr, Phillip H. S., Khan, Fahad Shahbaz, Khan, Salman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912288823312384
author Kumar, Komal
Ashraf, Tajamul
Thawakar, Omkar
Anwer, Rao Muhammad
Cholakkal, Hisham
Shah, Mubarak
Yang, Ming-Hsuan
Torr, Phillip H. S.
Khan, Fahad Shahbaz
Khan, Salman
author_facet Kumar, Komal
Ashraf, Tajamul
Thawakar, Omkar
Anwer, Rao Muhammad
Cholakkal, Hisham
Shah, Mubarak
Yang, Ming-Hsuan
Torr, Phillip H. S.
Khan, Fahad Shahbaz
Khan, Salman
contents Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the research community is now increasingly shifting focus toward post-training techniques to achieve further breakthroughs. While pretraining provides a broad linguistic foundation, post-training methods enable LLMs to refine their knowledge, improve reasoning, enhance factual accuracy, and align more effectively with user intents and ethical considerations. Fine-tuning, reinforcement learning, and test-time scaling have emerged as critical strategies for optimizing LLMs performance, ensuring robustness, and improving adaptability across various real-world tasks. This survey provides a systematic exploration of post-training methodologies, analyzing their role in refining LLMs beyond pretraining, addressing key challenges such as catastrophic forgetting, reward hacking, and inference-time trade-offs. We highlight emerging directions in model alignment, scalable adaptation, and inference-time reasoning, and outline future research directions. We also provide a public repository to continually track developments in this fast-evolving field: https://github.com/mbzuai-oryx/Awesome-LLM-Post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Kumar, Komal
Ashraf, Tajamul
Thawakar, Omkar
Anwer, Rao Muhammad
Cholakkal, Hisham
Shah, Mubarak
Yang, Ming-Hsuan
Torr, Phillip H. S.
Khan, Fahad Shahbaz
Khan, Salman
Computation and Language
Computer Vision and Pattern Recognition
Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the research community is now increasingly shifting focus toward post-training techniques to achieve further breakthroughs. While pretraining provides a broad linguistic foundation, post-training methods enable LLMs to refine their knowledge, improve reasoning, enhance factual accuracy, and align more effectively with user intents and ethical considerations. Fine-tuning, reinforcement learning, and test-time scaling have emerged as critical strategies for optimizing LLMs performance, ensuring robustness, and improving adaptability across various real-world tasks. This survey provides a systematic exploration of post-training methodologies, analyzing their role in refining LLMs beyond pretraining, addressing key challenges such as catastrophic forgetting, reward hacking, and inference-time trade-offs. We highlight emerging directions in model alignment, scalable adaptation, and inference-time reasoning, and outline future research directions. We also provide a public repository to continually track developments in this fast-evolving field: https://github.com/mbzuai-oryx/Awesome-LLM-Post-training.
title LLM Post-Training: A Deep Dive into Reasoning Large Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.21321