LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912288823312384 |
|---|---|
| author | Kumar, Komal Ashraf, Tajamul Thawakar, Omkar Anwer, Rao Muhammad Cholakkal, Hisham Shah, Mubarak Yang, Ming-Hsuan Torr, Phillip H. S. Khan, Fahad Shahbaz Khan, Salman |
| author_facet | Kumar, Komal Ashraf, Tajamul Thawakar, Omkar Anwer, Rao Muhammad Cholakkal, Hisham Shah, Mubarak Yang, Ming-Hsuan Torr, Phillip H. S. Khan, Fahad Shahbaz Khan, Salman |
| contents | Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the research community is now increasingly shifting focus toward post-training techniques to achieve further breakthroughs. While pretraining provides a broad linguistic foundation, post-training methods enable LLMs to refine their knowledge, improve reasoning, enhance factual accuracy, and align more effectively with user intents and ethical considerations. Fine-tuning, reinforcement learning, and test-time scaling have emerged as critical strategies for optimizing LLMs performance, ensuring robustness, and improving adaptability across various real-world tasks. This survey provides a systematic exploration of post-training methodologies, analyzing their role in refining LLMs beyond pretraining, addressing key challenges such as catastrophic forgetting, reward hacking, and inference-time trade-offs. We highlight emerging directions in model alignment, scalable adaptation, and inference-time reasoning, and outline future research directions. We also provide a public repository to continually track developments in this fast-evolving field: https://github.com/mbzuai-oryx/Awesome-LLM-Post-training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_21321 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLM Post-Training: A Deep Dive into Reasoning Large Language Models Kumar, Komal Ashraf, Tajamul Thawakar, Omkar Anwer, Rao Muhammad Cholakkal, Hisham Shah, Mubarak Yang, Ming-Hsuan Torr, Phillip H. S. Khan, Fahad Shahbaz Khan, Salman Computation and Language Computer Vision and Pattern Recognition Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the foundation for these models, yet the research community is now increasingly shifting focus toward post-training techniques to achieve further breakthroughs. While pretraining provides a broad linguistic foundation, post-training methods enable LLMs to refine their knowledge, improve reasoning, enhance factual accuracy, and align more effectively with user intents and ethical considerations. Fine-tuning, reinforcement learning, and test-time scaling have emerged as critical strategies for optimizing LLMs performance, ensuring robustness, and improving adaptability across various real-world tasks. This survey provides a systematic exploration of post-training methodologies, analyzing their role in refining LLMs beyond pretraining, addressing key challenges such as catastrophic forgetting, reward hacking, and inference-time trade-offs. We highlight emerging directions in model alignment, scalable adaptation, and inference-time reasoning, and outline future research directions. We also provide a public repository to continually track developments in this fast-evolving field: https://github.com/mbzuai-oryx/Awesome-LLM-Post-training. |
| title | LLM Post-Training: A Deep Dive into Reasoning Large Language Models |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2502.21321 |