Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Lingzhe, Jia, Tong, Zhai, Yunpeng, Fang, Liancheng, Zheng, Kening, Liu, Hongyi, Huang, Xiaosong, Yu, Philip S., Li, Ying
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910194196283392
author Zhang, Lingzhe
Jia, Tong
Zhai, Yunpeng
Fang, Liancheng
Zheng, Kening
Liu, Hongyi
Huang, Xiaosong
Yu, Philip S.
Li, Ying
author_facet Zhang, Lingzhe
Jia, Tong
Zhai, Yunpeng
Fang, Liancheng
Zheng, Kening
Liu, Hongyi
Huang, Xiaosong
Yu, Philip S.
Li, Ying
contents Reinforcement fine-tuning (RFT) has become a core paradigm for post-training large language models, yet its training process remains highly fragile. Existing efforts mainly improve reliability at the system level or address specific issues in individual subproblems by modifying RFT algorithms. Despite their effectiveness, they largely overlook the problem of failure management at the training-process level. When training goes wrong, practitioners still rely heavily on expert-driven manual inspection and correction, and automatic failure management for RFT remains largely unexplored. In this paper, we take a first step toward systematic failure management for reinforcement fine-tuning. To understand the empirical structure of RFT failures, we first construct RFT-FaultBench, the first benchmark for fine-grained failures in reinforcement fine-tuning, covering 5 fault families, 16 fault types, 779 training runs, 22,549 train-step records, and 1,457,288 trajectory-level records. Based on this benchmark, we conduct a comprehensive empirical study showing that RFT failures are both observable from training dynamics and distinguishable through their empirical fault fingerprints. Building on these findings, we propose RFT-FM, an automatic failure management framework for reinforcement fine-tuning that unifies anomaly detection, failure diagnosis, and auto remediation in a closed loop. Experimental results show that RFT-FaultBench is neither trivial nor saturated: it exhibits clear anomaly structure while still posing substantial challenges, especially under subtle fault settings. Moreover, RFT-FM shows strong capability in detecting, diagnosing, and mitigating RFT failures.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04431
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
Zhang, Lingzhe
Jia, Tong
Zhai, Yunpeng
Fang, Liancheng
Zheng, Kening
Liu, Hongyi
Huang, Xiaosong
Yu, Philip S.
Li, Ying
Software Engineering
Artificial Intelligence
Reinforcement fine-tuning (RFT) has become a core paradigm for post-training large language models, yet its training process remains highly fragile. Existing efforts mainly improve reliability at the system level or address specific issues in individual subproblems by modifying RFT algorithms. Despite their effectiveness, they largely overlook the problem of failure management at the training-process level. When training goes wrong, practitioners still rely heavily on expert-driven manual inspection and correction, and automatic failure management for RFT remains largely unexplored. In this paper, we take a first step toward systematic failure management for reinforcement fine-tuning. To understand the empirical structure of RFT failures, we first construct RFT-FaultBench, the first benchmark for fine-grained failures in reinforcement fine-tuning, covering 5 fault families, 16 fault types, 779 training runs, 22,549 train-step records, and 1,457,288 trajectory-level records. Based on this benchmark, we conduct a comprehensive empirical study showing that RFT failures are both observable from training dynamics and distinguishable through their empirical fault fingerprints. Building on these findings, we propose RFT-FM, an automatic failure management framework for reinforcement fine-tuning that unifies anomaly detection, failure diagnosis, and auto remediation in a closed loop. Experimental results show that RFT-FaultBench is neither trivial nor saturated: it exhibits clear anomaly structure while still posing substantial challenges, especially under subtle fault settings. Moreover, RFT-FM shows strong capability in detecting, diagnosing, and mitigating RFT failures.
title Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2605.04431