Saved in:
Bibliographic Details
Main Authors: Murray, Seoirse, Qi, Allison, Qian, Timothy, Schulman, John, Burns, Collin, Price, Sara
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.05910
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918324592443392
author Murray, Seoirse
Qi, Allison
Qian, Timothy
Schulman, John
Burns, Collin
Price, Sara
author_facet Murray, Seoirse
Qi, Allison
Qian, Timothy
Schulman, John
Burns, Collin
Price, Sara
contents LLM post-training involves many diverse datasets, each targeting a specific behavior. But these datasets encode incidental patterns alongside intended ones: correlations between formatting and content, narrow phrasings across diverse problems, and implicit associations arising from the discrete data curation process. These patterns are often invisible to developers yet salient to models, producing behaviors that surprise their creators, such as rejecting true facts presented in a particular question format. We call this chunky post-training: the model learns spurious correlations as a result of distinct chunks of post-training data. We introduce SURF, a black-box pipeline which surfaces these unintended behaviors at run time, and TURF, a tool that traces these failures back to specific post-training data. Applying these tools to frontier models (Claude 4.5, GPT-5.1, Grok 4.1, Gemini 3) and open models (Tülu 3), we show that chunky post-training produces miscalibrated behaviors, which often result from imbalanced or underspecified chunks of post-training data.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05910
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chunky Post-Training: Data Driven Failures of Generalization
Murray, Seoirse
Qi, Allison
Qian, Timothy
Schulman, John
Burns, Collin
Price, Sara
Machine Learning
LLM post-training involves many diverse datasets, each targeting a specific behavior. But these datasets encode incidental patterns alongside intended ones: correlations between formatting and content, narrow phrasings across diverse problems, and implicit associations arising from the discrete data curation process. These patterns are often invisible to developers yet salient to models, producing behaviors that surprise their creators, such as rejecting true facts presented in a particular question format. We call this chunky post-training: the model learns spurious correlations as a result of distinct chunks of post-training data. We introduce SURF, a black-box pipeline which surfaces these unintended behaviors at run time, and TURF, a tool that traces these failures back to specific post-training data. Applying these tools to frontier models (Claude 4.5, GPT-5.1, Grok 4.1, Gemini 3) and open models (Tülu 3), we show that chunky post-training produces miscalibrated behaviors, which often result from imbalanced or underspecified chunks of post-training data.
title Chunky Post-Training: Data Driven Failures of Generalization
topic Machine Learning
url https://arxiv.org/abs/2602.05910