Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Djuhera, Aladin, Kadhe, Swanand Ravindra, Zawad, Syed, Ahmed, Farhan, Ludwig, Heiko, Boche, Holger
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908816961961984
author Djuhera, Aladin
Kadhe, Swanand Ravindra
Zawad, Syed
Ahmed, Farhan
Ludwig, Heiko
Boche, Holger
author_facet Djuhera, Aladin
Kadhe, Swanand Ravindra
Zawad, Syed
Ahmed, Farhan
Ludwig, Heiko
Boche, Holger
contents Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible to the public, with limited information about their construction process. This lack of transparency has motivated the recent development of open-source post-training corpora. While training on these open alternatives can yield performance comparable to that of leading models, systematic comparisons remain challenging due to the significant computational cost of conducting them rigorously at scale, and are therefore largely absent. As a result, it remains unclear how specific samples, task types, or curation strategies influence downstream performance when assessing data quality. In this work, we conduct the first comprehensive side-by-side analysis of two prominent open post-training datasets: Tulu-3-SFT-Mix and SmolTalk. Using the Magpie framework, we annotate each sample with detailed quality metrics, including turn structure (single-turn vs. multi-turn), task category, input quality, and response quality, and we derive statistics that reveal structural and qualitative similarities and differences between the two datasets. Based on these insights, we design a principled curation recipe that produces a new data mixture, TuluTalk, which contains 14% fewer samples than either source dataset while matching or exceeding their performance on key benchmarks. Our findings offer actionable insights for constructing more effective post-training datasets that improve model performance within practical resource limits. To support future research, we publicly release both the annotated source datasets and our curated TuluTalk mixture.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06522
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
Djuhera, Aladin
Kadhe, Swanand Ravindra
Zawad, Syed
Ahmed, Farhan
Ludwig, Heiko
Boche, Holger
Computation and Language
Artificial Intelligence
Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and specialized skills. However, most post-training datasets used in leading open- and closed-source LLMs remain inaccessible to the public, with limited information about their construction process. This lack of transparency has motivated the recent development of open-source post-training corpora. While training on these open alternatives can yield performance comparable to that of leading models, systematic comparisons remain challenging due to the significant computational cost of conducting them rigorously at scale, and are therefore largely absent. As a result, it remains unclear how specific samples, task types, or curation strategies influence downstream performance when assessing data quality. In this work, we conduct the first comprehensive side-by-side analysis of two prominent open post-training datasets: Tulu-3-SFT-Mix and SmolTalk. Using the Magpie framework, we annotate each sample with detailed quality metrics, including turn structure (single-turn vs. multi-turn), task category, input quality, and response quality, and we derive statistics that reveal structural and qualitative similarities and differences between the two datasets. Based on these insights, we design a principled curation recipe that produces a new data mixture, TuluTalk, which contains 14% fewer samples than either source dataset while matching or exceeding their performance on key benchmarks. Our findings offer actionable insights for constructing more effective post-training datasets that improve model performance within practical resource limits. To support future research, we publicly release both the annotated source datasets and our curated TuluTalk mixture.
title Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.06522