PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yin, Shangjian, Liang, Shining, Ding, Wenbiao, Qian, Yuli, Shi, Zhouxing, Li, Hongzhi, Xie, Yutao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914460142141440
author Yin, Shangjian
Liang, Shining
Ding, Wenbiao
Qian, Yuli
Shi, Zhouxing
Li, Hongzhi
Xie, Yutao
author_facet Yin, Shangjian
Liang, Shining
Ding, Wenbiao
Qian, Yuli
Shi, Zhouxing
Li, Hongzhi
Xie, Yutao
contents High-quality instruction data is critical for LLM alignment, yet existing open-source datasets often lack efficiency, requiring hundreds of thousands of examples to approach proprietary performance. In this work, we find that beyond the widely recognized importance of prompt-response quality, prompt difficulty itself plays a critical role in driving alignment gains. Motivated by this observation, we introduce PiKa, a data-efficient family of expert-level alignment datasets that concentrates supervision on high-difficulty instructions. The PiKa-SFT dataset contains only 30k examples, an order of magnitude fewer than state-of-the-art open datasets like Magpie-Pro. Despite its small size, fine-tuning Llama-3-8B-Base on PiKa-SFT even outperforms the official Llama-3-8B-Instruct model trained on over 10M proprietary examples on widely used benchmarks such as AlpacaEval 2.0 and Arena-Hard. We also validate the generalizability of PiKa across the Qwen2.5 series (0.5B-7B), consistently surpassing their official instruction-tuned counterparts. Additionally, we provide 30k high-quality preference optimization examples to further enhance alignment. Our results demonstrate that promising alignment is achievable with significantly reduced data, democratizing access for resource-constrained research. Our code and data will be available at https://github.com/SJY8460/PiKa.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
Yin, Shangjian
Liang, Shining
Ding, Wenbiao
Qian, Yuli
Shi, Zhouxing
Li, Hongzhi
Xie, Yutao
Computation and Language
High-quality instruction data is critical for LLM alignment, yet existing open-source datasets often lack efficiency, requiring hundreds of thousands of examples to approach proprietary performance. In this work, we find that beyond the widely recognized importance of prompt-response quality, prompt difficulty itself plays a critical role in driving alignment gains. Motivated by this observation, we introduce PiKa, a data-efficient family of expert-level alignment datasets that concentrates supervision on high-difficulty instructions. The PiKa-SFT dataset contains only 30k examples, an order of magnitude fewer than state-of-the-art open datasets like Magpie-Pro. Despite its small size, fine-tuning Llama-3-8B-Base on PiKa-SFT even outperforms the official Llama-3-8B-Instruct model trained on over 10M proprietary examples on widely used benchmarks such as AlpacaEval 2.0 and Arena-Hard. We also validate the generalizability of PiKa across the Qwen2.5 series (0.5B-7B), consistently surpassing their official instruction-tuned counterparts. Additionally, we provide 30k high-quality preference optimization examples to further enhance alignment. Our results demonstrate that promising alignment is achievable with significantly reduced data, democratizing access for resource-constrained research. Our code and data will be available at https://github.com/SJY8460/PiKa.
title PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
topic Computation and Language
url https://arxiv.org/abs/2510.06670