AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Koh, Woosung, Oh, Wonbeen, Jang, Jaein, Lee, MinHyung, Kim, Hyeongjin, Kim, Ah Yeon, Kim, Joonkee, Lee, Junghyun, Kim, Taehyeon, Yun, Se-Young
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912628446593024
author Koh, Woosung
Oh, Wonbeen
Jang, Jaein
Lee, MinHyung
Kim, Hyeongjin
Kim, Ah Yeon
Kim, Joonkee
Lee, Junghyun
Kim, Taehyeon
Yun, Se-Young
author_facet Koh, Woosung
Oh, Wonbeen
Jang, Jaein
Lee, MinHyung
Kim, Hyeongjin
Kim, Ah Yeon
Kim, Joonkee
Lee, Junghyun
Kim, Taehyeon
Yun, Se-Young
contents Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (LMs). The self-improving mechanism often employs random observation (data) sampling. However, this results in trained observation imbalance; inefficiently over-training on solved examples while under-training on challenging ones. In response, we introduce Adaptive STaR (AdaSTaR), a novel algorithm that rectifies this by integrating two adaptive sampling principles: (1) Adaptive Sampling for Diversity: promoting balanced training across observations, and (2) Adaptive Sampling for Curriculum: dynamically adjusting data difficulty to match the model's evolving strength. Across six benchmarks, AdaSTaR achieves best test accuracy in all instances (6/6) and reduces training FLOPs by an average of 58.6% against an extensive list of baselines. These improvements in performance and efficiency generalize to different pre-trained LMs and larger models, paving the way for more efficient and effective self-improving LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Koh, Woosung
Oh, Wonbeen
Jang, Jaein
Lee, MinHyung
Kim, Hyeongjin
Kim, Ah Yeon
Kim, Joonkee
Lee, Junghyun
Kim, Taehyeon
Yun, Se-Young
Machine Learning
Artificial Intelligence
Computation and Language
Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (LMs). The self-improving mechanism often employs random observation (data) sampling. However, this results in trained observation imbalance; inefficiently over-training on solved examples while under-training on challenging ones. In response, we introduce Adaptive STaR (AdaSTaR), a novel algorithm that rectifies this by integrating two adaptive sampling principles: (1) Adaptive Sampling for Diversity: promoting balanced training across observations, and (2) Adaptive Sampling for Curriculum: dynamically adjusting data difficulty to match the model's evolving strength. Across six benchmarks, AdaSTaR achieves best test accuracy in all instances (6/6) and reduces training FLOPs by an average of 58.6% against an extensive list of baselines. These improvements in performance and efficiency generalize to different pre-trained LMs and larger models, paving the way for more efficient and effective self-improving LMs.
title AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.16322