AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912628446593024 |
|---|---|
| author | Koh, Woosung Oh, Wonbeen Jang, Jaein Lee, MinHyung Kim, Hyeongjin Kim, Ah Yeon Kim, Joonkee Lee, Junghyun Kim, Taehyeon Yun, Se-Young |
| author_facet | Koh, Woosung Oh, Wonbeen Jang, Jaein Lee, MinHyung Kim, Hyeongjin Kim, Ah Yeon Kim, Joonkee Lee, Junghyun Kim, Taehyeon Yun, Se-Young |
| contents | Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (LMs). The self-improving mechanism often employs random observation (data) sampling. However, this results in trained observation imbalance; inefficiently over-training on solved examples while under-training on challenging ones. In response, we introduce Adaptive STaR (AdaSTaR), a novel algorithm that rectifies this by integrating two adaptive sampling principles: (1) Adaptive Sampling for Diversity: promoting balanced training across observations, and (2) Adaptive Sampling for Curriculum: dynamically adjusting data difficulty to match the model's evolving strength. Across six benchmarks, AdaSTaR achieves best test accuracy in all instances (6/6) and reduces training FLOPs by an average of 58.6% against an extensive list of baselines. These improvements in performance and efficiency generalize to different pre-trained LMs and larger models, paving the way for more efficient and effective self-improving LMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16322 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners Koh, Woosung Oh, Wonbeen Jang, Jaein Lee, MinHyung Kim, Hyeongjin Kim, Ah Yeon Kim, Joonkee Lee, Junghyun Kim, Taehyeon Yun, Se-Young Machine Learning Artificial Intelligence Computation and Language Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (LMs). The self-improving mechanism often employs random observation (data) sampling. However, this results in trained observation imbalance; inefficiently over-training on solved examples while under-training on challenging ones. In response, we introduce Adaptive STaR (AdaSTaR), a novel algorithm that rectifies this by integrating two adaptive sampling principles: (1) Adaptive Sampling for Diversity: promoting balanced training across observations, and (2) Adaptive Sampling for Curriculum: dynamically adjusting data difficulty to match the model's evolving strength. Across six benchmarks, AdaSTaR achieves best test accuracy in all instances (6/6) and reduces training FLOPs by an average of 58.6% against an extensive list of baselines. These improvements in performance and efficiency generalize to different pre-trained LMs and larger models, paving the way for more efficient and effective self-improving LMs. |
| title | AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2505.16322 |