Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rao, Jun, Liu, Xuebo, Deng, Hexuan, Lin, Zepeng, Yu, Zixiong, Wei, Jiansheng, Meng, Xiaojun, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914480708911104
author Rao, Jun
Liu, Xuebo
Deng, Hexuan
Lin, Zepeng
Yu, Zixiong
Wei, Jiansheng
Meng, Xiaojun
Zhang, Min
author_facet Rao, Jun
Liu, Xuebo
Deng, Hexuan
Lin, Zepeng
Yu, Zixiong
Wei, Jiansheng
Meng, Xiaojun
Zhang, Min
contents In mathematical reasoning, data selection strategies predominantly rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. This misalignment limits the efficiency of Supervised Fine-Tuning and Reinforcement Learning. To bridge this gap, we introduce SAI-DPO (Self-Aware Iterative Data Persistent Optimization), a dynamic sampling framework that aligns training data with the model's intrinsic competence. SAI-DPO operationalizes two novel metrics: Knowledge Semantic Alignment for targeting domain weaknesses, and Self-Aware Difficulty, derived from pass rates and reasoning path characteristics, to gauge instance complexity relative to the model's current state. By iteratively recalibrating the data distribution based on real-time feedback, SAI-DPO dynamically aligns training samples with the model's evolving competence, ensuring the data remains strictly relevant to the model's current capability level. Extensive experiments on eight benchmarks (including AIME24 and AMC23) demonstrate that SAI-DPO outperforms static baselines at most nearly 6 points, achieving state-of-the-art efficiency with significantly less data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16176
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
Rao, Jun
Liu, Xuebo
Deng, Hexuan
Lin, Zepeng
Yu, Zixiong
Wei, Jiansheng
Meng, Xiaojun
Zhang, Min
Artificial Intelligence
Computation and Language
In mathematical reasoning, data selection strategies predominantly rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. This misalignment limits the efficiency of Supervised Fine-Tuning and Reinforcement Learning. To bridge this gap, we introduce SAI-DPO (Self-Aware Iterative Data Persistent Optimization), a dynamic sampling framework that aligns training data with the model's intrinsic competence. SAI-DPO operationalizes two novel metrics: Knowledge Semantic Alignment for targeting domain weaknesses, and Self-Aware Difficulty, derived from pass rates and reasoning path characteristics, to gauge instance complexity relative to the model's current state. By iteratively recalibrating the data distribution based on real-time feedback, SAI-DPO dynamically aligns training samples with the model's evolving competence, ensuring the data remains strictly relevant to the model's current capability level. Extensive experiments on eight benchmarks (including AIME24 and AMC23) demonstrate that SAI-DPO outperforms static baselines at most nearly 6 points, achieving state-of-the-art efficiency with significantly less data.
title Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.16176