Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Jun, Liu, Xuebo, Deng, Hexuan, Lin, Zepeng, Yu, Zixiong, Wei, Jiansheng, Meng, Xiaojun, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
by: Yu, Zixiong, et al.
Published: (2026)
by: Yu, Zixiong, et al.
Published: (2026)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
by: Deng, Hexuan, et al.
Published: (2025)
by: Deng, Hexuan, et al.
Published: (2025)
AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
by: Ke, Xiaopeng, et al.
Published: (2025)
by: Ke, Xiaopeng, et al.
Published: (2025)
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
by: Yuan, Bo, et al.
Published: (2026)
by: Yuan, Bo, et al.
Published: (2026)
APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training
by: Rao, Jun, et al.
Published: (2025)
by: Rao, Jun, et al.
Published: (2025)
DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models
by: Nie, Shuo, et al.
Published: (2026)
by: Nie, Shuo, et al.
Published: (2026)
NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
by: Rao, Jun, et al.
Published: (2025)
by: Rao, Jun, et al.
Published: (2025)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
AdaptGrad: Adaptive Sampling to Reduce Noise
by: Zhou, Linjiang, et al.
Published: (2024)
by: Zhou, Linjiang, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory
by: Zhao, Haokun, et al.
Published: (2025)
by: Zhao, Haokun, et al.
Published: (2025)
Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
SGIC: A Self-Guided Iterative Calibration Framework for RAG
by: Chen, Guanhua, et al.
Published: (2025)
by: Chen, Guanhua, et al.
Published: (2025)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
by: Deng, Hexuan, et al.
Published: (2026)
by: Deng, Hexuan, et al.
Published: (2026)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
by: Yu, Fengming, et al.
Published: (2025)
by: Yu, Fengming, et al.
Published: (2025)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
Machine Learning Potential Powered Insights into the Mechanical Stability of Amorphous Li-Si Alloys
by: Wei, Zixiong, et al.
Published: (2024)
by: Wei, Zixiong, et al.
Published: (2024)
History-Guided Iterative Visual Reasoning with Self-Correction
by: Yang, Xinglong, et al.
Published: (2026)
by: Yang, Xinglong, et al.
Published: (2026)
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
by: Pan, Ruwei, et al.
Published: (2026)
by: Pan, Ruwei, et al.
Published: (2026)
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
by: Ma, Xinyu, et al.
Published: (2024)
by: Ma, Xinyu, et al.
Published: (2024)
Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
by: LI, Xuying
Published: (2025)
by: LI, Xuying
Published: (2025)
TasTe: Teaching Large Language Models to Translate through Self-Reflection
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
Relaxation Dynamics in Persistent Epithelial Tissues
by: Li, Meng-Yuan, et al.
Published: (2024)
by: Li, Meng-Yuan, et al.
Published: (2024)
A Neural-Guided Dynamic Symbolic Network for Exploring Mathematical Expressions from Data
by: Li, Wenqiang, et al.
Published: (2023)
by: Li, Wenqiang, et al.
Published: (2023)
CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning
by: Wang, Zhuo, et al.
Published: (2026)
by: Wang, Zhuo, et al.
Published: (2026)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
by: Liu, Liangxin, et al.
Published: (2024)
by: Liu, Liangxin, et al.
Published: (2024)
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
by: Lin, Yunze
Published: (2025)
by: Lin, Yunze
Published: (2025)
A Probabilistic Approach for Demand-Aware Ride-Sharing Optimization
by: Lin, Qiulin, et al.
Published: (2019)
by: Lin, Qiulin, et al.
Published: (2019)
Iterative Reasoning Preference Optimization
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
PureSample: Neural Materials Learned by Sampling Microgeometry
by: Li, Zixuan, et al.
Published: (2025)
by: Li, Zixuan, et al.
Published: (2025)
Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
by: Mu, Yongyu, et al.
Published: (2026)
by: Mu, Yongyu, et al.
Published: (2026)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
by: Wan, Guangya, et al.
Published: (2024)
by: Wan, Guangya, et al.
Published: (2024)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
Similar Items
-
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
by: Yu, Zixiong, et al.
Published: (2026) -
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
by: Deng, Hexuan, et al.
Published: (2025) -
AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
by: Ke, Xiaopeng, et al.
Published: (2025) -
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
by: Yuan, Bo, et al.
Published: (2026) -
APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training
by: Rao, Jun, et al.
Published: (2025)