SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity
Fuente:
arXiv
Saved in:
| Main Authors: | Xi, Xiangyu, Kong, Deyang, Yang, Jian, Yang, Jiawei, Chen, Zhengyu, Wang, Wei, Wang, Jingang, Cai, Xunliang, Zhang, Shikun, Ye, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025)
by: Kong, Deyang, et al.
Published: (2025)
Autoformalizer with Tool Feedback
by: Guo, Qi, et al.
Published: (2025)
by: Guo, Qi, et al.
Published: (2025)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
OPE: Overcoming Information Saturation in Parallel Thinking via Outline-Guided Path Exploration
by: Guo, Qi, et al.
Published: (2026)
by: Guo, Qi, et al.
Published: (2026)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
by: Bai, Yang, et al.
Published: (2026)
by: Bai, Yang, et al.
Published: (2026)
Estimation of Heterogeneous Panel Data Models With Mixed Sampling Frequencies
by: Haoran Li, et al.
Published: (2026)
by: Haoran Li, et al.
Published: (2026)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025)
by: Yang, Kailai, et al.
Published: (2025)
Decentralized Robust Data-driven Predictive Control for Smoothing Mixed Traffic Flow
by: Shang, Xu, et al.
Published: (2024)
by: Shang, Xu, et al.
Published: (2024)
Scaling and Transferability of Annealing Strategies in Large Language Model Training
by: Wang, Siqi, et al.
Published: (2025)
by: Wang, Siqi, et al.
Published: (2025)
The Effects of Mixed Sample Data Augmentation are Class Dependent
by: Lee, Haeil, et al.
Published: (2023)
by: Lee, Haeil, et al.
Published: (2023)
RegMix: Data Mixture as Regression for Language Model Pre-training
by: Liu, Qian, et al.
Published: (2024)
by: Liu, Qian, et al.
Published: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
by: Li, Yudong, et al.
Published: (2026)
by: Li, Yudong, et al.
Published: (2026)
Quality-Diversity Generative Sampling for Learning with Synthetic Data
by: Chang, Allen, et al.
Published: (2023)
by: Chang, Allen, et al.
Published: (2023)
Mixed Matrix Completion in Complex Survey Sampling under Heterogeneous Missingness
by: Mao, Xiaojun, et al.
Published: (2024)
by: Mao, Xiaojun, et al.
Published: (2024)
Hierarchical Regularizers for Reverse Unrestricted Mixed Data Sampling Regressions
by: Hecq, Alain, et al.
Published: (2023)
by: Hecq, Alain, et al.
Published: (2023)
Analyzing Effects of Mixed Sample Data Augmentation on Model Interpretability
by: Won, Soyoun, et al.
Published: (2023)
by: Won, Soyoun, et al.
Published: (2023)
DynaMix: Generalizable Person Re-identification via Dynamic Relabeling and Mixed Data Sampling
by: Mamedov, Timur, et al.
Published: (2025)
by: Mamedov, Timur, et al.
Published: (2025)
Class-Aware PillarMix: Can Mixed Sample Data Augmentation Enhance 3D Object Detection with Radar Point Clouds?
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
SelectMix: Enhancing Label Noise Robustness through Targeted Sample Mixing
by: Liu, Qiuhao, et al.
Published: (2025)
by: Liu, Qiuhao, et al.
Published: (2025)
Large-Scale Diverse Synthesis for Mid-Training
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Free Performance Gain from Mixing Multiple Partially Labeled Samples in Multi-label Image Classification
by: Chong, Chak Fong, et al.
Published: (2024)
by: Chong, Chak Fong, et al.
Published: (2024)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
ALPCAH: Subspace Learning for Sample-wise Heteroscedastic Data
by: Cavazos, Javier Salazar, et al.
Published: (2025)
by: Cavazos, Javier Salazar, et al.
Published: (2025)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
by: Peng, Jiahui, et al.
Published: (2025)
by: Peng, Jiahui, et al.
Published: (2025)
When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
by: Deng, Mengyi, et al.
Published: (2025)
by: Deng, Mengyi, et al.
Published: (2025)
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
by: Wang, Guan, et al.
Published: (2023)
by: Wang, Guan, et al.
Published: (2023)
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
MedCutMix: A Data-Centric Approach to Improve Radiology Vision-Language Pre-training with Disease Awareness
by: Wang, Sinuo, et al.
Published: (2025)
by: Wang, Sinuo, et al.
Published: (2025)
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Revisiting Data Analysis with Pre-trained Foundation Models
by: Liang, Chen, et al.
Published: (2025)
by: Liang, Chen, et al.
Published: (2025)
ENTP: Enhancing Low-Quality SFT Data via Neural-Symbolic Text Purge-Mix
by: Yang, Zile, et al.
Published: (2025)
by: Yang, Zile, et al.
Published: (2025)
Generative Data Transformation: From Mixed to Unified Data
by: Zhang, Jiaqing, et al.
Published: (2026)
by: Zhang, Jiaqing, et al.
Published: (2026)
Mixed Cloud Control Testbed: Validating Vehicle-Road-Cloud Integration via Mixed Digital Twin
by: Dong, Jianghong, et al.
Published: (2022)
by: Dong, Jianghong, et al.
Published: (2022)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Robust Nonlinear Data-Driven Predictive Control for Mixed Vehicle Platoons via Koopman Operator and Reachability Analysis
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
Similar Items
-
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025) -
Autoformalizer with Tool Feedback
by: Guo, Qi, et al.
Published: (2025) -
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025) -
OPE: Overcoming Information Saturation in Parallel Thinking via Outline-Guided Path Exploration
by: Guo, Qi, et al.
Published: (2026) -
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)