Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Yiwei, Huang, Zhen, Mi, Tiantian, Si, Weiye, Zhou, Chenyang, Guo, Qipeng, Feng, Siyuan, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data Darwinism Part II: DataEvolve -- AI can Autonomously Evolve Pretraining Data Curation
von: Mi, Tiantian, et al.
Veröffentlicht: (2026)
von: Mi, Tiantian, et al.
Veröffentlicht: (2026)
daVinci-LLM:Towards the Science of Pretraining
von: Qin, Yiwei, et al.
Veröffentlicht: (2026)
von: Qin, Yiwei, et al.
Veröffentlicht: (2026)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
von: Guo, Xu, et al.
Veröffentlicht: (2026)
von: Guo, Xu, et al.
Veröffentlicht: (2026)
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
von: Zhou, Fan, et al.
Veröffentlicht: (2024)
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
von: Bi, Baolong, et al.
Veröffentlicht: (2025)
von: Bi, Baolong, et al.
Veröffentlicht: (2025)
Improving Continual Pre-training Through Seamless Data Packing
von: Yin, Ruicheng, et al.
Veröffentlicht: (2025)
von: Yin, Ruicheng, et al.
Veröffentlicht: (2025)
LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoran, et al.
Veröffentlicht: (2025)
DIVE: Diversified Iterative Self-Improvement
von: Qin, Yiwei, et al.
Veröffentlicht: (2025)
von: Qin, Yiwei, et al.
Veröffentlicht: (2025)
DataMan: Data Manager for Pre-training Large Language Models
von: Peng, Ru, et al.
Veröffentlicht: (2025)
von: Peng, Ru, et al.
Veröffentlicht: (2025)
Parallel Structures in Pre-training Data Yield In-Context Learning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Probing Language Models for Pre-training Data Detection
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Data-free Weight Compress and Denoise for Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
A Survey of Pre-trained Language Models for Processing Scientific Text
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
von: Ho, Xanh, et al.
Veröffentlicht: (2024)
PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
von: Liu, Jinhan, et al.
Veröffentlicht: (2026)
von: Liu, Jinhan, et al.
Veröffentlicht: (2026)
O1 Replication Journey: A Strategic Progress Report -- Part 1
von: Qin, Yiwei, et al.
Veröffentlicht: (2024)
von: Qin, Yiwei, et al.
Veröffentlicht: (2024)
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
von: Yamaguchi, Atsuki, et al.
Veröffentlicht: (2026)
von: Yamaguchi, Atsuki, et al.
Veröffentlicht: (2026)
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
Data Science and Technology Towards AGI Part I: Tiered Data Management
von: Wang, Yudong, et al.
Veröffentlicht: (2026)
von: Wang, Yudong, et al.
Veröffentlicht: (2026)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
von: Zeng, Hansi, et al.
Veröffentlicht: (2025)
von: Zeng, Hansi, et al.
Veröffentlicht: (2025)
LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction Tuning
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
von: Ye, Yangfan, et al.
Veröffentlicht: (2025)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
von: Theodoropoulos, Nikitas, et al.
Veröffentlicht: (2024)
von: Theodoropoulos, Nikitas, et al.
Veröffentlicht: (2024)
Inside the Black Box: Detecting Data Leakage in Pre-trained Language Encoders
von: Xin, Yuan, et al.
Veröffentlicht: (2024)
von: Xin, Yuan, et al.
Veröffentlicht: (2024)
The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity
von: Xi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Xi, Xiangyu, et al.
Veröffentlicht: (2025)
Data-efficient Performance Modeling via Pre-training
von: Liu, Chunting, et al.
Veröffentlicht: (2025)
von: Liu, Chunting, et al.
Veröffentlicht: (2025)
Understanding Data Temporality Impact on Large Language Models Pre-training
von: Pilchen, Hippolyte, et al.
Veröffentlicht: (2026)
von: Pilchen, Hippolyte, et al.
Veröffentlicht: (2026)
RegMix: Data Mixture as Regression for Language Model Pre-training
von: Liu, Qian, et al.
Veröffentlicht: (2024)
von: Liu, Qian, et al.
Veröffentlicht: (2024)
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
PGA-SciRE: Harnessing LLM on Data Augmentation for Enhancing Scientific Relation Extraction
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
On Predicting the Post-training Potential of Pre-trained LLMs
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2026)
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
von: Naous, Tarek, et al.
Veröffentlicht: (2025)
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
von: Chen, Shan, et al.
Veröffentlicht: (2024)
von: Chen, Shan, et al.
Veröffentlicht: (2024)
Influence-driven Curriculum Learning for Pre-training on Limited Data
von: Schoenegger, Loris, et al.
Veröffentlicht: (2025)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2025)
Scaling Agents via Continual Pre-training
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Data Darwinism Part II: DataEvolve -- AI can Autonomously Evolve Pretraining Data Curation
von: Mi, Tiantian, et al.
Veröffentlicht: (2026) -
daVinci-LLM:Towards the Science of Pretraining
von: Qin, Yiwei, et al.
Veröffentlicht: (2026) -
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
von: Guo, Xu, et al.
Veröffentlicht: (2026) -
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
von: Wang, Cheng, et al.
Veröffentlicht: (2024) -
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
von: Zhou, Fan, et al.
Veröffentlicht: (2024)