GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Simin, Glarou, Maria Ios, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
A Split-Client Approach to Second-Order Optimization
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
by: Fan, Simin, et al.
Published: (2026)
by: Fan, Simin, et al.
Published: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
CoBo: Collaborative Learning via Bilevel Optimization
by: Hashemi, Diba, et al.
Published: (2024)
by: Hashemi, Diba, et al.
Published: (2024)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
by: Li, Cheng, et al.
Published: (2025)
by: Li, Cheng, et al.
Published: (2025)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
by: Yen, Thomson, et al.
Published: (2025)
by: Yen, Thomson, et al.
Published: (2025)
Certified Robustness from Approximate Gaussian Mixture Structures in Pretrained Latent Spaces
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
by: Kosson, Atli, et al.
Published: (2024)
by: Kosson, Atli, et al.
Published: (2024)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Group-Adaptive Threshold Optimization for Robust AI-Generated Text Detection
by: Jung, Minseok, et al.
Published: (2025)
by: Jung, Minseok, et al.
Published: (2025)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Adaptive Heterogeneous Mixtures of Normalising Flows for Robust Variational Inference
by: Wiriyapong, Benjamin, et al.
Published: (2025)
by: Wiriyapong, Benjamin, et al.
Published: (2025)
Intrinsic User-Centric Interpretability through Global Mixture of Experts
by: Swamy, Vinitra, et al.
Published: (2024)
by: Swamy, Vinitra, et al.
Published: (2024)
Adaptive Group Robust Ensemble Knowledge Distillation
by: Kenfack, Patrik, et al.
Published: (2024)
by: Kenfack, Patrik, et al.
Published: (2024)
Adaptive Multi-task Learning for Multi-sector Portfolio Optimization
by: Fan, Qingliang, et al.
Published: (2025)
by: Fan, Qingliang, et al.
Published: (2025)
Optimizing the Training Diet: Data Mixture Search for Robust Time Series Forecasting
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
TapWeight: Reweighting Pretraining Objectives for Task-Adaptive Pretraining
by: Zhang, Ruiyi, et al.
Published: (2024)
by: Zhang, Ruiyi, et al.
Published: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Variance-Adaptive Muon: Accelerating LLM Pretraining with NSR-Modulated and Variance-Scaled Momentum
by: Li, Jingru, et al.
Published: (2026)
by: Li, Jingru, et al.
Published: (2026)
Stochastic Optimization with Random Search
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Robust Mixture Learning when Outliers Overwhelm Small Groups
by: Dmitriev, Daniil, et al.
Published: (2024)
by: Dmitriev, Daniil, et al.
Published: (2024)
Pretrained transformer efficiently learns low-dimensional target functions in-context
by: Oko, Kazusato, et al.
Published: (2024)
by: Oko, Kazusato, et al.
Published: (2024)
Similar Items
-
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
by: Zhou, Xinyu, et al.
Published: (2024) -
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024) -
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025) -
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)