Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jingwei, Gu, Xinran, Zhang, Jingzhao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
por: Gu, Xinran, et al.
Publicado: (2025)
por: Gu, Xinran, et al.
Publicado: (2025)
Research and Implementation of Data Enhancement Techniques for Graph Neural Networks
por: Gu, Jingzhao, et al.
Publicado: (2024)
por: Gu, Jingzhao, et al.
Publicado: (2024)
A Quadratic Synchronization Rule for Distributed Deep Learning
por: Gu, Xinran, et al.
Publicado: (2023)
por: Gu, Xinran, et al.
Publicado: (2023)
Understanding Nonlinear Implicit Bias via Region Counts in Input Space
por: Li, Jingwei, et al.
Publicado: (2025)
por: Li, Jingwei, et al.
Publicado: (2025)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
por: Liu, Siyuan, et al.
Publicado: (2026)
por: Liu, Siyuan, et al.
Publicado: (2026)
On the Condition Number Dependency in Bilevel Optimization
por: Chen, Lesi, et al.
Publicado: (2025)
por: Chen, Lesi, et al.
Publicado: (2025)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
por: Xu, Jing, et al.
Publicado: (2024)
por: Xu, Jing, et al.
Publicado: (2024)
Towards Black-Box Membership Inference Attack for Diffusion Models
por: Li, Jingwei, et al.
Publicado: (2024)
por: Li, Jingwei, et al.
Publicado: (2024)
Scaling Laws for Optimal Data Mixtures
por: Shukor, Mustafa, et al.
Publicado: (2025)
por: Shukor, Mustafa, et al.
Publicado: (2025)
Scalable Model Merging with Progressive Layer-wise Distillation
por: Xu, Jing, et al.
Publicado: (2025)
por: Xu, Jing, et al.
Publicado: (2025)
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
por: Wan, Weilin, et al.
Publicado: (2026)
por: Wan, Weilin, et al.
Publicado: (2026)
Similarity-Aware Mixture-of-Experts for Data-Efficient Continual Learning
por: Mclaughlin, Connor, et al.
Publicado: (2026)
por: Mclaughlin, Connor, et al.
Publicado: (2026)
Efficient Sampling on Riemannian Manifolds via Langevin MCMC
por: Cheng, Xiang, et al.
Publicado: (2024)
por: Cheng, Xiang, et al.
Publicado: (2024)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
por: Du, Zhixu, et al.
Publicado: (2023)
por: Du, Zhixu, et al.
Publicado: (2023)
Faster Gradient Methods for Highly-Smooth Stochastic Bilevel Optimization
por: Chen, Lesi, et al.
Publicado: (2025)
por: Chen, Lesi, et al.
Publicado: (2025)
Fast and Multiphase Rates for Nearest Neighbor Classifiers
por: Yang, Pengkun, et al.
Publicado: (2023)
por: Yang, Pengkun, et al.
Publicado: (2023)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
por: Chen, Lesi, et al.
Publicado: (2023)
por: Chen, Lesi, et al.
Publicado: (2023)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
por: He, Shwai, et al.
Publicado: (2025)
por: He, Shwai, et al.
Publicado: (2025)
Second-Order Min-Max Optimization with Lazy Hessians
por: Chen, Lesi, et al.
Publicado: (2024)
por: Chen, Lesi, et al.
Publicado: (2024)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
por: Ye, Jiasheng, et al.
Publicado: (2024)
por: Ye, Jiasheng, et al.
Publicado: (2024)
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
por: Jha, Nandan Kumar, et al.
Publicado: (2026)
por: Jha, Nandan Kumar, et al.
Publicado: (2026)
MetaKube: An Experience-Aware LLM Framework for Kubernetes Failure Diagnosis
por: Sun, Wei, et al.
Publicado: (2026)
por: Sun, Wei, et al.
Publicado: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
por: Sedova, Anastasiia, et al.
Publicado: (2026)
por: Sedova, Anastasiia, et al.
Publicado: (2026)
MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs
por: Gu, Yupu, et al.
Publicado: (2026)
por: Gu, Yupu, et al.
Publicado: (2026)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
por: Belenki, Lior, et al.
Publicado: (2025)
por: Belenki, Lior, et al.
Publicado: (2025)
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
por: Ding, Yuyang, et al.
Publicado: (2025)
por: Ding, Yuyang, et al.
Publicado: (2025)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
por: Benazir, Afsara, et al.
Publicado: (2026)
por: Benazir, Afsara, et al.
Publicado: (2026)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
por: Wu, Xiaojun, et al.
Publicado: (2025)
por: Wu, Xiaojun, et al.
Publicado: (2025)
U-CAN: Utility-Aware Contrastive Attenuation for Efficient Unlearning in Generative Recommendation
por: Wu, Zezheng, et al.
Publicado: (2026)
por: Wu, Zezheng, et al.
Publicado: (2026)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
por: Li, Bo, et al.
Publicado: (2026)
por: Li, Bo, et al.
Publicado: (2026)
Fast Conditional Mixing of MCMC Algorithms for Non-log-concave Distributions
por: Cheng, Xiang, et al.
Publicado: (2023)
por: Cheng, Xiang, et al.
Publicado: (2023)
D3: Diversity, Difficulty, and Dependability-Aware Data Selection for Sample-Efficient LLM Instruction Tuning
por: Zhang, Jia, et al.
Publicado: (2025)
por: Zhang, Jia, et al.
Publicado: (2025)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Functionally Constrained Algorithm Solves Convex Simple Bilevel Problems
por: Zhang, Huaqing, et al.
Publicado: (2024)
por: Zhang, Huaqing, et al.
Publicado: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
por: Wen, Kaiyue, et al.
Publicado: (2024)
por: Wen, Kaiyue, et al.
Publicado: (2024)
MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
por: Wen, Bingbing, et al.
Publicado: (2026)
por: Wen, Bingbing, et al.
Publicado: (2026)
ADO: Automatic Data Optimization for Inputs in LLM Prompts
por: Lin, Sam, et al.
Publicado: (2025)
por: Lin, Sam, et al.
Publicado: (2025)
Towards the Law of Capacity Gap in Distilling Language Models
por: Zhang, Chen, et al.
Publicado: (2023)
por: Zhang, Chen, et al.
Publicado: (2023)
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
por: Liu, Lei, et al.
Publicado: (2025)
por: Liu, Lei, et al.
Publicado: (2025)
Ejemplares similares
-
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
por: Gu, Xinran, et al.
Publicado: (2025) -
Research and Implementation of Data Enhancement Techniques for Graph Neural Networks
por: Gu, Jingzhao, et al.
Publicado: (2024) -
A Quadratic Synchronization Rule for Distributed Deep Learning
por: Gu, Xinran, et al.
Publicado: (2023) -
Understanding Nonlinear Implicit Bias via Region Counts in Input Space
por: Li, Jingwei, et al.
Publicado: (2025) -
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
por: Liu, Siyuan, et al.
Publicado: (2026)