DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
Fuente:
arXiv
Saved in:
| Main Authors: | Gualdoni, Eleonora, Laguna, Sonia, Bethune, Louis, Monteiro, Joao, Ablin, Pierre, Cuturi, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Sample and Map from a Single Convex Potential: Generation using Conjugate Moment Measures
by: Vesseron, Nina, et al.
Published: (2025)
by: Vesseron, Nina, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Scaling Categorical Flow Maps
by: Davis, Oscar, et al.
Published: (2026)
by: Davis, Oscar, et al.
Published: (2026)
Learning Unmasking Policies for Diffusion Language Models
by: Jazbec, Metod, et al.
Published: (2025)
by: Jazbec, Metod, et al.
Published: (2025)
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
Multivariate Conformal Prediction using Optimal Transport
by: Klein, Michal, et al.
Published: (2025)
by: Klein, Michal, et al.
Published: (2025)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
by: Sakamoto, Keitaro, et al.
Published: (2026)
by: Sakamoto, Keitaro, et al.
Published: (2026)
HyperTransport: Amortized Conditioning of T2I Generative Models
by: Maiorca, Valentino, et al.
Published: (2026)
by: Maiorca, Valentino, et al.
Published: (2026)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
by: Jeha, Paul, et al.
Published: (2026)
by: Jeha, Paul, et al.
Published: (2026)
Learning Elastic Costs to Shape Monge Displacements
by: Klein, Michal, et al.
Published: (2023)
by: Klein, Michal, et al.
Published: (2023)
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
by: Filippova, Anastasiia, et al.
Published: (2026)
by: Filippova, Anastasiia, et al.
Published: (2026)
Amortizing Maximum Inner Product Search with Learned Support Functions
by: Olausson, Theo X., et al.
Published: (2026)
by: Olausson, Theo X., et al.
Published: (2026)
Why do objects have many names? A study on word informativeness in language use and lexical systems
by: Gualdoni, Eleonora, et al.
Published: (2024)
by: Gualdoni, Eleonora, et al.
Published: (2024)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
On a Neural Implementation of Brenier's Polar Factorization
by: Vesseron, Nina, et al.
Published: (2024)
by: Vesseron, Nina, et al.
Published: (2024)
Infeasible Deterministic, Stochastic, and Variance-Reduction Algorithms for Optimization under Orthogonality Constraints
by: Ablin, Pierre, et al.
Published: (2023)
by: Ablin, Pierre, et al.
Published: (2023)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
by: Zhou, Huichi, et al.
Published: (2025)
by: Zhou, Huichi, et al.
Published: (2025)
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
by: Rodriguez, Pau, et al.
Published: (2025)
by: Rodriguez, Pau, et al.
Published: (2025)
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023)
by: Castin, Valérie, et al.
Published: (2023)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
The Design Space of Tri-Modal Masked Diffusion Models
by: Bethune, Louis, et al.
Published: (2026)
by: Bethune, Louis, et al.
Published: (2026)
Estimating Player Performance in Different Contexts Using Fine-tuned Large Events Models
by: Mendes-Neves, Tiago, et al.
Published: (2024)
by: Mendes-Neves, Tiago, et al.
Published: (2024)
Deep Sturm--Liouville: From Sample-Based to 1D Regularization with Learnable Orthogonal Basis Functions
by: Vigouroux, David, et al.
Published: (2025)
by: Vigouroux, David, et al.
Published: (2025)
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
by: Tran, Van-Tuan, et al.
Published: (2026)
by: Tran, Van-Tuan, et al.
Published: (2026)
Counter-Dyna: Data-Efficient RL-Based HVAC Control using Counterfactual Building Models
by: de Vargas, Jan Marco Ruiz, et al.
Published: (2026)
by: de Vargas, Jan Marco Ruiz, et al.
Published: (2026)
Topic Modeling with Fine-tuning LLMs and Bag of Sentences
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Activated LoRA: Fine-tuned LLMs for Intrinsics
by: Greenewald, Kristjan, et al.
Published: (2025)
by: Greenewald, Kristjan, et al.
Published: (2025)
Functional-level Uncertainty Quantification for Calibrated Fine-tuning on LLMs
by: Niu, Ruijia, et al.
Published: (2024)
by: Niu, Ruijia, et al.
Published: (2024)
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
by: Ye, Zhenzhang, et al.
Published: (2024)
by: Ye, Zhenzhang, et al.
Published: (2024)
MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations
by: Heurtebise, Ambroise, et al.
Published: (2025)
by: Heurtebise, Ambroise, et al.
Published: (2025)
A Specialized Semismooth Newton Method for Kernel-Based Optimal Transport
by: Lin, Tianyi, et al.
Published: (2023)
by: Lin, Tianyi, et al.
Published: (2023)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
Similar Items
-
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025) -
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
by: Kirchhof, Michael, et al.
Published: (2024) -
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026) -
Sample and Map from a Single Convex Potential: Generation using Conjugate Moment Measures
by: Vesseron, Nina, et al.
Published: (2025) -
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)