Saved in:
| Main Authors: | Pagliardini, Matteo, Ablin, Pierre, Grangier, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.03137 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023)
by: Castin, Valérie, et al.
Published: (2023)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
A Primal-Dual Approach to Solving Variational Inequalities with General Constraints
by: Chavdarova, Tatjana, et al.
Published: (2022)
by: Chavdarova, Tatjana, et al.
Published: (2022)
Infeasible Deterministic, Stochastic, and Variance-Reduction Algorithms for Optimization under Orthogonality Constraints
by: Ablin, Pierre, et al.
Published: (2023)
by: Ablin, Pierre, et al.
Published: (2023)
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
by: Ye, Zhenzhang, et al.
Published: (2024)
by: Ye, Zhenzhang, et al.
Published: (2024)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
by: Sakamoto, Keitaro, et al.
Published: (2026)
by: Sakamoto, Keitaro, et al.
Published: (2026)
MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations
by: Heurtebise, Ambroise, et al.
Published: (2025)
by: Heurtebise, Ambroise, et al.
Published: (2025)
Optimization without Retraction on the Random Generalized Stiefel Manifold
by: Vary, Simon, et al.
Published: (2024)
by: Vary, Simon, et al.
Published: (2024)
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)
by: González, Ramón Calvo, et al.
Published: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
A framework for bilevel optimization that enables stochastic and global variance reduction algorithms
by: Dagréou, Mathieu, et al.
Published: (2022)
by: Dagréou, Mathieu, et al.
Published: (2022)
A Lower Bound and a Near-Optimal Algorithm for Bilevel Empirical Risk Minimization
by: Dagréou, Mathieu, et al.
Published: (2023)
by: Dagréou, Mathieu, et al.
Published: (2023)
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
by: Castin, Valérie, et al.
Published: (2026)
by: Castin, Valérie, et al.
Published: (2026)
ANO : Faster is Better in Noisy Landscape
by: Kegreisz, Adrien
Published: (2025)
by: Kegreisz, Adrien
Published: (2025)
FERRET: Private Deep Learning Faster And Better Than DPSGD
by: Zagardo, David
Published: (2025)
by: Zagardo, David
Published: (2025)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
by: Filippova, Anastasiia, et al.
Published: (2026)
by: Filippova, Anastasiia, et al.
Published: (2026)
Partial Parameter Updates for Efficient Distributed Training
by: Filippova, Anastasiia, et al.
Published: (2025)
by: Filippova, Anastasiia, et al.
Published: (2025)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Playing Markov Games Without Observing Payoffs
by: Ablin, Daniel, et al.
Published: (2025)
by: Ablin, Daniel, et al.
Published: (2025)
Multivector Neurons: Better and Faster O(n)-Equivariant Clifford Graph Neural Networks
by: Liu, Cong, et al.
Published: (2024)
by: Liu, Cong, et al.
Published: (2024)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026)
by: Gualdoni, Eleonora, et al.
Published: (2026)
Faster Predictive Coding Networks via Better Initialization
by: Pinchetti, Luca, et al.
Published: (2026)
by: Pinchetti, Luca, et al.
Published: (2026)
FREE: Faster and Better Data-Free Meta-Learning
by: Wei, Yongxian, et al.
Published: (2024)
by: Wei, Yongxian, et al.
Published: (2024)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition
by: Basharin, Artem, et al.
Published: (2024)
by: Basharin, Artem, et al.
Published: (2024)
Sparser, Better, Faster, Stronger: Sparsity Detection for Efficient Automatic Differentiation
by: Hill, Adrian, et al.
Published: (2025)
by: Hill, Adrian, et al.
Published: (2025)
Similar Items
-
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025) -
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024) -
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024) -
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024) -
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)