Stabilizing Native Low-Rank LLM Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Janson, Paul, Oyallon, Edouard, Belilovsky, Eugene |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
von: Nabli, Adel, et al.
Veröffentlicht: (2024)
von: Nabli, Adel, et al.
Veröffentlicht: (2024)
Model Parallelism With Subnetwork Data Parallelism
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
PETRA: Parallel End-to-end Training with Reversible Architectures
von: Rivaud, Stéphane, et al.
Veröffentlicht: (2024)
von: Rivaud, Stéphane, et al.
Veröffentlicht: (2024)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
von: Thérien, Benjamin, et al.
Veröffentlicht: (2024)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2024)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
von: Obeidi, Yazan, et al.
Veröffentlicht: (2026)
von: Obeidi, Yazan, et al.
Veröffentlicht: (2026)
WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
PyLO: Towards Accessible Learned Optimizers in PyTorch
von: Janson, Paul, et al.
Veröffentlicht: (2025)
von: Janson, Paul, et al.
Veröffentlicht: (2025)
Cyclic Data Parallelism for Efficient Parallelism of Deep Neural Networks
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
DISCO: learning to DISCover an evolution Operator for multi-physics-agnostic prediction
von: Morel, Rudy, et al.
Veröffentlicht: (2025)
von: Morel, Rudy, et al.
Veröffentlicht: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
von: Miahi, Erfan, et al.
Veröffentlicht: (2026)
von: Miahi, Erfan, et al.
Veröffentlicht: (2026)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2025)
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2025)
Efficient Refusal Ablation in LLM through Optimal Transport
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2026)
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2026)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
von: Legate, Gwen, et al.
Veröffentlicht: (2025)
von: Legate, Gwen, et al.
Veröffentlicht: (2025)
Unbiased Approximate Vector-Jacobian Products for Efficient Backpropagation
von: Bakong, Killian, et al.
Veröffentlicht: (2026)
von: Bakong, Killian, et al.
Veröffentlicht: (2026)
Communication Efficient LLM Pre-training with SparseLoCo
von: Sarfi, Amir, et al.
Veröffentlicht: (2025)
von: Sarfi, Amir, et al.
Veröffentlicht: (2025)
Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
von: Davari, MohammadReza, et al.
Veröffentlicht: (2023)
von: Davari, MohammadReza, et al.
Veröffentlicht: (2023)
Test-time Generalization for Physics through Neural Operator Splitting
von: Serrano, Louis, et al.
Veröffentlicht: (2026)
von: Serrano, Louis, et al.
Veröffentlicht: (2026)
When Data Falls Short: Grokking Below the Critical Threshold
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
von: Singh, Vaibhav, et al.
Veröffentlicht: (2024)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2024)
Celo2: Towards Learned Optimization Free Lunch
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2026)
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2026)
Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
On Non-Linear operators for Geometric Deep Learning
von: Sergeant-Perthuis, Grégoire, et al.
Veröffentlicht: (2022)
von: Sergeant-Perthuis, Grégoire, et al.
Veröffentlicht: (2022)
Celo: Training Versatile Learned Optimizers on a Compute Diet
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2025)
von: Moudgil, Abhinav, et al.
Veröffentlicht: (2025)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
von: Hameed, Humza Wajid, et al.
Veröffentlicht: (2024)
von: Hameed, Humza Wajid, et al.
Veröffentlicht: (2024)
Towards motion from video diffusion models
von: Janson, Paul, et al.
Veröffentlicht: (2024)
von: Janson, Paul, et al.
Veröffentlicht: (2024)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
von: Xie, Tianhao, et al.
Veröffentlicht: (2023)
von: Xie, Tianhao, et al.
Veröffentlicht: (2023)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2024)
von: Nanfack, Geraldin, et al.
Veröffentlicht: (2024)
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
von: Davari, MohammadReza, et al.
Veröffentlicht: (2025)
von: Davari, MohammadReza, et al.
Veröffentlicht: (2025)
LoQT: Low-Rank Adapters for Quantized Pretraining
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
von: Loeschcke, Sebastian, et al.
Veröffentlicht: (2024)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2026)
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2026)
Less is More: Undertraining Experts Improves Model Upcycling
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
von: Horoi, Stefan, et al.
Veröffentlicht: (2025)
Meta-learning Optimizers for Communication-Efficient Learning
von: Joseph, Charles-Étienne, et al.
Veröffentlicht: (2023)
von: Joseph, Charles-Étienne, et al.
Veröffentlicht: (2023)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
Non-Uniform Parameter-Wise Model Merging
von: Camacho, Albert Manuel Orozco, et al.
Veröffentlicht: (2024)
von: Camacho, Albert Manuel Orozco, et al.
Veröffentlicht: (2024)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
AdaFisher: Adaptive Second Order Optimization via Fisher Information
von: Gomes, Damien Martins, et al.
Veröffentlicht: (2024)
von: Gomes, Damien Martins, et al.
Veröffentlicht: (2024)
Accelerating Training with Neuron Interaction and Nowcasting Networks
von: Knyazev, Boris, et al.
Veröffentlicht: (2024)
von: Knyazev, Boris, et al.
Veröffentlicht: (2024)
Test Time Adaptation Using Adaptive Quantile Recalibration
von: Mehrbod, Paria, et al.
Veröffentlicht: (2025)
von: Mehrbod, Paria, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
von: Nabli, Adel, et al.
Veröffentlicht: (2024) -
Model Parallelism With Subnetwork Data Parallelism
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025) -
PETRA: Parallel End-to-end Training with Reversible Architectures
von: Rivaud, Stéphane, et al.
Veröffentlicht: (2024) -
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
von: Thérien, Benjamin, et al.
Veröffentlicht: (2024) -
Heterogeneous Low-Bandwidth Pre-Training of LLMs
von: Obeidi, Yazan, et al.
Veröffentlicht: (2026)