Locking Pretrained Weights via Deep Low-Rank Residual Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Sakamoto, Keitaro, Ablin, Pierre, Danieli, Federico, Cuturi, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026)
by: Gualdoni, Eleonora, et al.
Published: (2026)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Scaling Categorical Flow Maps
by: Davis, Oscar, et al.
Published: (2026)
by: Davis, Oscar, et al.
Published: (2026)
Learning Elastic Costs to Shape Monge Displacements
by: Klein, Michal, et al.
Published: (2023)
by: Klein, Michal, et al.
Published: (2023)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Learning Unmasking Policies for Diffusion Language Models
by: Jazbec, Metod, et al.
Published: (2025)
by: Jazbec, Metod, et al.
Published: (2025)
On a Neural Implementation of Brenier's Polar Factorization
by: Vesseron, Nina, et al.
Published: (2024)
by: Vesseron, Nina, et al.
Published: (2024)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
by: Wang, Yiming, et al.
Published: (2023)
by: Wang, Yiming, et al.
Published: (2023)
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)
by: Janson, Paul, et al.
Published: (2026)
How Smooth Is Attention?
by: Castin, Valérie, et al.
Published: (2023)
by: Castin, Valérie, et al.
Published: (2023)
The AdEMAMix Optimizer: Better, Faster, Older
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
Sample and Map from a Single Convex Potential: Generation using Conjugate Moment Measures
by: Vesseron, Nina, et al.
Published: (2025)
by: Vesseron, Nina, et al.
Published: (2025)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
by: Ramapuram, Jason, et al.
Published: (2024)
by: Ramapuram, Jason, et al.
Published: (2024)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
by: Loeschcke, Sebastian, et al.
Published: (2024)
by: Loeschcke, Sebastian, et al.
Published: (2024)
Low-Rank Compression of Pretrained Models via Randomized Subspace Iteration
by: Pourkamali-Anaraki, Farhad
Published: (2026)
by: Pourkamali-Anaraki, Farhad
Published: (2026)
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
by: Ye, Zhenzhang, et al.
Published: (2024)
by: Ye, Zhenzhang, et al.
Published: (2024)
MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations
by: Heurtebise, Ambroise, et al.
Published: (2025)
by: Heurtebise, Ambroise, et al.
Published: (2025)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
The Coupling Within: Flow Matching via Distilled Normalizing Flows
by: Berthelot, David, et al.
Published: (2026)
by: Berthelot, David, et al.
Published: (2026)
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
Multivariate Conformal Prediction using Optimal Transport
by: Klein, Michal, et al.
Published: (2025)
by: Klein, Michal, et al.
Published: (2025)
Contrasting Multiple Representations with the Multi-Marginal Matching Gap
by: Piran, Zoe, et al.
Published: (2024)
by: Piran, Zoe, et al.
Published: (2024)
GENOT: Entropic (Gromov) Wasserstein Flow Matching with Applications to Single-Cell Genomics
by: Klein, Dominik, et al.
Published: (2023)
by: Klein, Dominik, et al.
Published: (2023)
Revisiting Weight Regularization for Low-Rank Continual Learning
by: Zheng, Yaoyue, et al.
Published: (2026)
by: Zheng, Yaoyue, et al.
Published: (2026)
Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining
by: Zhang, Haochen, et al.
Published: (2025)
by: Zhang, Haochen, et al.
Published: (2025)
A Specialized Semismooth Newton Method for Kernel-Based Optimal Transport
by: Lin, Tianyi, et al.
Published: (2023)
by: Lin, Tianyi, et al.
Published: (2023)
Lillama: Large Language Models Compression via Low-Rank Feature Distillation
by: Sy, Yaya, et al.
Published: (2024)
by: Sy, Yaya, et al.
Published: (2024)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Similar Items
-
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026) -
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025) -
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025) -
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026) -
Careful with that Scalpel: Improving Gradient Surgery with an EMA
by: Hsieh, Yu-Guan, et al.
Published: (2024)