Partial Parameter Updates for Efficient Distributed Training
Fuente:
arXiv
Saved in:
| Main Authors: | Filippova, Anastasiia, Katharopoulos, Angelos, Grangier, David, Collobert, Ronan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
by: Filippova, Anastasiia, et al.
Published: (2026)
by: Filippova, Anastasiia, et al.
Published: (2026)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Time-series attribution maps with regularized contrastive learning
by: Schneider, Steffen, et al.
Published: (2025)
by: Schneider, Steffen, et al.
Published: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Pretraining with hierarchical memories: separating long-tail and common knowledge
by: Pouransari, Hadi, et al.
Published: (2025)
by: Pouransari, Hadi, et al.
Published: (2025)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
by: Sabry, Mohammed, et al.
Published: (2024)
by: Sabry, Mohammed, et al.
Published: (2024)
Efficient Knowledge Deletion from Trained Models through Layer-wise Partial Machine Unlearning
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
by: Gogineni, Vinay Chakravarthi, et al.
Published: (2024)
Dynamic Layer Tying for Parameter-Efficient Transformers
by: Hay, Tamir David, et al.
Published: (2024)
by: Hay, Tamir David, et al.
Published: (2024)
Task-Aware Parameter-Efficient Fine-Tuning of Large Pre-Trained Models at the Edge
by: Hu, Senkang, et al.
Published: (2025)
by: Hu, Senkang, et al.
Published: (2025)
From Small to Large: Generalization Bounds for Transformers on Variable-Size Inputs
by: Alokhina, Anastasiia, et al.
Published: (2025)
by: Alokhina, Anastasiia, et al.
Published: (2025)
Test-time Adaptation of Tiny Recursive Models
by: McGovern, Ronan Killian
Published: (2025)
by: McGovern, Ronan Killian
Published: (2025)
Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting
by: Pan, Licheng, et al.
Published: (2025)
by: Pan, Licheng, et al.
Published: (2025)
Efficient Post-Training Augmentation for Adaptive Inference in Heterogeneous and Distributed IoT Environments
by: Sponner, Max, et al.
Published: (2024)
by: Sponner, Max, et al.
Published: (2024)
Dion: Distributed Orthonormalized Updates
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Ambient Physics: Training Neural PDE Solvers with Partial Observations
by: Majid, Harris Abdul, et al.
Published: (2026)
by: Majid, Harris Abdul, et al.
Published: (2026)
Discovering Data Manifold Geometry via Non-Contracting Flows
by: Vigouroux, David, et al.
Published: (2026)
by: Vigouroux, David, et al.
Published: (2026)
MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
by: Iacob, Alex, et al.
Published: (2025)
by: Iacob, Alex, et al.
Published: (2025)
Adversarial Latent-State Training for Robust Policies in Partially Observable Domains
by: Ahuja, Angad Singh
Published: (2026)
by: Ahuja, Angad Singh
Published: (2026)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient Update
by: Mao, Liyuan, et al.
Published: (2024)
by: Mao, Liyuan, et al.
Published: (2024)
RapidGNN: Energy and Communication-Efficient Distributed Training on Large-Scale Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025)
by: Niam, Arefin, et al.
Published: (2025)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Visual Analytics for Explainable and Trustworthy Artificial Intelligence
by: Chatzimparmpas, Angelos
Published: (2025)
by: Chatzimparmpas, Angelos
Published: (2025)
Neural Operators as Efficient Function Interpolators
by: Niarchos, Vasilis, et al.
Published: (2026)
by: Niarchos, Vasilis, et al.
Published: (2026)
PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
by: Woo, Sunghyeon, et al.
Published: (2025)
by: Woo, Sunghyeon, et al.
Published: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
by: Sedova, Anastasiia, et al.
Published: (2023)
by: Sedova, Anastasiia, et al.
Published: (2023)
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
by: Wu, Bo, et al.
Published: (2025)
by: Wu, Bo, et al.
Published: (2025)
Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation
by: Lin, Luxi, et al.
Published: (2026)
by: Lin, Luxi, et al.
Published: (2026)
MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
by: Lopez-Piqueres, Javier, et al.
Published: (2025)
by: Lopez-Piqueres, Javier, et al.
Published: (2025)
Distributional Training Data Attribution: What do Influence Functions Sample?
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
by: Wang, Runqian, et al.
Published: (2024)
by: Wang, Runqian, et al.
Published: (2024)
Online Feedback Efficient Active Target Discovery in Partially Observable Environments
by: Sarkar, Anindya, et al.
Published: (2025)
by: Sarkar, Anindya, et al.
Published: (2025)
CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
by: Kanatas, Angelos-Nikolaos, et al.
Published: (2025)
Parameter-Efficient Continual Fine-Tuning: A Survey
by: Coleman, Eric Nuertey, et al.
Published: (2025)
by: Coleman, Eric Nuertey, et al.
Published: (2025)
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
by: Bao, Wenjie, et al.
Published: (2025)
by: Bao, Wenjie, et al.
Published: (2025)
Similar Items
-
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024) -
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025) -
Stochastic KV Routing: Enabling Adaptive Depth-Wise Cache Sharing
by: Filippova, Anastasiia, et al.
Published: (2026) -
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026) -
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)