Model Breadcrumbs: Scaling Multi-Task Model Merging with Sparse Masks
Fuente:
arXiv
Saved in:
| Main Authors: | Davari, MohammadReza, Belilovsky, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
by: Davari, MohammadReza, et al.
Published: (2025)
by: Davari, MohammadReza, et al.
Published: (2025)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026)
by: Miahi, Erfan, et al.
Published: (2026)
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)
by: Nanfack, Geraldin, et al.
Published: (2025)
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024)
by: Horoi, Stefan, et al.
Published: (2024)
From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation
by: Nanfack, Geraldin, et al.
Published: (2024)
by: Nanfack, Geraldin, et al.
Published: (2024)
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)
by: Janson, Paul, et al.
Published: (2026)
When Data Falls Short: Grokking Below the Critical Threshold
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent
by: Wei, Yongxian, et al.
Published: (2025)
by: Wei, Yongxian, et al.
Published: (2025)
AdaMerging: Adaptive Model Merging for Multi-Task Learning
by: Yang, Enneng, et al.
Published: (2023)
by: Yang, Enneng, et al.
Published: (2023)
SeriesGAN: Time Series Generation via Adversarial and Autoregressive Learning
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
AVATAR: Adversarial Autoencoders with Autoregressive Refinement for Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
ChronoGAN: Supervised and Embedded Generative Adversarial Networks for Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
Celo: Training Versatile Learned Optimizers on a Compute Diet
by: Moudgil, Abhinav, et al.
Published: (2025)
by: Moudgil, Abhinav, et al.
Published: (2025)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
DragD3D: Realistic Mesh Editing with Rigidity Control Driven by 2D Diffusion Priors
by: Xie, Tianhao, et al.
Published: (2023)
by: Xie, Tianhao, et al.
Published: (2023)
SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
by: Chapagain, Santosh, et al.
Published: (2026)
by: Chapagain, Santosh, et al.
Published: (2026)
Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
TIMED: Adversarial and Autoregressive Refinement of Diffusion-Based Time Series Generation
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
by: EskandariNasab, MohammadReza, et al.
Published: (2025)
Representation Surgery for Multi-Task Model Merging
by: Yang, Enneng, et al.
Published: (2024)
by: Yang, Enneng, et al.
Published: (2024)
Fisher Mask Nodes for Language Model Merging
by: K, Thennal D, et al.
Published: (2024)
by: K, Thennal D, et al.
Published: (2024)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
Enhancing Multivariate Time Series-based Solar Flare Prediction with Multifaceted Preprocessing and Contrastive Learning
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
by: EskandariNasab, MohammadReza, et al.
Published: (2024)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
by: Aakanksha, et al.
Published: (2024)
by: Aakanksha, et al.
Published: (2024)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
MOMA: Masked Orthogonal Matrix Alignment for Zero-Additional-Parameter Model Merging
by: Kong, Fanshuang, et al.
Published: (2024)
by: Kong, Fanshuang, et al.
Published: (2024)
Heterogeneous Low-Bandwidth Pre-Training of LLMs
by: Obeidi, Yazan, et al.
Published: (2026)
by: Obeidi, Yazan, et al.
Published: (2026)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces
by: Marczak, Daniel, et al.
Published: (2025)
by: Marczak, Daniel, et al.
Published: (2025)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Superpose Task-specific Features for Model Merging
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Pareto Merging: Multi-Objective Optimization for Preference-Aware Model Merging
by: Chen, Weiyu, et al.
Published: (2024)
by: Chen, Weiyu, et al.
Published: (2024)
Similar Items
-
Rethinking Prompt Optimization: Reinforcement, Diversification, and Migration in Blackbox LLMs
by: Davari, MohammadReza, et al.
Published: (2025) -
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024) -
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025) -
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026) -
FairDropout: Using Example-Tied Dropout to Enhance Generalization of Minority Groups
by: Nanfack, Geraldin, et al.
Published: (2025)