Saved in:
| Main Authors: | Yu, Shi Jie, Choi, Sehyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.18580 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Weight Averaging for Model Merging
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging
by: Choi, Minsik, et al.
Published: (2026)
by: Choi, Minsik, et al.
Published: (2026)
Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
by: Joung, Youngju, et al.
Published: (2025)
by: Joung, Youngju, et al.
Published: (2025)
Weight Weaving: Parameter Pooling for Data-Free Model Merging
by: Chaves, Levy, et al.
Published: (2025)
by: Chaves, Levy, et al.
Published: (2025)
SeWA: Selective Weight Average via Probabilistic Masking
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Sample Weight Averaging for Stable Prediction
by: Yu, Han, et al.
Published: (2025)
by: Yu, Han, et al.
Published: (2025)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
by: Li, Wenshuo, et al.
Published: (2024)
by: Li, Wenshuo, et al.
Published: (2024)
Generalizing the Geometry of Model Merging Through Frechet Averages
by: da Silva, Marvin F., et al.
Published: (2026)
by: da Silva, Marvin F., et al.
Published: (2026)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Adaptive Stochastic Weight Averaging
by: Demir, Caglar, et al.
Published: (2024)
by: Demir, Caglar, et al.
Published: (2024)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
OLR-WAA: Adaptive and Drift-Resilient Online Regression with Dynamic Weighted Averaging
by: Abu-Shaira, Mohammad, et al.
Published: (2025)
by: Abu-Shaira, Mohammad, et al.
Published: (2025)
Parameter Averaging in Link Prediction
by: Sapkota, Rupesh, et al.
Published: (2025)
by: Sapkota, Rupesh, et al.
Published: (2025)
NegMerge: Sign-Consensual Weight Merging for Machine Unlearning
by: Kim, Hyo Seo, et al.
Published: (2024)
by: Kim, Hyo Seo, et al.
Published: (2024)
Class-Wise Federated Averaging for Efficient Personalization
by: Lee, Gyuejeong, et al.
Published: (2024)
by: Lee, Gyuejeong, et al.
Published: (2024)
When, Where and Why to Average Weights?
by: Ajroldi, Niccolò, et al.
Published: (2025)
by: Ajroldi, Niccolò, et al.
Published: (2025)
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
by: Gauthier-Caron, Thomas, et al.
Published: (2024)
Optimizing the Optimal Weighted Average: Efficient Distributed Sparse Classification
by: Lu, Fred, et al.
Published: (2024)
by: Lu, Fred, et al.
Published: (2024)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
by: Miao, Ruijie, et al.
Published: (2025)
by: Miao, Ruijie, et al.
Published: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
Efficient Estimation for Longitudinal Networks via Adaptive Merging
by: Zhang, Haoran, et al.
Published: (2022)
by: Zhang, Haoran, et al.
Published: (2022)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Implicit Contrastive Representation Learning with Guided Stop-gradient
by: Lee, Byeongchan, et al.
Published: (2025)
by: Lee, Byeongchan, et al.
Published: (2025)
OLC-WA: Drift Aware Tuning-Free Online Classification with Weighted Average
by: Shaira, Mohammad Abu, et al.
Published: (2025)
by: Shaira, Mohammad Abu, et al.
Published: (2025)
Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging
by: Shen, Li, et al.
Published: (2024)
by: Shen, Li, et al.
Published: (2024)
How to Weight Multitask Finetuning? Fast Previews via Bayesian Model-Merging
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
by: Maldonado, Hugo Monzón, et al.
Published: (2024)
Interaction-Aware Influence Functions for Group Attribution
by: Heo, Jaeseung, et al.
Published: (2026)
by: Heo, Jaeseung, et al.
Published: (2026)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Trainable Weight Averaging: Accelerating Training and Improving Generalization
by: Li, Tao, et al.
Published: (2022)
by: Li, Tao, et al.
Published: (2022)
Average Certified Radius is a Poor Metric for Randomized Smoothing
by: Sun, Chenhao, et al.
Published: (2024)
by: Sun, Chenhao, et al.
Published: (2024)
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Multi-Task Model Merging via Adaptive Weight Disentanglement
by: Xiong, Feng, et al.
Published: (2024)
by: Xiong, Feng, et al.
Published: (2024)
Orthogonal Model Merging
by: Yang, Sihan, et al.
Published: (2026)
by: Yang, Sihan, et al.
Published: (2026)
MergeIT: From Selection to Merging for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
FedMerge: Federated Personalization via Model Merging
by: Chen, Shutong, et al.
Published: (2025)
by: Chen, Shutong, et al.
Published: (2025)
Dataless Knowledge Fusion by Merging Weights of Language Models
by: Jin, Xisen, et al.
Published: (2022)
by: Jin, Xisen, et al.
Published: (2022)
Similar Items
-
Revisiting Weight Averaging for Model Merging
by: Choi, Jiho, et al.
Published: (2024) -
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024) -
Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging
by: Choi, Minsik, et al.
Published: (2026) -
Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
by: Joung, Youngju, et al.
Published: (2025) -
Weight Weaving: Parameter Pooling for Data-Free Model Merging
by: Chaves, Levy, et al.
Published: (2025)