Merging Text Transformer Models from Different Initializations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Verma, Neha, Elbayad, Maha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Merging Feed-Forward Sublayers for Compressed Transformers
von: Verma, Neha, et al.
Veröffentlicht: (2025)
von: Verma, Neha, et al.
Veröffentlicht: (2025)
Merge to Mix: Mixing Datasets via Model Merging
von: Tao, Zhixu Silvia, et al.
Veröffentlicht: (2025)
von: Tao, Zhixu Silvia, et al.
Veröffentlicht: (2025)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
Arcee's MergeKit: A Toolkit for Merging Large Language Models
von: Goddard, Charles, et al.
Veröffentlicht: (2024)
von: Goddard, Charles, et al.
Veröffentlicht: (2024)
Model Merging for Knowledge Editing
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
von: Shenaj, Donald, et al.
Veröffentlicht: (2025)
von: Shenaj, Donald, et al.
Veröffentlicht: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
What Matters for Model Merging at Scale?
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
Dynamic Model Merging Made Slim
von: Du, Guodong, et al.
Veröffentlicht: (2026)
von: Du, Guodong, et al.
Veröffentlicht: (2026)
Channel Merging: Preserving Specialization for Merged Experts
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Pourreza, Mohammadreza, et al.
Veröffentlicht: (2024)
Fisher Mask Nodes for Language Model Merging
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
STAR: Spectral Truncation and Rescale for Model Merging
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2025)
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2025)
Model Merging by Uncertainty-Based Gradient Matching
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
von: Daheim, Nico, et al.
Veröffentlicht: (2023)
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
von: Gauthier-Caron, Thomas, et al.
Veröffentlicht: (2024)
von: Gauthier-Caron, Thomas, et al.
Veröffentlicht: (2024)
Optimal Brain Iterative Merging: Mitigating Interference in LLM Merging
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
Model Assembly Learning with Heterogeneous Layer Weight Merging
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
Merge-of-Thought Distillation
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
Data-driven Clustering and Merging of Adapters for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model Merging
von: Zhang, Haobo, et al.
Veröffentlicht: (2025)
von: Zhang, Haobo, et al.
Veröffentlicht: (2025)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
Linear Model Merging Unlocks Simple and Scalable Multimodal Data Mixture Optimization
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
von: Berasi, Davide, et al.
Veröffentlicht: (2026)
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
von: Hu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2025)
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Training-free LLM Merging for Multi-task Learning
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
von: Luo, Yingfeng, et al.
Veröffentlicht: (2025)
DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
von: Verma, Neha, et al.
Veröffentlicht: (2025)
von: Verma, Neha, et al.
Veröffentlicht: (2025)
Context-Aware Initialization for Reducing Generative Path Length in Diffusion Language Models
von: Miao, Tongyuan, et al.
Veröffentlicht: (2025)
von: Miao, Tongyuan, et al.
Veröffentlicht: (2025)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
von: Li, Pingzhi, et al.
Veröffentlicht: (2023)
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
von: Tang, Yao, et al.
Veröffentlicht: (2026)
von: Tang, Yao, et al.
Veröffentlicht: (2026)
Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
von: Glocker, Kevin, et al.
Veröffentlicht: (2025)
von: Glocker, Kevin, et al.
Veröffentlicht: (2025)
Transformer Explainer: Interactive Learning of Text-Generative Models
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
EvoMerge: Neuroevolution for Large Language Models
von: Jiang, Yushu
Veröffentlicht: (2024)
von: Jiang, Yushu
Veröffentlicht: (2024)
Random Initialization of Gated Sparse Adapters
von: Retault, Vi, et al.
Veröffentlicht: (2025)
von: Retault, Vi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Merging Feed-Forward Sublayers for Compressed Transformers
von: Verma, Neha, et al.
Veröffentlicht: (2025) -
Merge to Mix: Mixing Datasets via Model Merging
von: Tao, Zhixu Silvia, et al.
Veröffentlicht: (2025) -
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024) -
Arcee's MergeKit: A Toolkit for Merging Large Language Models
von: Goddard, Charles, et al.
Veröffentlicht: (2024) -
Model Merging for Knowledge Editing
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)