NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Si, Chongjie, Lv, Kangtao, Jiang, Jingjing, Wang, Yadao, Wang, Yongwei, Yang, Xiaokang, Su, Wenbo, Zheng, Bo, Shen, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Weight Spectra Induced Efficient Model Adaptation
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
von: Lv, Kangtao, et al.
Veröffentlicht: (2025)
Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
von: Lv, Kangtao, et al.
Veröffentlicht: (2026)
von: Lv, Kangtao, et al.
Veröffentlicht: (2026)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
von: Liu, Langming, et al.
Veröffentlicht: (2026)
von: Liu, Langming, et al.
Veröffentlicht: (2026)
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
Tendency-driven Mutual Exclusivity for Weakly Supervised Incremental Semantic Segmentation
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
von: Chen, Haibin, et al.
Veröffentlicht: (2025)
von: Chen, Haibin, et al.
Veröffentlicht: (2025)
ACE-Merging: Data-Free Model Merging with Adaptive Covariance Estimation
von: Xu, Bo, et al.
Veröffentlicht: (2026)
von: Xu, Bo, et al.
Veröffentlicht: (2026)
HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs
von: Cao, Huangsen, et al.
Veröffentlicht: (2024)
von: Cao, Huangsen, et al.
Veröffentlicht: (2024)
Appeal: Allow Mislabeled Samples the Chance to be Rectified in Partial Label Learning
von: Si, Chongjie, et al.
Veröffentlicht: (2023)
von: Si, Chongjie, et al.
Veröffentlicht: (2023)
Revisiting Sparsity Constraint Under High-Rank Property in Partial Multi-Label Learning
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Why Can Accurate Models Be Learned from Inaccurate Annotations?
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
von: Wang, Zijing, et al.
Veröffentlicht: (2026)
von: Wang, Zijing, et al.
Veröffentlicht: (2026)
Generalized Tensor-based Parameter-Efficient Fine-Tuning via Lie Group Transformations
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Transport and Merge: Cross-Architecture Merging for Large Language Models
von: Cui, Chenhang, et al.
Veröffentlicht: (2026)
von: Cui, Chenhang, et al.
Veröffentlicht: (2026)
BARD: budget-aware reasoning distillation
von: Niu, Lujie, et al.
Veröffentlicht: (2025)
von: Niu, Lujie, et al.
Veröffentlicht: (2025)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
von: Du, Yiyang, et al.
Veröffentlicht: (2025)
Training-free LLM Merging for Multi-task Learning
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
SuperMerge: An Approach For Gradient-Based Model Merging
von: Yang, Haoyu, et al.
Veröffentlicht: (2024)
von: Yang, Haoyu, et al.
Veröffentlicht: (2024)
Multi-Stage Evolutionary Model Merging with Meta Data Driven Curriculum Learning for Sentiment-Specialized Large Language Modeling
von: Inoshita, Keito, et al.
Veröffentlicht: (2026)
von: Inoshita, Keito, et al.
Veröffentlicht: (2026)
E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring
von: Wang, Wenjun, et al.
Veröffentlicht: (2026)
von: Wang, Wenjun, et al.
Veröffentlicht: (2026)
SeMe: Training-Free Language Model Merging via Semantic Alignment
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
von: Wang, Zheng, et al.
Veröffentlicht: (2024)
von: Wang, Zheng, et al.
Veröffentlicht: (2024)
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
von: Si, Wai Man, et al.
Veröffentlicht: (2026)
von: Si, Wai Man, et al.
Veröffentlicht: (2026)
MTA: A Merge-then-Adapt Framework for Personalized Large Language Model
von: Li, Xiaopeng, et al.
Veröffentlicht: (2025)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2025)
Hyper Adversarial Tuning for Boosting Adversarial Robustness of Pretrained Large Vision Models
von: Lv, Kangtao, et al.
Veröffentlicht: (2024)
von: Lv, Kangtao, et al.
Veröffentlicht: (2024)
FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
LEWIS (LayEr WIse Sparsity) -- A Training Free Guided Model Merging Approach
von: Chopra, Hetarth, et al.
Veröffentlicht: (2025)
von: Chopra, Hetarth, et al.
Veröffentlicht: (2025)
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
von: Wu, Yu-Hang, et al.
Veröffentlicht: (2026)
von: Wu, Yu-Hang, et al.
Veröffentlicht: (2026)
ProgCo: Program Helps Self-Correction of Large Language Models
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2025)
On the Limits of Model Merging for Multilinguality in Pre-Training
von: Aycock, Seth, et al.
Veröffentlicht: (2026)
von: Aycock, Seth, et al.
Veröffentlicht: (2026)
AdaMuon: Adaptive Muon Optimizer
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
Training-free LLM-generated Text Detection by Mining Token Probability Sequences
von: Xu, Yihuai, et al.
Veröffentlicht: (2024)
von: Xu, Yihuai, et al.
Veröffentlicht: (2024)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
von: Si, Chongjie, et al.
Veröffentlicht: (2025) -
Weight Spectra Induced Efficient Model Adaptation
von: Si, Chongjie, et al.
Veröffentlicht: (2025) -
How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
von: Lv, Kangtao, et al.
Veröffentlicht: (2025) -
Unveiling the Mystery of Weight in Large Foundation Models: Gaussian Distribution Never Fades
von: Si, Chongjie, et al.
Veröffentlicht: (2025) -
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
von: Lv, Kangtao, et al.
Veröffentlicht: (2026)