Enregistré dans:
| Auteurs principaux: | Qiu, Haiyun, Wu, Xingyu, Feng, Liang, Tan, Kay Chen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.06552 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards Adaptive Continual Model Merging via Manifold-Aware Expert Evolution
par: Qiu, Haiyun, et autres
Publié: (2026)
par: Qiu, Haiyun, et autres
Publié: (2026)
HM3: Hierarchical Multi-Objective Model Merging for Pretrained Models
par: Zhou, Yu, et autres
Publié: (2024)
par: Zhou, Yu, et autres
Publié: (2024)
Design Principle Transfer in Neural Architecture Search via Large Language Models
par: Zhou, Xun, et autres
Publié: (2024)
par: Zhou, Xun, et autres
Publié: (2024)
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
par: Zhang, Dengming, et autres
Publié: (2025)
par: Zhang, Dengming, et autres
Publié: (2025)
FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models
par: Huang, Chenyu, et autres
Publié: (2026)
par: Huang, Chenyu, et autres
Publié: (2026)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
par: Zhou, Yu, et autres
Publié: (2024)
par: Zhou, Yu, et autres
Publié: (2024)
Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models
par: Wang, Yuxiao, et autres
Publié: (2025)
par: Wang, Yuxiao, et autres
Publié: (2025)
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts
par: Morrison, Jacob, et autres
Publié: (2026)
par: Morrison, Jacob, et autres
Publié: (2026)
RainSeer: Fine-Grained Rainfall Reconstruction via Physics-Guided Modeling
par: Chen, Lin, et autres
Publié: (2025)
par: Chen, Lin, et autres
Publié: (2025)
Diversity-Aware Policy Optimization for Large Language Model Reasoning
par: Yao, Jian, et autres
Publié: (2025)
par: Yao, Jian, et autres
Publié: (2025)
LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery
par: Wu, Xingyu, et autres
Publié: (2025)
par: Wu, Xingyu, et autres
Publié: (2025)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
par: Nguyen, Dung V., et autres
Publié: (2025)
par: Nguyen, Dung V., et autres
Publié: (2025)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
par: Lu, Zhenyi, et autres
Publié: (2024)
par: Lu, Zhenyi, et autres
Publié: (2024)
CAMEx: Curvature-aware Merging of Experts
par: Nguyen, Dung V., et autres
Publié: (2025)
par: Nguyen, Dung V., et autres
Publié: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
par: Miao, Ruijie, et autres
Publié: (2025)
par: Miao, Ruijie, et autres
Publié: (2025)
Large Language Model-Enhanced Algorithm Selection: Towards Comprehensive Algorithm Representation
par: Wu, Xingyu, et autres
Publié: (2023)
par: Wu, Xingyu, et autres
Publié: (2023)
How Multimodal Integration Boost the Performance of LLM for Optimization: Case Study on Capacitated Vehicle Routing Problems
par: Huang, Yuxiao, et autres
Publié: (2024)
par: Huang, Yuxiao, et autres
Publié: (2024)
Vanishing Feature: Diagnosing Model Merging and Beyond
par: Qu, Xingyu, et autres
Publié: (2024)
par: Qu, Xingyu, et autres
Publié: (2024)
Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE
par: Chen, Yuanteng, et autres
Publié: (2026)
par: Chen, Yuanteng, et autres
Publié: (2026)
MIN-Merging: Merge the Important Neurons for Model Merging
par: Liang, Yunfei
Publié: (2025)
par: Liang, Yunfei
Publié: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
par: Park, Sejik
Publié: (2024)
par: Park, Sejik
Publié: (2024)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
par: Chen, Zhaoyang, et autres
Publié: (2025)
par: Chen, Zhaoyang, et autres
Publié: (2025)
FedMerge: Federated Personalization via Model Merging
par: Chen, Shutong, et autres
Publié: (2025)
par: Chen, Shutong, et autres
Publié: (2025)
Why Do More Experts Fail? A Theoretical Analysis of Model Merging
par: Wang, Zijing, et autres
Publié: (2025)
par: Wang, Zijing, et autres
Publié: (2025)
Soft Merging of Experts with Adaptive Routing
par: Muqeeth, Mohammed, et autres
Publié: (2023)
par: Muqeeth, Mohammed, et autres
Publié: (2023)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
par: Li, Lujun, et autres
Publié: (2025)
par: Li, Lujun, et autres
Publié: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
par: Zhao, Adrian, et autres
Publié: (2026)
par: Zhao, Adrian, et autres
Publié: (2026)
MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
par: Qiu, Zihuan, et autres
Publié: (2025)
par: Qiu, Zihuan, et autres
Publié: (2025)
Channel Merging: Preserving Specialization for Merged Experts
par: Zhang, Mingyang, et autres
Publié: (2024)
par: Zhang, Mingyang, et autres
Publié: (2024)
Unlock the Power of Algorithm Features: A Generalization Analysis for Algorithm Selection
par: Wu, Xingyu, et autres
Publié: (2024)
par: Wu, Xingyu, et autres
Publié: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
par: Krajewski, Jakub, et autres
Publié: (2024)
par: Krajewski, Jakub, et autres
Publié: (2024)
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
par: Tang, Anke, et autres
Publié: (2024)
par: Tang, Anke, et autres
Publié: (2024)
Superpose Task-specific Features for Model Merging
par: Qiu, Haiquan, et autres
Publié: (2025)
par: Qiu, Haiquan, et autres
Publié: (2025)
CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
par: Feng, Jinyuan, et autres
Publié: (2025)
par: Feng, Jinyuan, et autres
Publié: (2025)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
par: Bertolissi, Ryo, et autres
Publié: (2025)
par: Bertolissi, Ryo, et autres
Publié: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
par: Hui, Tingfeng, et autres
Publié: (2024)
par: Hui, Tingfeng, et autres
Publié: (2024)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
par: Zhao, Yushu, et autres
Publié: (2025)
par: Zhao, Yushu, et autres
Publié: (2025)
Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging
par: Wang, Yuanyi, et autres
Publié: (2026)
par: Wang, Yuanyi, et autres
Publié: (2026)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
par: Jain, Swayambhoo, et autres
Publié: (2024)
par: Jain, Swayambhoo, et autres
Publié: (2024)
Can Muon Fine-tune Adam-Pretrained Models?
par: Qu, Xingyu, et autres
Publié: (2026)
par: Qu, Xingyu, et autres
Publié: (2026)
Documents similaires
-
Towards Adaptive Continual Model Merging via Manifold-Aware Expert Evolution
par: Qiu, Haiyun, et autres
Publié: (2026) -
HM3: Hierarchical Multi-Objective Model Merging for Pretrained Models
par: Zhou, Yu, et autres
Publié: (2024) -
Design Principle Transfer in Neural Architecture Search via Large Language Models
par: Zhou, Xun, et autres
Publié: (2024) -
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
par: Zhang, Dengming, et autres
Publié: (2025) -
FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models
par: Huang, Chenyu, et autres
Publié: (2026)