Less is More: Undertraining Experts Improves Model Upcycling
Fuente:
arXiv
Saved in:
| Main Authors: | Horoi, Stefan, Wolf, Guy, Belilovsky, Eugene, Dziugaite, Gintare Karolina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024)
by: Horoi, Stefan, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
by: Ohib, Riyasat, et al.
Published: (2024)
by: Ohib, Riyasat, et al.
Published: (2024)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
by: Dwivedi, Chaitanya, et al.
Published: (2026)
by: Dwivedi, Chaitanya, et al.
Published: (2026)
Upcycling Large Language Models into Mixture of Experts
by: He, Ethan, et al.
Published: (2024)
by: He, Ethan, et al.
Published: (2024)
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
by: Zhang, Michael S., et al.
Published: (2025)
by: Zhang, Michael S., et al.
Published: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Transformer Multivariate Forecasting: Less is More?
by: Xu, Jingjing, et al.
Published: (2023)
by: Xu, Jingjing, et al.
Published: (2023)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
by: Tejankar, Ajinkya, et al.
Published: (2024)
by: Tejankar, Ajinkya, et al.
Published: (2024)
Less is More: Recursive Reasoning with Tiny Networks
by: Jolicoeur-Martineau, Alexia
Published: (2025)
by: Jolicoeur-Martineau, Alexia
Published: (2025)
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
Less is More: on the Over-Globalizing Problem in Graph Transformers
by: Xing, Yujie, et al.
Published: (2024)
by: Xing, Yujie, et al.
Published: (2024)
Cut Less, Fold More: Model Compression through the Lens of Projection Geometry
by: Saukh, Olga, et al.
Published: (2026)
by: Saukh, Olga, et al.
Published: (2026)
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
by: Zhao, Lifan, et al.
Published: (2025)
by: Zhao, Lifan, et al.
Published: (2025)
Towards a General Recipe for Combinatorial Optimization with Multi-Filter GNNs
by: Wenkel, Frederik, et al.
Published: (2024)
by: Wenkel, Frederik, et al.
Published: (2024)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
by: Shen, Yuhao, et al.
Published: (2026)
by: Shen, Yuhao, et al.
Published: (2026)
Less Is More -- On the Importance of Sparsification for Transformers and Graph Neural Networks for TSP
by: Lischka, Attila, et al.
Published: (2024)
by: Lischka, Attila, et al.
Published: (2024)
Less is More: Pseudo-Label Filtering for Continual Test-Time Adaptation
by: Tan, Jiayao, et al.
Published: (2024)
by: Tan, Jiayao, et al.
Published: (2024)
Similar Items
-
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024) -
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024) -
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025) -
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024) -
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)