Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Haoyuan, Zheng, Haisheng, He, Zhuolun, Yu, Bei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
von: He, Zhuolun, et al.
Veröffentlicht: (2023)
von: He, Zhuolun, et al.
Veröffentlicht: (2023)
Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design Automation
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
von: Liu, Yilun, et al.
Veröffentlicht: (2024)
Architect of the Bits World: Masked Autoregressive Modeling for Circuit Generation Guided by Truth Table
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
Sparsity and Superposition in Mixture of Experts
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
von: Xie, Zhen-Hao, et al.
Veröffentlicht: (2026)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
Parameter Efficient Instruction Tuning: An Empirical Study
von: He, Pengfei
Veröffentlicht: (2024)
von: He, Pengfei
Veröffentlicht: (2024)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
von: Srinivas, Sakhinana Sagar, et al.
Veröffentlicht: (2024)
von: Srinivas, Sakhinana Sagar, et al.
Veröffentlicht: (2024)
Mixture of Experts Approaches in Dense Retrieval Tasks
von: Sokli, Effrosyni, et al.
Veröffentlicht: (2025)
von: Sokli, Effrosyni, et al.
Veröffentlicht: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Investigating Mixture of Experts in Dense Retrieval
von: Sokli, Effrosyni, et al.
Veröffentlicht: (2024)
von: Sokli, Effrosyni, et al.
Veröffentlicht: (2024)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
von: Liu, Ningyuan, et al.
Veröffentlicht: (2025)
von: Liu, Ningyuan, et al.
Veröffentlicht: (2025)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
von: Kunwar, Pradip, et al.
Veröffentlicht: (2025)
von: Kunwar, Pradip, et al.
Veröffentlicht: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Low-Resource Fine-Tuning for Multi-Task Structured Information Extraction with a Billion-Parameter Instruction-Tuned Model
von: Chih, Yu Cheng, et al.
Veröffentlicht: (2025)
von: Chih, Yu Cheng, et al.
Veröffentlicht: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fang, Zhiyuan, et al.
Veröffentlicht: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
von: Li, Houyi, et al.
Veröffentlicht: (2025)
von: Li, Houyi, et al.
Veröffentlicht: (2025)
Unsupervised Text Representation Learning via Instruction-Tuning for Zero-Shot Dense Retrieval
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2024)
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2024)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
von: Raje, Arian, et al.
Veröffentlicht: (2026)
von: Raje, Arian, et al.
Veröffentlicht: (2026)
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
von: Guo, Chenyou, et al.
Veröffentlicht: (2026)
von: Guo, Chenyou, et al.
Veröffentlicht: (2026)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
von: Du, Zhiying, et al.
Veröffentlicht: (2025)
von: Du, Zhiying, et al.
Veröffentlicht: (2025)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2026)
von: Chowdhury, Mohammed Nowaz Rabbani, et al.
Veröffentlicht: (2026)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
von: Cao, Jie, et al.
Veröffentlicht: (2025)
von: Cao, Jie, et al.
Veröffentlicht: (2025)
Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets
von: Bilican, Ahmet, et al.
Veröffentlicht: (2025)
von: Bilican, Ahmet, et al.
Veröffentlicht: (2025)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
von: Ming, Rui, et al.
Veröffentlicht: (2025)
von: Ming, Rui, et al.
Veröffentlicht: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024) -
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
von: He, Zhuolun, et al.
Veröffentlicht: (2023) -
Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design Automation
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025) -
Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025) -
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
von: Liu, Yilun, et al.
Veröffentlicht: (2024)