Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Haoyuan, Zheng, Haisheng, He, Zhuolun, Yu, Bei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
por: Hui, Tingfeng, et al.
Publicado: (2024)
por: Hui, Tingfeng, et al.
Publicado: (2024)
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
por: He, Zhuolun, et al.
Publicado: (2023)
por: He, Zhuolun, et al.
Publicado: (2023)
Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design Automation
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
por: Liu, Yilun, et al.
Publicado: (2024)
por: Liu, Yilun, et al.
Publicado: (2024)
Architect of the Bits World: Masked Autoregressive Modeling for Circuit Generation Guided by Truth Table
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
por: Zhang, Buze, et al.
Publicado: (2026)
por: Zhang, Buze, et al.
Publicado: (2026)
Sparsity and Superposition in Mixture of Experts
por: Chaudhari, Marmik, et al.
Publicado: (2025)
por: Chaudhari, Marmik, et al.
Publicado: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
por: Nakamura, Taishi, et al.
Publicado: (2025)
por: Nakamura, Taishi, et al.
Publicado: (2025)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
por: Xie, Zhen-Hao, et al.
Publicado: (2026)
por: Xie, Zhen-Hao, et al.
Publicado: (2026)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
por: Liu, Yilun, et al.
Publicado: (2025)
por: Liu, Yilun, et al.
Publicado: (2025)
Parameter Efficient Instruction Tuning: An Empirical Study
por: He, Pengfei
Publicado: (2024)
por: He, Pengfei
Publicado: (2024)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
por: Abnar, Samira, et al.
Publicado: (2025)
por: Abnar, Samira, et al.
Publicado: (2025)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
por: Srinivas, Sakhinana Sagar, et al.
Publicado: (2024)
por: Srinivas, Sakhinana Sagar, et al.
Publicado: (2024)
Mixture of Experts Approaches in Dense Retrieval Tasks
por: Sokli, Effrosyni, et al.
Publicado: (2025)
por: Sokli, Effrosyni, et al.
Publicado: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
por: Park, Jongseok, et al.
Publicado: (2026)
por: Park, Jongseok, et al.
Publicado: (2026)
Investigating Mixture of Experts in Dense Retrieval
por: Sokli, Effrosyni, et al.
Publicado: (2024)
por: Sokli, Effrosyni, et al.
Publicado: (2024)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
por: Deng, Guanzhi, et al.
Publicado: (2026)
por: Deng, Guanzhi, et al.
Publicado: (2026)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
por: Liu, Ningyuan, et al.
Publicado: (2025)
por: Liu, Ningyuan, et al.
Publicado: (2025)
TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
por: Kunwar, Pradip, et al.
Publicado: (2025)
por: Kunwar, Pradip, et al.
Publicado: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
por: He, Yifei, et al.
Publicado: (2025)
por: He, Yifei, et al.
Publicado: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
por: Panda, Ashwinee, et al.
Publicado: (2025)
por: Panda, Ashwinee, et al.
Publicado: (2025)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
por: Guo, Yongxin, et al.
Publicado: (2024)
por: Guo, Yongxin, et al.
Publicado: (2024)
Low-Resource Fine-Tuning for Multi-Task Structured Information Extraction with a Billion-Parameter Instruction-Tuned Model
por: Chih, Yu Cheng, et al.
Publicado: (2025)
por: Chih, Yu Cheng, et al.
Publicado: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
por: Fang, Zhiyuan, et al.
Publicado: (2025)
por: Fang, Zhiyuan, et al.
Publicado: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
por: Li, Houyi, et al.
Publicado: (2025)
por: Li, Houyi, et al.
Publicado: (2025)
Unsupervised Text Representation Learning via Instruction-Tuning for Zero-Shot Dense Retrieval
por: Zeng, Qiuhai, et al.
Publicado: (2024)
por: Zeng, Qiuhai, et al.
Publicado: (2024)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
por: Raje, Arian, et al.
Publicado: (2026)
por: Raje, Arian, et al.
Publicado: (2026)
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
por: Wang, Tuowei, et al.
Publicado: (2025)
por: Wang, Tuowei, et al.
Publicado: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
por: Kim, Junhyuck, et al.
Publicado: (2026)
por: Kim, Junhyuck, et al.
Publicado: (2026)
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
por: Guo, Chenyou, et al.
Publicado: (2026)
por: Guo, Chenyou, et al.
Publicado: (2026)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
por: Xu, Yu, et al.
Publicado: (2026)
por: Xu, Yu, et al.
Publicado: (2026)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
por: Du, Zhiying, et al.
Publicado: (2025)
por: Du, Zhiying, et al.
Publicado: (2025)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2026)
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2026)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2024)
por: Yan, Fanqi, et al.
Publicado: (2024)
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
por: Cao, Jie, et al.
Publicado: (2025)
por: Cao, Jie, et al.
Publicado: (2025)
Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets
por: Bilican, Ahmet, et al.
Publicado: (2025)
por: Bilican, Ahmet, et al.
Publicado: (2025)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
por: Ding, Yifeng, et al.
Publicado: (2024)
por: Ding, Yifeng, et al.
Publicado: (2024)
One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
por: Ming, Rui, et al.
Publicado: (2025)
por: Ming, Rui, et al.
Publicado: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
por: Obando-Ceron, Johan, et al.
Publicado: (2024)
Ejemplares similares
-
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
por: Hui, Tingfeng, et al.
Publicado: (2024) -
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
por: He, Zhuolun, et al.
Publicado: (2023) -
Divergent Thoughts toward One Goal: LLM-based Multi-Agent Collaboration System for Electronic Design Automation
por: Wu, Haoyuan, et al.
Publicado: (2025) -
Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
por: Wu, Haoyuan, et al.
Publicado: (2025) -
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
por: Liu, Yilun, et al.
Publicado: (2024)