Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Tong, Dong, Daize, Qu, Xiaoye, Ruan, Jiacheng, Chen, Wenliang, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2026)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2026)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking
von: Tang, Tianwen, et al.
Veröffentlicht: (2024)
von: Tang, Tianwen, et al.
Veröffentlicht: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
von: Fan, Chenghao, et al.
Veröffentlicht: (2025)
SEE: Continual Fine-tuning with Sequential Ensemble of Experts
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
von: Gui, Runquan, et al.
Veröffentlicht: (2026)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
von: Tan, Zhen, et al.
Veröffentlicht: (2024)
von: Tan, Zhen, et al.
Veröffentlicht: (2024)
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
von: He, Haoze, et al.
Veröffentlicht: (2026)
von: He, Haoze, et al.
Veröffentlicht: (2026)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
von: Chen, Yicheng, et al.
Veröffentlicht: (2025)
von: Chen, Yicheng, et al.
Veröffentlicht: (2025)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
von: Zeng, Runjia, et al.
Veröffentlicht: (2025)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
von: Renduchintala, H S V N S Kowndinya, et al.
Veröffentlicht: (2024)
von: Renduchintala, H S V N S Kowndinya, et al.
Veröffentlicht: (2024)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
von: Li, Dengchun, et al.
Veröffentlicht: (2024)
von: Li, Dengchun, et al.
Veröffentlicht: (2024)
Seal-Tools: Self-Instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark
von: Wu, Mengsong, et al.
Veröffentlicht: (2024)
von: Wu, Mengsong, et al.
Veröffentlicht: (2024)
Aurora:Activating Chinese chat capability for Mixtral-8x7B sparse Mixture-of-Experts through Instruction-Tuning
von: Wang, Rongsheng, et al.
Veröffentlicht: (2023)
von: Wang, Rongsheng, et al.
Veröffentlicht: (2023)
Mixture of Neuron Experts
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
Is Fine-Tuning an Effective Solution? Reassessing Knowledge Editing for Unstructured Data
von: Xiong, Hao, et al.
Veröffentlicht: (2025)
von: Xiong, Hao, et al.
Veröffentlicht: (2025)
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
von: Yang, Yixin, et al.
Veröffentlicht: (2025)
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert Controllers
von: Li, Sijia, et al.
Veröffentlicht: (2023)
von: Li, Sijia, et al.
Veröffentlicht: (2023)
DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
von: Deng, Guanzhi, et al.
Veröffentlicht: (2026)
Timo: Towards Better Temporal Reasoning for Language Models
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
von: Su, Zhaochen, et al.
Veröffentlicht: (2024)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
von: Bai, Jun, et al.
Veröffentlicht: (2025)
von: Bai, Jun, et al.
Veröffentlicht: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyi, et al.
Veröffentlicht: (2024)
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
von: Kadasi, Pritam, et al.
Veröffentlicht: (2025)
von: Kadasi, Pritam, et al.
Veröffentlicht: (2025)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
iDAT: inverse Distillation Adapter-Tuning
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2024)
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
von: Li, Yafu, et al.
Veröffentlicht: (2025)
von: Li, Yafu, et al.
Veröffentlicht: (2025)
LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review
von: Yuan, Cheng, et al.
Veröffentlicht: (2025)
von: Yuan, Cheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024) -
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024) -
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2026) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025) -
MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking
von: Tang, Tianwen, et al.
Veröffentlicht: (2024)