Aurora:Activating Chinese chat capability for Mixtral-8x7B sparse Mixture-of-Experts through Instruction-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Rongsheng, Chen, Haoming, Zhou, Ruizhe, Duan, Yaofei, Cai, Kunyan, Ma, Han, Cui, Jiaxi, Li, Jian, Pang, Patrick Cheong-Iao, Wang, Yapeng, Tan, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning
by: Wang, Rongsheng, et al.
Published: (2024)
by: Wang, Rongsheng, et al.
Published: (2024)
Mixtral of Experts
by: Jiang, Albert Q., et al.
Published: (2024)
by: Jiang, Albert Q., et al.
Published: (2024)
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
by: Wang, Zhong-Qiu, et al.
Published: (2025)
by: Wang, Zhong-Qiu, et al.
Published: (2025)
Enhancing Exploratory Learning through Exploratory Search with the Emergence of Large Language Models
by: Luo, Yiming, et al.
Published: (2024)
by: Luo, Yiming, et al.
Published: (2024)
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
by: Eo, Sugyeong, et al.
Published: (2025)
by: Eo, Sugyeong, et al.
Published: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
by: Xie, Zhen-Hao, et al.
Published: (2026)
by: Xie, Zhen-Hao, et al.
Published: (2026)
Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
by: Ge, Chendi, et al.
Published: (2025)
by: Ge, Chendi, et al.
Published: (2025)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
by: Wang, Yaoxiang, et al.
Published: (2025)
by: Wang, Yaoxiang, et al.
Published: (2025)
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2023)
by: Gou, Yunhao, et al.
Published: (2023)
AnyTaskTune: Advanced Domain-Specific Solutions through Task-Fine-Tuning
by: Cui, Jiaxi, et al.
Published: (2024)
by: Cui, Jiaxi, et al.
Published: (2024)
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
by: Vooturi, Dharma Teja, et al.
Published: (2026)
by: Vooturi, Dharma Teja, et al.
Published: (2026)
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
by: Imani, HamidReza, et al.
Published: (2024)
by: Imani, HamidReza, et al.
Published: (2024)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
by: Wu, Haoyuan, et al.
Published: (2024)
by: Wu, Haoyuan, et al.
Published: (2024)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
by: Shen, Ying, et al.
Published: (2024)
by: Shen, Ying, et al.
Published: (2024)
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
by: Wang, Qingyue, et al.
Published: (2025)
by: Wang, Qingyue, et al.
Published: (2025)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
by: He, Haoze, et al.
Published: (2026)
by: He, Haoze, et al.
Published: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
by: Liu, Baihui, et al.
Published: (2026)
by: Liu, Baihui, et al.
Published: (2026)
DynMoLE: Boosting Mixture of LoRA Experts Fine-Tuning with a Hybrid Routing Mechanism
by: Li, Dengchun, et al.
Published: (2025)
by: Li, Dengchun, et al.
Published: (2025)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
by: Zeng, Runjia, et al.
Published: (2025)
by: Zeng, Runjia, et al.
Published: (2025)
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts
by: Siddiky, Md Nurul Absar
Published: (2026)
by: Siddiky, Md Nurul Absar
Published: (2026)
Theory of Mixture-of-Experts for Mobile Edge Computing
by: Li, Hongbo, et al.
Published: (2024)
by: Li, Hongbo, et al.
Published: (2024)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
by: Tangtartharakul, Gene, et al.
Published: (2026)
by: Tangtartharakul, Gene, et al.
Published: (2026)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
Phased Instruction Fine-Tuning for Large Language Models
by: Pang, Wei, et al.
Published: (2024)
by: Pang, Wei, et al.
Published: (2024)
Enhanced Bloom's Educational Taxonomy for Fostering Information Literacy in the Era of Large Language Models
by: Luo, Yiming, et al.
Published: (2025)
by: Luo, Yiming, et al.
Published: (2025)
A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models
by: Sun, Mengyang, et al.
Published: (2025)
by: Sun, Mengyang, et al.
Published: (2025)
Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models
by: Deng, Guanzhi, et al.
Published: (2026)
by: Deng, Guanzhi, et al.
Published: (2026)
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
Quality Assessment for AI Generated Images with Instruction Tuning
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
FetalFlex: Anatomy-Guided Diffusion Model for Flexible Control on Fetal Ultrasound Image Synthesis
by: Duan, Yaofei, et al.
Published: (2025)
by: Duan, Yaofei, et al.
Published: (2025)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
by: Li, Dengchun, et al.
Published: (2024)
by: Li, Dengchun, et al.
Published: (2024)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
Similar Items
-
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning
by: Wang, Rongsheng, et al.
Published: (2024) -
Mixtral of Experts
by: Jiang, Albert Q., et al.
Published: (2024) -
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
by: Wang, Zhong-Qiu, et al.
Published: (2025) -
Enhancing Exploratory Learning through Exploratory Search with the Emergence of Large Language Models
by: Luo, Yiming, et al.
Published: (2024) -
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
by: Cui, Yiming, et al.
Published: (2024)