LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Tong, Qu, Xiaoye, Dong, Daize, Ruan, Jiacheng, Tong, Jingqi, He, Conghui, Cheng, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
LLaMA Pro: Progressive LLaMA with Block Expansion
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
di: Aftab, Danyal, et al.
Pubblicazione: (2024)
di: Aftab, Danyal, et al.
Pubblicazione: (2024)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
di: Yang, Zhuoyi, et al.
Pubblicazione: (2025)
di: Yang, Zhuoyi, et al.
Pubblicazione: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
Adapting LLaMA Decoder to Vision Transformer
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
BanglaLlama: LLaMA for Bangla Language
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning
di: Jahangir, Md. Zihad Bin, et al.
Pubblicazione: (2025)
di: Jahangir, Md. Zihad Bin, et al.
Pubblicazione: (2025)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
di: Yuan, Fei, et al.
Pubblicazione: (2023)
di: Yuan, Fei, et al.
Pubblicazione: (2023)
Multimodal Medical Disease Classification with LLaMA II
di: Gapp, Christian, et al.
Pubblicazione: (2024)
di: Gapp, Christian, et al.
Pubblicazione: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
di: Ma, Chi, et al.
Pubblicazione: (2024)
di: Ma, Chi, et al.
Pubblicazione: (2024)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
PUMA: Secure Inference of LLaMA-7B in Five Minutes
di: Dong, Ye, et al.
Pubblicazione: (2023)
di: Dong, Ye, et al.
Pubblicazione: (2023)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
di: Cui, Yiming, et al.
Pubblicazione: (2023)
di: Cui, Yiming, et al.
Pubblicazione: (2023)
Faster Speech-LLaMA Inference with Multi-token Prediction
di: Raj, Desh, et al.
Pubblicazione: (2024)
di: Raj, Desh, et al.
Pubblicazione: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
di: Gonçalves, José, et al.
Pubblicazione: (2025)
di: Gonçalves, José, et al.
Pubblicazione: (2025)
LLaMA-Based Models for Aspect-Based Sentiment Analysis
di: Šmíd, Jakub, et al.
Pubblicazione: (2025)
di: Šmíd, Jakub, et al.
Pubblicazione: (2025)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
di: Kang, Boyi, et al.
Pubblicazione: (2025)
di: Kang, Boyi, et al.
Pubblicazione: (2025)
Me LLaMA: Foundation Large Language Models for Medical Applications
di: Xie, Qianqian, et al.
Pubblicazione: (2024)
di: Xie, Qianqian, et al.
Pubblicazione: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
di: Xing, Bohao, et al.
Pubblicazione: (2024)
di: Xing, Bohao, et al.
Pubblicazione: (2024)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024)
di: Zhang, Di, et al.
Pubblicazione: (2024)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
di: Zhao, Jun, et al.
Pubblicazione: (2024)
di: Zhao, Jun, et al.
Pubblicazione: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
di: Li, Xianhang, et al.
Pubblicazione: (2024)
di: Li, Xianhang, et al.
Pubblicazione: (2024)
LLaMA based Punctuation Restoration With Forward Pass Only Decoding
di: Pang, Yutong, et al.
Pubblicazione: (2024)
di: Pang, Yutong, et al.
Pubblicazione: (2024)
CUBETESTERAI: Automated JUnit Test Generation using the LLaMA Model
di: Gorla, Daniele, et al.
Pubblicazione: (2025)
di: Gorla, Daniele, et al.
Pubblicazione: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
di: Huang, Wei, et al.
Pubblicazione: (2024)
di: Huang, Wei, et al.
Pubblicazione: (2024)
Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Experts
di: Zhu, Tong, et al.
Pubblicazione: (2024)
di: Zhu, Tong, et al.
Pubblicazione: (2024)
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
di: Wang, Zhengyi, et al.
Pubblicazione: (2024)
di: Wang, Zhengyi, et al.
Pubblicazione: (2024)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
di: Zhang, Jihai, et al.
Pubblicazione: (2024)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
Efficient Training of Robust Traditional Chinese LLaMA-1B on a Single Consumer GPU: Continual Pre-training, SFT, and DPO
di: Chih, Yu-Cheng, et al.
Pubblicazione: (2025)
di: Chih, Yu-Cheng, et al.
Pubblicazione: (2025)
LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
di: Shi, Kaize, et al.
Pubblicazione: (2023)
di: Shi, Kaize, et al.
Pubblicazione: (2023)
FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models
di: Zhang, Zishuai, et al.
Pubblicazione: (2025)
di: Zhang, Zishuai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
di: Qu, Xiaoye, et al.
Pubblicazione: (2024) -
LLaMA Pro: Progressive LLaMA with Block Expansion
di: Wu, Chengyue, et al.
Pubblicazione: (2024) -
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
di: Ma, Mingrui, et al.
Pubblicazione: (2024) -
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
di: Dialameh, Maryam, et al.
Pubblicazione: (2025) -
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
di: Aftab, Danyal, et al.
Pubblicazione: (2024)