LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qu, Xiaoye, Dong, Daize, Hu, Xuyang, Zhu, Tong, Sun, Weigao, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
von: Zhu, Tong, et al.
Veröffentlicht: (2024)
LLaMA Pro: Progressive LLaMA with Block Expansion
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
von: Ma, Mingrui, et al.
Veröffentlicht: (2024)
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
von: Aftab, Danyal, et al.
Veröffentlicht: (2024)
von: Aftab, Danyal, et al.
Veröffentlicht: (2024)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoyi, et al.
Veröffentlicht: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Chu, Xiangxiang, et al.
Veröffentlicht: (2024)
Adapting LLaMA Decoder to Vision Transformer
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
von: Wang, Jiahao, et al.
Veröffentlicht: (2024)
BanglaLlama: LLaMA for Bangla Language
von: Zehady, Abdullah Khan, et al.
Veröffentlicht: (2024)
von: Zehady, Abdullah Khan, et al.
Veröffentlicht: (2024)
LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning
von: Jahangir, Md. Zihad Bin, et al.
Veröffentlicht: (2025)
von: Jahangir, Md. Zihad Bin, et al.
Veröffentlicht: (2025)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
von: Yuan, Fei, et al.
Veröffentlicht: (2023)
von: Yuan, Fei, et al.
Veröffentlicht: (2023)
Multimodal Medical Disease Classification with LLaMA II
von: Gapp, Christian, et al.
Veröffentlicht: (2024)
von: Gapp, Christian, et al.
Veröffentlicht: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
von: Andersland, Michael
Veröffentlicht: (2024)
von: Andersland, Michael
Veröffentlicht: (2024)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
von: Zou, Haosheng, et al.
Veröffentlicht: (2025)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
PUMA: Secure Inference of LLaMA-7B in Five Minutes
von: Dong, Ye, et al.
Veröffentlicht: (2023)
von: Dong, Ye, et al.
Veröffentlicht: (2023)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
von: Cui, Yiming, et al.
Veröffentlicht: (2023)
von: Cui, Yiming, et al.
Veröffentlicht: (2023)
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
von: Ma, Chi, et al.
Veröffentlicht: (2024)
von: Ma, Chi, et al.
Veröffentlicht: (2024)
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2023)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2023)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
LLaMA-Based Models for Aspect-Based Sentiment Analysis
von: Šmíd, Jakub, et al.
Veröffentlicht: (2025)
von: Šmíd, Jakub, et al.
Veröffentlicht: (2025)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Me LLaMA: Foundation Large Language Models for Medical Applications
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
von: Dréano, Sören, et al.
Veröffentlicht: (2025)
von: Dréano, Sören, et al.
Veröffentlicht: (2025)
LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
von: Shi, Kaize, et al.
Veröffentlicht: (2023)
von: Shi, Kaize, et al.
Veröffentlicht: (2023)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
von: Li, Xianhang, et al.
Veröffentlicht: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
von: Xing, Bohao, et al.
Veröffentlicht: (2024)
LLaMA based Punctuation Restoration With Forward Pass Only Decoding
von: Pang, Yutong, et al.
Veröffentlicht: (2024)
von: Pang, Yutong, et al.
Veröffentlicht: (2024)
CUBETESTERAI: Automated JUnit Test Generation using the LLaMA Model
von: Gorla, Daniele, et al.
Veröffentlicht: (2025)
von: Gorla, Daniele, et al.
Veröffentlicht: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
von: Wang, Zhengyi, et al.
Veröffentlicht: (2024)
von: Wang, Zhengyi, et al.
Veröffentlicht: (2024)
LLaMA-NAS: Efficient Neural Architecture Search for Large Language Models
von: Sarah, Anthony, et al.
Veröffentlicht: (2024)
von: Sarah, Anthony, et al.
Veröffentlicht: (2024)
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
von: Chen, Nuo, et al.
Veröffentlicht: (2023)
The Uniqueness of LLaMA3-70B Series with Per-Channel Quantization
von: Qin, Minghai
Veröffentlicht: (2024)
von: Qin, Minghai
Veröffentlicht: (2024)
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Analysis
von: Buscemi, Alessio, et al.
Veröffentlicht: (2024)
von: Buscemi, Alessio, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
von: Zhu, Tong, et al.
Veröffentlicht: (2024) -
LLaMA Pro: Progressive LLaMA with Block Expansion
von: Wu, Chengyue, et al.
Veröffentlicht: (2024) -
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025) -
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
von: Ma, Mingrui, et al.
Veröffentlicht: (2024) -
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
von: Aftab, Danyal, et al.
Veröffentlicht: (2024)