Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Chi, Huang, Mincong, Wang, Chao, Wang, Yujie, Yu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOYU: A Theoretical Study on Massive Over-activation Yielded Uplifts in LLMs
by: Ma, Chi, et al.
Published: (2024)
by: Ma, Chi, et al.
Published: (2024)
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
by: Ma, Chi, et al.
Published: (2024)
by: Ma, Chi, et al.
Published: (2024)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
by: Yang, Zhuoyi, et al.
Published: (2025)
by: Yang, Zhuoyi, et al.
Published: (2025)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
BanglaLlama: LLaMA for Bangla Language
by: Zehady, Abdullah Khan, et al.
Published: (2024)
by: Zehady, Abdullah Khan, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023)
by: Kavehzadeh, Parsa, et al.
Published: (2023)
Less Is More: Generating Time Series with LLaMA-Style Autoregression in Simple Factorized Latent Spaces
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
by: Cui, Yiming, et al.
Published: (2023)
by: Cui, Yiming, et al.
Published: (2023)
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
by: Kim, Bo-Kyeong, et al.
Published: (2024)
by: Kim, Bo-Kyeong, et al.
Published: (2024)
Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model
by: Kim, Dongki, et al.
Published: (2025)
by: Kim, Dongki, et al.
Published: (2025)
Evaluating LLaMA 3.2 for Software Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2025)
by: Gonçalves, José, et al.
Published: (2025)
The Uniqueness of LLaMA3-70B Series with Per-Channel Quantization
by: Qin, Minghai
Published: (2024)
by: Qin, Minghai
Published: (2024)
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
by: Dréano, Sören, et al.
Published: (2025)
by: Dréano, Sören, et al.
Published: (2025)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
by: Xia, Mengzhou, et al.
Published: (2023)
by: Xia, Mengzhou, et al.
Published: (2023)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
by: Zou, Haosheng, et al.
Published: (2025)
by: Zou, Haosheng, et al.
Published: (2025)
LLaMA Pro: Progressive LLaMA with Block Expansion
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
by: Wang, Zhengyi, et al.
Published: (2024)
by: Wang, Zhengyi, et al.
Published: (2024)
Enhancing Document-Level Question Answering via Multi-Hop Retrieval-Augmented Generation with LLaMA 3
by: Huang, Xinyue, et al.
Published: (2025)
by: Huang, Xinyue, et al.
Published: (2025)
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
Fine-tuning LLaMA 2 interference: a comparative study of language implementations for optimal efficiency
by: Hossain, Sazzad, et al.
Published: (2025)
by: Hossain, Sazzad, et al.
Published: (2025)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024)
by: Huang, Mincong, et al.
Published: (2024)
Instruction Finetuning LLaMA-3-8B Model Using LoRA for Financial Named Entity Recognition
by: Lian, Zhiming
Published: (2026)
by: Lian, Zhiming
Published: (2026)
Zero-Shot End-to-End Relation Extraction in Chinese: A Comparative Study of Gemini, LLaMA and ChatGPT
by: Du, Shaoshuai, et al.
Published: (2025)
by: Du, Shaoshuai, et al.
Published: (2025)
I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2
by: McLaughlin, Oliver, et al.
Published: (2025)
by: McLaughlin, Oliver, et al.
Published: (2025)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
QQQ: Quality Quattuor-Bit Quantization for Large Language Models
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
The Role of Model Architecture and Scale in Predicting Molecular Properties: Insights from Fine-Tuning RoBERTa, BART, and LLaMA
by: Youngmin, Lee, et al.
Published: (2024)
by: Youngmin, Lee, et al.
Published: (2024)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
Secure Federated Learning Across Heterogeneous Cloud and High-Performance Computing Resources -- A Case Study on Federated Fine-tuning of LLaMA 2
by: Li, Zilinghan, et al.
Published: (2024)
by: Li, Zilinghan, et al.
Published: (2024)
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
by: Aftab, Danyal, et al.
Published: (2024)
by: Aftab, Danyal, et al.
Published: (2024)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
by: Kang, Boyi, et al.
Published: (2025)
by: Kang, Boyi, et al.
Published: (2025)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
by: Yuan, Fei, et al.
Published: (2023)
by: Yuan, Fei, et al.
Published: (2023)
Adapting LLaMA Decoder to Vision Transformer
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
QuickLLaMA: Query-aware Inference Acceleration for Large Language Models
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
Similar Items
-
MOYU: A Theoretical Study on Massive Over-activation Yielded Uplifts in LLMs
by: Ma, Chi, et al.
Published: (2024) -
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
by: Ma, Chi, et al.
Published: (2024) -
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025) -
LogLLaMA: Transformer-based log anomaly detection with LLaMA
by: Yang, Zhuoyi, et al.
Published: (2025) -
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
by: Huang, Wei, et al.
Published: (2024)