Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Nuo, Wu, Ning, Liang, Shining, Gong, Ming, Shou, Linjun, Zhang, Dongmei, Li, Jia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLaMA Pro: Progressive LLaMA with Block Expansion
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
di: Dialameh, Maryam, et al.
Pubblicazione: (2025)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
di: Di Palma, Dario, et al.
Pubblicazione: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
di: Zhu, Tong, et al.
Pubblicazione: (2024)
di: Zhu, Tong, et al.
Pubblicazione: (2024)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
di: Yang, Zhuoyi, et al.
Pubblicazione: (2025)
di: Yang, Zhuoyi, et al.
Pubblicazione: (2025)
Are Bigger Encoders Always Better in Vision Large Models?
di: Li, Bozhou, et al.
Pubblicazione: (2024)
di: Li, Bozhou, et al.
Pubblicazione: (2024)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
di: Qu, Xiaoye, et al.
Pubblicazione: (2024)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
di: Yuan, Fei, et al.
Pubblicazione: (2023)
di: Yuan, Fei, et al.
Pubblicazione: (2023)
360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
di: Zou, Haosheng, et al.
Pubblicazione: (2025)
di: Zou, Haosheng, et al.
Pubblicazione: (2025)
BanglaLlama: LLaMA for Bangla Language
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
di: Zehady, Abdullah Khan, et al.
Pubblicazione: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
di: Andersland, Michael
Pubblicazione: (2024)
di: Andersland, Michael
Pubblicazione: (2024)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
LLaMA-Based Models for Aspect-Based Sentiment Analysis
di: Šmíd, Jakub, et al.
Pubblicazione: (2025)
di: Šmíd, Jakub, et al.
Pubblicazione: (2025)
Me LLaMA: Foundation Large Language Models for Medical Applications
di: Xie, Qianqian, et al.
Pubblicazione: (2024)
di: Xie, Qianqian, et al.
Pubblicazione: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
di: Gema, Aryo Pradipta, et al.
Pubblicazione: (2023)
LLaMA based Punctuation Restoration With Forward Pass Only Decoding
di: Pang, Yutong, et al.
Pubblicazione: (2024)
di: Pang, Yutong, et al.
Pubblicazione: (2024)
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
di: Cui, Yiming, et al.
Pubblicazione: (2023)
di: Cui, Yiming, et al.
Pubblicazione: (2023)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
di: Zhao, Jun, et al.
Pubblicazione: (2024)
di: Zhao, Jun, et al.
Pubblicazione: (2024)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
di: Kang, Boyi, et al.
Pubblicazione: (2025)
di: Kang, Boyi, et al.
Pubblicazione: (2025)
Selected Languages are All You Need for Cross-lingual Truthfulness Transfer
di: Liu, Weihao, et al.
Pubblicazione: (2024)
di: Liu, Weihao, et al.
Pubblicazione: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
di: Li, Xianhang, et al.
Pubblicazione: (2024)
di: Li, Xianhang, et al.
Pubblicazione: (2024)
Llamazip: Leveraging LLaMA for Lossless Text Compression and Training Dataset Detection
di: Dréano, Sören, et al.
Pubblicazione: (2025)
di: Dréano, Sören, et al.
Pubblicazione: (2025)
ChatGPT vs Gemini vs LLaMA on Multilingual Sentiment Analysis
di: Buscemi, Alessio, et al.
Pubblicazione: (2024)
di: Buscemi, Alessio, et al.
Pubblicazione: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
di: Robertson, Sean, et al.
Pubblicazione: (2023)
di: Robertson, Sean, et al.
Pubblicazione: (2023)
Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets
di: Azime, Israel Abebe, et al.
Pubblicazione: (2024)
di: Azime, Israel Abebe, et al.
Pubblicazione: (2024)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
di: Fang, Qingkai, et al.
Pubblicazione: (2024)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
di: Zhang, Di, et al.
Pubblicazione: (2024)
di: Zhang, Di, et al.
Pubblicazione: (2024)
MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation
di: Izzati, Fathinah, et al.
Pubblicazione: (2025)
di: Izzati, Fathinah, et al.
Pubblicazione: (2025)
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
di: Kim, Bo-Kyeong, et al.
Pubblicazione: (2024)
di: Kim, Bo-Kyeong, et al.
Pubblicazione: (2024)
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
di: Vargas, Francisco, et al.
Pubblicazione: (2025)
di: Vargas, Francisco, et al.
Pubblicazione: (2025)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
di: Ma, Mingrui, et al.
Pubblicazione: (2024)
Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations
di: Chen, Nuo, et al.
Pubblicazione: (2023)
di: Chen, Nuo, et al.
Pubblicazione: (2023)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
di: Zou, Bo, et al.
Pubblicazione: (2024)
di: Zou, Bo, et al.
Pubblicazione: (2024)
LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
di: Shi, Kaize, et al.
Pubblicazione: (2023)
di: Shi, Kaize, et al.
Pubblicazione: (2023)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
di: Liu, Weihao, et al.
Pubblicazione: (2025)
di: Liu, Weihao, et al.
Pubblicazione: (2025)
Adapting LLaMA Decoder to Vision Transformer
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
di: Syromiatnikov, Mykyta, et al.
Pubblicazione: (2025)
Inverse Scaling: When Bigger Isn't Better
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
di: McKenzie, Ian R., et al.
Pubblicazione: (2023)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLaMA Pro: Progressive LLaMA with Block Expansion
di: Wu, Chengyue, et al.
Pubblicazione: (2024) -
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
di: Dialameh, Maryam, et al.
Pubblicazione: (2025) -
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
di: Di Palma, Dario, et al.
Pubblicazione: (2025) -
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
di: Zhu, Tong, et al.
Pubblicazione: (2024) -
LogLLaMA: Transformer-based log anomaly detection with LLaMA
di: Yang, Zhuoyi, et al.
Pubblicazione: (2025)