Training Dynamics of a 1.7B LLaMa Model: A Data-Efficient Approach
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Miles Q., Fung, Benjamin C. M., Huang, Shih-Chia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Effectiveness of Incremental Training of Large Language Models
por: Li, Miles Q., et al.
Publicado: (2024)
por: Li, Miles Q., et al.
Publicado: (2024)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
por: Santos, Matheus L. O., et al.
Publicado: (2023)
por: Santos, Matheus L. O., et al.
Publicado: (2023)
A Deep User Interface for Exploring LLaMa
por: Perumal, Divya, et al.
Publicado: (2025)
por: Perumal, Divya, et al.
Publicado: (2025)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
por: Allard, Marc-Antoine, et al.
Publicado: (2024)
por: Allard, Marc-Antoine, et al.
Publicado: (2024)
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
Security Concerns for Large Language Models: A Survey
por: Li, Miles Q., et al.
Publicado: (2025)
por: Li, Miles Q., et al.
Publicado: (2025)
Evaluating LLMs for Quotation Attribution in Literary Texts: A Case Study of LLaMa3
por: Michel, Gaspard, et al.
Publicado: (2024)
por: Michel, Gaspard, et al.
Publicado: (2024)
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning
por: Huang, Shih-Cheng, et al.
Publicado: (2022)
por: Huang, Shih-Cheng, et al.
Publicado: (2022)
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation
por: Chiang, Yuan, et al.
Publicado: (2024)
por: Chiang, Yuan, et al.
Publicado: (2024)
Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework
por: Statkiewicz, Grzegorz, et al.
Publicado: (2026)
por: Statkiewicz, Grzegorz, et al.
Publicado: (2026)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
por: Di Palma, Dario, et al.
Publicado: (2025)
por: Di Palma, Dario, et al.
Publicado: (2025)
Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
por: Mansha, Imran
Publicado: (2025)
por: Mansha, Imran
Publicado: (2025)
LLaDA2.0: Scaling Up Diffusion Language Models to 100B
por: Bie, Tiwei, et al.
Publicado: (2025)
por: Bie, Tiwei, et al.
Publicado: (2025)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
por: Zhang, Di, et al.
Publicado: (2024)
por: Zhang, Di, et al.
Publicado: (2024)
Me LLaMA: Foundation Large Language Models for Medical Applications
por: Xie, Qianqian, et al.
Publicado: (2024)
por: Xie, Qianqian, et al.
Publicado: (2024)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
por: Huang, Wenxuan, et al.
Publicado: (2024)
por: Huang, Wenxuan, et al.
Publicado: (2024)
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
por: Huang, Sterling, et al.
Publicado: (2026)
por: Huang, Sterling, et al.
Publicado: (2026)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
por: Zhu, Fengqi, et al.
Publicado: (2025)
por: Zhu, Fengqi, et al.
Publicado: (2025)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
por: Ge, Albert, et al.
Publicado: (2025)
por: Ge, Albert, et al.
Publicado: (2025)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
por: Yuan, Fei, et al.
Publicado: (2023)
por: Yuan, Fei, et al.
Publicado: (2023)
LLaSA: Large Language and Structured Data Assistant
por: Xu, Yao, et al.
Publicado: (2024)
por: Xu, Yao, et al.
Publicado: (2024)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
por: Lan, Zhibin, et al.
Publicado: (2024)
por: Lan, Zhibin, et al.
Publicado: (2024)
How to Train Data-Efficient LLMs
por: Sachdeva, Noveen, et al.
Publicado: (2024)
por: Sachdeva, Noveen, et al.
Publicado: (2024)
Crowdsourcing with Enhanced Data Quality Assurance: An Efficient Approach to Mitigate Resource Scarcity Challenges in Training Large Language Models for Healthcare
por: Barai, P., et al.
Publicado: (2024)
por: Barai, P., et al.
Publicado: (2024)
LLaMandement: Large Language Models for Summarization of French Legislative Proposals
por: Gesnouin, Joseph, et al.
Publicado: (2024)
por: Gesnouin, Joseph, et al.
Publicado: (2024)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
por: Zhao, Jun, et al.
Publicado: (2024)
por: Zhao, Jun, et al.
Publicado: (2024)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
por: Wu, Wei, et al.
Publicado: (2026)
por: Wu, Wei, et al.
Publicado: (2026)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
por: Hinck, Musashi, et al.
Publicado: (2024)
por: Hinck, Musashi, et al.
Publicado: (2024)
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
por: Yu, Longxuan, et al.
Publicado: (2026)
por: Yu, Longxuan, et al.
Publicado: (2026)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
por: He, Guangxin, et al.
Publicado: (2025)
por: He, Guangxin, et al.
Publicado: (2025)
StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
por: Zeng, Jing-Yi, et al.
Publicado: (2025)
por: Zeng, Jing-Yi, et al.
Publicado: (2025)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
por: Li, Shujie, et al.
Publicado: (2024)
por: Li, Shujie, et al.
Publicado: (2024)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
por: Zhang, Wenqiao, et al.
Publicado: (2024)
por: Zhang, Wenqiao, et al.
Publicado: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
por: Weller, Orion, et al.
Publicado: (2023)
por: Weller, Orion, et al.
Publicado: (2023)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
por: Xu, Yuanjian, et al.
Publicado: (2026)
por: Xu, Yuanjian, et al.
Publicado: (2026)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
por: Sun, Wei, et al.
Publicado: (2025)
por: Sun, Wei, et al.
Publicado: (2025)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
por: Tian, Junfeng, et al.
Publicado: (2024)
por: Tian, Junfeng, et al.
Publicado: (2024)
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
por: Syromiatnikov, Mykyta, et al.
Publicado: (2025)
por: Syromiatnikov, Mykyta, et al.
Publicado: (2025)
Data Efficacy for Language Model Training
por: Dai, Yalun, et al.
Publicado: (2025)
por: Dai, Yalun, et al.
Publicado: (2025)
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
por: Li, Miles Q., et al.
Publicado: (2026)
por: Li, Miles Q., et al.
Publicado: (2026)
Ejemplares similares
-
On the Effectiveness of Incremental Training of Large Language Models
por: Li, Miles Q., et al.
Publicado: (2024) -
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
por: Santos, Matheus L. O., et al.
Publicado: (2023) -
A Deep User Interface for Exploring LLaMa
por: Perumal, Divya, et al.
Publicado: (2025) -
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
por: Allard, Marc-Antoine, et al.
Publicado: (2024) -
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
por: Zhang, Di, et al.
Publicado: (2024)