Training Dynamics of a 1.7B LLaMa Model: A Data-Efficient Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Miles Q., Fung, Benjamin C. M., Huang, Shih-Chia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Effectiveness of Incremental Training of Large Language Models
von: Li, Miles Q., et al.
Veröffentlicht: (2024)
von: Li, Miles Q., et al.
Veröffentlicht: (2024)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
von: Santos, Matheus L. O., et al.
Veröffentlicht: (2023)
von: Santos, Matheus L. O., et al.
Veröffentlicht: (2023)
A Deep User Interface for Exploring LLaMa
von: Perumal, Divya, et al.
Veröffentlicht: (2025)
von: Perumal, Divya, et al.
Veröffentlicht: (2025)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
von: Allard, Marc-Antoine, et al.
Veröffentlicht: (2024)
von: Allard, Marc-Antoine, et al.
Veröffentlicht: (2024)
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
Security Concerns for Large Language Models: A Survey
von: Li, Miles Q., et al.
Veröffentlicht: (2025)
von: Li, Miles Q., et al.
Veröffentlicht: (2025)
Evaluating LLMs for Quotation Attribution in Literary Texts: A Case Study of LLaMa3
von: Michel, Gaspard, et al.
Veröffentlicht: (2024)
von: Michel, Gaspard, et al.
Veröffentlicht: (2024)
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning
von: Huang, Shih-Cheng, et al.
Veröffentlicht: (2022)
von: Huang, Shih-Cheng, et al.
Veröffentlicht: (2022)
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation
von: Chiang, Yuan, et al.
Veröffentlicht: (2024)
von: Chiang, Yuan, et al.
Veröffentlicht: (2024)
Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework
von: Statkiewicz, Grzegorz, et al.
Veröffentlicht: (2026)
von: Statkiewicz, Grzegorz, et al.
Veröffentlicht: (2026)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
von: Di Palma, Dario, et al.
Veröffentlicht: (2025)
Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
von: Mansha, Imran
Veröffentlicht: (2025)
von: Mansha, Imran
Veröffentlicht: (2025)
LLaDA2.0: Scaling Up Diffusion Language Models to 100B
von: Bie, Tiwei, et al.
Veröffentlicht: (2025)
von: Bie, Tiwei, et al.
Veröffentlicht: (2025)
LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
von: Zhang, Di, et al.
Veröffentlicht: (2024)
von: Zhang, Di, et al.
Veröffentlicht: (2024)
Me LLaMA: Foundation Large Language Models for Medical Applications
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
von: Huang, Sterling, et al.
Veröffentlicht: (2026)
von: Huang, Sterling, et al.
Veröffentlicht: (2026)
LLaDA-MoE: A Sparse MoE Diffusion Language Model
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
von: Zhu, Fengqi, et al.
Veröffentlicht: (2025)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
von: Ge, Albert, et al.
Veröffentlicht: (2025)
von: Ge, Albert, et al.
Veröffentlicht: (2025)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
von: Yuan, Fei, et al.
Veröffentlicht: (2023)
von: Yuan, Fei, et al.
Veröffentlicht: (2023)
LLaSA: Large Language and Structured Data Assistant
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
How to Train Data-Efficient LLMs
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
Crowdsourcing with Enhanced Data Quality Assurance: An Efficient Approach to Mitigate Resource Scarcity Challenges in Training Large Language Models for Healthcare
von: Barai, P., et al.
Veröffentlicht: (2024)
von: Barai, P., et al.
Veröffentlicht: (2024)
LLaMandement: Large Language Models for Summarization of French Legislative Proposals
von: Gesnouin, Joseph, et al.
Veröffentlicht: (2024)
von: Gesnouin, Joseph, et al.
Veröffentlicht: (2024)
LLaMA Beyond English: An Empirical Study on Language Capability Transfer
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs
von: Yu, Longxuan, et al.
Veröffentlicht: (2026)
von: Yu, Longxuan, et al.
Veröffentlicht: (2026)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
von: He, Guangxin, et al.
Veröffentlicht: (2025)
von: He, Guangxin, et al.
Veröffentlicht: (2025)
StatLLaMA: Multi-Stage training for domain-optimized statistical large language models
von: Zeng, Jing-Yi, et al.
Veröffentlicht: (2025)
von: Zeng, Jing-Yi, et al.
Veröffentlicht: (2025)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
von: Li, Shujie, et al.
Veröffentlicht: (2024)
von: Li, Shujie, et al.
Veröffentlicht: (2024)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
von: Sun, Wei, et al.
Veröffentlicht: (2025)
von: Sun, Wei, et al.
Veröffentlicht: (2025)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
von: Tian, Junfeng, et al.
Veröffentlicht: (2024)
von: Tian, Junfeng, et al.
Veröffentlicht: (2024)
Empowering Smaller Models: Tuning LLaMA and Gemma with Chain-of-Thought for Ukrainian Exam Tasks
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Data Efficacy for Language Model Training
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On the Effectiveness of Incremental Training of Large Language Models
von: Li, Miles Q., et al.
Veröffentlicht: (2024) -
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
von: Santos, Matheus L. O., et al.
Veröffentlicht: (2023) -
A Deep User Interface for Exploring LLaMa
von: Perumal, Divya, et al.
Veröffentlicht: (2025) -
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
von: Allard, Marc-Antoine, et al.
Veröffentlicht: (2024) -
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
von: Zhang, Di, et al.
Veröffentlicht: (2024)