Dense vs Sparse Pretraining at Tiny Scale: Active-Parameter vs Total-Parameter Matching
Fuente:
arXiv
Salvato in:
| Autore principale: | Wael, Abdalrahman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
di: Jin, Tian, et al.
Pubblicazione: (2025)
di: Jin, Tian, et al.
Pubblicazione: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
di: Ablin, Pierre, et al.
Pubblicazione: (2025)
di: Ablin, Pierre, et al.
Pubblicazione: (2025)
Merging by Matching Models in Task Parameter Subspaces
di: Tam, Derek, et al.
Pubblicazione: (2023)
di: Tam, Derek, et al.
Pubblicazione: (2023)
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
di: Ficek, Aleksander, et al.
Pubblicazione: (2024)
di: Ficek, Aleksander, et al.
Pubblicazione: (2024)
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
di: Song, Yixin, et al.
Pubblicazione: (2024)
di: Song, Yixin, et al.
Pubblicazione: (2024)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
di: Wang, Xinyi, et al.
Pubblicazione: (2024)
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
di: Choudhary, Anand, et al.
Pubblicazione: (2025)
di: Choudhary, Anand, et al.
Pubblicazione: (2025)
Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG
di: Klila, Jaafer, et al.
Pubblicazione: (2026)
di: Klila, Jaafer, et al.
Pubblicazione: (2026)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
di: Lab, Mind, et al.
Pubblicazione: (2026)
di: Lab, Mind, et al.
Pubblicazione: (2026)
Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge
di: Šliogeris, Vytenis, et al.
Pubblicazione: (2025)
di: Šliogeris, Vytenis, et al.
Pubblicazione: (2025)
DropLoRA: Sparse Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
di: Zhang, Haojie
Pubblicazione: (2025)
di: Zhang, Haojie
Pubblicazione: (2025)
Diffusion-Pretrained Dense and Contextual Embeddings
di: Eslami, Sedigheh, et al.
Pubblicazione: (2026)
di: Eslami, Sedigheh, et al.
Pubblicazione: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
di: Roberts, Nicholas, et al.
Pubblicazione: (2025)
di: Roberts, Nicholas, et al.
Pubblicazione: (2025)
Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
di: Lin, Zekai, et al.
Pubblicazione: (2026)
di: Lin, Zekai, et al.
Pubblicazione: (2026)
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
di: Tao, Leitian, et al.
Pubblicazione: (2025)
di: Tao, Leitian, et al.
Pubblicazione: (2025)
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
di: Snell, Charlie, et al.
Pubblicazione: (2024)
di: Snell, Charlie, et al.
Pubblicazione: (2024)
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
di: Tang, Qiaoyu, et al.
Pubblicazione: (2024)
di: Tang, Qiaoyu, et al.
Pubblicazione: (2024)
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models
di: Chang, Aofei, et al.
Pubblicazione: (2024)
di: Chang, Aofei, et al.
Pubblicazione: (2024)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
di: AbdElhameed, Marwan, et al.
Pubblicazione: (2024)
di: AbdElhameed, Marwan, et al.
Pubblicazione: (2024)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
di: Song, Jonghyun, et al.
Pubblicazione: (2025)
di: Song, Jonghyun, et al.
Pubblicazione: (2025)
Towards Efficient Active Learning in NLP via Pretrained Representations
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
Ayn: A Tiny yet Competitive Indian Legal Language Model Pretrained from Scratch
di: Niyogi, Mitodru, et al.
Pubblicazione: (2024)
di: Niyogi, Mitodru, et al.
Pubblicazione: (2024)
Language Models Improve When Pretraining Data Matches Target Tasks
di: Mizrahi, David, et al.
Pubblicazione: (2025)
di: Mizrahi, David, et al.
Pubblicazione: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
di: Li, Guanchen, et al.
Pubblicazione: (2024)
di: Li, Guanchen, et al.
Pubblicazione: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
di: Chen, Hao, et al.
Pubblicazione: (2026)
di: Chen, Hao, et al.
Pubblicazione: (2026)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
di: Sen, Tanmay, et al.
Pubblicazione: (2024)
di: Sen, Tanmay, et al.
Pubblicazione: (2024)
Scaling Laws for Mixture Pretraining Under Data Constraints
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
di: Song, Chenyang, et al.
Pubblicazione: (2026)
di: Song, Chenyang, et al.
Pubblicazione: (2026)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
di: Liu, Yong, et al.
Pubblicazione: (2024)
di: Liu, Yong, et al.
Pubblicazione: (2024)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
di: Wu, Chen, et al.
Pubblicazione: (2025)
di: Wu, Chen, et al.
Pubblicazione: (2025)
Exploring Activation Patterns of Parameters in Language Models
di: Wang, Yudong, et al.
Pubblicazione: (2024)
di: Wang, Yudong, et al.
Pubblicazione: (2024)
Is Parameter Collision Hindering Continual Learning in LLMs?
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
di: Lee, Harrison, et al.
Pubblicazione: (2023)
di: Lee, Harrison, et al.
Pubblicazione: (2023)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
di: Sabry, Mohammed, et al.
Pubblicazione: (2024)
di: Sabry, Mohammed, et al.
Pubblicazione: (2024)
Optimal Stopping vs Best-of-$N$ for Inference Time Optimization
di: Kalayci, Yusuf, et al.
Pubblicazione: (2025)
di: Kalayci, Yusuf, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
di: Jin, Tian, et al.
Pubblicazione: (2025) -
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
di: Ablin, Pierre, et al.
Pubblicazione: (2025) -
Merging by Matching Models in Task Parameter Subspaces
di: Tam, Derek, et al.
Pubblicazione: (2023) -
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning
di: Ficek, Aleksander, et al.
Pubblicazione: (2024) -
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
di: Song, Yixin, et al.
Pubblicazione: (2024)