Deep Fusion: Efficient Network Training via Pre-trained Initializations
Fuente:
arXiv
Saved in:
| Main Authors: | Mazzawi, Hanna, Gonzalvo, Xavi, Wunder, Michael, Jerome, Sammy, Dherin, Benoit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning without training: The implicit dynamics of in-context learning
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Transmuting prompts into weights
by: Mazzawi, Hanna, et al.
Published: (2025)
by: Mazzawi, Hanna, et al.
Published: (2025)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
by: Adila, Dyah, et al.
Published: (2026)
by: Adila, Dyah, et al.
Published: (2026)
Learning by solving differential equations
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
by: Mazzawi, Hanna, et al.
Published: (2024)
by: Mazzawi, Hanna, et al.
Published: (2024)
How iteration order influences convergence and stability in deep learning
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
A Margin-based Multiclass Generalization Bound via Geometric Complexity
by: Munn, Michael, et al.
Published: (2024)
by: Munn, Michael, et al.
Published: (2024)
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
by: Munn, Michael, et al.
Published: (2024)
by: Munn, Michael, et al.
Published: (2024)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
by: Goldwaser, Adrian, et al.
Published: (2025)
by: Goldwaser, Adrian, et al.
Published: (2025)
On residual network depth
by: Dherin, Benoit, et al.
Published: (2025)
by: Dherin, Benoit, et al.
Published: (2025)
Corridor Geometry in Gradient-Based Optimization
by: Dherin, Benoit, et al.
Published: (2024)
by: Dherin, Benoit, et al.
Published: (2024)
Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and Beyond
by: Axiotis, Kyriakos, et al.
Published: (2024)
by: Axiotis, Kyriakos, et al.
Published: (2024)
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-Training of Deep Networks
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
Depth-Aware Initialization for Stable and Efficient Neural Network Training
by: Pandey, Vijay
Published: (2025)
by: Pandey, Vijay
Published: (2025)
Constraint-based Pre-training: From Structured Constraints to Scalable Model Initialization
by: Feng, Fu, et al.
Published: (2026)
by: Feng, Fu, et al.
Published: (2026)
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
by: Lin, Xi Victoria, et al.
Published: (2024)
by: Lin, Xi Victoria, et al.
Published: (2024)
DeepRTE: Pre-trained Attention-based Neural Network for Radiative Transfer
by: Zhu, Yekun, et al.
Published: (2025)
by: Zhu, Yekun, et al.
Published: (2025)
Designing Pre-training Datasets from Unlabeled Data for EEG Classification with Transformers
by: Bary, Tim, et al.
Published: (2024)
by: Bary, Tim, et al.
Published: (2024)
CPT: Efficient Deep Neural Network Training via Cyclic Precision
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
NetMamba: Efficient Network Traffic Classification via Pre-training Unidirectional Mamba
by: Wang, Tongze, et al.
Published: (2024)
by: Wang, Tongze, et al.
Published: (2024)
Experience-Efficient Model-Free Deep Reinforcement Learning Using Pre-Training
by: Yang, Ruoxing
Published: (2025)
by: Yang, Ruoxing
Published: (2025)
GP2F: Cross-Domain Graph Prompting with Adaptive Fusion of Pre-trained Graph Neural Networks
by: He, Dongxiao, et al.
Published: (2026)
by: He, Dongxiao, et al.
Published: (2026)
The logic of rational graph neural networks
by: Khalife, Sammy
Published: (2023)
by: Khalife, Sammy
Published: (2023)
Language verY Rare for All
by: Merad, Ibrahim, et al.
Published: (2024)
by: Merad, Ibrahim, et al.
Published: (2024)
Effects of Initialization Biases on Deep Neural Network Training Dynamics
by: Pellegrino, Nicholas, et al.
Published: (2025)
by: Pellegrino, Nicholas, et al.
Published: (2025)
Pre-training Tensor-Train Networks Facilitates Machine Learning with Variational Quantum Circuits
by: Qi, Jun, et al.
Published: (2023)
by: Qi, Jun, et al.
Published: (2023)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
by: Bogdanov, Sergei, et al.
Published: (2024)
by: Bogdanov, Sergei, et al.
Published: (2024)
NetMamba+: A Framework of Pre-trained Models for Efficient and Accurate Network Traffic Classification
by: Wang, Tongze, et al.
Published: (2026)
by: Wang, Tongze, et al.
Published: (2026)
PreNeT: Leveraging Computational Features to Predict Deep Neural Network Training Time
by: Pourali, Alireza, et al.
Published: (2024)
by: Pourali, Alireza, et al.
Published: (2024)
Memory-Efficient Training for Text-Dependent SV with Independent Pre-trained Models
by: Farokh, Seyed Ali, et al.
Published: (2024)
by: Farokh, Seyed Ali, et al.
Published: (2024)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
Initialization Matters: Unraveling the Impact of Pre-Training on Federated Learning
by: Jhunjhunwala, Divyansh, et al.
Published: (2025)
by: Jhunjhunwala, Divyansh, et al.
Published: (2025)
DeepRV: Accelerating Spatiotemporal Inference with Pre-trained Neural Priors
by: Navott, Jhonathan, et al.
Published: (2025)
by: Navott, Jhonathan, et al.
Published: (2025)
Deep Neural Network Initialization with Sparsity Inducing Activations
by: Price, Ilan, et al.
Published: (2024)
by: Price, Ilan, et al.
Published: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
by: Thapa, Krishu K, et al.
Published: (2025)
by: Thapa, Krishu K, et al.
Published: (2025)
PreND: Enhancing Intrinsic Motivation in Reinforcement Learning through Pre-trained Network Distillation
by: Davoodabadi, Mohammadamin, et al.
Published: (2024)
by: Davoodabadi, Mohammadamin, et al.
Published: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
Efficient Training of Deep Neural Operator Networks via Randomized Sampling
by: Karumuri, Sharmila, et al.
Published: (2024)
by: Karumuri, Sharmila, et al.
Published: (2024)
Pruning-based Data Selection and Network Fusion for Efficient Deep Learning
by: Kousar, Humaira, et al.
Published: (2025)
by: Kousar, Humaira, et al.
Published: (2025)
Similar Items
-
Learning without training: The implicit dynamics of in-context learning
by: Dherin, Benoit, et al.
Published: (2025) -
Transmuting prompts into weights
by: Mazzawi, Hanna, et al.
Published: (2025) -
Grow, Don't Overwrite: Fine-tuning Without Forgetting
by: Adila, Dyah, et al.
Published: (2026) -
Learning by solving differential equations
by: Dherin, Benoit, et al.
Published: (2025) -
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
by: Mazzawi, Hanna, et al.
Published: (2024)