Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Krishnakumar, Arjun, Sukthanker, Rhea Sanjay, Mahadik, Hannan Javed, Kadlecová, Gabriela, Moroshan, Vladyslav, Carstensen, Timur, Hutter, Frank, Klein, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight-Entanglement Meets Gradient-Based Neural Architecture Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2023)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2023)
TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
by: Moroshan, Vladyslav, et al.
Published: (2025)
by: Moroshan, Vladyslav, et al.
Published: (2025)
Compressing Large Language Models with Automated Sub-Network Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
Multi-objective Differentiable Neural Architecture Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
Quickly Tuning Foundation Models for Image Segmentation
by: Das, Breenda, et al.
Published: (2025)
by: Das, Breenda, et al.
Published: (2025)
Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
by: Carstensen, Timur, et al.
Published: (2025)
by: Carstensen, Timur, et al.
Published: (2025)
Selective Rotary Position Embedding
by: Movahedi, Sajad, et al.
Published: (2025)
by: Movahedi, Sajad, et al.
Published: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
by: Panigrahi, Abhishek, et al.
Published: (2024)
by: Panigrahi, Abhishek, et al.
Published: (2024)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
by: Siems, Julien, et al.
Published: (2025)
by: Siems, Julien, et al.
Published: (2025)
Can LLMs Beat Classical Hyperparameter Optimization Algorithms? A Study on autoresearch
by: Ferreira, Fabio, et al.
Published: (2026)
by: Ferreira, Fabio, et al.
Published: (2026)
confopt: A Library for Implementation and Evaluation of Gradient-based One-Shot NAS Methods
by: Jha, Abhash Kumar, et al.
Published: (2025)
by: Jha, Abhash Kumar, et al.
Published: (2025)
On the absence of shock waves and vacuum birefringence in Born--Infeld electrodynamics
by: Kadlecová, Hedvika
Published: (2021)
by: Kadlecová, Hedvika
Published: (2021)
Surprisingly Strong Performance Prediction with Neural Graph Features
by: Kadlecová, Gabriela, et al.
Published: (2024)
by: Kadlecová, Gabriela, et al.
Published: (2024)
Where It all Begins
by: Turner, Dorothy B.
Published: (1973)
by: Turner, Dorothy B.
Published: (1973)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
Collaboration: Where Does It Begin?
by: Small, Ruth V.
Published: (2002)
by: Small, Ruth V.
Published: (2002)
Finding Stable Subnetworks at Initialization with Dataset Distillation
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
by: Bhaskar, Adithya, et al.
Published: (2024)
by: Bhaskar, Adithya, et al.
Published: (2024)
Where Δμ Begins, Gödel Ends
by: Ednyashev, Sanal
Published: (2025)
by: Ednyashev, Sanal
Published: (2025)
Hallucination Begins Where Saliency Drops
by: Zhang, Xiaofeng, et al.
Published: (2026)
by: Zhang, Xiaofeng, et al.
Published: (2026)
Technology--Where Do We Begin?
by: Kostecki, Sister Gladys
Published: (1973)
by: Kostecki, Sister Gladys
Published: (1973)
Source data for the graphs presented in "Synthesis, Anthelmintic Activity, and Mechanism of Action of 5-Aryl-1H-indoles"
by: Kadlecová, Alena, et al.
Published: (2026)
by: Kadlecová, Alena, et al.
Published: (2026)
Source data for the graphs presented in "Advanced screening methods for assessing motility and hatching in plant-parasitic nematodes"
by: Kadlecová, Alena, et al.
Published: (2026)
by: Kadlecová, Alena, et al.
Published: (2026)
Library Research Guides: Where to Begin. A Selected, Annotated Bibliography, B-82.
by: Holt, Dorothy, Comp.
Published: (1982)
by: Holt, Dorothy, Comp.
Published: (1982)
Beyond Random Augmentations: Pretraining with Hard Views
by: Ferreira, Fabio, et al.
Published: (2023)
by: Ferreira, Fabio, et al.
Published: (2023)
Efficient Fault-Tolerant Search by Fast Indexing of Subnetworks
by: Bilò, Davide, et al.
Published: (2024)
by: Bilò, Davide, et al.
Published: (2024)
Active Short Circuit and Safe Discharge Mechanisms in Multi-Phase Inverters During Critical Failures
by: Pimpale, Siddhesh, et al.
Published: (2025)
by: Pimpale, Siddhesh, et al.
Published: (2025)
Do Linear Probes Generalize Better in Persona Coordinates?
by: Mahadik, Prasad, et al.
Published: (2026)
by: Mahadik, Prasad, et al.
Published: (2026)
Neural Subnetwork Ensembles
by: Whitaker, Tim
Published: (2023)
by: Whitaker, Tim
Published: (2023)
Stochastic Subnetwork Annealing: A Regularization Technique for Fine Tuning Pruned Subnetworks
by: Whitaker, Tim, et al.
Published: (2024)
by: Whitaker, Tim, et al.
Published: (2024)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026)
by: Liu, Langming, et al.
Published: (2026)
FedSI: Federated Subnetwork Inference for Efficient Uncertainty Quantification
by: Chen, Hui, et al.
Published: (2024)
by: Chen, Hui, et al.
Published: (2024)
REDS: Resource-Efficient Deep Subnetworks for Dynamic Resource Constraints
by: Corti, Francesco, et al.
Published: (2023)
by: Corti, Francesco, et al.
Published: (2023)
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
by: Kukreja, Dikshant, et al.
Published: (2026)
by: Kukreja, Dikshant, et al.
Published: (2026)
Follow‐Up After Preterm Birth: Where Evidence Ends and Uncertainty Begins
by: Ilari Kuitunen
Published: (2026)
by: Ilari Kuitunen
Published: (2026)
Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How
by: Arango, Sebastian Pineda, et al.
Published: (2023)
by: Arango, Sebastian Pineda, et al.
Published: (2023)
Instilling Inductive Biases with Subnetworks
by: Zhang, Enyan, et al.
Published: (2023)
by: Zhang, Enyan, et al.
Published: (2023)
PlotPick: AI-powered batch extraction of numerical data from scientific figures
by: Carstensen, Tommy
Published: (2026)
by: Carstensen, Tommy
Published: (2026)
Social Media in der Arbeitswelt
by: Carstensen, Tanja
Published: (2018)
by: Carstensen, Tanja
Published: (2018)
Similar Items
-
Weight-Entanglement Meets Gradient-Based Neural Architecture Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2023) -
TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
by: Moroshan, Vladyslav, et al.
Published: (2025) -
Compressing Large Language Models with Automated Sub-Network Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024) -
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024) -
Multi-objective Differentiable Neural Architecture Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)