When is Warmstarting Effective for Scaling Language Models?
Fuente:
arXiv
Saved in:
| Main Authors: | Mallik, Neeratyoy, Janowski, Maciej, Hog, Johannes, Rakotoarison, Herilalaina, Grabocka, Josif, Hutter, Frank, Klein, Aaron |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Warmstarting for Scaling Language Models
by: Mallik, Neeratyoy, et al.
Published: (2024)
by: Mallik, Neeratyoy, et al.
Published: (2024)
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization
by: Rakotoarison, Herilalaina, et al.
Published: (2024)
by: Rakotoarison, Herilalaina, et al.
Published: (2024)
Ensembling Finetuned Language Models for Text Classification
by: Arango, Sebastian Pineda, et al.
Published: (2024)
by: Arango, Sebastian Pineda, et al.
Published: (2024)
Regularized Neural Ensemblers
by: Arango, Sebastian Pineda, et al.
Published: (2024)
by: Arango, Sebastian Pineda, et al.
Published: (2024)
An Open-Source Training Dataset for Foundation Models for Black-box Optimization
by: Klein, Aaron, et al.
Published: (2026)
by: Klein, Aaron, et al.
Published: (2026)
Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and How
by: Arango, Sebastian Pineda, et al.
Published: (2023)
by: Arango, Sebastian Pineda, et al.
Published: (2023)
Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
by: Carstensen, Timur, et al.
Published: (2025)
by: Carstensen, Timur, et al.
Published: (2025)
Fast Benchmarking of Asynchronous Multi-Fidelity Optimization on Zero-Cost Benchmarks
by: Watanabe, Shuhei, et al.
Published: (2024)
by: Watanabe, Shuhei, et al.
Published: (2024)
Transformers Can Do Bayesian Inference
by: Müller, Samuel, et al.
Published: (2021)
by: Müller, Samuel, et al.
Published: (2021)
LMEMs for post-hoc analysis of HPO Benchmarking
by: Geburek, Anton, et al.
Published: (2024)
by: Geburek, Anton, et al.
Published: (2024)
Hierarchical Transformers are Efficient Meta-Reinforcement Learners
by: Shala, Gresa, et al.
Published: (2024)
by: Shala, Gresa, et al.
Published: (2024)
Interpretable Mesomorphic Networks for Tabular Data
by: Kadra, Arlind, et al.
Published: (2023)
by: Kadra, Arlind, et al.
Published: (2023)
How Usable is Automated Feature Engineering for Tabular Data?
by: Schäfer, Bastian, et al.
Published: (2025)
by: Schäfer, Bastian, et al.
Published: (2025)
nanoTabPFN: A Lightweight and Educational Reimplementation of TabPFN
by: Pfefferle, Alexander, et al.
Published: (2025)
by: Pfefferle, Alexander, et al.
Published: (2025)
Multi-objective Differentiable Neural Architecture Search
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
Learning to Order: Task Sequencing as In-Context Optimization
by: Kobiolka, Jan, et al.
Published: (2026)
by: Kobiolka, Jan, et al.
Published: (2026)
End-to-End Compression for Tabular Foundation Models
by: Zabërgja, Guri, et al.
Published: (2026)
by: Zabërgja, Guri, et al.
Published: (2026)
Tabular Data: Is Deep Learning all you need?
by: Zabërgja, Guri, et al.
Published: (2024)
by: Zabërgja, Guri, et al.
Published: (2024)
POP: Prior-Fitted First-Order Optimization Policies
by: Kobiolka, Jan, et al.
Published: (2026)
by: Kobiolka, Jan, et al.
Published: (2026)
Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
by: Abdalla, M. H. I., et al.
Published: (2025)
by: Abdalla, M. H. I., et al.
Published: (2025)
Lightweight Correlation-Aware Table Compression
by: Stoian, Mihail, et al.
Published: (2024)
by: Stoian, Mihail, et al.
Published: (2024)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks
by: Benjamins, Carolin, et al.
Published: (2025)
by: Benjamins, Carolin, et al.
Published: (2025)
Large Language Models Engineer Too Many Simple Features For Tabular Data
by: Küken, Jaris, et al.
Published: (2024)
by: Küken, Jaris, et al.
Published: (2024)
Multi-Objective Hierarchical Optimization with Large Language Models
by: Schwanke, Andrej, et al.
Published: (2026)
by: Schwanke, Andrej, et al.
Published: (2026)
Improving LLM-based Global Optimization with Search Space Partitioning
by: Schwanke, Andrej, et al.
Published: (2025)
by: Schwanke, Andrej, et al.
Published: (2025)
Causal Data Augmentation for Robust Fine-Tuning of Tabular Foundation Models
by: Bühler, Magnus, et al.
Published: (2026)
by: Bühler, Magnus, et al.
Published: (2026)
Towards Scaling Laws for Symbolic Regression
by: Otte, David, et al.
Published: (2025)
by: Otte, David, et al.
Published: (2025)
Isotuning With Applications To Scale-Free Online Learning
by: Orseau, Laurent, et al.
Published: (2021)
by: Orseau, Laurent, et al.
Published: (2021)
c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperparameter Optimization
by: Watanabe, Shuhei, et al.
Published: (2022)
by: Watanabe, Shuhei, et al.
Published: (2022)
A Human-in-the-Loop Fairness-Aware Model Selection Framework for Complex Fairness Objective Landscapes
by: Robertson, Jake, et al.
Published: (2024)
by: Robertson, Jake, et al.
Published: (2024)
Birth of the Intelligentsia – 1750–1831
by: Janowski, Maciej
Published: (2021)
by: Janowski, Maciej
Published: (2021)
A General Framework for User-Guided Bayesian Optimization
by: Hvarfner, Carl, et al.
Published: (2023)
by: Hvarfner, Carl, et al.
Published: (2023)
Early Stopping Tabular In-Context Learning
by: Küken, Jaris, et al.
Published: (2025)
by: Küken, Jaris, et al.
Published: (2025)
Bayes' Power for Explaining In-Context Learning Generalizations
by: Müller, Samuel, et al.
Published: (2024)
by: Müller, Samuel, et al.
Published: (2024)
Auction-Based Online Policy Adaptation for Evolving Objectives
by: Shabadi, Guruprerana, et al.
Published: (2026)
by: Shabadi, Guruprerana, et al.
Published: (2026)
Channel Estimation by Infinite Width Convolutional Networks
by: Mallik, Mohammed, et al.
Published: (2025)
by: Mallik, Mohammed, et al.
Published: (2025)
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
by: Wu, Boqian, et al.
Published: (2026)
by: Wu, Boqian, et al.
Published: (2026)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
Transfer Learning for Finetuning Large Language Models
by: Strangmann, Tobias, et al.
Published: (2024)
by: Strangmann, Tobias, et al.
Published: (2024)
Similar Items
-
Warmstarting for Scaling Language Models
by: Mallik, Neeratyoy, et al.
Published: (2024) -
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization
by: Rakotoarison, Herilalaina, et al.
Published: (2024) -
Ensembling Finetuned Language Models for Text Classification
by: Arango, Sebastian Pineda, et al.
Published: (2024) -
Regularized Neural Ensemblers
by: Arango, Sebastian Pineda, et al.
Published: (2024) -
An Open-Source Training Dataset for Foundation Models for Black-box Optimization
by: Klein, Aaron, et al.
Published: (2026)