How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sengupta, Ayan, Goel, Yash, Chakraborty, Tanmoy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915305869017088
author Sengupta, Ayan
Goel, Yash
Chakraborty, Tanmoy
author_facet Sengupta, Ayan
Goel, Yash
Chakraborty, Tanmoy
contents Neural scaling laws have revolutionized the design and optimization of large-scale AI models by revealing predictable relationships between model size, dataset volume, and computational resources. Early research established power-law relationships in model performance, leading to compute-optimal scaling strategies. However, recent studies highlighted their limitations across architectures, modalities, and deployment contexts. Sparse models, mixture-of-experts, retrieval-augmented learning, and multimodal models often deviate from traditional scaling patterns. Moreover, scaling behaviors vary across domains such as vision, reinforcement learning, and fine-tuning, underscoring the need for more nuanced approaches. In this survey, we synthesize insights from over 50 studies, examining the theoretical foundations, empirical findings, and practical implications of scaling laws. We also explore key challenges, including data efficiency, inference scaling, and architecture-specific constraints, advocating for adaptive scaling strategies tailored to real-world applications. We suggest that while scaling laws provide a useful guide, they do not always generalize across all architectures and training strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
Sengupta, Ayan
Goel, Yash
Chakraborty, Tanmoy
Computation and Language
Machine Learning
Neural scaling laws have revolutionized the design and optimization of large-scale AI models by revealing predictable relationships between model size, dataset volume, and computational resources. Early research established power-law relationships in model performance, leading to compute-optimal scaling strategies. However, recent studies highlighted their limitations across architectures, modalities, and deployment contexts. Sparse models, mixture-of-experts, retrieval-augmented learning, and multimodal models often deviate from traditional scaling patterns. Moreover, scaling behaviors vary across domains such as vision, reinforcement learning, and fine-tuning, underscoring the need for more nuanced approaches. In this survey, we synthesize insights from over 50 studies, examining the theoretical foundations, empirical findings, and practical implications of scaling laws. We also explore key challenges, including data efficiency, inference scaling, and architecture-specific constraints, advocating for adaptive scaling strategies tailored to real-world applications. We suggest that while scaling laws provide a useful guide, they do not always generalize across all architectures and training strategies.
title How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.12051