How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
Fuente:
arXiv
Saved in:
| Main Authors: | Sengupta, Ayan, Goel, Yash, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
by: Goel, Yash, et al.
Published: (2025)
by: Goel, Yash, et al.
Published: (2025)
Persona-aware Generative Model for Code-mixed Language
by: Sengupta, Ayan, et al.
Published: (2023)
by: Sengupta, Ayan, et al.
Published: (2023)
Compression Laws for Large Language Models
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
The Art of Scaling Test-Time Compute for Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)
by: Agarwal, Aradhye, et al.
Published: (2025)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)
by: Agarwal, Aradhye, et al.
Published: (2025)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
by: Sengupta, Ayan, et al.
Published: (2024)
by: Sengupta, Ayan, et al.
Published: (2024)
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
by: Dutta, Subhabrata, et al.
Published: (2024)
by: Dutta, Subhabrata, et al.
Published: (2024)
Latent Performance Profiling of Large Language Models
by: Chakraborty, Tanmoy, et al.
Published: (2026)
by: Chakraborty, Tanmoy, et al.
Published: (2026)
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
by: Ananthanarayanan, Samhruth, et al.
Published: (2026)
Neural Neural Scaling Laws
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
by: Chatterjee, Anwoy, et al.
Published: (2025)
by: Chatterjee, Anwoy, et al.
Published: (2025)
On the Generalization vs Fidelity Paradox in Knowledge Distillation
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
by: Ramesh, Suhas Kamasetty, et al.
Published: (2025)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
by: Joshi, Maithili, et al.
Published: (2025)
by: Joshi, Maithili, et al.
Published: (2025)
An Analysis of Embedding Layers and Similarity Scores using Siamese Neural Networks
by: Bingi, Yash, et al.
Published: (2023)
by: Bingi, Yash, et al.
Published: (2023)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Temporally Consistent Factuality Probing for Large Language Models
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
by: Singh, Joykirat, et al.
Published: (2025)
by: Singh, Joykirat, et al.
Published: (2025)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers
by: Sengupta, Ayan, et al.
Published: (2026)
by: Sengupta, Ayan, et al.
Published: (2026)
Step-by-Step Unmasking for Parameter-Efficient Fine-tuning of Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2024)
by: Agarwal, Aradhye, et al.
Published: (2024)
A Survey : Neural Networks for AMR-to-Text
by: Hao, Hongyu, et al.
Published: (2022)
by: Hao, Hongyu, et al.
Published: (2022)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
by: Bajpai, Ashutosh, et al.
Published: (2024)
by: Bajpai, Ashutosh, et al.
Published: (2024)
Multilingual Language Models Encode Script Over Linguistic Structure
by: Verma, Aastha A K, et al.
Published: (2026)
by: Verma, Aastha A K, et al.
Published: (2026)
PerSEval: Assessing Personalization in Text Summarizers
by: Dasgupta, Sourish, et al.
Published: (2024)
by: Dasgupta, Sourish, et al.
Published: (2024)
Scaling Laws for Precision
by: Kumar, Tanishq, et al.
Published: (2024)
by: Kumar, Tanishq, et al.
Published: (2024)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
by: Schwethelm, Kristian, et al.
Published: (2026)
by: Schwethelm, Kristian, et al.
Published: (2026)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
by: Kartik, Kartik, et al.
Published: (2024)
by: Kartik, Kartik, et al.
Published: (2024)
A Survey of Early Exit Deep Neural Networks in NLP
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Relative-Based Scaling Law for Neural Language Models
by: Yue, Baoqing, et al.
Published: (2025)
by: Yue, Baoqing, et al.
Published: (2025)
Scaling Laws For Mixed Quantization
by: Cao, Zeyu, et al.
Published: (2024)
by: Cao, Zeyu, et al.
Published: (2024)
Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
by: Chatterjee, Parthiv, et al.
Published: (2025)
by: Chatterjee, Parthiv, et al.
Published: (2025)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Scaling Law for Quantization-Aware Training
by: Chen, Mengzhao, et al.
Published: (2025)
by: Chen, Mengzhao, et al.
Published: (2025)
Reconciling Kaplan and Chinchilla Scaling Laws
by: Pearce, Tim, et al.
Published: (2024)
by: Pearce, Tim, et al.
Published: (2024)
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Similar Items
-
Position: Enough of Scaling LLMs! Lets Focus on Downscaling
by: Goel, Yash, et al.
Published: (2025) -
Persona-aware Generative Model for Code-mixed Language
by: Sengupta, Ayan, et al.
Published: (2023) -
Compression Laws for Large Language Models
by: Sengupta, Ayan, et al.
Published: (2025) -
The Art of Scaling Test-Time Compute for Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025) -
First Finish Search: Efficient Test-Time Scaling in Large Language Models
by: Agarwal, Aradhye, et al.
Published: (2025)