Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alnemari, Mohammed, Qureshi, Rizwan, Begrazadah, Nader |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026)
von: Mysore, Naveen
Veröffentlicht: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
von: Dai, Yanning, et al.
Veröffentlicht: (2026)
von: Dai, Yanning, et al.
Veröffentlicht: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
von: Ho, Siu Hang, et al.
Veröffentlicht: (2025)
von: Ho, Siu Hang, et al.
Veröffentlicht: (2025)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
von: Agrawal, Vidhi, et al.
Veröffentlicht: (2026)
von: Agrawal, Vidhi, et al.
Veröffentlicht: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
von: Merin, Aur Shalev
Veröffentlicht: (2026)
von: Merin, Aur Shalev
Veröffentlicht: (2026)
Is Cambodia the World's Largest Cashew Producer?
von: Chaya, Veasna, et al.
Veröffentlicht: (2024)
von: Chaya, Veasna, et al.
Veröffentlicht: (2024)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
Revisiting GAN with Bayes-Optimal Discrimination
von: Naeini, Mohammadreza Tavasoli, et al.
Veröffentlicht: (2025)
von: Naeini, Mohammadreza Tavasoli, et al.
Veröffentlicht: (2025)
JacNet: Learning Functions with Structured Jacobians
von: Lorraine, Jonathan, et al.
Veröffentlicht: (2024)
von: Lorraine, Jonathan, et al.
Veröffentlicht: (2024)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
von: Lau, Matthew, et al.
Veröffentlicht: (2023)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
von: Di Santi, Eduardo, et al.
Veröffentlicht: (2025)
von: Di Santi, Eduardo, et al.
Veröffentlicht: (2025)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
von: Rahimiasl, Mohammadmahdi, et al.
Veröffentlicht: (2025)
von: Rahimiasl, Mohammadmahdi, et al.
Veröffentlicht: (2025)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
von: Sharma, Paras
Veröffentlicht: (2026)
von: Sharma, Paras
Veröffentlicht: (2026)
GraphNNK -- Graph Classification and Interpretability
von: Bolevic, Zeljko, et al.
Veröffentlicht: (2026)
von: Bolevic, Zeljko, et al.
Veröffentlicht: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026)
von: Mysore, Naveen
Veröffentlicht: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
von: Hospodarchuk, Dmytro
Veröffentlicht: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
von: Berthier, Louis, et al.
Veröffentlicht: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
von: BC, Samiksha
Veröffentlicht: (2025)
von: BC, Samiksha
Veröffentlicht: (2025)
CellARC: Measuring Intelligence with Cellular Automata
von: Lžičař, Miroslav
Veröffentlicht: (2025)
von: Lžičař, Miroslav
Veröffentlicht: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
von: Johnson, David R., et al.
Veröffentlicht: (2025)
von: Johnson, David R., et al.
Veröffentlicht: (2025)
Towards understanding how attention mechanism works in deep learning
von: Ruan, Tianyu, et al.
Veröffentlicht: (2024)
von: Ruan, Tianyu, et al.
Veröffentlicht: (2024)
Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
von: Kumar, Yash
Veröffentlicht: (2025)
von: Kumar, Yash
Veröffentlicht: (2025)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
von: K, Prasanth K, et al.
Veröffentlicht: (2025)
GRALIS: A Unified Canonical Framework for Linear Attribution Methods via Riesz Representation
von: Fanale, Raimondo
Veröffentlicht: (2026)
von: Fanale, Raimondo
Veröffentlicht: (2026)
KAN vs LSTM Performance in Time Series Forecasting
von: Rather, Tabish Ali, et al.
Veröffentlicht: (2025)
von: Rather, Tabish Ali, et al.
Veröffentlicht: (2025)
On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
von: Štrupl, Miroslav, et al.
Veröffentlicht: (2025)
von: Štrupl, Miroslav, et al.
Veröffentlicht: (2025)
A Boltzmann-machine-enhanced Transformer For DNA Sequence Classification
von: Cao, Zhixuan, et al.
Veröffentlicht: (2026)
von: Cao, Zhixuan, et al.
Veröffentlicht: (2026)
Model Input-Output Configuration Search with Embedded Feature Selection for Sensor Time-series and Image Classification
von: Hoang, Anh T., et al.
Veröffentlicht: (2023)
von: Hoang, Anh T., et al.
Veröffentlicht: (2023)
Predicting Traffic Accident Severity with Deep Neural Networks
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
von: Bibb, Meghan, et al.
Veröffentlicht: (2025)
An in-depth look at approximation via deep and narrow neural networks
von: Dommel, Joris, et al.
Veröffentlicht: (2025)
von: Dommel, Joris, et al.
Veröffentlicht: (2025)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
Label Smoothing is a Pragmatic Information Bottleneck
von: Kudo, Sota
Veröffentlicht: (2025)
von: Kudo, Sota
Veröffentlicht: (2025)
Hybrid Imbalanced Regression Through Unified Data-Level and Algorithm-Level Balancing
von: Shahbazi, Shermin, et al.
Veröffentlicht: (2026)
von: Shahbazi, Shermin, et al.
Veröffentlicht: (2026)
Tricks and Plug-ins for Gradient Boosting with Transformers
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
iLTM: Integrated Large Tabular Model
von: Bonet, David, et al.
Veröffentlicht: (2025)
von: Bonet, David, et al.
Veröffentlicht: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025) -
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
von: Mysore, Naveen
Veröffentlicht: (2026) -
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
von: Mukherjee, Debdeep, et al.
Veröffentlicht: (2025) -
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
von: Dai, Yanning, et al.
Veröffentlicht: (2026) -
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
von: Ho, Siu Hang, et al.
Veröffentlicht: (2025)