Scaling Laws for Precision
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Tanishq, Ankner, Zachary, Spector, Benjamin F., Bordelon, Blake, Muennighoff, Niklas, Paul, Mansheej, Pehlevan, Cengiz, Ré, Christopher, Raghunathan, Aditi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Dynamical Model of Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
How Feature Learning Can Improve Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
Grokking as the Transition from Lazy to Rich Training Dynamics
di: Kumar, Tanishq, et al.
Pubblicazione: (2023)
di: Kumar, Tanishq, et al.
Pubblicazione: (2023)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
Do Mice Grok? Glimpses of Hidden Progress During Overtraining in Sensory Cortex
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
Transfer Learning in Infinite Width Feature Learning Networks
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
di: Lauditi, Clarissa, et al.
Pubblicazione: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
di: Lauditi, Clarissa, et al.
Pubblicazione: (2026)
di: Lauditi, Clarissa, et al.
Pubblicazione: (2026)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
Infinite Limits of Multi-head Transformer Dynamics
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
di: Bordelon, Blake, et al.
Pubblicazione: (2024)
Hyperparameter Transfer with Mixture-of-Expert Layers
di: Jiang, Tianze, et al.
Pubblicazione: (2026)
di: Jiang, Tianze, et al.
Pubblicazione: (2026)
Critique-out-Loud Reward Models
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
di: Ankner, Zachary, et al.
Pubblicazione: (2024)
Dynamically Learning to Integrate in Recurrent Neural Networks
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
di: Bordelon, Blake, et al.
Pubblicazione: (2025)
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
di: Halder, Indranil, et al.
Pubblicazione: (2026)
di: Halder, Indranil, et al.
Pubblicazione: (2026)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
di: Atanasov, Alexander, et al.
Pubblicazione: (2025)
di: Atanasov, Alexander, et al.
Pubblicazione: (2025)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
di: Bordelon, Blake, et al.
Pubblicazione: (2026)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
di: Halder, Indranil, et al.
Pubblicazione: (2025)
di: Halder, Indranil, et al.
Pubblicazione: (2025)
Mode-Conditioning Unlocks Superior Test-Time Scaling
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
No Free Lunch From Random Feature Ensembles: Scaling Laws and Near-Optimality Conditions
di: Ruben, Benjamin S., et al.
Pubblicazione: (2024)
di: Ruben, Benjamin S., et al.
Pubblicazione: (2024)
Understanding Finetuning for Factual Knowledge Extraction
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
di: Ghosal, Gaurav, et al.
Pubblicazione: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
di: Ruben, Benjamin S., et al.
Pubblicazione: (2023)
di: Ruben, Benjamin S., et al.
Pubblicazione: (2023)
Self-Trained Verification for Training- and Test-Time Self-Improvement
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
di: Wu, Chen Henry, et al.
Pubblicazione: (2026)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
di: Kim, Taeyoun, et al.
Pubblicazione: (2024)
Mitigating Bias in RAG: Controlling the Embedder
di: Kim, Taeyoun, et al.
Pubblicazione: (2025)
di: Kim, Taeyoun, et al.
Pubblicazione: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
di: Blakeney, Cody, et al.
Pubblicazione: (2024)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
Just read twice: closing the recall gap for recurrent language models
di: Arora, Simran, et al.
Pubblicazione: (2024)
di: Arora, Simran, et al.
Pubblicazione: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Don't be lazy: CompleteP enables compute-efficient deep transformers
di: Dey, Nolan, et al.
Pubblicazione: (2025)
di: Dey, Nolan, et al.
Pubblicazione: (2025)
Summary statistics of learning link changing neural representations to behavior
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2025)
di: Zavatone-Veth, Jacob A., et al.
Pubblicazione: (2025)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
di: Maini, Pratyush, et al.
Pubblicazione: (2023)
di: Maini, Pratyush, et al.
Pubblicazione: (2023)
Universal One-third Time Scaling in Learning Peaked Distributions
di: Liu, Yizhou, et al.
Pubblicazione: (2026)
di: Liu, Yizhou, et al.
Pubblicazione: (2026)
Learning richness modulates equality reasoning in neural networks
di: Tong, William L., et al.
Pubblicazione: (2025)
di: Tong, William L., et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Dynamical Model of Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024) -
How Feature Learning Can Improve Neural Scaling Laws
di: Bordelon, Blake, et al.
Pubblicazione: (2024) -
Grokking as the Transition from Lazy to Rich Training Dynamics
di: Kumar, Tanishq, et al.
Pubblicazione: (2023) -
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
di: Bordelon, Blake, et al.
Pubblicazione: (2025) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
di: Bordelon, Blake, et al.
Pubblicazione: (2026)