Prescriptive Scaling Laws for Data Constrained Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Lovelace, Justin, Belardi, Christian, Kundurthy, Srivatsa, Sudhakar, Shriya, Weinberger, Kilian Q. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
di: Lovelace, Justin, et al.
Pubblicazione: (2026)
Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling
di: Belardi, Christian, et al.
Pubblicazione: (2026)
di: Belardi, Christian, et al.
Pubblicazione: (2026)
Diffusion Guided Language Modeling
di: Lovelace, Justin, et al.
Pubblicazione: (2024)
di: Lovelace, Justin, et al.
Pubblicazione: (2024)
IncDSI: Incrementally Updatable Document Retrieval
di: Kishore, Varsha, et al.
Pubblicazione: (2023)
di: Kishore, Varsha, et al.
Pubblicazione: (2023)
Pre-training Limited Memory Language Models with Internal and External Knowledge
di: Zhao, Linxi, et al.
Pubblicazione: (2025)
di: Zhao, Linxi, et al.
Pubblicazione: (2025)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
di: Kundurthy, Srivatsa, et al.
Pubblicazione: (2026)
Sample-Efficient Diffusion for Text-To-Speech Synthesis
di: Lovelace, Justin, et al.
Pubblicazione: (2024)
di: Lovelace, Justin, et al.
Pubblicazione: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
Scaling Law for Quantization-Aware Training
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
di: Chen, Mengzhao, et al.
Pubblicazione: (2025)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
Scaling Multi-Hop Training Data via Graph-Constrained Path Selection
di: Chen, Pengyu, et al.
Pubblicazione: (2026)
di: Chen, Pengyu, et al.
Pubblicazione: (2026)
Scaling Laws for Post Training Quantized Large Language Models
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
Scaling Law for Language Models Training Considering Batch Size
di: Shuai, Xian, et al.
Pubblicazione: (2024)
di: Shuai, Xian, et al.
Pubblicazione: (2024)
gzip Predicts Data-dependent Scaling Laws
di: Pandey, Rohan
Pubblicazione: (2024)
di: Pandey, Rohan
Pubblicazione: (2024)
SpeechOp: Inference-Time Task Composition for Generative Speech Processing
di: Lovelace, Justin, et al.
Pubblicazione: (2025)
di: Lovelace, Justin, et al.
Pubblicazione: (2025)
Exploring Scaling Laws for Local SGD in Large Language Model Training
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
di: He, Qiaozhi, et al.
Pubblicazione: (2024)
Learning from Synthetic Data Improves Multi-hop Reasoning
di: Kabra, Anmol, et al.
Pubblicazione: (2026)
di: Kabra, Anmol, et al.
Pubblicazione: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
di: Sedova, Anastasiia, et al.
Pubblicazione: (2026)
Scaling Laws for Floating Point Quantization Training
di: Sun, Xingwu, et al.
Pubblicazione: (2025)
di: Sun, Xingwu, et al.
Pubblicazione: (2025)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
di: Zhang, Hanlin, et al.
Pubblicazione: (2026)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
di: Bethune, Louis, et al.
Pubblicazione: (2025)
di: Bethune, Louis, et al.
Pubblicazione: (2025)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
Scaling Laws for Precision
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
di: Kumar, Tanishq, et al.
Pubblicazione: (2024)
Scaling Data-Constrained Language Models
di: Muennighoff, Niklas, et al.
Pubblicazione: (2023)
di: Muennighoff, Niklas, et al.
Pubblicazione: (2023)
Neural Neural Scaling Laws
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
di: Hu, Michael Y., et al.
Pubblicazione: (2026)
Scaling Laws For Mixed Quantization
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
di: Cao, Zeyu, et al.
Pubblicazione: (2024)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Training LLMs for Multi-Step Tool Orchestration with Constrained Data Synthesis and Graduated Rewards
di: Jiayang, Cheng, et al.
Pubblicazione: (2026)
di: Jiayang, Cheng, et al.
Pubblicazione: (2026)
Music Transcription with (Almost) No Supervision
di: Shin, Saebyeol, et al.
Pubblicazione: (2026)
di: Shin, Saebyeol, et al.
Pubblicazione: (2026)
Unified Scaling Laws for Compressed Representations
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
di: Panferov, Andrei, et al.
Pubblicazione: (2025)
Parallel Scaling Law for Language Models
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
di: Chen, Mouxiang, et al.
Pubblicazione: (2025)
Reconciling Kaplan and Chinchilla Scaling Laws
di: Pearce, Tim, et al.
Pubblicazione: (2024)
di: Pearce, Tim, et al.
Pubblicazione: (2024)
Scaling Laws for Multilingual Language Models
di: He, Yifei, et al.
Pubblicazione: (2024)
di: He, Yifei, et al.
Pubblicazione: (2024)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
di: Liu, Lei, et al.
Pubblicazione: (2025)
di: Liu, Lei, et al.
Pubblicazione: (2025)
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
di: Gong, Albert, et al.
Pubblicazione: (2025)
di: Gong, Albert, et al.
Pubblicazione: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
di: Sadhukhan, Ranajoy, et al.
Pubblicazione: (2025)
Unraveling the Mystery of Scaling Laws: Part I
di: Su, Hui, et al.
Pubblicazione: (2024)
di: Su, Hui, et al.
Pubblicazione: (2024)
Compression Scaling Laws:Unifying Sparsity and Quantization
di: Frantar, Elias, et al.
Pubblicazione: (2025)
di: Frantar, Elias, et al.
Pubblicazione: (2025)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
di: Mayilvahanan, Prasanna, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
di: Lovelace, Justin, et al.
Pubblicazione: (2026) -
Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling
di: Belardi, Christian, et al.
Pubblicazione: (2026) -
Diffusion Guided Language Modeling
di: Lovelace, Justin, et al.
Pubblicazione: (2024) -
IncDSI: Incrementally Updatable Document Retrieval
di: Kishore, Varsha, et al.
Pubblicazione: (2023) -
Pre-training Limited Memory Language Models with Internal and External Knowledge
di: Zhao, Linxi, et al.
Pubblicazione: (2025)