Gespeichert in:
| Hauptverfasser: | Li, Ruihang, Wei, Yixuan, Zhang, Miaosen, Yu, Nenghai, Hu, Han, Peng, Houwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2408.08310 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
von: Ni, Bolin, et al.
Veröffentlicht: (2024)
von: Ni, Bolin, et al.
Veröffentlicht: (2024)
Common 7B Language Models Already Possess Strong Math Capabilities
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
von: Liu, Fengze, et al.
Veröffentlicht: (2026)
von: Liu, Fengze, et al.
Veröffentlicht: (2026)
LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
von: Zhang, Sikui, et al.
Veröffentlicht: (2025)
von: Zhang, Sikui, et al.
Veröffentlicht: (2025)
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
Scaling Laws of Synthetic Data for Language Models
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
Uncovering Scaling Laws for Large Language Models via Inverse Problems
von: Verma, Arun, et al.
Veröffentlicht: (2025)
von: Verma, Arun, et al.
Veröffentlicht: (2025)
Scaling Laws for Code: A More Data-Hungry Regime
von: Luo, Xianzhen, et al.
Veröffentlicht: (2025)
von: Luo, Xianzhen, et al.
Veröffentlicht: (2025)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
von: Alrashed, Sultan, et al.
Veröffentlicht: (2025)
von: Alrashed, Sultan, et al.
Veröffentlicht: (2025)
Flora: Effortless Context Construction to Arbitrary Length and Scale
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
The Scaling Laws of Skills in LLM Agent Systems
von: Chen, Charles, et al.
Veröffentlicht: (2026)
von: Chen, Charles, et al.
Veröffentlicht: (2026)
Temporal Scaling Law for Large Language Models
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
Neural Neural Scaling Laws
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
What Scales in Cross-Entropy Scaling Law?
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
von: Zeng, Liang, et al.
Veröffentlicht: (2025)
von: Zeng, Liang, et al.
Veröffentlicht: (2025)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
gzip Predicts Data-dependent Scaling Laws
von: Pandey, Rohan
Veröffentlicht: (2024)
von: Pandey, Rohan
Veröffentlicht: (2024)
Prescriptive Scaling Laws for Data Constrained Training
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
Scaling Laws for Code: Every Programming Language Matters
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
von: Bi, Ziqian, et al.
Veröffentlicht: (2025)
von: Bi, Ziqian, et al.
Veröffentlicht: (2025)
Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
von: Yang, Wen, et al.
Veröffentlicht: (2025)
von: Yang, Wen, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Scaling Laws for Multilingual Language Models
von: He, Yifei, et al.
Veröffentlicht: (2024)
von: He, Yifei, et al.
Veröffentlicht: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
Scaling Laws for Floating Point Quantization Training
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
von: Sun, Xingwu, et al.
Veröffentlicht: (2025)
Relative Scaling Laws for LLMs
von: Held, William, et al.
Veröffentlicht: (2025)
von: Held, William, et al.
Veröffentlicht: (2025)
Scaling Laws For Mixed Quantization
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
von: Cao, Zeyu, et al.
Veröffentlicht: (2024)
Scaling Laws for Precision
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Scaling Laws for Mixture Pretraining Under Data Constraints
von: Sedova, Anastasiia, et al.
Veröffentlicht: (2026)
von: Sedova, Anastasiia, et al.
Veröffentlicht: (2026)
Scaling Law for Quantization-Aware Training
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Scaling Laws for Linear Complexity Language Models
von: Shen, Xuyang, et al.
Veröffentlicht: (2024)
von: Shen, Xuyang, et al.
Veröffentlicht: (2024)
Scale-Free Graph-Language Models
von: Lu, Jianglin, et al.
Veröffentlicht: (2025)
von: Lu, Jianglin, et al.
Veröffentlicht: (2025)
Scaling Laws in Scientific Discovery with AI and Robot Scientists
von: Zhang, Pengsong, et al.
Veröffentlicht: (2025)
von: Zhang, Pengsong, et al.
Veröffentlicht: (2025)
The Limits of Data Scaling: Sub-token Utilization and Acoustic Saturation in Multilingual ASR
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
von: Liang, Siyu, et al.
Veröffentlicht: (2025)
Scaling Parameter-Constrained Language Models with Quality Data
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
von: Chang, Ernie, et al.
Veröffentlicht: (2024)
Exploring Training and Inference Scaling Laws in Generative Retrieval
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
von: Bethune, Louis, et al.
Veröffentlicht: (2025)
von: Bethune, Louis, et al.
Veröffentlicht: (2025)
Scaling Laws For Dense Retrieval
von: Fang, Yan, et al.
Veröffentlicht: (2024)
von: Fang, Yan, et al.
Veröffentlicht: (2024)
Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
von: Lourie, Nicholas, et al.
Veröffentlicht: (2025)
von: Lourie, Nicholas, et al.
Veröffentlicht: (2025)
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
von: Ni, Bolin, et al.
Veröffentlicht: (2024) -
Common 7B Language Models Already Possess Strong Math Capabilities
von: Li, Chen, et al.
Veröffentlicht: (2024) -
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
von: Liu, Fengze, et al.
Veröffentlicht: (2026) -
LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
von: Zhang, Sikui, et al.
Veröffentlicht: (2025) -
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)