Training Optimal Large Diffusion Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915598388166656 |
|---|---|
| author | Ni, Jinjie Liu, Qian Du, Chao Dou, Longxu Yan, Hang Wang, Zili Pang, Tianyu Shieh, Michael Qizhe |
| author_facet | Ni, Jinjie Liu, Qian Du, Chao Dou, Longxu Yan, Hang Wang, Zili Pang, Tianyu Shieh, Michael Qizhe |
| contents | We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03280 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Training Optimal Large Diffusion Language Models Ni, Jinjie Liu, Qian Du, Chao Dou, Longxu Yan, Hang Wang, Zili Pang, Tianyu Shieh, Michael Qizhe Machine Learning Artificial Intelligence Computation and Language We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community. |
| title | Training Optimal Large Diffusion Language Models |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2510.03280 |