Training Optimal Large Diffusion Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ni, Jinjie, Liu, Qian, Du, Chao, Dou, Longxu, Yan, Hang, Wang, Zili, Pang, Tianyu, Shieh, Michael Qizhe
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915598388166656
author Ni, Jinjie
Liu, Qian
Du, Chao
Dou, Longxu
Yan, Hang
Wang, Zili
Pang, Tianyu
Shieh, Michael Qizhe
author_facet Ni, Jinjie
Liu, Qian
Du, Chao
Dou, Longxu
Yan, Hang
Wang, Zili
Pang, Tianyu
Shieh, Michael Qizhe
contents We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training Optimal Large Diffusion Language Models
Ni, Jinjie
Liu, Qian
Du, Chao
Dou, Longxu
Yan, Hang
Wang, Zili
Pang, Tianyu
Shieh, Michael Qizhe
Machine Learning
Artificial Intelligence
Computation and Language
We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community.
title Training Optimal Large Diffusion Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.03280