Paris 2.0: A Decentralized Diffusion Model for Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rouzbayani, Ali, Roy, Bidhan, Villagra, Marcos, Jiang, Zhiying
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913170453430272
author Rouzbayani, Ali
Roy, Bidhan
Villagra, Marcos
Jiang, Zhiying
author_facet Rouzbayani, Ali
Roy, Bidhan
Villagra, Marcos
Jiang, Zhiying
contents We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentralized training, and Paris 2.0 closes it. In low-resolution text-to-video training, against a monolithic model trained on the same data under a matched total compute budget, Paris 2.0 cuts Frechet Video Distance (FVD) from 561.04 to 279.01, a ~2.0x improvement, and lifts CLIP text-video similarity and aesthetic score.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26064
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Paris 2.0: A Decentralized Diffusion Model for Video Generation
Rouzbayani, Ali
Roy, Bidhan
Villagra, Marcos
Jiang, Zhiying
Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.2.11
We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentralized training, and Paris 2.0 closes it. In low-resolution text-to-video training, against a monolithic model trained on the same data under a matched total compute budget, Paris 2.0 cuts Frechet Video Distance (FVD) from 561.04 to 279.01, a ~2.0x improvement, and lifts CLIP text-video similarity and aesthetic score.
title Paris 2.0: A Decentralized Diffusion Model for Video Generation
topic Computer Vision and Pattern Recognition
Machine Learning
I.2.10; I.2.11
url https://arxiv.org/abs/2605.26064