WorldScore: A Unified Evaluation Benchmark for World Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917112877940736 |
|---|---|
| author | Duan, Haoyi Yu, Hong-Xing Chen, Sirui Fei-Fei, Li Wu, Jiajun |
| author_facet | Duan, Haoyi Yu, Hong-Xing Chen, Sirui Fei-Fei, Li Wu, Jiajun |
| contents | We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifications, enabling unified evaluation of diverse approaches from 3D and 4D scene generation to video generation models. The WorldScore benchmark encompasses a curated dataset of 3,000 test examples that span diverse worlds: static and dynamic, indoor and outdoor, photorealistic and stylized. The WorldScore metrics evaluate generated worlds through three key aspects: controllability, quality, and dynamics. Through extensive evaluation of 19 representative models, including both open-source and closed-source ones, we reveal key insights and challenges for each category of models. Our dataset, evaluation code, and leaderboard can be found at https://haoyi-duan.github.io/WorldScore/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_00983 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WorldScore: A Unified Evaluation Benchmark for World Generation Duan, Haoyi Yu, Hong-Xing Chen, Sirui Fei-Fei, Li Wu, Jiajun Graphics Artificial Intelligence Computer Vision and Pattern Recognition We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifications, enabling unified evaluation of diverse approaches from 3D and 4D scene generation to video generation models. The WorldScore benchmark encompasses a curated dataset of 3,000 test examples that span diverse worlds: static and dynamic, indoor and outdoor, photorealistic and stylized. The WorldScore metrics evaluate generated worlds through three key aspects: controllability, quality, and dynamics. Through extensive evaluation of 19 representative models, including both open-source and closed-source ones, we reveal key insights and challenges for each category of models. Our dataset, evaluation code, and leaderboard can be found at https://haoyi-duan.github.io/WorldScore/ |
| title | WorldScore: A Unified Evaluation Benchmark for World Generation |
| topic | Graphics Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.00983 |