Scaling Properties of Diffusion Models for Perceptual Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917840007725056 |
|---|---|
| author | Ravishankar, Rahul Patel, Zeeshan Rajasegaran, Jathushan Malik, Jitendra |
| author_facet | Ravishankar, Rahul Patel, Zeeshan Rajasegaran, Jathushan Malik, Jitendra |
| contents | In this paper, we argue that iterative computation with diffusion models offers a powerful paradigm for not only generation but also visual perception tasks. We unify tasks such as depth estimation, optical flow, and amodal segmentation under the framework of image-to-image translation, and show how diffusion models benefit from scaling training and test-time compute for these perceptual tasks. Through a careful analysis of these scaling properties, we formulate compute-optimal training and inference recipes to scale diffusion models for visual perception tasks. Our models achieve competitive performance to state-of-the-art methods using significantly less data and compute. To access our code and models, see https://scaling-diffusion-perception.github.io . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_08034 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Scaling Properties of Diffusion Models for Perceptual Tasks Ravishankar, Rahul Patel, Zeeshan Rajasegaran, Jathushan Malik, Jitendra Computer Vision and Pattern Recognition Artificial Intelligence In this paper, we argue that iterative computation with diffusion models offers a powerful paradigm for not only generation but also visual perception tasks. We unify tasks such as depth estimation, optical flow, and amodal segmentation under the framework of image-to-image translation, and show how diffusion models benefit from scaling training and test-time compute for these perceptual tasks. Through a careful analysis of these scaling properties, we formulate compute-optimal training and inference recipes to scale diffusion models for visual perception tasks. Our models achieve competitive performance to state-of-the-art methods using significantly less data and compute. To access our code and models, see https://scaling-diffusion-perception.github.io . |
| title | Scaling Properties of Diffusion Models for Perceptual Tasks |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2411.08034 |