Generative Models: What Do They Know? Do They Know Things? Let's Find Out!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Du, Xiaodan, Kolkin, Nicholas, Shakhnarovich, Greg, Bhattad, Anand
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914974339694592
author Du, Xiaodan
Kolkin, Nicholas
Shakhnarovich, Greg
Bhattad, Anand
author_facet Du, Xiaodan
Kolkin, Nicholas
Shakhnarovich, Greg
Bhattad, Anand
contents Generative models excel at mimicking real scenes, suggesting they might inherently encode important intrinsic scene properties. In this paper, we aim to explore the following key questions: (1) What intrinsic knowledge do generative models like GANs, Autoregressive models, and Diffusion models encode? (2) Can we establish a general framework to recover intrinsic representations from these models, regardless of their architecture or model type? (3) How minimal can the required learnable parameters and labeled data be to successfully recover this knowledge? (4) Is there a direct link between the quality of a generative model and the accuracy of the recovered scene intrinsics? Our findings indicate that a small Low-Rank Adaptators (LoRA) can recover intrinsic images-depth, normals, albedo and shading-across different generators (Autoregressive, GANs and Diffusion) while using the same decoder head that generates the image. As LoRA is lightweight, we introduce very few learnable parameters (as few as 0.04% of Stable Diffusion model weights for a rank of 2), and we find that as few as 250 labeled images are enough to generate intrinsic images with these LoRA modules. Finally, we also show a positive correlation between the generative model's quality and the accuracy of the recovered intrinsics through control experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17137
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
Du, Xiaodan
Kolkin, Nicholas
Shakhnarovich, Greg
Bhattad, Anand
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Generative models excel at mimicking real scenes, suggesting they might inherently encode important intrinsic scene properties. In this paper, we aim to explore the following key questions: (1) What intrinsic knowledge do generative models like GANs, Autoregressive models, and Diffusion models encode? (2) Can we establish a general framework to recover intrinsic representations from these models, regardless of their architecture or model type? (3) How minimal can the required learnable parameters and labeled data be to successfully recover this knowledge? (4) Is there a direct link between the quality of a generative model and the accuracy of the recovered scene intrinsics? Our findings indicate that a small Low-Rank Adaptators (LoRA) can recover intrinsic images-depth, normals, albedo and shading-across different generators (Autoregressive, GANs and Diffusion) while using the same decoder head that generates the image. As LoRA is lightweight, we introduce very few learnable parameters (as few as 0.04% of Stable Diffusion model weights for a rank of 2), and we find that as few as 250 labeled images are enough to generate intrinsic images with these LoRA modules. Finally, we also show a positive correlation between the generative model's quality and the accuracy of the recovered intrinsics through control experiments.
title Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2311.17137