Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Haoyang, Sun, Dongfang, Ma, Caoyuan, Wang, Shiqin, Zhang, Kewei, Wang, Zheng, Wang, Zhixiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915505631133696
author Chen, Haoyang
Sun, Dongfang
Ma, Caoyuan
Wang, Shiqin
Zhang, Kewei
Wang, Zheng
Wang, Zhixiang
author_facet Chen, Haoyang
Sun, Dongfang
Ma, Caoyuan
Wang, Shiqin
Zhang, Kewei
Wang, Zheng
Wang, Zhixiang
contents We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn rough sketches. Built on optimization-based alignment of diffusion models, our approach avoids large-scale paired training data and mitigates generalization issues. To address the challenge of integrating multiple abstract concepts in real-world scenarios, we design a Sequence-Aware Sketch-Guided Diffusion framework with three loss terms for concept-wise sequential optimization, following the natural order of subjective readouts. Experiments on two datasets demonstrate that our method achieves state-of-the-art performance in image quality as well as spatial and semantic alignment with target scenes. User studies with 40 participants further confirm that our approach is consistently preferred. Our project page is at: subjective-camera.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2506_23711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion
Chen, Haoyang
Sun, Dongfang
Ma, Caoyuan
Wang, Shiqin
Zhang, Kewei
Wang, Zheng
Wang, Zhixiang
Computer Vision and Pattern Recognition
We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn rough sketches. Built on optimization-based alignment of diffusion models, our approach avoids large-scale paired training data and mitigates generalization issues. To address the challenge of integrating multiple abstract concepts in real-world scenarios, we design a Sequence-Aware Sketch-Guided Diffusion framework with three loss terms for concept-wise sequential optimization, following the natural order of subjective readouts. Experiments on two datasets demonstrate that our method achieves state-of-the-art performance in image quality as well as spatial and semantic alignment with target scenes. User studies with 40 participants further confirm that our approach is consistently preferred. Our project page is at: subjective-camera.github.io
title Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23711