Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Hyundo, Choi, Suhyung, Hwang, Inwoo, Zhang, Byoung-Tak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Locality-aware Concept Bottleneck Model
by: Jeon, Sujin, et al.
Published: (2025)
by: Jeon, Sujin, et al.
Published: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
by: Choi, Won-Seok, et al.
Published: (2025)
by: Choi, Won-Seok, et al.
Published: (2025)
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
On the Consistency of Video Large Language Models in Temporal Comprehension
by: Jung, Minjoon, et al.
Published: (2024)
by: Jung, Minjoon, et al.
Published: (2024)
Unlearning for One-Step Generative Models via Unbalanced Optimal Transport
by: Choi, Hyundo, et al.
Published: (2026)
by: Choi, Hyundo, et al.
Published: (2026)
Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene
by: Lim, Donggeun, et al.
Published: (2025)
by: Lim, Donggeun, et al.
Published: (2025)
Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes
by: Song, Yeon-Ji, et al.
Published: (2026)
by: Song, Yeon-Ji, et al.
Published: (2026)
EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing
by: Hwang, Inwoo, et al.
Published: (2026)
by: Hwang, Inwoo, et al.
Published: (2026)
SnapMoGen: Human Motion Generation from Expressive Texts
by: Guo, Chuan, et al.
Published: (2025)
by: Guo, Chuan, et al.
Published: (2025)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026)
by: Jung, Minjoon, et al.
Published: (2026)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
Less is More: Improving Motion Diffusion Models with Sparse Keyframes
by: Bae, Jinseok, et al.
Published: (2025)
by: Bae, Jinseok, et al.
Published: (2025)
Setting the Stage: Text-Driven Scene-Consistent Image Generation
by: Xie, Cong, et al.
Published: (2025)
by: Xie, Cong, et al.
Published: (2025)
Consistent Image Layout Editing with Diffusion Models
by: Xia, Tao, et al.
Published: (2025)
by: Xia, Tao, et al.
Published: (2025)
One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single Image
by: Wang, Pengfei, et al.
Published: (2026)
by: Wang, Pengfei, et al.
Published: (2026)
Fine-Grained Causal Dynamics Learning with Quantization for Improving Robustness in Reinforcement Learning
by: Hwang, Inwoo, et al.
Published: (2024)
by: Hwang, Inwoo, et al.
Published: (2024)
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
by: Hwang, Inwoo, et al.
Published: (2026)
by: Hwang, Inwoo, et al.
Published: (2026)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization
by: Tang, Xiang, et al.
Published: (2025)
by: Tang, Xiang, et al.
Published: (2025)
4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models
by: Yu, Heng, et al.
Published: (2024)
by: Yu, Heng, et al.
Published: (2024)
MOSAIC: Generating Consistent, Privacy-Preserving Scenes from Multiple Depth Views in Multi-Room Environments
by: Liu, Zhixuan, et al.
Published: (2025)
by: Liu, Zhixuan, et al.
Published: (2025)
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
by: Liu, Lijuan, et al.
Published: (2025)
by: Liu, Lijuan, et al.
Published: (2025)
Consistent Diffusion: Denoising Diffusion Model with Data-Consistent Training for Image Restoration
by: Cheng, Xinlong, et al.
Published: (2024)
by: Cheng, Xinlong, et al.
Published: (2024)
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
by: Lee, JoungBin, et al.
Published: (2025)
by: Lee, JoungBin, et al.
Published: (2025)
Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
by: Xu, Bicheng, et al.
Published: (2024)
by: Xu, Bicheng, et al.
Published: (2024)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
by: Liang, Hanwen, et al.
Published: (2024)
by: Liang, Hanwen, et al.
Published: (2024)
Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
by: Liu, Tianqi, et al.
Published: (2025)
by: Liu, Tianqi, et al.
Published: (2025)
PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization
by: Wu, Dongli, et al.
Published: (2026)
by: Wu, Dongli, et al.
Published: (2026)
MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
Human-Aware 3D Scene Generation with Spatially-constrained Diffusion Models
by: Hong, Xiaolin, et al.
Published: (2024)
by: Hong, Xiaolin, et al.
Published: (2024)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
A Survey on Human Interaction Motion Generation
by: Sui, Kewei, et al.
Published: (2025)
by: Sui, Kewei, et al.
Published: (2025)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
Intrinsic Image Decomposition for Robust Self-supervised Monocular Depth Estimation on Reflective Surfaces
by: Choi, Wonhyeok, et al.
Published: (2025)
by: Choi, Wonhyeok, et al.
Published: (2025)
Similar Items
-
Locality-aware Concept Bottleneck Model
by: Jeon, Sujin, et al.
Published: (2025) -
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024) -
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
by: Choi, Won-Seok, et al.
Published: (2025) -
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025) -
On the Consistency of Video Large Language Models in Temporal Comprehension
by: Jung, Minjoon, et al.
Published: (2024)