Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Bicheng, Yan, Qi, Liao, Renjie, Wang, Lele, Sigal, Leonid |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026)
by: Rahman, Tanzila, et al.
Published: (2026)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
by: Xian, Jia Jun Cheng, et al.
Published: (2025)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
QDM: Quadtree-Based Region-Adaptive Sparse Diffusion Models for Efficient Image Super-Resolution
by: Yang, Donglin, et al.
Published: (2025)
by: Yang, Donglin, et al.
Published: (2025)
Factorized Video Autoencoders for Efficient Generative Modelling
by: Suhail, Mohammed, et al.
Published: (2024)
by: Suhail, Mohammed, et al.
Published: (2024)
Fréchet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos
by: Liu, Jiahe, et al.
Published: (2024)
by: Liu, Jiahe, et al.
Published: (2024)
StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
by: Wu, Zike, et al.
Published: (2025)
by: Wu, Zike, et al.
Published: (2025)
MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation
by: Fu, Yuxiang, et al.
Published: (2025)
by: Fu, Yuxiang, et al.
Published: (2025)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
by: Sakai, Shunsuke, et al.
Published: (2025)
by: Sakai, Shunsuke, et al.
Published: (2025)
SGEdit: Bridging LLM with Text2Image Generative Model for Scene Graph-based Image Editing
by: Zhang, Zhiyuan, et al.
Published: (2024)
by: Zhang, Zhiyuan, et al.
Published: (2024)
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
by: Wu, Weijia, et al.
Published: (2023)
by: Wu, Weijia, et al.
Published: (2023)
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
by: Chinchure, Aditya, et al.
Published: (2023)
by: Chinchure, Aditya, et al.
Published: (2023)
Image Synthesis with Graph Conditioning: CLIP-Guided Diffusion Models for Scene Graphs
by: Mishra, Rameshwar, et al.
Published: (2024)
by: Mishra, Rameshwar, et al.
Published: (2024)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attacks on Breast Ultrasound Images
by: Medghalchi, Yasamin, et al.
Published: (2024)
by: Medghalchi, Yasamin, et al.
Published: (2024)
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
Multi-scale Latent Point Consistency Models for 3D Shape Generation
by: Du, Bi'an, et al.
Published: (2024)
by: Du, Bi'an, et al.
Published: (2024)
GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
by: Ruiz, Antonio, et al.
Published: (2025)
by: Ruiz, Antonio, et al.
Published: (2025)
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2024)
by: Zhai, Guangyao, et al.
Published: (2024)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
by: Bhattacharyya, Apratim, et al.
Published: (2025)
by: Bhattacharyya, Apratim, et al.
Published: (2025)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
by: Chen, Chaofeng, et al.
Published: (2024)
by: Chen, Chaofeng, et al.
Published: (2024)
Controllable 3D Outdoor Scene Generation via Scene Graphs
by: Liu, Yuheng, et al.
Published: (2025)
by: Liu, Yuheng, et al.
Published: (2025)
MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
by: Huang, Xiufeng, et al.
Published: (2025)
by: Huang, Xiufeng, et al.
Published: (2025)
Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent
by: Qin, Ziyuan, et al.
Published: (2024)
by: Qin, Ziyuan, et al.
Published: (2024)
Test-Time Consistency in Vision Language Models
by: Chou, Shih-Han, et al.
Published: (2025)
by: Chou, Shih-Han, et al.
Published: (2025)
MM-R$^3$: On (In-)Consistency of Vision-Language Models (VLMs)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
Supplementing Missing Visions via Dialog for Scene Graph Generations
by: Zhao, Zhenghao, et al.
Published: (2022)
by: Zhao, Zhenghao, et al.
Published: (2022)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
by: Kouzelis, Theodoros, et al.
Published: (2025)
by: Kouzelis, Theodoros, et al.
Published: (2025)
DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis
by: Tang, Jiapeng, et al.
Published: (2023)
by: Tang, Jiapeng, et al.
Published: (2023)
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
by: Chen, Mu, et al.
Published: (2025)
by: Chen, Mu, et al.
Published: (2025)
Tinted Frames: Question Framing Blinds Vision-Language Models
by: Fan, Wan-Cyuan, et al.
Published: (2026)
by: Fan, Wan-Cyuan, et al.
Published: (2026)
JeDi: Joint-Image Diffusion Models for Finetuning-Free Personalized Text-to-Image Generation
by: Zeng, Yu, et al.
Published: (2024)
by: Zeng, Yu, et al.
Published: (2024)
Informative Scene Graph Generation via Debiasing
by: Gao, Lianli, et al.
Published: (2023)
by: Gao, Lianli, et al.
Published: (2023)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)
by: Lee, Hyundo, et al.
Published: (2025)
Using Powerful Prior Knowledge of Diffusion Model in Deep Unfolding Networks for Image Compressive Sensing
by: Liao, Chen, et al.
Published: (2025)
by: Liao, Chen, et al.
Published: (2025)
Scene Graph Generation via Conditional Random Fields
by: Cong, Weilin, et al.
Published: (2018)
by: Cong, Weilin, et al.
Published: (2018)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
Similar Items
-
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
by: Rahman, Tanzila, et al.
Published: (2026) -
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
by: Xian, Jia Jun Cheng, et al.
Published: (2025) -
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024) -
QDM: Quadtree-Based Region-Adaptive Sparse Diffusion Models for Efficient Image Super-Resolution
by: Yang, Donglin, et al.
Published: (2025) -
Factorized Video Autoencoders for Efficient Generative Modelling
by: Suhail, Mohammed, et al.
Published: (2024)