Evaluating Compositional Scene Understanding in Multimodal Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shuhao, Lee, Andrew Jun, Wang, Anna, Momennejad, Ida, Bihl, Trevor, Lu, Hongjing, Webb, Taylor W. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
by: Lee, Andrew Jun, et al.
Published: (2025)
by: Lee, Andrew Jun, et al.
Published: (2025)
Hierarchical Abstraction Enables Human-Like 3D Object Recognition in Deep Learning Models
by: Fu, Shuhao, et al.
Published: (2025)
by: Fu, Shuhao, et al.
Published: (2025)
Compositional Scene Understanding through Inverse Generative Modeling
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Structured Generative Models for Scene Understanding
by: Williams, Christopher K. I.
Published: (2023)
by: Williams, Christopher K. I.
Published: (2023)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
by: Du, Sinan, et al.
Published: (2025)
by: Du, Sinan, et al.
Published: (2025)
Compositional Feature Augmentation for Unbiased Scene Graph Generation
by: Li, Lin, et al.
Published: (2023)
by: Li, Lin, et al.
Published: (2023)
GaussianGraph: 3D Gaussian-based Scene Graph Generation for Open-world Scene Understanding
by: Wang, Xihan, et al.
Published: (2025)
by: Wang, Xihan, et al.
Published: (2025)
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind
by: Li, Qingmei, et al.
Published: (2025)
by: Li, Qingmei, et al.
Published: (2025)
FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
by: Yang, Zhifei, et al.
Published: (2026)
by: Yang, Zhifei, et al.
Published: (2026)
Adaptive Visual Scene Understanding: Incremental Scene Graph Generation
by: Khandelwal, Naitik, et al.
Published: (2023)
by: Khandelwal, Naitik, et al.
Published: (2023)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025)
by: Wu, Yanhao, et al.
Published: (2025)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
by: Li, Teng, et al.
Published: (2025)
by: Li, Teng, et al.
Published: (2025)
SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation
by: Luo, Jun, et al.
Published: (2026)
by: Luo, Jun, et al.
Published: (2026)
SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding
by: Zhu, Dekai, et al.
Published: (2025)
by: Zhu, Dekai, et al.
Published: (2025)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
Generative Learning of Differentiable Object Models for Compositional Interpretation of Complex Scenes
by: Nowinowski, Antoni, et al.
Published: (2025)
by: Nowinowski, Antoni, et al.
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
SceneTransporter: Optimal Transport-Guided Compositional Latent Diffusion for Single-Image Structured 3D Scene Generation
by: Wang, Ling, et al.
Published: (2026)
by: Wang, Ling, et al.
Published: (2026)
SceneGPT: A Language Model for 3D Scene Understanding
by: Chandhok, Shivam
Published: (2024)
by: Chandhok, Shivam
Published: (2024)
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
by: Feng, Jun, et al.
Published: (2025)
by: Feng, Jun, et al.
Published: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
by: Kim, Seok-Young, et al.
Published: (2026)
by: Kim, Seok-Young, et al.
Published: (2026)
Unified 3D Scene Understanding Through Physical World Modeling
by: Lee, Wanhee, et al.
Published: (2026)
by: Lee, Wanhee, et al.
Published: (2026)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
by: Li, Haoxin, et al.
Published: (2025)
by: Li, Haoxin, et al.
Published: (2025)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
by: Berman, Nimrod, et al.
Published: (2025)
by: Berman, Nimrod, et al.
Published: (2025)
Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
by: Pan, Jiadong, et al.
Published: (2026)
by: Pan, Jiadong, et al.
Published: (2026)
Syn-Mediverse: A Multimodal Synthetic Dataset for Intelligent Scene Understanding of Healthcare Facilities
by: Mohan, Rohit, et al.
Published: (2023)
by: Mohan, Rohit, et al.
Published: (2023)
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
by: Tong, Yujun, et al.
Published: (2026)
by: Tong, Yujun, et al.
Published: (2026)
Similar Items
-
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
by: Lee, Andrew Jun, et al.
Published: (2025) -
Hierarchical Abstraction Enables Human-Like 3D Object Recognition in Deep Learning Models
by: Fu, Shuhao, et al.
Published: (2025) -
Compositional Scene Understanding through Inverse Generative Modeling
by: Wang, Yanbo, et al.
Published: (2025) -
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
by: Liu, Pei, et al.
Published: (2025) -
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
by: Xiong, Haomiao, et al.
Published: (2025)