Evaluating the Generation of Spatial Relations in Text and Image Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Sim, Shang Hong, Lee, Clarence, Tan, Alvin, Tan, Cheston |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
by: Agrawal, Palaash, et al.
Published: (2023)
by: Agrawal, Palaash, et al.
Published: (2023)
Stencil: Subject-Driven Generation with Context Guidance
by: Chen, Gordon, et al.
Published: (2025)
by: Chen, Gordon, et al.
Published: (2025)
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
by: Jia, Yuhao, et al.
Published: (2024)
by: Jia, Yuhao, et al.
Published: (2024)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
by: Tan, Yaoteng, et al.
Published: (2026)
by: Tan, Yaoteng, et al.
Published: (2026)
Human-like compositional learning of visually-grounded concepts using synthetic environments
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation
by: Tan, Binhong, et al.
Published: (2026)
by: Tan, Binhong, et al.
Published: (2026)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
by: Tan, Chaoyi, et al.
Published: (2025)
by: Tan, Chaoyi, et al.
Published: (2025)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
by: Li, Niantong, et al.
Published: (2026)
by: Li, Niantong, et al.
Published: (2026)
Aggregating Diverse Cue Experts for AI-Generated Image Detection
by: Tan, Lei, et al.
Published: (2026)
by: Tan, Lei, et al.
Published: (2026)
Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation
by: Li, Binglei, et al.
Published: (2026)
by: Li, Binglei, et al.
Published: (2026)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
by: Chinchure, Aditya, et al.
Published: (2023)
by: Chinchure, Aditya, et al.
Published: (2023)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
by: Liu, Ziyue, et al.
Published: (2025)
by: Liu, Ziyue, et al.
Published: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
by: Jeon, Jinwoo, et al.
Published: (2025)
by: Jeon, Jinwoo, et al.
Published: (2025)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
by: Meng, Chutian, et al.
Published: (2024)
by: Meng, Chutian, et al.
Published: (2024)
Clinical-grade Multi-Organ Pathology Report Generation for Multi-scale Whole Slide Images via a Semantically Guided Medical Text Foundation Model
by: Tan, Jing Wei, et al.
Published: (2024)
by: Tan, Jing Wei, et al.
Published: (2024)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
by: Teotia, Revant, et al.
Published: (2025)
by: Teotia, Revant, et al.
Published: (2025)
Text in the Dark: Extremely Low-Light Text Image Enhancement
by: Lin, Che-Tsung, et al.
Published: (2024)
by: Lin, Che-Tsung, et al.
Published: (2024)
Seeing Beyond Haze: Generative Nighttime Image Dehazing
by: Lin, Beibei, et al.
Published: (2025)
by: Lin, Beibei, et al.
Published: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
by: Trusca, Maria Mihaela, et al.
Published: (2024)
by: Trusca, Maria Mihaela, et al.
Published: (2024)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
Human Image Generation: A Comprehensive Survey
by: Jia, Zhen, et al.
Published: (2022)
by: Jia, Zhen, et al.
Published: (2022)
A Novel Evaluation Framework for Image2Text Generation
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
by: Tan, Xiaofeng, et al.
Published: (2024)
by: Tan, Xiaofeng, et al.
Published: (2024)
GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
by: Butt, Muhammad Atif, et al.
Published: (2025)
by: Butt, Muhammad Atif, et al.
Published: (2025)
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
by: Conwell, Colin, et al.
Published: (2024)
by: Conwell, Colin, et al.
Published: (2024)
Similar Items
-
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
by: Agrawal, Palaash, et al.
Published: (2023) -
Stencil: Subject-Driven Generation with Context Guidance
by: Chen, Gordon, et al.
Published: (2025) -
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
by: Jia, Yuhao, et al.
Published: (2024) -
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022) -
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
by: Tan, Yaoteng, et al.
Published: (2026)