Evaluating the Generation of Spatial Relations in Text and Image Generative Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Sim, Shang Hong, Lee, Clarence, Tan, Alvin, Tan, Cheston |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
por: Agrawal, Palaash, et al.
Publicado: (2023)
por: Agrawal, Palaash, et al.
Publicado: (2023)
Stencil: Subject-Driven Generation with Context Guidance
por: Chen, Gordon, et al.
Publicado: (2025)
por: Chen, Gordon, et al.
Publicado: (2025)
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
por: Jia, Yuhao, et al.
Publicado: (2024)
por: Jia, Yuhao, et al.
Publicado: (2024)
Inferring Past Human Actions in Homes with Abductive Reasoning
por: Tan, Clement, et al.
Publicado: (2022)
por: Tan, Clement, et al.
Publicado: (2022)
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
por: Tan, Yaoteng, et al.
Publicado: (2026)
por: Tan, Yaoteng, et al.
Publicado: (2026)
Human-like compositional learning of visually-grounded concepts using synthetic environments
por: Lin, Zijun, et al.
Publicado: (2025)
por: Lin, Zijun, et al.
Publicado: (2025)
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
por: Lin, Zijun, et al.
Publicado: (2025)
por: Lin, Zijun, et al.
Publicado: (2025)
DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation
por: Tan, Binhong, et al.
Publicado: (2026)
por: Tan, Binhong, et al.
Publicado: (2026)
Generating Multimodal Images with GAN: Integrating Text, Image, and Style
por: Tan, Chaoyi, et al.
Publicado: (2025)
por: Tan, Chaoyi, et al.
Publicado: (2025)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
por: Nagar, Aishik, et al.
Publicado: (2024)
por: Nagar, Aishik, et al.
Publicado: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
por: Li, Niantong, et al.
Publicado: (2026)
por: Li, Niantong, et al.
Publicado: (2026)
Aggregating Diverse Cue Experts for AI-Generated Image Detection
por: Tan, Lei, et al.
Publicado: (2026)
por: Tan, Lei, et al.
Publicado: (2026)
Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation
por: Li, Binglei, et al.
Publicado: (2026)
por: Li, Binglei, et al.
Publicado: (2026)
Personalized Reward Modeling for Text-to-Image Generation
por: Lee, Jeongeun, et al.
Publicado: (2025)
por: Lee, Jeongeun, et al.
Publicado: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
por: Marioriyad, Arash, et al.
Publicado: (2024)
por: Marioriyad, Arash, et al.
Publicado: (2024)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
por: Rahman, Tanzila, et al.
Publicado: (2024)
por: Rahman, Tanzila, et al.
Publicado: (2024)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
por: Tan, Zhiyu, et al.
Publicado: (2024)
por: Tan, Zhiyu, et al.
Publicado: (2024)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
por: Chinchure, Aditya, et al.
Publicado: (2023)
por: Chinchure, Aditya, et al.
Publicado: (2023)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
por: Rawal, Ishaan Singh, et al.
Publicado: (2023)
por: Rawal, Ishaan Singh, et al.
Publicado: (2023)
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
por: Liu, Ziyue, et al.
Publicado: (2025)
por: Liu, Ziyue, et al.
Publicado: (2025)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
por: Kasaei, Seyed Amir, et al.
Publicado: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
por: Jeon, Jinwoo, et al.
Publicado: (2025)
por: Jeon, Jinwoo, et al.
Publicado: (2025)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
por: Meng, Chutian, et al.
Publicado: (2024)
por: Meng, Chutian, et al.
Publicado: (2024)
Clinical-grade Multi-Organ Pathology Report Generation for Multi-scale Whole Slide Images via a Semantically Guided Medical Text Foundation Model
por: Tan, Jing Wei, et al.
Publicado: (2024)
por: Tan, Jing Wei, et al.
Publicado: (2024)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
por: Teotia, Revant, et al.
Publicado: (2025)
por: Teotia, Revant, et al.
Publicado: (2025)
Text in the Dark: Extremely Low-Light Text Image Enhancement
por: Lin, Che-Tsung, et al.
Publicado: (2024)
por: Lin, Che-Tsung, et al.
Publicado: (2024)
Seeing Beyond Haze: Generative Nighttime Image Dehazing
por: Lin, Beibei, et al.
Publicado: (2025)
por: Lin, Beibei, et al.
Publicado: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
por: Cui, Tianyu, et al.
Publicado: (2025)
por: Cui, Tianyu, et al.
Publicado: (2025)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
por: Wang, Wenhao, et al.
Publicado: (2025)
por: Wang, Wenhao, et al.
Publicado: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
por: Zhou, Sashuai, et al.
Publicado: (2026)
por: Zhou, Sashuai, et al.
Publicado: (2026)
Debiasing Text-to-Image Diffusion Models
por: He, Ruifei, et al.
Publicado: (2024)
por: He, Ruifei, et al.
Publicado: (2024)
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control
por: Trusca, Maria Mihaela, et al.
Publicado: (2024)
por: Trusca, Maria Mihaela, et al.
Publicado: (2024)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
por: Eldesokey, Abdelrahman, et al.
Publicado: (2026)
por: Eldesokey, Abdelrahman, et al.
Publicado: (2026)
Human Image Generation: A Comprehensive Survey
por: Jia, Zhen, et al.
Publicado: (2022)
por: Jia, Zhen, et al.
Publicado: (2022)
A Novel Evaluation Framework for Image2Text Generation
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
por: Tan, Xiaofeng, et al.
Publicado: (2024)
por: Tan, Xiaofeng, et al.
Publicado: (2024)
GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
por: Butt, Muhammad Atif, et al.
Publicado: (2025)
por: Butt, Muhammad Atif, et al.
Publicado: (2025)
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
por: Conwell, Colin, et al.
Publicado: (2024)
por: Conwell, Colin, et al.
Publicado: (2024)
Ejemplares similares
-
STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning
por: Agrawal, Palaash, et al.
Publicado: (2023) -
Stencil: Subject-Driven Generation with Context Guidance
por: Chen, Gordon, et al.
Publicado: (2025) -
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
por: Jia, Yuhao, et al.
Publicado: (2024) -
Inferring Past Human Actions in Homes with Abductive Reasoning
por: Tan, Clement, et al.
Publicado: (2022) -
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
por: Tan, Yaoteng, et al.
Publicado: (2026)