Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chatterjee, Agneet, Stan, Gabriela Ben Melech, Aflalo, Estelle, Paul, Sayak, Ghosh, Dhruba, Gokhale, Tejas, Schmidt, Ludwig, Hajishirzi, Hannaneh, Lal, Vasudev, Baral, Chitta, Yang, Yezhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024)
Learning from Reasoning Failures via Synthetic Data Generation
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2025)
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2025)
FastRM: An efficient and automatic explainability framework for multimodal generative models
von: Stan, Gabriela Ben-Melech, et al.
Veröffentlicht: (2024)
von: Stan, Gabriela Ben-Melech, et al.
Veröffentlicht: (2024)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
von: Aflalo, Estelle, et al.
Veröffentlicht: (2024)
von: Aflalo, Estelle, et al.
Veröffentlicht: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
Dual Caption Preference Optimization for Diffusion Models
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
von: Saeidi, Amir, et al.
Veröffentlicht: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
von: Luo, Yiran, et al.
Veröffentlicht: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2025)
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
von: Varshney, Neeraj, et al.
Veröffentlicht: (2024)
von: Varshney, Neeraj, et al.
Veröffentlicht: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
von: Ghosh, Dhruba, et al.
Veröffentlicht: (2026)
von: Ghosh, Dhruba, et al.
Veröffentlicht: (2026)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
von: Rajput, Krishna Singh, et al.
Veröffentlicht: (2025)
von: Rajput, Krishna Singh, et al.
Veröffentlicht: (2025)
L-MAGIC: Language Model Assisted Generation of Images with Coherence
von: Cai, Zhipeng, et al.
Veröffentlicht: (2024)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
von: Kusumba, Abhiram, et al.
Veröffentlicht: (2025)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
von: Zhao, Bowen, et al.
Veröffentlicht: (2024)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
The EarlyBird Gets the WORM: Heuristically Accelerating EarlyBird Convergence
von: Vasudev, Adithya
Veröffentlicht: (2024)
von: Vasudev, Adithya
Veröffentlicht: (2024)
Improving Shift Invariance in Convolutional Neural Networks with Translation Invariant Polyphase Sampling
von: Saha, Sourajit, et al.
Veröffentlicht: (2024)
von: Saha, Sourajit, et al.
Veröffentlicht: (2024)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
von: Madasu, Avinash, et al.
Veröffentlicht: (2023)
von: Madasu, Avinash, et al.
Veröffentlicht: (2023)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
von: Alqurnawi, Yahia, et al.
Veröffentlicht: (2026)
von: Alqurnawi, Yahia, et al.
Veröffentlicht: (2026)
Data or Language Supervision: What Makes CLIP Better than DINO?
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
von: Merrill, William, et al.
Veröffentlicht: (2025)
von: Merrill, William, et al.
Veröffentlicht: (2025)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
von: Kim, Joongwon, et al.
Veröffentlicht: (2024)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
von: Wang, Yike, et al.
Veröffentlicht: (2025)
von: Wang, Yike, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024) -
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
von: Chatterjee, Agneet, et al.
Veröffentlicht: (2024) -
Learning from Reasoning Failures via Synthetic Data Generation
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2025) -
FastRM: An efficient and automatic explainability framework for multimodal generative models
von: Stan, Gabriela Ben-Melech, et al.
Veröffentlicht: (2024) -
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
von: Aflalo, Estelle, et al.
Veröffentlicht: (2024)