Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Chatterjee, Agneet, Stan, Gabriela Ben Melech, Aflalo, Estelle, Paul, Sayak, Ghosh, Dhruba, Gokhale, Tejas, Schmidt, Ludwig, Hajishirzi, Hannaneh, Lal, Vasudev, Baral, Chitta, Yang, Yezhou |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
por: Chatterjee, Agneet, et al.
Publicado: (2024)
por: Chatterjee, Agneet, et al.
Publicado: (2024)
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
por: Chatterjee, Agneet, et al.
Publicado: (2024)
por: Chatterjee, Agneet, et al.
Publicado: (2024)
Learning from Reasoning Failures via Synthetic Data Generation
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
FastRM: An efficient and automatic explainability framework for multimodal generative models
por: Stan, Gabriela Ben-Melech, et al.
Publicado: (2024)
por: Stan, Gabriela Ben-Melech, et al.
Publicado: (2024)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
por: Aflalo, Estelle, et al.
Publicado: (2024)
por: Aflalo, Estelle, et al.
Publicado: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
por: Patel, Maitreya, et al.
Publicado: (2023)
por: Patel, Maitreya, et al.
Publicado: (2023)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
por: Malaviya, Vatsal, et al.
Publicado: (2025)
por: Malaviya, Vatsal, et al.
Publicado: (2025)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
por: Liu, Xiangrui, et al.
Publicado: (2025)
por: Liu, Xiangrui, et al.
Publicado: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
por: Fallah, Forouzan, et al.
Publicado: (2025)
por: Fallah, Forouzan, et al.
Publicado: (2025)
Dual Caption Preference Optimization for Diffusion Models
por: Saeidi, Amir, et al.
Publicado: (2025)
por: Saeidi, Amir, et al.
Publicado: (2025)
Chimera: Compositional Image Generation using Part-based Concepting
por: Singh, Shivam, et al.
Publicado: (2025)
por: Singh, Shivam, et al.
Publicado: (2025)
Grounding Stylistic Domain Generalization with Quantitative Domain Shift Measures and Synthetic Scene Images
por: Luo, Yiran, et al.
Publicado: (2024)
por: Luo, Yiran, et al.
Publicado: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
por: Patel, Maitreya, et al.
Publicado: (2024)
por: Patel, Maitreya, et al.
Publicado: (2024)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2024)
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
por: Yilmaz, Nilay, et al.
Publicado: (2025)
por: Yilmaz, Nilay, et al.
Publicado: (2025)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
por: Chatterjee, Agneet, et al.
Publicado: (2025)
por: Chatterjee, Agneet, et al.
Publicado: (2025)
ActionCOMET: A Zero-shot Approach to Learn Image-specific Commonsense Concepts about Actions
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
por: Varshney, Neeraj, et al.
Publicado: (2024)
por: Varshney, Neeraj, et al.
Publicado: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
por: Patel, Maitreya, et al.
Publicado: (2024)
por: Patel, Maitreya, et al.
Publicado: (2024)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
por: Ghosh, Dhruba, et al.
Publicado: (2026)
por: Ghosh, Dhruba, et al.
Publicado: (2026)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
por: Ratzlaff, Neale, et al.
Publicado: (2024)
por: Ratzlaff, Neale, et al.
Publicado: (2024)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
por: Pathiraja, Bimsara, et al.
Publicado: (2025)
por: Pathiraja, Bimsara, et al.
Publicado: (2025)
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective
por: Rajput, Krishna Singh, et al.
Publicado: (2025)
por: Rajput, Krishna Singh, et al.
Publicado: (2025)
L-MAGIC: Language Model Assisted Generation of Images with Coherence
por: Cai, Zhipeng, et al.
Publicado: (2024)
por: Cai, Zhipeng, et al.
Publicado: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
por: Vani, Sameep, et al.
Publicado: (2025)
por: Vani, Sameep, et al.
Publicado: (2025)
EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
por: Kusumba, Abhiram, et al.
Publicado: (2025)
por: Kusumba, Abhiram, et al.
Publicado: (2025)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
por: Zhao, Bowen, et al.
Publicado: (2024)
por: Zhao, Bowen, et al.
Publicado: (2024)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
por: Anvekar, Tejas, et al.
Publicado: (2025)
por: Anvekar, Tejas, et al.
Publicado: (2025)
The EarlyBird Gets the WORM: Heuristically Accelerating EarlyBird Convergence
por: Vasudev, Adithya
Publicado: (2024)
por: Vasudev, Adithya
Publicado: (2024)
Improving Shift Invariance in Convolutional Neural Networks with Translation Invariant Polyphase Sampling
por: Saha, Sourajit, et al.
Publicado: (2024)
por: Saha, Sourajit, et al.
Publicado: (2024)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
por: Madasu, Avinash, et al.
Publicado: (2023)
por: Madasu, Avinash, et al.
Publicado: (2023)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
por: Alqurnawi, Yahia, et al.
Publicado: (2026)
por: Alqurnawi, Yahia, et al.
Publicado: (2026)
Data or Language Supervision: What Makes CLIP Better than DINO?
por: Liu, Yiming, et al.
Publicado: (2025)
por: Liu, Yiming, et al.
Publicado: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
por: Sampat, Shailaja Keyur, et al.
Publicado: (2024)
Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training
por: Merrill, William, et al.
Publicado: (2025)
por: Merrill, William, et al.
Publicado: (2025)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
por: Kim, Joongwon, et al.
Publicado: (2024)
por: Kim, Joongwon, et al.
Publicado: (2024)
ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
por: Wang, Yike, et al.
Publicado: (2025)
por: Wang, Yike, et al.
Publicado: (2025)
Ejemplares similares
-
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
por: Chatterjee, Agneet, et al.
Publicado: (2024) -
REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
por: Chatterjee, Agneet, et al.
Publicado: (2024) -
Learning from Reasoning Failures via Synthetic Data Generation
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025) -
FastRM: An efficient and automatic explainability framework for multimodal generative models
por: Stan, Gabriela Ben-Melech, et al.
Publicado: (2024) -
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
por: Aflalo, Estelle, et al.
Publicado: (2024)