Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Vatsa, Mayank, Bharati, Aparna, Singh, Richa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
by: Roy, Susim, et al.
Published: (2025)
by: Roy, Susim, et al.
Published: (2025)
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
by: Singh, Jaisidh, et al.
Published: (2024)
by: Singh, Jaisidh, et al.
Published: (2024)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025)
by: Khan, Misaal, et al.
Published: (2025)
Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional Noises
by: Jain, Anubhooti, et al.
Published: (2024)
by: Jain, Anubhooti, et al.
Published: (2024)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
Optimizing Skin Lesion Classification via Multimodal Data and Auxiliary Task Integration
by: Khurshid, Mahapara, et al.
Published: (2024)
by: Khurshid, Mahapara, et al.
Published: (2024)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security
by: Gorthi, Nidheesh, et al.
Published: (2025)
by: Gorthi, Nidheesh, et al.
Published: (2025)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
Harmonizing Geometry and Uncertainty: Diffusion with Hyperspheres
by: Dosi, Muskan, et al.
Published: (2025)
by: Dosi, Muskan, et al.
Published: (2025)
Unbiased Model Prediction Without Using Protected Attribute Information
by: Majumdar, Puspita, et al.
Published: (2026)
by: Majumdar, Puspita, et al.
Published: (2026)
HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
I Am Big, You Are Little; I Am Right, You Are Wrong
by: Kelly, David A., et al.
Published: (2025)
by: Kelly, David A., et al.
Published: (2025)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
by: Song, Shezheng, et al.
Published: (2026)
by: Song, Shezheng, et al.
Published: (2026)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
by: Bao, Zhipeng, et al.
Published: (2023)
by: Bao, Zhipeng, et al.
Published: (2023)
Low-Resolution Chest X-ray Classification via Knowledge Distillation and Multi-task Learning
by: Akhter, Yasmeena, et al.
Published: (2024)
by: Akhter, Yasmeena, et al.
Published: (2024)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation
by: Ogezi, Michael, et al.
Published: (2024)
by: Ogezi, Michael, et al.
Published: (2024)
DreamText: High Fidelity Scene Text Synthesis
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
Plug-and-Play Multi-Concept Adaptive Blending for High-Fidelity Text-to-Image Synthesis
by: Woo, Young-Beom
Published: (2025)
by: Woo, Young-Beom
Published: (2025)
Evaluating Design Video Generation: Metrics for Compositional Fidelity
by: Deganutti, Adrienne, et al.
Published: (2026)
by: Deganutti, Adrienne, et al.
Published: (2026)
Poze: Sports Technique Feedback under Data Constraints
by: Singh, Agamdeep, et al.
Published: (2024)
by: Singh, Agamdeep, et al.
Published: (2024)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
by: Elmaaroufi, Karim, et al.
Published: (2025)
by: Elmaaroufi, Karim, et al.
Published: (2025)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
by: Kwon, Jihoon, et al.
Published: (2025)
by: Kwon, Jihoon, et al.
Published: (2025)
A Generative Approach to High Fidelity 3D Reconstruction from Text Data
by: R, Venkat Kumar, et al.
Published: (2025)
by: R, Venkat Kumar, et al.
Published: (2025)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
by: Poletti, Silvia, et al.
Published: (2026)
by: Poletti, Silvia, et al.
Published: (2026)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
Compositional Text-to-Image Generation with Dense Blob Representations
by: Nie, Weili, et al.
Published: (2024)
by: Nie, Weili, et al.
Published: (2024)
CrossMed: A Multimodal Cross-Task Benchmark for Compositional Generalization in Medical Imaging
by: Singh, Pooja, et al.
Published: (2025)
by: Singh, Pooja, et al.
Published: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
by: Cao, Yushe, et al.
Published: (2025)
by: Cao, Yushe, et al.
Published: (2025)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
by: Malik, Sameer, et al.
Published: (2025)
by: Malik, Sameer, et al.
Published: (2025)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026)
by: Shen, Yuxiang, et al.
Published: (2026)
Similar Items
-
TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
by: Roy, Susim, et al.
Published: (2025) -
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
by: Singh, Jaisidh, et al.
Published: (2024) -
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025) -
Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional Noises
by: Jain, Anubhooti, et al.
Published: (2024) -
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
by: Vatsa, Mayank, et al.
Published: (2025)