Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Xuyang, Huo, Jiayan, Liang, Yingyu, Shi, Zhenmei, Song, Zhao, Zhang, Jiahao, Zhuang, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
by: Chen, Yubin, et al.
Published: (2025)
by: Chen, Yubin, et al.
Published: (2025)
Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
by: Cao, Yuefan, et al.
Published: (2025)
by: Cao, Yuefan, et al.
Published: (2025)
Theoretical Guarantees for High Order Trajectory Refinement in Generative Flows
by: Gong, Chengyue, et al.
Published: (2025)
by: Gong, Chengyue, et al.
Published: (2025)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
by: Liang, Yingyu, et al.
Published: (2025)
by: Liang, Yingyu, et al.
Published: (2025)
High-Order Matching for One-Step Shortcut Diffusion Models
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
by: Wu, Jiaqi, et al.
Published: (2026)
by: Wu, Jiaqi, et al.
Published: (2026)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
by: Khangaonkar, Om, et al.
Published: (2026)
by: Khangaonkar, Om, et al.
Published: (2026)
Prompt Refinement with Image Pivot for Text-to-Image Generation
by: Zhan, Jingtao, et al.
Published: (2024)
by: Zhan, Jingtao, et al.
Published: (2024)
Force Matching with Relativistic Constraints: A Physics-Inspired Approach to Stable and Efficient Generative Modeling
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
by: Jeon, Jinwoo, et al.
Published: (2025)
by: Jeon, Jinwoo, et al.
Published: (2025)
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
by: Zhen, Zongwei, et al.
Published: (2025)
by: Zhen, Zongwei, et al.
Published: (2025)
Source Prompt Disentangled Inversion for Boosting Image Editability with Diffusion Models
by: Li, Ruibin, et al.
Published: (2024)
by: Li, Ruibin, et al.
Published: (2024)
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
by: Zhuang, Jingyu, et al.
Published: (2024)
by: Zhuang, Jingyu, et al.
Published: (2024)
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
by: Jung, Chaeyoung, et al.
Published: (2026)
by: Jung, Chaeyoung, et al.
Published: (2026)
Linearized Attention Cannot Enter the Kernel Regime at Any Practical Width
by: Miñoza, Jose Marie Antonio, et al.
Published: (2026)
by: Miñoza, Jose Marie Antonio, et al.
Published: (2026)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
by: Wu, Chen, et al.
Published: (2024)
by: Wu, Chen, et al.
Published: (2024)
Verify Claimed Text-to-Image Models via Boundary-Aware Prompt Optimization
by: Zhao, Zidong, et al.
Published: (2026)
by: Zhao, Zidong, et al.
Published: (2026)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
Image Super-Resolution with Text Prompt Diffusion
by: Chen, Zheng, et al.
Published: (2023)
by: Chen, Zheng, et al.
Published: (2023)
HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models
by: Manukyan, Hayk, et al.
Published: (2023)
by: Manukyan, Hayk, et al.
Published: (2023)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
by: Li, Mingzhe, et al.
Published: (2025)
by: Li, Mingzhe, et al.
Published: (2025)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
by: Kumar, Divake, et al.
Published: (2026)
by: Kumar, Divake, et al.
Published: (2026)
Similar Items
-
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025) -
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
by: Guo, Xuyang, et al.
Published: (2025) -
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025) -
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
by: Guo, Xuyang, et al.
Published: (2025) -
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
by: Chen, Yubin, et al.
Published: (2025)