DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Qirui, Chen, Daoyuan, Huang, Yilun, Lin, Xika, Shen, Ying, Li, Yaliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
by: Li, Zhiwen, et al.
Published: (2025)
by: Li, Zhiwen, et al.
Published: (2025)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
by: Han, Yucheng, et al.
Published: (2024)
by: Han, Yucheng, et al.
Published: (2024)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
by: Li, Yuyi, et al.
Published: (2025)
by: Li, Yuyi, et al.
Published: (2025)
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
by: Gani, Hanan, et al.
Published: (2023)
by: Gani, Hanan, et al.
Published: (2023)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation
by: Ye, Jinyan, et al.
Published: (2026)
by: Ye, Jinyan, et al.
Published: (2026)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
by: Chen, Lifeng, et al.
Published: (2025)
by: Chen, Lifeng, et al.
Published: (2025)
VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
by: Li, Zhiwen, et al.
Published: (2026)
by: Li, Zhiwen, et al.
Published: (2026)
Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts
by: Liu, Yufan, et al.
Published: (2025)
by: Liu, Yufan, et al.
Published: (2025)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
Can Score-Based Generative Modeling Effectively Handle Medical Image Classification?
by: Sarker, Sushmita, et al.
Published: (2025)
by: Sarker, Sushmita, et al.
Published: (2025)
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
by: Zhuang, Jingyu, et al.
Published: (2024)
by: Zhuang, Jingyu, et al.
Published: (2024)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
by: Wang, Juntong, et al.
Published: (2025)
by: Wang, Juntong, et al.
Published: (2025)
TIPO: Text to Image with Text Presampling for Prompt Optimization
by: Yeh, Shih-Ying, et al.
Published: (2024)
by: Yeh, Shih-Ying, et al.
Published: (2024)
Text-guided Foundation Model Adaptation for Long-Tailed Medical Image Classification
by: Li, Sirui, et al.
Published: (2024)
by: Li, Sirui, et al.
Published: (2024)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
by: Wang, Luozhou, et al.
Published: (2023)
by: Wang, Luozhou, et al.
Published: (2023)
Unified Prompt Attack Against Text-to-Image Generation Models
by: Peng, Duo, et al.
Published: (2025)
by: Peng, Duo, et al.
Published: (2025)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
by: Jo, Kyungmin, et al.
Published: (2025)
by: Jo, Kyungmin, et al.
Published: (2025)
Personalized Image Filter: Mastering Your Photographic Style
by: Zhu, Chengxuan, et al.
Published: (2025)
by: Zhu, Chengxuan, et al.
Published: (2025)
Is Your Text-to-Image Model Robust to Caption Noise?
by: Yu, Weichen, et al.
Published: (2024)
by: Yu, Weichen, et al.
Published: (2024)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
by: Wu, Chen, et al.
Published: (2024)
by: Wu, Chen, et al.
Published: (2024)
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)
by: Li, Yanyu, et al.
Published: (2024)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
by: Zhong, Jianping, et al.
Published: (2026)
by: Zhong, Jianping, et al.
Published: (2026)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
EPIC: Efficient Prompt Interaction for Text-Image Classification
by: Yu, Xinyao, et al.
Published: (2025)
by: Yu, Xinyao, et al.
Published: (2025)
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
by: Thakkar, Parth, et al.
Published: (2025)
by: Thakkar, Parth, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Similar Items
-
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
by: Jiao, Qirui, et al.
Published: (2024) -
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
by: Jiao, Qirui, et al.
Published: (2024) -
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024) -
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026) -
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
by: Li, Zhiwen, et al.
Published: (2025)