The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Xiaofeng, Zhang, Courville, Aaron, Drozdzal, Michal, Romero-Soriano, Adriana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024)
por: Mañas, Oscar, et al.
Publicado: (2024)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
por: Hall, Melissa, et al.
Publicado: (2023)
por: Hall, Melissa, et al.
Publicado: (2023)
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
por: Assouel, Rim, et al.
Publicado: (2026)
por: Assouel, Rim, et al.
Publicado: (2026)
Increasing the Utility of Synthetic Images through Chamfer Guidance
por: Dall'Asen, Nicola, et al.
Publicado: (2025)
por: Dall'Asen, Nicola, et al.
Publicado: (2025)
Entropy Rectifying Guidance for Diffusion and Flow Models
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)
Object-centric Binding in Contrastive Language-Image Pretraining
por: Assouel, Rim, et al.
Publicado: (2025)
por: Assouel, Rim, et al.
Publicado: (2025)
Consistency-diversity-realism Pareto fronts of conditional image generative models
por: Astolfi, Pietro, et al.
Publicado: (2024)
por: Astolfi, Pietro, et al.
Publicado: (2024)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
por: Yuan, Jianhao, et al.
Publicado: (2026)
por: Yuan, Jianhao, et al.
Publicado: (2026)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Feedback-guided Data Synthesis for Imbalanced Classification
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
Improving the Physics of Video Generation with VJEPA-2 Reward Signal
por: Yuan, Jianhao, et al.
Publicado: (2025)
por: Yuan, Jianhao, et al.
Publicado: (2025)
Multi-Modal Language Models as Text-to-Image Model Evaluators
por: Chen, Jiahui, et al.
Publicado: (2025)
por: Chen, Jiahui, et al.
Publicado: (2025)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
Boosting Latent Diffusion with Perceptual Objectives
por: Berrada, Tariq, et al.
Publicado: (2024)
por: Berrada, Tariq, et al.
Publicado: (2024)
Controlling Multimodal LLMs via Reward-guided Decoding
por: Mañas, Oscar, et al.
Publicado: (2025)
por: Mañas, Oscar, et al.
Publicado: (2025)
Content-Rich AIGC Video Quality Assessment via Intricate Text Alignment and Motion-Aware Consistency
por: Sun, Shangkun, et al.
Publicado: (2025)
por: Sun, Shangkun, et al.
Publicado: (2025)
Bias Analysis in Unconditional Image Generative Models
por: Zhang, Xiaofeng, et al.
Publicado: (2025)
por: Zhang, Xiaofeng, et al.
Publicado: (2025)
On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2024)
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2024)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
por: Teotia, Revant, et al.
Publicado: (2025)
por: Teotia, Revant, et al.
Publicado: (2025)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
por: Lavoie, Samuel, et al.
Publicado: (2025)
por: Lavoie, Samuel, et al.
Publicado: (2025)
EvalGIM: A Library for Evaluating Generative Image Models
por: Hall, Melissa, et al.
Publicado: (2024)
por: Hall, Melissa, et al.
Publicado: (2024)
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
por: Ma, Xu, et al.
Publicado: (2026)
por: Ma, Xu, et al.
Publicado: (2026)
ConsiStyle: Style Diversity in Training-Free Consistent T2I Generation
por: Mazuz, Yohai, et al.
Publicado: (2025)
por: Mazuz, Yohai, et al.
Publicado: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024)
por: Lavoie, Samuel, et al.
Publicado: (2024)
Rethinking Interactive Image Segmentation with Low Latency, High Quality, and Diverse Prompts
por: Liu, Qin, et al.
Publicado: (2024)
por: Liu, Qin, et al.
Publicado: (2024)
Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate Rains
por: Ran, Wu, et al.
Publicado: (2024)
por: Ran, Wu, et al.
Publicado: (2024)
Consistency-guided Prompt Learning for Vision-Language Models
por: Roy, Shuvendu, et al.
Publicado: (2023)
por: Roy, Shuvendu, et al.
Publicado: (2023)
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
por: Yang, Kaixing, et al.
Publicado: (2025)
por: Yang, Kaixing, et al.
Publicado: (2025)
SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
por: Nguyen, Bac, et al.
Publicado: (2024)
por: Nguyen, Bac, et al.
Publicado: (2024)
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
por: Urbanek, Jack, et al.
Publicado: (2023)
por: Urbanek, Jack, et al.
Publicado: (2023)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
por: Ren, Weiming, et al.
Publicado: (2024)
por: Ren, Weiming, et al.
Publicado: (2024)
Inspiring the Next Generation of Segment Anything Models: Comprehensively Evaluate SAM and SAM 2 with Diverse Prompts Towards Context-Dependent Concepts under Different Scenes
por: Zhao, Xiaoqi, et al.
Publicado: (2024)
por: Zhao, Xiaoqi, et al.
Publicado: (2024)
MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation
por: Yang, Kaixing, et al.
Publicado: (2025)
por: Yang, Kaixing, et al.
Publicado: (2025)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
por: He, Wen-Jue, et al.
Publicado: (2025)
por: He, Wen-Jue, et al.
Publicado: (2025)
OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
por: Zhang, Jinlu, et al.
Publicado: (2025)
por: Zhang, Jinlu, et al.
Publicado: (2025)
Augmented Conditioning Is Enough For Effective Training Image Generation
por: Chen, Jiahui, et al.
Publicado: (2025)
por: Chen, Jiahui, et al.
Publicado: (2025)
Cross Prompting Consistency with Segment Anything Model for Semi-supervised Medical Image Segmentation
por: Miao, Juzheng, et al.
Publicado: (2024)
por: Miao, Juzheng, et al.
Publicado: (2024)
Consistency-Preserving Diverse Video Generation
por: Liu, Xinshuang, et al.
Publicado: (2026)
por: Liu, Xinshuang, et al.
Publicado: (2026)
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
por: Park, Mingue, et al.
Publicado: (2025)
por: Park, Mingue, et al.
Publicado: (2025)
Diversity Covariance-Aware Prompt Learning for Vision-Language Models
por: Dong, Songlin, et al.
Publicado: (2025)
por: Dong, Songlin, et al.
Publicado: (2025)
Ejemplares similares
-
Improving Text-to-Image Consistency via Automatic Prompt Optimization
por: Mañas, Oscar, et al.
Publicado: (2024) -
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
por: Hall, Melissa, et al.
Publicado: (2023) -
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
por: Assouel, Rim, et al.
Publicado: (2026) -
Increasing the Utility of Synthetic Images through Chamfer Guidance
por: Dall'Asen, Nicola, et al.
Publicado: (2025) -
Entropy Rectifying Guidance for Diffusion and Flow Models
por: Ifriqi, Tariq Berrada, et al.
Publicado: (2025)