Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Saichandran, Ketan Suhaas, Thomas, Xavier, Kaushik, Prakhar, Ghadiyaram, Deepti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hand Shape and Gesture Recognition using Multiscale Template Matching, Background Subtraction and Binary Image Analysis
by: Saichandran, Ketan Suhaas
Published: (2024)
by: Saichandran, Ketan Suhaas
Published: (2024)
Improving Physical Object State Representation in Text-to-Image Generative Systems
by: Chen, Tianle, et al.
Published: (2025)
by: Chen, Tianle, et al.
Published: (2025)
A Comparative Analysis of U-Net-based models for Segmentation of Cardiac MRI
by: Saichandran, Ketan Suhaas
Published: (2024)
by: Saichandran, Ketan Suhaas
Published: (2024)
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
by: Lim, Youngsun, et al.
Published: (2026)
by: Lim, Youngsun, et al.
Published: (2026)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
A Bayesian Approach to OOD Robustness in Image Classification
by: Kaushik, Prakhar, et al.
Published: (2024)
by: Kaushik, Prakhar, et al.
Published: (2024)
$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
by: Kim, Dahye, et al.
Published: (2024)
by: Kim, Dahye, et al.
Published: (2024)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
by: Jiao, Qirui, et al.
Published: (2025)
by: Jiao, Qirui, et al.
Published: (2025)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
by: Kaushik, Prakhar, et al.
Published: (2024)
by: Kaushik, Prakhar, et al.
Published: (2024)
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025)
by: Kim, Dahye, et al.
Published: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
by: Chen, Lifeng, et al.
Published: (2025)
by: Chen, Lifeng, et al.
Published: (2025)
DebiasPI: Inference-time Debiasing by Prompt Iteration of a Text-to-Image Generative Model
by: Bonna, Sarah, et al.
Published: (2025)
by: Bonna, Sarah, et al.
Published: (2025)
DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models
by: Tian, Mengyuan, et al.
Published: (2026)
by: Tian, Mengyuan, et al.
Published: (2026)
Swift Sampling: Selecting Temporal Surprises via Taylor Series
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)
by: Chen, Tianle, et al.
Published: (2026)
Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
by: Chen, Hongxu, et al.
Published: (2025)
by: Chen, Hongxu, et al.
Published: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
by: Zou, Hao, et al.
Published: (2025)
by: Zou, Hao, et al.
Published: (2025)
Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Instant Preference Alignment for Text-to-Image Diffusion Models
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
by: Chen, Wenting, et al.
Published: (2024)
by: Chen, Wenting, et al.
Published: (2024)
Face-MakeUp: Multimodal Facial Prompts for Text-to-Image Generation
by: Dai, Dawei, et al.
Published: (2025)
by: Dai, Dawei, et al.
Published: (2025)
FairQueue: Rethinking Prompt Learning for Fair Text-to-Image Generation
by: Teo, Christopher T. H, et al.
Published: (2024)
by: Teo, Christopher T. H, et al.
Published: (2024)
Generating Accurate and Detailed Captions for High-Resolution Images
by: Lee, Hankyeol, et al.
Published: (2025)
by: Lee, Hankyeol, et al.
Published: (2025)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
by: Chatterjee, Abhiroop, et al.
Published: (2025)
by: Chatterjee, Abhiroop, et al.
Published: (2025)
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models
by: Liu, Bingchen, et al.
Published: (2024)
by: Liu, Bingchen, et al.
Published: (2024)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
by: Zhang, Huixuan, et al.
Published: (2025)
by: Zhang, Huixuan, et al.
Published: (2025)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
by: Singla, Pratham, et al.
Published: (2025)
by: Singla, Pratham, et al.
Published: (2025)
REMAP: Regularized Matching and Partial Alignment of Video Embeddings
by: Chandra, Soumyadeep, et al.
Published: (2025)
by: Chandra, Soumyadeep, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
by: Zhang, Qian, et al.
Published: (2025)
by: Zhang, Qian, et al.
Published: (2025)
Similar Items
-
Hand Shape and Gesture Recognition using Multiscale Template Matching, Background Subtraction and Binary Image Analysis
by: Saichandran, Ketan Suhaas
Published: (2024) -
Improving Physical Object State Representation in Text-to-Image Generative Systems
by: Chen, Tianle, et al.
Published: (2025) -
A Comparative Analysis of U-Net-based models for Segmentation of Cardiac MRI
by: Saichandran, Ketan Suhaas
Published: (2024) -
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
by: Lim, Youngsun, et al.
Published: (2026) -
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)