Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Subin, Mo, Sangwoo, Rizve, Mamshad Nayeem, Xu, Yiran, Liu, Difan, Shin, Jinwoo, Hinz, Tobias |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
por: Pillai, Manu S, et al.
Publicado: (2024)
por: Pillai, Manu S, et al.
Publicado: (2024)
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
por: Dave, Ishan Rajendrakumar, et al.
Publicado: (2024)
por: Dave, Ishan Rajendrakumar, et al.
Publicado: (2024)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
por: Zhu, Zixin, et al.
Publicado: (2025)
por: Zhu, Zixin, et al.
Publicado: (2025)
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
por: Venkataramanan, Shashanka, et al.
Publicado: (2023)
FontAdapter: Instant Font Adaptation in Visual Text Generation
por: Koo, Myungkyu, et al.
Publicado: (2025)
por: Koo, Myungkyu, et al.
Publicado: (2025)
VidLA: Video-Language Alignment at Scale
por: Rizve, Mamshad Nayeem, et al.
Publicado: (2024)
por: Rizve, Mamshad Nayeem, et al.
Publicado: (2024)
Open Vocabulary Multi-Label Video Classification
por: Gupta, Rohit, et al.
Publicado: (2024)
por: Gupta, Rohit, et al.
Publicado: (2024)
Discovering and Mitigating Visual Biases through Keyword Explanation
por: Kim, Younghyun, et al.
Publicado: (2023)
por: Kim, Younghyun, et al.
Publicado: (2023)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
por: Swetha, Sirnam, et al.
Publicado: (2024)
por: Swetha, Sirnam, et al.
Publicado: (2024)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
por: Ahmed, Sabbir, et al.
Publicado: (2025)
por: Ahmed, Sabbir, et al.
Publicado: (2025)
Restoration-Aligned Generative Flow Models for Blind Motion Deblurring
por: Kim, Insoo, et al.
Publicado: (2026)
por: Kim, Insoo, et al.
Publicado: (2026)
DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
por: Zhang, Peiying, et al.
Publicado: (2025)
por: Zhang, Peiying, et al.
Publicado: (2025)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
por: Lee, Kyungmin, et al.
Publicado: (2024)
por: Lee, Kyungmin, et al.
Publicado: (2024)
Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales
por: Kim, Myeongsoo, et al.
Publicado: (2026)
por: Kim, Myeongsoo, et al.
Publicado: (2026)
Iterative Prompt Refinement for Safer Text-to-Image Generation
por: Jeon, Jinwoo, et al.
Publicado: (2025)
por: Jeon, Jinwoo, et al.
Publicado: (2025)
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
por: Kim, Subin, et al.
Publicado: (2025)
por: Kim, Subin, et al.
Publicado: (2025)
Rethinking Layered Graphic Design Generation with a Top-Down Approach
por: Chen, Jingye, et al.
Publicado: (2025)
por: Chen, Jingye, et al.
Publicado: (2025)
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
por: Varghese, Subin, et al.
Publicado: (2024)
por: Varghese, Subin, et al.
Publicado: (2024)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
por: Shin, Youngwoo, et al.
Publicado: (2026)
por: Shin, Youngwoo, et al.
Publicado: (2026)
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
por: Jeon, Byungwoo, et al.
Publicado: (2026)
por: Jeon, Byungwoo, et al.
Publicado: (2026)
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
por: Kim, Sanghyun, et al.
Publicado: (2024)
por: Kim, Sanghyun, et al.
Publicado: (2024)
Personalized Residuals for Concept-Driven Text-to-Image Generation
por: Ham, Cusuh, et al.
Publicado: (2024)
por: Ham, Cusuh, et al.
Publicado: (2024)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
por: Lee, Kyungmin, et al.
Publicado: (2024)
por: Lee, Kyungmin, et al.
Publicado: (2024)
Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-Reasoning
por: Park, Subin, et al.
Publicado: (2026)
por: Park, Subin, et al.
Publicado: (2026)
Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion
por: Kim, Sanghyun, et al.
Publicado: (2024)
por: Kim, Sanghyun, et al.
Publicado: (2024)
Text-guided Visual Prompt DINO for Generic Segmentation
por: Guan, Yuchen, et al.
Publicado: (2025)
por: Guan, Yuchen, et al.
Publicado: (2025)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
por: Kara, Ozgur, et al.
Publicado: (2025)
por: Kara, Ozgur, et al.
Publicado: (2025)
SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
por: Li, Zhengang, et al.
Publicado: (2024)
por: Li, Zhengang, et al.
Publicado: (2024)
FairQueue: Rethinking Prompt Learning for Fair Text-to-Image Generation
por: Teo, Christopher T. H, et al.
Publicado: (2024)
por: Teo, Christopher T. H, et al.
Publicado: (2024)
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
por: Kim, Bryan Sangwoo, et al.
Publicado: (2026)
por: Kim, Bryan Sangwoo, et al.
Publicado: (2026)
Multimodal Dataset Distillation Made Simple by Prototype-Guided Data Synthesis
por: Choi, Junhyeok, et al.
Publicado: (2026)
por: Choi, Junhyeok, et al.
Publicado: (2026)
Adversarial Robustification via Text-to-Image Diffusion Models
por: Choi, Daewon, et al.
Publicado: (2024)
por: Choi, Daewon, et al.
Publicado: (2024)
A Review of Image Retrieval Techniques: Data Augmentation and Adversarial Learning Approaches
por: Jinwoo, Kim
Publicado: (2024)
por: Jinwoo, Kim
Publicado: (2024)
Dynamic Prompt Optimizing for Text-to-Image Generation
por: Mo, Wenyi, et al.
Publicado: (2024)
por: Mo, Wenyi, et al.
Publicado: (2024)
CoAPT: Context Attribute words for Prompt Tuning
por: Lee, Gun, et al.
Publicado: (2024)
por: Lee, Gun, et al.
Publicado: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
por: Shin, Chaehun, et al.
Publicado: (2024)
por: Shin, Chaehun, et al.
Publicado: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
por: Fallah, Forouzan, et al.
Publicado: (2025)
por: Fallah, Forouzan, et al.
Publicado: (2025)
Text2VR: Automated instruction Generation in Virtual Reality using Large language Models for Assembly Task
por: Peter, Subin Raj
Publicado: (2025)
por: Peter, Subin Raj
Publicado: (2025)
Extreme Blind Image Restoration via Prompt-Conditioned Information Bottleneck
por: Kim, Hongeun, et al.
Publicado: (2025)
por: Kim, Hongeun, et al.
Publicado: (2025)
Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization
por: Peng, Jiangweizhi, et al.
Publicado: (2024)
por: Peng, Jiangweizhi, et al.
Publicado: (2024)
Ejemplares similares
-
GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers
por: Pillai, Manu S, et al.
Publicado: (2024) -
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
por: Dave, Ishan Rajendrakumar, et al.
Publicado: (2024) -
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
por: Zhu, Zixin, et al.
Publicado: (2025) -
Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video
por: Venkataramanan, Shashanka, et al.
Publicado: (2023) -
FontAdapter: Instant Font Adaptation in Visual Text Generation
por: Koo, Myungkyu, et al.
Publicado: (2025)