Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xie, Dian, Shao, Shitong, Bai, Lichen, Zhou, Zikai, Cheng, Bojun, Yang, Shuo, Wu, Jun, Xie, Zeke |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Optimizing Few-Step Generation with Adaptive Matching Distillation
par: Bai, Lichen, et autres
Publié: (2026)
par: Bai, Lichen, et autres
Publié: (2026)
CoRe^2: Collect, Reflect and Refine to Generate Better and Faster
par: Shao, Shitong, et autres
Publié: (2025)
par: Shao, Shitong, et autres
Publié: (2025)
IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis
par: Shao, Shitong, et autres
Publié: (2024)
par: Shao, Shitong, et autres
Publié: (2024)
Exploring Data-Free LoRA Transferability for Video Diffusion Models
par: Wang, Yuchen, et autres
Publié: (2026)
par: Wang, Yuchen, et autres
Publié: (2026)
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
par: Shao, Shitong, et autres
Publié: (2024)
par: Shao, Shitong, et autres
Publié: (2024)
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
par: Sun, Zening, et autres
Publié: (2026)
par: Sun, Zening, et autres
Publié: (2026)
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
par: Shao, Shitong, et autres
Publié: (2026)
par: Shao, Shitong, et autres
Publié: (2026)
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
par: Li, Haopeng, et autres
Publié: (2026)
par: Li, Haopeng, et autres
Publié: (2026)
Golden Noise for Diffusion Models: A Learning Framework
par: Zhou, Zikai, et autres
Publié: (2024)
par: Zhou, Zikai, et autres
Publié: (2024)
Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-Reflection
par: Bai, Lichen, et autres
Publié: (2024)
par: Bai, Lichen, et autres
Publié: (2024)
Reflective Flow Sampling Enhancement
par: Zhou, Zikai, et autres
Publié: (2026)
par: Zhou, Zikai, et autres
Publié: (2026)
Efficient Video Diffusion Models: Advancements and Challenges
par: Shao, Shitong, et autres
Publié: (2026)
par: Shao, Shitong, et autres
Publié: (2026)
FastLightGen: Fast and Light Video Generation with Fewer Steps and Parameters
par: Shao, Shitong, et autres
Publié: (2026)
par: Shao, Shitong, et autres
Publié: (2026)
Weak-to-Strong Diffusion with Reflection
par: Bai, Lichen, et autres
Publié: (2025)
par: Bai, Lichen, et autres
Publié: (2025)
Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization
par: Qi, Zipeng, et autres
Publié: (2024)
par: Qi, Zipeng, et autres
Publié: (2024)
Rethinking Centered Kernel Alignment in Knowledge Distillation
par: Zhou, Zikai, et autres
Publié: (2024)
par: Zhou, Zikai, et autres
Publié: (2024)
Alignment of Diffusion Models: Fundamentals, Challenges, and Future
par: Liu, Buhua, et autres
Publié: (2024)
par: Liu, Buhua, et autres
Publié: (2024)
Elucidating the Design Space of Dataset Condensation
par: Shao, Shitong, et autres
Publié: (2024)
par: Shao, Shitong, et autres
Publié: (2024)
MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis
par: Shao, Shitong, et autres
Publié: (2025)
par: Shao, Shitong, et autres
Publié: (2025)
Accelerating Diffusion Model Training under Minimal Budgets: A Condensation-Based Perspective
par: Huang, Rui, et autres
Publié: (2025)
par: Huang, Rui, et autres
Publié: (2025)
Fused attention mechanism-based ore sorting network
par: Zhen, Junjiang, et autres
Publié: (2024)
par: Zhen, Junjiang, et autres
Publié: (2024)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
par: Yan, Junkai, et autres
Publié: (2024)
par: Yan, Junkai, et autres
Publié: (2024)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
par: Sun, Peng, et autres
Publié: (2026)
par: Sun, Peng, et autres
Publié: (2026)
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
par: Wu, Jiong, et autres
Publié: (2025)
par: Wu, Jiong, et autres
Publié: (2025)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
par: Yi, Hongwei, et autres
Publié: (2025)
par: Yi, Hongwei, et autres
Publié: (2025)
Conditional Text-to-Image Generation with Reference Guidance
par: Kim, Taewook, et autres
Publié: (2024)
par: Kim, Taewook, et autres
Publié: (2024)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
par: Zhou, Dewei, et autres
Publié: (2024)
par: Zhou, Dewei, et autres
Publié: (2024)
Robust Remote Sensing Image-Text Retrieval with Noisy Correspondence
par: Song, Qiya, et autres
Publié: (2026)
par: Song, Qiya, et autres
Publié: (2026)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
par: Meng, Zichong, et autres
Publié: (2024)
par: Meng, Zichong, et autres
Publié: (2024)
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
par: Cheng, Jiaxin, et autres
Publié: (2024)
par: Cheng, Jiaxin, et autres
Publié: (2024)
Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation
par: Luo, Minxing, et autres
Publié: (2025)
par: Luo, Minxing, et autres
Publié: (2025)
Text-Image Conditioned 3D Generation
par: Cen, Jiazhong, et autres
Publié: (2026)
par: Cen, Jiazhong, et autres
Publié: (2026)
KB-DMGen: Knowledge-Based Global Guidance and Dynamic Pose Masking for Human Image Generation
par: Liu, Shibang, et autres
Publié: (2025)
par: Liu, Shibang, et autres
Publié: (2025)
MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation
par: Wang, Muyao, et autres
Publié: (2026)
par: Wang, Muyao, et autres
Publié: (2026)
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
par: Li, Niantong, et autres
Publié: (2026)
par: Li, Niantong, et autres
Publié: (2026)
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
par: Tapli, Merve, et autres
Publié: (2026)
par: Tapli, Merve, et autres
Publié: (2026)
PIG: Prompt Images Guidance for Night-Time Scene Parsing
par: Xie, Zhifeng, et autres
Publié: (2024)
par: Xie, Zhifeng, et autres
Publié: (2024)
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
par: Li, Yaqi, et autres
Publié: (2025)
par: Li, Yaqi, et autres
Publié: (2025)
Tuning-Free Image Customization with Image and Text Guidance
par: Li, Pengzhi, et autres
Publié: (2024)
par: Li, Pengzhi, et autres
Publié: (2024)
Power Line Aerial Image Restoration under dverse Weather: Datasets and Baselines
par: Yang, Sai, et autres
Publié: (2024)
par: Yang, Sai, et autres
Publié: (2024)
Documents similaires
-
Optimizing Few-Step Generation with Adaptive Matching Distillation
par: Bai, Lichen, et autres
Publié: (2026) -
CoRe^2: Collect, Reflect and Refine to Generate Better and Faster
par: Shao, Shitong, et autres
Publié: (2025) -
IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis
par: Shao, Shitong, et autres
Publié: (2024) -
Exploring Data-Free LoRA Transferability for Video Diffusion Models
par: Wang, Yuchen, et autres
Publié: (2026) -
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
par: Shao, Shitong, et autres
Publié: (2024)