Self-Evaluation Unlocks Any-Step Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Xin, Qi, Xiaojuan, Li, Zhengqi, Zhang, Kai, Zhang, Richard, Lin, Zhe, Shechtman, Eli, Wang, Tianyu, Nitzan, Yotam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024)
by: Nitzan, Yotam, et al.
Published: (2024)
Long-Context State-Space Video World Models
by: Po, Ryan, et al.
Published: (2025)
by: Po, Ryan, et al.
Published: (2025)
Jump Cut Smoothing for Talking Heads
by: Wang, Xiaojuan, et al.
Published: (2024)
by: Wang, Xiaojuan, et al.
Published: (2024)
Learning an Image Editing Model without Image Editing Pairs
by: Kumari, Nupur, et al.
Published: (2025)
by: Kumari, Nupur, et al.
Published: (2025)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
Customizing Text-to-Image Diffusion with Object Viewpoint Control
by: Kumari, Nupur, et al.
Published: (2024)
by: Kumari, Nupur, et al.
Published: (2024)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
by: Mo, Sicheng, et al.
Published: (2025)
by: Mo, Sicheng, et al.
Published: (2025)
Identifying Prompted Artist Names from Generated Images
by: Su, Grace, et al.
Published: (2025)
by: Su, Grace, et al.
Published: (2025)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
Image Neural Field Diffusion Models
by: Chen, Yinbo, et al.
Published: (2024)
by: Chen, Yinbo, et al.
Published: (2024)
VLM-Guided Adaptive Negative Prompting for Creative Generation
by: Golan, Shelly, et al.
Published: (2025)
by: Golan, Shelly, et al.
Published: (2025)
Fine-grained Defocus Blur Control for Generative Image Models
by: Shrivastava, Ayush, et al.
Published: (2025)
by: Shrivastava, Ayush, et al.
Published: (2025)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Editable Image Elements for Controllable Synthesis
by: Mu, Jiteng, et al.
Published: (2024)
by: Mu, Jiteng, et al.
Published: (2024)
Improved Distribution Matching Distillation for Fast Image Synthesis
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
Generative Video Motion Editing with 3D Point Tracks
by: Lee, Yao-Chih, et al.
Published: (2025)
by: Lee, Yao-Chih, et al.
Published: (2025)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
by: Xue, Haotian, et al.
Published: (2025)
by: Xue, Haotian, et al.
Published: (2025)
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026)
by: Bai, Xingjian, et al.
Published: (2026)
AnyTrans: Translate AnyText in the Image with Large Scale Models
by: Qian, Zhipeng, et al.
Published: (2024)
by: Qian, Zhipeng, et al.
Published: (2024)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
TurboEdit: Instant text-based image editing
by: Wu, Zongze, et al.
Published: (2024)
by: Wu, Zongze, et al.
Published: (2024)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation
by: Sun, Yanan, et al.
Published: (2024)
by: Sun, Yanan, et al.
Published: (2024)
Generative Image Dynamics
by: Li, Zhengqi, et al.
Published: (2023)
by: Li, Zhengqi, et al.
Published: (2023)
NewMove: Customizing text-to-video models with novel motions
by: Materzynska, Joanna, et al.
Published: (2023)
by: Materzynska, Joanna, et al.
Published: (2023)
ObjectMover: Generative Object Movement with Video Prior
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
Vision Foundation Models as Generalist Tokenizers for Image Generation
by: Zheng, Anlin, et al.
Published: (2026)
by: Zheng, Anlin, et al.
Published: (2026)
Structure-Guided Image Completion with Image-level and Object-level Semantic Discriminators
by: Zheng, Haitian, et al.
Published: (2022)
by: Zheng, Haitian, et al.
Published: (2022)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
One-step Diffusion with Distribution Matching Distillation
by: Yin, Tianwei, et al.
Published: (2023)
by: Yin, Tianwei, et al.
Published: (2023)
GAP: Gaussianize Any Point Clouds with Text Guidance
by: Zhang, Weiqi, et al.
Published: (2025)
by: Zhang, Weiqi, et al.
Published: (2025)
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
by: Zhang, Tianyu, et al.
Published: (2025)
by: Zhang, Tianyu, et al.
Published: (2025)
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
by: Gandikota, Rohit, et al.
Published: (2025)
by: Gandikota, Rohit, et al.
Published: (2025)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
by: Zheng, Anlin, et al.
Published: (2025)
by: Zheng, Anlin, et al.
Published: (2025)
ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control
by: Golan, Shelly, et al.
Published: (2026)
by: Golan, Shelly, et al.
Published: (2026)
LidarPainter: One-Step Away From Any Lidar View To Novel Guidance
by: Ji, Yuzhou, et al.
Published: (2025)
by: Ji, Yuzhou, et al.
Published: (2025)
AnyText: Multilingual Visual Text Generation And Editing
by: Tuo, Yuxiang, et al.
Published: (2023)
by: Tuo, Yuxiang, et al.
Published: (2023)
Similar Items
-
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024) -
Long-Context State-Space Video World Models
by: Po, Ryan, et al.
Published: (2025) -
Jump Cut Smoothing for Talking Heads
by: Wang, Xiaojuan, et al.
Published: (2024) -
Learning an Image Editing Model without Image Editing Pairs
by: Kumari, Nupur, et al.
Published: (2025) -
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
by: Huang, Xun, et al.
Published: (2025)