Few-Step Distillation for Text-to-Image Generation: A Practical Guide
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pu, Yifan, Han, Yizeng, Tang, Zhiwei, Tang, Jiasheng, Wang, Fan, Zhuang, Bohan, Huang, Gao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025)
Dynamic Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
von: Dao, Quan, et al.
Veröffentlicht: (2024)
von: Dao, Quan, et al.
Veröffentlicht: (2024)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
von: Wang, Jiangshan, et al.
Veröffentlicht: (2024)
von: Wang, Jiangshan, et al.
Veröffentlicht: (2024)
Exploring contextual modeling with linear complexity for point cloud segmentation
von: Chng, Yong Xien, et al.
Veröffentlicht: (2024)
von: Chng, Yong Xien, et al.
Veröffentlicht: (2024)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
von: Pu, Yifan, et al.
Veröffentlicht: (2024)
von: Pu, Yifan, et al.
Veröffentlicht: (2024)
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2024)
Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2024)
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2024)
InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
von: Gong, Yiming, et al.
Veröffentlicht: (2025)
von: Gong, Yiming, et al.
Veröffentlicht: (2025)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
von: Liu, Akide, et al.
Veröffentlicht: (2025)
von: Liu, Akide, et al.
Veröffentlicht: (2025)
Mask Grounding for Referring Image Segmentation
von: Chng, Yong Xien, et al.
Veröffentlicht: (2023)
von: Chng, Yong Xien, et al.
Veröffentlicht: (2023)
Bridging the Divide: Reconsidering Softmax and Linear Attention
von: Han, Dongchen, et al.
Veröffentlicht: (2024)
von: Han, Dongchen, et al.
Veröffentlicht: (2024)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
GSVA: Generalized Segmentation via Multimodal Large Language Models
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
von: Xia, Zhuofan, et al.
Veröffentlicht: (2023)
SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2026)
von: Wu, Zhuguanyu, et al.
Veröffentlicht: (2026)
Demystify Mamba in Vision: A Linear Attention Perspective
von: Han, Dongchen, et al.
Veröffentlicht: (2024)
von: Han, Dongchen, et al.
Veröffentlicht: (2024)
Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection
von: Wu, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2026)
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
von: Gu, Youping, et al.
Veröffentlicht: (2025)
von: Gu, Youping, et al.
Veröffentlicht: (2025)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2024)
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2024)
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
von: Inferix Team, et al.
Veröffentlicht: (2025)
von: Inferix Team, et al.
Veröffentlicht: (2025)
Optimizing Few-Step Generation with Adaptive Matching Distillation
von: Bai, Lichen, et al.
Veröffentlicht: (2026)
von: Bai, Lichen, et al.
Veröffentlicht: (2026)
DistillKac: Few-Step Image Generation via Damped Wave Equations
von: Han, Weiqiao, et al.
Veröffentlicht: (2025)
von: Han, Weiqiao, et al.
Veröffentlicht: (2025)
Semantic-Guided Diffusion Model for Single-Step Image Super-Resolution
von: Liu, Zihang, et al.
Veröffentlicht: (2025)
von: Liu, Zihang, et al.
Veröffentlicht: (2025)
DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
Learning Multi-dimensional Human Preference for Text-to-Image Generation
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
Distilling ODE Solvers of Diffusion Models into Smaller Steps
von: Kim, Sanghwan, et al.
Veröffentlicht: (2023)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2023)
Pseudo-triplet Guided Few-shot Composed Image Retrieval
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
von: Hou, Bohan, et al.
Veröffentlicht: (2024)
Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
von: Chen, Guanjie, et al.
Veröffentlicht: (2025)
von: Chen, Guanjie, et al.
Veröffentlicht: (2025)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
von: Huang-Menders, Alexander, et al.
Veröffentlicht: (2025)
von: Huang-Menders, Alexander, et al.
Veröffentlicht: (2025)
MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis
von: Shao, Shitong, et al.
Veröffentlicht: (2025)
von: Shao, Shitong, et al.
Veröffentlicht: (2025)
Few-Step Diffusion via Score identity Distillation
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Mingyuan, et al.
Veröffentlicht: (2025)
Learning Few-Step Diffusion Models by Trajectory Distribution Matching
von: Luo, Yihong, et al.
Veröffentlicht: (2025)
von: Luo, Yihong, et al.
Veröffentlicht: (2025)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
von: Yang, Le, et al.
Veröffentlicht: (2024)
von: Yang, Le, et al.
Veröffentlicht: (2024)
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
von: Chang, Shuning, et al.
Veröffentlicht: (2024)
von: Chang, Shuning, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025) -
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025) -
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2025) -
Dynamic Diffusion Transformer
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024) -
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)