FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jingfeng, Cheng, Wang, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
by: Yao, Jingfeng, et al.
Published: (2023)
by: Yao, Jingfeng, et al.
Published: (2023)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025)
by: Zeng, Lunbin, et al.
Published: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024)
by: Song, Yuehao, et al.
Published: (2024)
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
by: Zou, Ya, et al.
Published: (2025)
by: Zou, Ya, et al.
Published: (2025)
Faster Diffusion Action Segmentation
by: Wang, Shuaibing, et al.
Published: (2024)
by: Wang, Shuaibing, et al.
Published: (2024)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
by: Xiong, Haijun, et al.
Published: (2025)
by: Xiong, Haijun, et al.
Published: (2025)
Point Transformer V3: Simpler, Faster, Stronger
by: Wu, Xiaoyang, et al.
Published: (2023)
by: Wu, Xiaoyang, et al.
Published: (2023)
EVA-X: A Foundation Model for General Chest X-ray Analysis with Self-supervised Learning
by: Yao, Jingfeng, et al.
Published: (2024)
by: Yao, Jingfeng, et al.
Published: (2024)
LKCell: Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels
by: Cui, Ziwei, et al.
Published: (2024)
by: Cui, Ziwei, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
Faster Projected GAN: Towards Faster Few-Shot Image Generation
by: Wang, Chuang, et al.
Published: (2024)
by: Wang, Chuang, et al.
Published: (2024)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Faster Diffusion via Temporal Attention Decomposition
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Faster and Better 3D Splatting via Group Training
by: Wang, Chengbo, et al.
Published: (2024)
by: Wang, Chengbo, et al.
Published: (2024)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
by: Mei, Kangfu, et al.
Published: (2023)
by: Mei, Kangfu, et al.
Published: (2023)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
by: Xiong, Haijun, et al.
Published: (2024)
by: Xiong, Haijun, et al.
Published: (2024)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
by: Zhang, Wenrui, et al.
Published: (2025)
by: Zhang, Wenrui, et al.
Published: (2025)
Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance
by: Lin, Liting, et al.
Published: (2024)
by: Lin, Liting, et al.
Published: (2024)
Occupancy as Set of Points
by: Shi, Yiang, et al.
Published: (2024)
by: Shi, Yiang, et al.
Published: (2024)
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
by: Gupta, Samarth, et al.
Published: (2025)
by: Gupta, Samarth, et al.
Published: (2025)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
by: Zhang, Shuai, et al.
Published: (2025)
by: Zhang, Shuai, et al.
Published: (2025)
Visual Generation Tuning
by: Guo, Jiahao, et al.
Published: (2025)
by: Guo, Jiahao, et al.
Published: (2025)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)
by: Chen, Shaoyu, et al.
Published: (2022)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Similar Items
-
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
by: Yao, Jingfeng, et al.
Published: (2023) -
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025) -
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024) -
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
by: Yao, Jingfeng, et al.
Published: (2025) -
Towards Scalable Pre-training of Visual Tokenizers for Generation
by: Yao, Jingfeng, et al.
Published: (2025)