SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Dongting, Chen, Jierun, Huang, Xijie, Coskun, Huseyin, Sahni, Arpit, Gupta, Aarush, Goyal, Anujraaj, Lahiri, Dishani, Singh, Rajesh, Idelbayev, Yerlan, Cao, Junli, Li, Yanyu, Cheng, Kwang-Ting, Chan, S. -H. Gary, Gong, Mingming, Tulyakov, Sergey, Kag, Anil, Xu, Yanwu, Ren, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026)
by: Hu, Dongting, et al.
Published: (2026)
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
by: Wu, Yushu, et al.
Published: (2024)
by: Wu, Yushu, et al.
Published: (2024)
AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation
by: Kag, Anil, et al.
Published: (2024)
by: Kag, Anil, et al.
Published: (2024)
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)
by: Li, Yanyu, et al.
Published: (2024)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
Scalable Ranked Preference Optimization for Text-to-Image Generation
by: Karthik, Shyamgopal, et al.
Published: (2024)
by: Karthik, Shyamgopal, et al.
Published: (2024)
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
by: Park, Dogyun, et al.
Published: (2026)
by: Park, Dogyun, et al.
Published: (2026)
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
Efficient Training with Denoised Neural Weights
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
SF-V: Single Forward Video Generation Model
by: Zhang, Zhixing, et al.
Published: (2024)
by: Zhang, Zhixing, et al.
Published: (2024)
E$^{2}$GAN: Efficient Training of Efficient GANs for Image-to-Image Translation
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
LayerComposer: Multi-Human Personalized Generation via Layered Canvas
by: Qian, Guocheng Gordon, et al.
Published: (2025)
by: Qian, Guocheng Gordon, et al.
Published: (2025)
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
by: Chen, Yunuo, et al.
Published: (2025)
by: Chen, Yunuo, et al.
Published: (2025)
BodyMAP -- Jointly Predicting Body Mesh and 3D Applied Pressure Map for People in Bed
by: Tandon, Abhishek, et al.
Published: (2024)
by: Tandon, Abhishek, et al.
Published: (2024)
Improving Progressive Generation with Decomposable Flow Matching
by: Haji-Ali, Moayed, et al.
Published: (2025)
by: Haji-Ali, Moayed, et al.
Published: (2025)
Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
Lightweight Predictive 3D Gaussian Splats
by: Cao, Junli, et al.
Published: (2024)
by: Cao, Junli, et al.
Published: (2024)
Recovering the state and dynamics of autonomous system with partial states solution using neural networks
by: Kag, Vijay
Published: (2024)
by: Kag, Vijay
Published: (2024)
OmniPatch: A Universal Adversarial Patch for ViT-CNN Cross-Architecture Transfer in Semantic Segmentation
by: Aggarwal, Aarush, et al.
Published: (2026)
by: Aggarwal, Aarush, et al.
Published: (2026)
MF-VITON: High-Fidelity Mask-Free Virtual Try-On with Minimal Input
by: Wan, Zhenchen, et al.
Published: (2025)
by: Wan, Zhenchen, et al.
Published: (2025)
Research Paper: Customer Segmentation and Personalization in Insurance: A Business Analysis Approach
by: Rajesh Goyal
Published: (2024)
by: Rajesh Goyal
Published: (2024)
KUZATUV KAMERALARI VIDEO OQIMLARIDA HARAKATNI ANIQLASH USULLARI TAHLILI
by: Tashmanov, Yerlan
Published: (2025)
by: Tashmanov, Yerlan
Published: (2025)
A Multi-Branched Radial Basis Network Approach to Predicting Complex Chaotic Behaviours
by: Sinha, Aarush
Published: (2024)
by: Sinha, Aarush
Published: (2024)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
GMLM: Bridging Graph Neural Networks and Language Models for Heterophilic Node Classification
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
Joint stereo 3D object detection and implicit surface reconstruction
by: Li, Shichao, et al.
Published: (2021)
by: Li, Shichao, et al.
Published: (2021)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
A novel framework for generalization of deep hidden physics models
by: Kag, Vijay, et al.
Published: (2024)
by: Kag, Vijay, et al.
Published: (2024)
Physics-informed neural network for modeling dynamic linear elasticity
by: Kag, Vijay, et al.
Published: (2023)
by: Kag, Vijay, et al.
Published: (2023)
SAM3D: Zero-Shot Semi-Automatic Segmentation in 3D Medical Images with the Segment Anything Model
by: Chan, Trevor J., et al.
Published: (2024)
by: Chan, Trevor J., et al.
Published: (2024)
Zombie firms, state subsidies and aggregate productivity
by: Xijie Gao
Published: (2025)
by: Xijie Gao
Published: (2025)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
CFD-DEM modeling of fracture initiation with polymer injection in granular media
by: Kazidenov, Daniyar, et al.
Published: (2024)
by: Kazidenov, Daniyar, et al.
Published: (2024)
Similar Items
-
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026) -
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
by: Wu, Yushu, et al.
Published: (2024) -
AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation
by: Kag, Anil, et al.
Published: (2024) -
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
by: Sui, Yang, et al.
Published: (2024) -
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)