SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yushu, Zhang, Zhixing, Li, Yanyu, Xu, Yanwu, Kag, Anil, Sui, Yang, Coskun, Huseyin, Ma, Ke, Lebedev, Aleksei, Hu, Ju, Metaxas, Dimitris, Wang, Yanzhi, Tulyakov, Sergey, Ren, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026)
by: Hu, Dongting, et al.
Published: (2026)
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
by: Hu, Dongting, et al.
Published: (2024)
by: Hu, Dongting, et al.
Published: (2024)
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
SF-V: Single Forward Video Generation Model
by: Zhang, Zhixing, et al.
Published: (2024)
by: Zhang, Zhixing, et al.
Published: (2024)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
by: Park, Dogyun, et al.
Published: (2026)
by: Park, Dogyun, et al.
Published: (2026)
Scalable Ranked Preference Optimization for Text-to-Image Generation
by: Karthik, Shyamgopal, et al.
Published: (2024)
by: Karthik, Shyamgopal, et al.
Published: (2024)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)
by: Li, Yanyu, et al.
Published: (2024)
AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation
by: Kag, Anil, et al.
Published: (2024)
by: Kag, Anil, et al.
Published: (2024)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
by: Park, Dogyun, et al.
Published: (2025)
by: Park, Dogyun, et al.
Published: (2025)
Lightweight Predictive 3D Gaussian Splats
by: Cao, Junli, et al.
Published: (2024)
by: Cao, Junli, et al.
Published: (2024)
BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models
by: Deng, Fei, et al.
Published: (2026)
by: Deng, Fei, et al.
Published: (2026)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
by: Chen, Yunuo, et al.
Published: (2025)
by: Chen, Yunuo, et al.
Published: (2025)
Five Seconds to End Corruption and Political Dynasty - Will they sign? (Sworn Affidavit of Declaration + Instructions)
by: Munoz, Jonelle Peter
Published: (2025)
by: Munoz, Jonelle Peter
Published: (2025)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
Efficient Training with Denoised Neural Weights
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
by: Bang, Jaehun, et al.
Published: (2026)
by: Bang, Jaehun, et al.
Published: (2026)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
by: Zhao, Yang, et al.
Published: (2023)
by: Zhao, Yang, et al.
Published: (2023)
One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
by: Haji-Ali, Moayed, et al.
Published: (2026)
by: Haji-Ali, Moayed, et al.
Published: (2026)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
Five shades of KMS: Statistical properties in the spectral geometry of Cuntz--Krieger algebras
by: Gerontogiannis, Dimitris Michail, et al.
Published: (2026)
by: Gerontogiannis, Dimitris Michail, et al.
Published: (2026)
Scientific and Technical Information Services in Indonesia; An Approach to Development in 1974-79 Under the Second Five-Year Plan.
by: Gray, J. C.
Published: (1972)
by: Gray, J. C.
Published: (1972)
Five Things Right, Five Wrong
by: Morley, Gabriel
Published: (2005)
by: Morley, Gabriel
Published: (2005)
Second Thoughts: How 1-second subslots transform CEX-DEX Arbitrage on Ethereum
by: Adadurov, Aleksei, et al.
Published: (2026)
by: Adadurov, Aleksei, et al.
Published: (2026)
Five Years into the Past...Five Years into the Future.
by: Tenopir, Carol
Published: (1988)
by: Tenopir, Carol
Published: (1988)
Beyond Explicit Edges: Robust Reasoning over Noisy and Sparse Knowledge Graphs
by: Gao, Hang, et al.
Published: (2026)
by: Gao, Hang, et al.
Published: (2026)
Central Nervous System Tumors in Xeroderma Pigmentosum: Five Cases and Review of the Literature
by: Farrah S. Bakr, et al.
Published: (2026)
by: Farrah S. Bakr, et al.
Published: (2026)
PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation
by: Mo, Wenyi, et al.
Published: (2025)
by: Mo, Wenyi, et al.
Published: (2025)
Recovering the state and dynamics of autonomous system with partial states solution using neural networks
by: Kag, Vijay
Published: (2024)
by: Kag, Vijay
Published: (2024)
Educational Mobility of Second-generation Turks
by: Schnell, Philipp
Published: (2025)
by: Schnell, Philipp
Published: (2025)
Five poems
by: John F. Sherry
Published: (2025)
by: John F. Sherry
Published: (2025)
Five Women
by: Gell, Marilyn
Published: (1975)
by: Gell, Marilyn
Published: (1975)
E$^{2}$GAN: Efficient Training of Efficient GANs for Image-to-Image Translation
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025)
by: Dafnis, Konstantinos M., et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
Similar Items
-
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026) -
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
by: Hu, Dongting, et al.
Published: (2024) -
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
by: Wu, Yushu, et al.
Published: (2025) -
SF-V: Single Forward Video Generation Model
by: Zhang, Zhixing, et al.
Published: (2024) -
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)