Neodragon: Mobile Video Generation using Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Karnewar, Animesh, Korzhenkov, Denis, Lelekas, Ioannis, Karjauv, Adil, Fathima, Noor, Xiong, Hanwen, Vaidyanathan, Vancheeswaran, Zeng, Will, Esteves, Rafael, Singhal, Tushar, Porikli, Fatih, Ghafoorian, Mohsen, Habibian, Amirhossein |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026)
by: Korzhenkov, Denis, et al.
Published: (2026)
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)
by: Karjauv, Adil, et al.
Published: (2024)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024)
by: Yahia, Haitam Ben, et al.
Published: (2024)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025)
by: Bhowmik, Aritra, et al.
Published: (2025)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)
by: Habibian, Amirhossein, et al.
Published: (2023)
HLA: Hadamard Linear Attention
by: Ackermann, Hanno, et al.
Published: (2026)
by: Ackermann, Hanno, et al.
Published: (2026)
On Sampling Strategies for Spectral Model Sharding
by: Korzhenkov, Denis, et al.
Published: (2024)
by: Korzhenkov, Denis, et al.
Published: (2024)
Multi-Scale Local Speculative Decoding for Image Generation
by: Peruzzo, Elia, et al.
Published: (2026)
by: Peruzzo, Elia, et al.
Published: (2026)
A Mutual Information Perspective on Federated Contrastive Learning
by: Louizos, Christos, et al.
Published: (2024)
by: Louizos, Christos, et al.
Published: (2024)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
by: Zanjani, Farhad G., et al.
Published: (2026)
by: Zanjani, Farhad G., et al.
Published: (2026)
LaFAM: Unsupervised Feature Attribution with Label-free Activation Maps
by: Karjauv, Aray, et al.
Published: (2024)
by: Karjauv, Aray, et al.
Published: (2024)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
by: Chen, Yanlong, et al.
Published: (2026)
by: Chen, Yanlong, et al.
Published: (2026)
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
by: Petersen, Jens, et al.
Published: (2025)
by: Petersen, Jens, et al.
Published: (2025)
GOEmbed: Gradient Origin Embeddings for Representation Agnostic 3D Feature Learning
by: Karnewar, Animesh, et al.
Published: (2023)
by: Karnewar, Animesh, et al.
Published: (2023)
FastCAD: Real-Time CAD Retrieval and Alignment from Scans and Videos
by: Langer, Florian, et al.
Published: (2024)
by: Langer, Florian, et al.
Published: (2024)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
by: Porikli, Sedat, et al.
Published: (2025)
by: Porikli, Sedat, et al.
Published: (2025)
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
by: Yasarla, Rajeev, et al.
Published: (2025)
by: Yasarla, Rajeev, et al.
Published: (2025)
Resolving the Identity Crisis in Text-to-Image Generation
by: Borse, Shubhankar, et al.
Published: (2025)
by: Borse, Shubhankar, et al.
Published: (2025)
Segmentation-Free Guidance for Text-to-Image Diffusion Models
by: Azarian, Kambiz, et al.
Published: (2024)
by: Azarian, Kambiz, et al.
Published: (2024)
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models
by: Farhadzadeh, Farzad, et al.
Published: (2025)
by: Farhadzadeh, Farzad, et al.
Published: (2025)
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
by: Garrepalli, Risheek, et al.
Published: (2024)
by: Garrepalli, Risheek, et al.
Published: (2024)
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
by: Farhadzadeh, Farzad, et al.
Published: (2025)
by: Farhadzadeh, Farzad, et al.
Published: (2025)
Hybrid Gaussian Splatting for Novel Urban View Synthesis
by: Omran, Mohamed, et al.
Published: (2025)
by: Omran, Mohamed, et al.
Published: (2025)
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
by: Omran, Mohamed, et al.
Published: (2025)
by: Omran, Mohamed, et al.
Published: (2025)
Improving Optical Flow and Stereo Depth Estimation by Leveraging Uncertainty-Based Learning Difficulties
by: Jeong, Jisoo, et al.
Published: (2025)
by: Jeong, Jisoo, et al.
Published: (2025)
DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions
by: Shi, Yunxiao, et al.
Published: (2024)
by: Shi, Yunxiao, et al.
Published: (2024)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
Imagining the Unseen: Generative Location Modeling for Object Placement
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
by: Kadambi, Shreya, et al.
Published: (2025)
by: Kadambi, Shreya, et al.
Published: (2025)
HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories
by: Hedlin, Eric, et al.
Published: (2024)
by: Hedlin, Eric, et al.
Published: (2024)
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
Planar Gaussian Splatting
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
Neural Mesh Fusion: Unsupervised 3D Planar Surface Understanding
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
by: Li, Zhuojin, et al.
Published: (2026)
by: Li, Zhuojin, et al.
Published: (2026)
Designing a Water Pipeline to Last 100 Years or More
by: Ahmad Habibian, et al.
Published: (2024)
by: Ahmad Habibian, et al.
Published: (2024)
FouRA: Fourier Low Rank Adaptation
by: Borse, Shubhankar, et al.
Published: (2024)
by: Borse, Shubhankar, et al.
Published: (2024)
Similar Items
-
PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
by: Korzhenkov, Denis, et al.
Published: (2026) -
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024) -
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025) -
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024) -
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)