GenTron: Diffusion Transformers for Image and Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shoufa, Xu, Mengmeng, Ren, Jiawei, Cong, Yuren, He, Sen, Xie, Yanping, Sinha, Animesh, Luo, Ping, Xiang, Tao, Perez-Rua, Juan-Manuel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023)
by: Cong, Yuren, et al.
Published: (2023)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
by: Simon, Christian, et al.
Published: (2023)
by: Simon, Christian, et al.
Published: (2023)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Learning Flow Fields in Attention for Controllable Person Image Generation
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
by: Liu, Zhiheng, et al.
Published: (2026)
by: Liu, Zhiheng, et al.
Published: (2026)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
by: Sun, Peize, et al.
Published: (2024)
by: Sun, Peize, et al.
Published: (2024)
Move Anything with Layered Scene Diffusion
by: Ren, Jiawei, et al.
Published: (2024)
by: Ren, Jiawei, et al.
Published: (2024)
SPAN: Learning Similarity between Scene Graphs and Images with Transformers
by: Cong, Yuren, et al.
Published: (2023)
by: Cong, Yuren, et al.
Published: (2023)
Faster Diffusion via Temporal Attention Decomposition
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
GenCompositor: Generative Video Compositing with Diffusion Transformer
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
Context Diffusion: In-Context Aware Image Generation
by: Najdenkoska, Ivona, et al.
Published: (2023)
by: Najdenkoska, Ivona, et al.
Published: (2023)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Scaling Zero-Shot Reference-to-Video Generation
by: Zhou, Zijian, et al.
Published: (2025)
by: Zhou, Zijian, et al.
Published: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
WavFlow: Audio Generation in Waveform Space
by: Zhou, Feiyan, et al.
Published: (2026)
by: Zhou, Feiyan, et al.
Published: (2026)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
A Time-Series Data Augmentation Model through Diffusion and Transformer Integration
by: Zhang, Yuren, et al.
Published: (2025)
by: Zhang, Yuren, et al.
Published: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
SimulTron: On-Device Simultaneous Speech to Speech Translation
by: Agranovich, Alex, et al.
Published: (2024)
by: Agranovich, Alex, et al.
Published: (2024)
GenAI Distortion: The Effect of GenAI Fluency and Positive Affect
by: Yang, Xiantong, et al.
Published: (2024)
by: Yang, Xiantong, et al.
Published: (2024)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
The Tron Technical Challenge: History of Visual Effects and Computer Graphics
by: Elio Quiroga Rodríguez
Published: (2025)
by: Elio Quiroga Rodríguez
Published: (2025)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
SerialGen: Personalized Image Generation by First Standardization Then Personalization
by: Xie, Cong, et al.
Published: (2024)
by: Xie, Cong, et al.
Published: (2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024)
by: Liu, Shaowei, et al.
Published: (2024)
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
by: Zhu, Hanxin, et al.
Published: (2026)
by: Zhu, Hanxin, et al.
Published: (2026)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
WorldAfford: Affordance Grounding based on Natural Language Instructions
by: Chen, Changmao, et al.
Published: (2024)
by: Chen, Changmao, et al.
Published: (2024)
A Beam-Segmenting Polar Format Algorithm Based on Double PCS for Video SAR Persistent Imaging
by: Jiang, Jiawei, et al.
Published: (2023)
by: Jiang, Jiawei, et al.
Published: (2023)
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025)
by: Karnewar, Animesh, et al.
Published: (2025)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
Parameter extraction for a superconducting thermal switch (hTron) SPICE model
by: Karam, Valentin, et al.
Published: (2024)
by: Karam, Valentin, et al.
Published: (2024)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
by: Yan, Feng, et al.
Published: (2024)
by: Yan, Feng, et al.
Published: (2024)
Tron [Película] / Steven Lisberger, director y guionista ; Donald Kushner, productor
by: Lisberger, Steven
Published: (1982)
by: Lisberger, Steven
Published: (1982)
Drift Analysis with Fitness Levels for Elitist Evolutionary Algorithms
by: He, Jun, et al.
Published: (2023)
by: He, Jun, et al.
Published: (2023)
Similar Items
-
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023) -
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025) -
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
by: Simon, Christian, et al.
Published: (2023) -
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024) -
Learning Flow Fields in Attention for Controllable Person Image Generation
by: Zhou, Zijian, et al.
Published: (2024)