GenTron: Diffusion Transformers for Image and Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Shoufa, Xu, Mengmeng, Ren, Jiawei, Cong, Yuren, He, Sen, Xie, Yanping, Sinha, Animesh, Luo, Ping, Xiang, Tao, Perez-Rua, Juan-Manuel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
di: Cong, Yuren, et al.
Pubblicazione: (2023)
di: Cong, Yuren, et al.
Pubblicazione: (2023)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
di: Qiu, Haonan, et al.
Pubblicazione: (2025)
di: Qiu, Haonan, et al.
Pubblicazione: (2025)
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
di: Simon, Christian, et al.
Pubblicazione: (2023)
di: Simon, Christian, et al.
Pubblicazione: (2023)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
Learning Flow Fields in Attention for Controllable Person Image Generation
di: Zhou, Zijian, et al.
Pubblicazione: (2024)
di: Zhou, Zijian, et al.
Pubblicazione: (2024)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
di: Liu, Zhiheng, et al.
Pubblicazione: (2026)
di: Liu, Zhiheng, et al.
Pubblicazione: (2026)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
di: Sun, Peize, et al.
Pubblicazione: (2024)
di: Sun, Peize, et al.
Pubblicazione: (2024)
Move Anything with Layered Scene Diffusion
di: Ren, Jiawei, et al.
Pubblicazione: (2024)
di: Ren, Jiawei, et al.
Pubblicazione: (2024)
SPAN: Learning Similarity between Scene Graphs and Images with Transformers
di: Cong, Yuren, et al.
Pubblicazione: (2023)
di: Cong, Yuren, et al.
Pubblicazione: (2023)
Faster Diffusion via Temporal Attention Decomposition
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
di: Liu, Haozhe, et al.
Pubblicazione: (2024)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
di: Liu, Zhiheng, et al.
Pubblicazione: (2025)
di: Liu, Zhiheng, et al.
Pubblicazione: (2025)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
di: Yang, Jingyuan, et al.
Pubblicazione: (2024)
di: Yang, Jingyuan, et al.
Pubblicazione: (2024)
GenCompositor: Generative Video Compositing with Diffusion Transformer
di: Yang, Shuzhou, et al.
Pubblicazione: (2025)
di: Yang, Shuzhou, et al.
Pubblicazione: (2025)
Context Diffusion: In-Context Aware Image Generation
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2023)
di: Najdenkoska, Ivona, et al.
Pubblicazione: (2023)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
di: Ji, Yatai, et al.
Pubblicazione: (2024)
di: Ji, Yatai, et al.
Pubblicazione: (2024)
Scaling Zero-Shot Reference-to-Video Generation
di: Zhou, Zijian, et al.
Pubblicazione: (2025)
di: Zhou, Zijian, et al.
Pubblicazione: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
di: Hao, Yuren, et al.
Pubblicazione: (2025)
di: Hao, Yuren, et al.
Pubblicazione: (2025)
WavFlow: Audio Generation in Waveform Space
di: Zhou, Feiyan, et al.
Pubblicazione: (2026)
di: Zhou, Feiyan, et al.
Pubblicazione: (2026)
PixelFlow: Pixel-Space Generative Models with Flow
di: Chen, Shoufa, et al.
Pubblicazione: (2025)
di: Chen, Shoufa, et al.
Pubblicazione: (2025)
A Time-Series Data Augmentation Model through Diffusion and Transformer Integration
di: Zhang, Yuren, et al.
Pubblicazione: (2025)
di: Zhang, Yuren, et al.
Pubblicazione: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
di: Xi, Haocheng, et al.
Pubblicazione: (2025)
SimulTron: On-Device Simultaneous Speech to Speech Translation
di: Agranovich, Alex, et al.
Pubblicazione: (2024)
di: Agranovich, Alex, et al.
Pubblicazione: (2024)
GenAI Distortion: The Effect of GenAI Fluency and Positive Affect
di: Yang, Xiantong, et al.
Pubblicazione: (2024)
di: Yang, Xiantong, et al.
Pubblicazione: (2024)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
The Tron Technical Challenge: History of Visual Effects and Computer Graphics
di: Elio Quiroga Rodríguez
Pubblicazione: (2025)
di: Elio Quiroga Rodríguez
Pubblicazione: (2025)
DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
di: Zhong, Yufeng, et al.
Pubblicazione: (2025)
di: Zhong, Yufeng, et al.
Pubblicazione: (2025)
SerialGen: Personalized Image Generation by First Standardization Then Personalization
di: Xie, Cong, et al.
Pubblicazione: (2024)
di: Xie, Cong, et al.
Pubblicazione: (2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
di: Liu, Shaowei, et al.
Pubblicazione: (2024)
di: Liu, Shaowei, et al.
Pubblicazione: (2024)
GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
di: Zhu, Hanxin, et al.
Pubblicazione: (2026)
di: Zhu, Hanxin, et al.
Pubblicazione: (2026)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
di: Chen, Junsong, et al.
Pubblicazione: (2024)
di: Chen, Junsong, et al.
Pubblicazione: (2024)
WorldAfford: Affordance Grounding based on Natural Language Instructions
di: Chen, Changmao, et al.
Pubblicazione: (2024)
di: Chen, Changmao, et al.
Pubblicazione: (2024)
A Beam-Segmenting Polar Format Algorithm Based on Double PCS for Video SAR Persistent Imaging
di: Jiang, Jiawei, et al.
Pubblicazione: (2023)
di: Jiang, Jiawei, et al.
Pubblicazione: (2023)
Neodragon: Mobile Video Generation using Diffusion Transformer
di: Karnewar, Animesh, et al.
Pubblicazione: (2025)
di: Karnewar, Animesh, et al.
Pubblicazione: (2025)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
di: Liu, Zhiheng, et al.
Pubblicazione: (2025)
di: Liu, Zhiheng, et al.
Pubblicazione: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
di: Huang, Zhijian, et al.
Pubblicazione: (2024)
di: Huang, Zhijian, et al.
Pubblicazione: (2024)
Parameter extraction for a superconducting thermal switch (hTron) SPICE model
di: Karam, Valentin, et al.
Pubblicazione: (2024)
di: Karam, Valentin, et al.
Pubblicazione: (2024)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
di: Yan, Feng, et al.
Pubblicazione: (2024)
di: Yan, Feng, et al.
Pubblicazione: (2024)
Tron [Película] / Steven Lisberger, director y guionista ; Donald Kushner, productor
di: Lisberger, Steven
Pubblicazione: (1982)
di: Lisberger, Steven
Pubblicazione: (1982)
Drift Analysis with Fitness Levels for Elitist Evolutionary Algorithms
di: He, Jun, et al.
Pubblicazione: (2023)
di: He, Jun, et al.
Pubblicazione: (2023)
Documenti analoghi
-
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
di: Cong, Yuren, et al.
Pubblicazione: (2023) -
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
di: Qiu, Haonan, et al.
Pubblicazione: (2025) -
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
di: Simon, Christian, et al.
Pubblicazione: (2023) -
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
di: Liu, Haozhe, et al.
Pubblicazione: (2024) -
Learning Flow Fields in Attention for Controllable Person Image Generation
di: Zhou, Zijian, et al.
Pubblicazione: (2024)