LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Hongjie, Ma, Chih-Yao, Liu, Yen-Cheng, Hou, Ji, Xu, Tao, Wang, Jialiang, Juefei-Xu, Felix, Luo, Yaqiao, Zhang, Peizhao, Hou, Tingbo, Vajda, Peter, Jha, Niraj K., Dai, Xiaoliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LinMU: Multimodal Understanding Made Linear
di: Wang, Hongjie, et al.
Pubblicazione: (2026)
di: Wang, Hongjie, et al.
Pubblicazione: (2026)
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
di: Song, Kunpeng, et al.
Pubblicazione: (2024)
di: Song, Kunpeng, et al.
Pubblicazione: (2024)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
di: Liang, Feng, et al.
Pubblicazione: (2025)
di: Liang, Feng, et al.
Pubblicazione: (2025)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
di: Ma, Xu, et al.
Pubblicazione: (2025)
di: Ma, Xu, et al.
Pubblicazione: (2025)
Populate-A-Scene: Affordance-Aware Human Video Generation
di: Shan, Mengyi, et al.
Pubblicazione: (2025)
di: Shan, Mengyi, et al.
Pubblicazione: (2025)
StreamDiT: Real-Time Streaming Text-to-Video Generation
di: Kodaira, Akio, et al.
Pubblicazione: (2025)
di: Kodaira, Akio, et al.
Pubblicazione: (2025)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
di: Zhang, Haochen, et al.
Pubblicazione: (2026)
di: Zhang, Haochen, et al.
Pubblicazione: (2026)
MoCha: Towards Movie-Grade Talking Character Synthesis
di: Wei, Cong, et al.
Pubblicazione: (2025)
di: Wei, Cong, et al.
Pubblicazione: (2025)
An Analysis on Quantizing Diffusion Transformers
di: Yang, Yuewei, et al.
Pubblicazione: (2024)
di: Yang, Yuewei, et al.
Pubblicazione: (2024)
Pixel-Space Post-Training of Latent Diffusion Models
di: Zhang, Christina, et al.
Pubblicazione: (2024)
di: Zhang, Christina, et al.
Pubblicazione: (2024)
FBNetV5: Neural Architecture Search for Multiple Tasks in One Run
di: Wu, Bichen, et al.
Pubblicazione: (2021)
di: Wu, Bichen, et al.
Pubblicazione: (2021)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
di: You, Haoran, et al.
Pubblicazione: (2022)
di: You, Haoran, et al.
Pubblicazione: (2022)
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
di: Wimbauer, Felix, et al.
Pubblicazione: (2023)
di: Wimbauer, Felix, et al.
Pubblicazione: (2023)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
di: Cai, Yuanhao, et al.
Pubblicazione: (2025)
di: Cai, Yuanhao, et al.
Pubblicazione: (2025)
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers
di: Wang, Hongjie, et al.
Pubblicazione: (2023)
di: Wang, Hongjie, et al.
Pubblicazione: (2023)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2025)
di: Hansen-Estruch, Philippe, et al.
Pubblicazione: (2025)
AVID: Any-Length Video Inpainting with Diffusion Model
di: Zhang, Zhixing, et al.
Pubblicazione: (2023)
di: Zhang, Zhixing, et al.
Pubblicazione: (2023)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
di: Zhao, Yang, et al.
Pubblicazione: (2023)
di: Zhao, Yang, et al.
Pubblicazione: (2023)
Teaching LLMs to See Graphs: Unifying Text and Structural Reasoning
di: Vajda, Dario
Pubblicazione: (2026)
di: Vajda, Dario
Pubblicazione: (2026)
LinFusion: 1 GPU, 1 Minute, 16K Image
di: Liu, Songhua, et al.
Pubblicazione: (2024)
di: Liu, Songhua, et al.
Pubblicazione: (2024)
Undecidability of Polynomial Inequalities in Subset Densities and Additive Energies
di: Li, Yaqiao
Pubblicazione: (2025)
di: Li, Yaqiao
Pubblicazione: (2025)
TopAY: Efficient Trajectory Planning for Differential Drive Mobile Manipulators via Topological Paths Search and Arc Length-Yaw Parameterization
di: Xu, Long, et al.
Pubblicazione: (2025)
di: Xu, Long, et al.
Pubblicazione: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
FLIP: Real-Time and Resilient Formation Planning for Large-Scale DIstributed Swarms via Point Cloud Registration
di: Zhou, Yuan, et al.
Pubblicazione: (2026)
di: Zhou, Yuan, et al.
Pubblicazione: (2026)
INGeo: Accelerating Instant Neural Scene Reconstruction with Noisy Geometry Priors
di: Li, Chaojian, et al.
Pubblicazione: (2022)
di: Li, Chaojian, et al.
Pubblicazione: (2022)
PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models
di: Deng, Fei, et al.
Pubblicazione: (2024)
di: Deng, Fei, et al.
Pubblicazione: (2024)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
di: Liang, Feng, et al.
Pubblicazione: (2023)
di: Liang, Feng, et al.
Pubblicazione: (2023)
Ethical-Lens: Curbing Malicious Usages of Open-Source Text-to-Image Models
di: Cai, Yuzhu, et al.
Pubblicazione: (2024)
di: Cai, Yuzhu, et al.
Pubblicazione: (2024)
Number Adaptive Formation Flight Planning via Affine Deformable Guidance in Narrow Environments
di: Zhou, Yuan, et al.
Pubblicazione: (2025)
di: Zhou, Yuan, et al.
Pubblicazione: (2025)
Transfer between Modalities with MetaQueries
di: Pan, Xichen, et al.
Pubblicazione: (2025)
di: Pan, Xichen, et al.
Pubblicazione: (2025)
Movie Gen: A Cast of Media Foundation Models
di: Polyak, Adam, et al.
Pubblicazione: (2024)
di: Polyak, Adam, et al.
Pubblicazione: (2024)
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
di: Chen, Leon Liangyu, et al.
Pubblicazione: (2026)
di: Chen, Leon Liangyu, et al.
Pubblicazione: (2026)
Exploring MLLM-Diffusion Information Transfer with MetaCanvas
di: Lin, Han, et al.
Pubblicazione: (2025)
di: Lin, Han, et al.
Pubblicazione: (2025)
Graph Similarity Computation via Interpretable Neural Node Alignment
di: Wang, Jingjing, et al.
Pubblicazione: (2024)
di: Wang, Jingjing, et al.
Pubblicazione: (2024)
Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models
di: Wang, Hongjie, et al.
Pubblicazione: (2024)
di: Wang, Hongjie, et al.
Pubblicazione: (2024)
Learning Interpretable Differentiable Logic Networks for Time-Series Classification
di: Yue, Chang, et al.
Pubblicazione: (2025)
di: Yue, Chang, et al.
Pubblicazione: (2025)
Uncertainty-Aware Transformers: Conformal Prediction for Language Models
di: Vellore, Abhiram, et al.
Pubblicazione: (2026)
di: Vellore, Abhiram, et al.
Pubblicazione: (2026)
TAD-SIE: Sample Size Estimation for Clinical Randomized Controlled Trials using a Trend-Adaptive Design with a Synthetic-Intervention-Based Estimator
di: Lala, Sayeri, et al.
Pubblicazione: (2024)
di: Lala, Sayeri, et al.
Pubblicazione: (2024)
METRIK: Measurement-Efficient Randomized Controlled Trials using Transformers with Input Masking
di: Lala, Sayeri, et al.
Pubblicazione: (2024)
di: Lala, Sayeri, et al.
Pubblicazione: (2024)
Learning Interpretable Differentiable Logic Networks
di: Yue, Chang, et al.
Pubblicazione: (2024)
di: Yue, Chang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LinMU: Multimodal Understanding Made Linear
di: Wang, Hongjie, et al.
Pubblicazione: (2026) -
Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation
di: Song, Kunpeng, et al.
Pubblicazione: (2024) -
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
di: Liang, Feng, et al.
Pubblicazione: (2025) -
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
di: Ma, Xu, et al.
Pubblicazione: (2025) -
Populate-A-Scene: Affordance-Aware Human Video Generation
di: Shan, Mengyi, et al.
Pubblicazione: (2025)