Loong: Generating Minute-level Long Videos with Autoregressive Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuqing, Xiong, Tianwei, Zhou, Daquan, Lin, Zhijie, Zhao, Yang, Kang, Bingyi, Feng, Jiashi, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
by: Xu, Lin, et al.
Published: (2024)
by: Xu, Lin, et al.
Published: (2024)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Magic 1-For-1: Generating One Minute Video Clips within One Minute
by: Yi, Hongwei, et al.
Published: (2025)
by: Yi, Hongwei, et al.
Published: (2025)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
Classification Done Right for Vision-Language Pre-Training
by: Huang, Zilong, et al.
Published: (2024)
by: Huang, Zilong, et al.
Published: (2024)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation
by: Wang, Weimin, et al.
Published: (2024)
by: Wang, Weimin, et al.
Published: (2024)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
by: Ren, Zhongwei, et al.
Published: (2026)
by: Ren, Zhongwei, et al.
Published: (2026)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
Minute-Long Videos with Dual Parallelisms
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
by: Chen, Zhekai, et al.
Published: (2026)
by: Chen, Zhekai, et al.
Published: (2026)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
Long Context Tuning for Video Generation
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
by: Zhou, Ziqin, et al.
Published: (2025)
by: Zhou, Ziqin, et al.
Published: (2025)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
by: Li, Zongyi, et al.
Published: (2024)
by: Li, Zongyi, et al.
Published: (2024)
Magic-Me: Identity-Specific Video Customized Diffusion
by: Ma, Ze, et al.
Published: (2024)
by: Ma, Ze, et al.
Published: (2024)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
by: Zhou, Yupeng, et al.
Published: (2026)
by: Zhou, Yupeng, et al.
Published: (2026)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
Depth Anything 3: Recovering the Visual Space from Any Views
by: Lin, Haotong, et al.
Published: (2025)
by: Lin, Haotong, et al.
Published: (2025)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
by: Ma, Fan, et al.
Published: (2023)
by: Ma, Fan, et al.
Published: (2023)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
by: Ouyang, Junyi, et al.
Published: (2026)
by: Ouyang, Junyi, et al.
Published: (2026)
Similar Items
-
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
by: Xiong, Tianwei, et al.
Published: (2024) -
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026) -
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025) -
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025) -
Parallelized Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2024)