Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Hangjie, Chen, Weihua, Cen, Jun, Yu, Hu, Liang, Jingyun, Chang, Shuning, Lin, Zhihui, Feng, Tao, Liu, Pengwei, Xing, Jiazheng, Luo, Hao, Tang, Jiasheng, Wang, Fan, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LumosFlow: Motion-Guided Long Video Generation
by: Chen, Jiahao, et al.
Published: (2025)
by: Chen, Jiahao, et al.
Published: (2025)
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
by: Liu, Ropeway, et al.
Published: (2025)
by: Liu, Ropeway, et al.
Published: (2025)
AMD: Autoregressive Motion Diffusion
by: Han, Bo, et al.
Published: (2023)
by: Han, Bo, et al.
Published: (2023)
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
ConCLVD: Controllable Chinese Landscape Video Generation via Diffusion Model
by: Liu, Dingming, et al.
Published: (2024)
by: Liu, Dingming, et al.
Published: (2024)
CAFE: Channel-Autoregressive Factorized Encoding for Robust Biosignal Spatial Super-Resolution
by: Liu, Hongjun, et al.
Published: (2026)
by: Liu, Hongjun, et al.
Published: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation
by: Zhu, Jian, et al.
Published: (2026)
by: Zhu, Jian, et al.
Published: (2026)
Bridging Discrete and Continuous: A Multimodal Strategy for Complex Emotion Detection
by: Jia, Jiehui, et al.
Published: (2024)
by: Jia, Jiehui, et al.
Published: (2024)
LayerT2V: A Unified Multi-Layer Video Generation Framework
by: Li, Guangzhao, et al.
Published: (2025)
by: Li, Guangzhao, et al.
Published: (2025)
Joint Optimization of Buffer Delay and HARQ for Video Communications
by: Cheng, Baoping, et al.
Published: (2024)
by: Cheng, Baoping, et al.
Published: (2024)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
by: Chen, Zhikai, et al.
Published: (2024)
by: Chen, Zhikai, et al.
Published: (2024)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
by: Bai, Hayes, et al.
Published: (2026)
by: Bai, Hayes, et al.
Published: (2026)
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
by: Mao, Yuxin, et al.
Published: (2025)
by: Mao, Yuxin, et al.
Published: (2025)
SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
by: Yang, Chenyu, et al.
Published: (2025)
by: Yang, Chenyu, et al.
Published: (2025)
Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok
by: Zha, Mingyue, et al.
Published: (2026)
by: Zha, Mingyue, et al.
Published: (2026)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
by: Tang, Guowei, et al.
Published: (2026)
by: Tang, Guowei, et al.
Published: (2026)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
by: Lei, Ke, et al.
Published: (2026)
by: Lei, Ke, et al.
Published: (2026)
EV-NVC: Efficient Variable bitrate Neural Video Compression
by: Hu, Yongcun, et al.
Published: (2025)
by: Hu, Yongcun, et al.
Published: (2025)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
by: Zhang, Jiaxu, et al.
Published: (2025)
by: Zhang, Jiaxu, et al.
Published: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
by: Zhou, Yang-Hao, et al.
Published: (2026)
by: Zhou, Yang-Hao, et al.
Published: (2026)
Emotional Cues Extraction and Fusion for Multi-modal Emotion Prediction and Recognition in Conversation
by: Shi, Haoxiang, et al.
Published: (2024)
by: Shi, Haoxiang, et al.
Published: (2024)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
by: Cheng, Zhi-Qi, et al.
Published: (2024)
by: Cheng, Zhi-Qi, et al.
Published: (2024)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
by: Wu, Yuheng, et al.
Published: (2026)
by: Wu, Yuheng, et al.
Published: (2026)
Diffusion Model-Based Size Variable Virtual Try-On Technology and Evaluation Method
by: Zhang, Shufang, et al.
Published: (2025)
by: Zhang, Shufang, et al.
Published: (2025)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
by: Gong, Haisong, et al.
Published: (2025)
by: Gong, Haisong, et al.
Published: (2025)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
by: Mao, Qingyang, et al.
Published: (2025)
by: Mao, Qingyang, et al.
Published: (2025)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
by: Yan, Mingxuan, et al.
Published: (2024)
by: Yan, Mingxuan, et al.
Published: (2024)
Temporally Aligned Audio for Video with Autoregression
by: Viertola, Ilpo, et al.
Published: (2024)
by: Viertola, Ilpo, et al.
Published: (2024)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
by: Cai, Haonan, et al.
Published: (2026)
by: Cai, Haonan, et al.
Published: (2026)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
by: Zheng, Guangting, et al.
Published: (2025)
by: Zheng, Guangting, et al.
Published: (2025)
Efficient Geometry Compression and Communication for 3D Gaussian Splatting Point Clouds
by: Xie, Liang, et al.
Published: (2025)
by: Xie, Liang, et al.
Published: (2025)
FlexCache: Flexible Approximate Cache System for Video Diffusion
by: Sun, Desen, et al.
Published: (2024)
by: Sun, Desen, et al.
Published: (2024)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
by: Dai, Yusheng, et al.
Published: (2026)
by: Dai, Yusheng, et al.
Published: (2026)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Similar Items
-
LumosFlow: Motion-Guided Long Video Generation
by: Chen, Jiahao, et al.
Published: (2025) -
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
by: Xing, Jiazheng, et al.
Published: (2026) -
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
by: Liu, Ropeway, et al.
Published: (2025) -
AMD: Autoregressive Motion Diffusion
by: Han, Bo, et al.
Published: (2023) -
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
by: Xing, Jiazheng, et al.
Published: (2026)