AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Mai, Ziyang, Zhang, Yuyao, Tai, Yu-Wing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
by: Zhang, Yuyao, et al.
Published: (2025)
by: Zhang, Yuyao, et al.
Published: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing
by: Zhang, Yuyao, et al.
Published: (2026)
by: Zhang, Yuyao, et al.
Published: (2026)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Global-Local Stepwise Generative Network for Ultra High-Resolution Image Restoration
by: Feng, Xin, et al.
Published: (2022)
by: Feng, Xin, et al.
Published: (2022)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRF
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
UltraGen: High-Resolution Video Generation with Hierarchical Attention
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
PixelWizard: Towards Efficient High-Fidelity Video Generation at Ultra-Large Spatial Resolution
by: Li, Wenxue, et al.
Published: (2026)
by: Li, Wenxue, et al.
Published: (2026)
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
by: Zhou, Junwei, et al.
Published: (2025)
by: Zhou, Junwei, et al.
Published: (2025)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos
by: Yu, Rongyi, et al.
Published: (2026)
by: Yu, Rongyi, et al.
Published: (2026)
SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
by: Huang-Menders, Alexander, et al.
Published: (2025)
by: Huang-Menders, Alexander, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
by: Yang, Xinran, et al.
Published: (2025)
by: Yang, Xinran, et al.
Published: (2025)
Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation
by: Xie, Ziyang, et al.
Published: (2025)
by: Xie, Ziyang, et al.
Published: (2025)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
by: Ma, Yongjia, et al.
Published: (2025)
by: Ma, Yongjia, et al.
Published: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
by: Huang, Xiaohu, et al.
Published: (2024)
by: Huang, Xiaohu, et al.
Published: (2024)
Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
by: Li, Ruibin, et al.
Published: (2026)
by: Li, Ruibin, et al.
Published: (2026)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
by: Jin, Wonjoon, et al.
Published: (2026)
by: Jin, Wonjoon, et al.
Published: (2026)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025)
by: Zhong, Yangyang, et al.
Published: (2025)
Multimodal Generation of Animatable 3D Human Models with AvatarForge
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation
by: Feng, Runyang, et al.
Published: (2025)
by: Feng, Runyang, et al.
Published: (2025)
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
by: Liu, Jianmeng, et al.
Published: (2024)
by: Liu, Jianmeng, et al.
Published: (2024)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
by: Nan, Kepan, et al.
Published: (2024)
by: Nan, Kepan, et al.
Published: (2024)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
MMG-Vid: Maximizing Marginal Gains at Segment-level and Token-level for Efficient Video LLMs
by: Ma, Junpeng, et al.
Published: (2025)
by: Ma, Junpeng, et al.
Published: (2025)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Local-Global Temporal Difference Learning for Satellite Video Super-Resolution
by: Xiao, Yi, et al.
Published: (2023)
by: Xiao, Yi, et al.
Published: (2023)
BachVid: Training-Free Video Generation with Consistent Background and Character
by: Yan, Han, et al.
Published: (2025)
by: Yan, Han, et al.
Published: (2025)
Similar Items
-
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
by: Zhang, Yuyao, et al.
Published: (2025) -
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025) -
HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing
by: Zhang, Yuyao, et al.
Published: (2026) -
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025) -
Global-Local Stepwise Generative Network for Ultra High-Resolution Image Restoration
by: Feng, Xin, et al.
Published: (2022)