Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Xunzhi, Chen, Yabo, Zhang, Guiyu, Wang, Zhongyu, Gao, Zhe, Xiang, Quanming, Shang, Gonghu, Liu, Junqi, Huang, Haibin, Gao, Yang, Zhang, Chi, Fan, Qi, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026)
by: Xiang, Xunzhi, et al.
Published: (2026)
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
by: Zhang, Guiyu, et al.
Published: (2026)
by: Zhang, Guiyu, et al.
Published: (2026)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
by: Zhang, Guiyu, et al.
Published: (2025)
by: Zhang, Guiyu, et al.
Published: (2025)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
by: Ma, Junyuan, et al.
Published: (2026)
by: Ma, Junyuan, et al.
Published: (2026)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
TeleStyle: Content-Preserving Style Transfer in Images and Videos
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
by: Gao, Xiangxiang, et al.
Published: (2024)
by: Gao, Xiangxiang, et al.
Published: (2024)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
by: Tian, Jiahao, et al.
Published: (2026)
by: Tian, Jiahao, et al.
Published: (2026)
Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation
by: Zhang, Yunyao, et al.
Published: (2026)
by: Zhang, Yunyao, et al.
Published: (2026)
Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image
by: Wen, Tao, et al.
Published: (2025)
by: Wen, Tao, et al.
Published: (2025)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
by: Cheng, Tianle, et al.
Published: (2025)
by: Cheng, Tianle, et al.
Published: (2025)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
by: Gao, Zhe, et al.
Published: (2026)
by: Gao, Zhe, et al.
Published: (2026)
Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
by: Feng, Yujie, et al.
Published: (2026)
by: Feng, Yujie, et al.
Published: (2026)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
by: Gao, Shouwei, et al.
Published: (2026)
by: Gao, Shouwei, et al.
Published: (2026)
ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents
by: Guo, Shuhan, et al.
Published: (2026)
by: Guo, Shuhan, et al.
Published: (2026)
Point2Insert: Video Object Insertion via Sparse Point Guidance
by: Zhou, Yu, et al.
Published: (2026)
by: Zhou, Yu, et al.
Published: (2026)
Geometry-Aware Implicit Memory for Video World Models
by: Wei, Zhengxuan, et al.
Published: (2026)
by: Wei, Zhengxuan, et al.
Published: (2026)
Micro, Macro & Mezzo Geoinformation
Published: (2020)
Published: (2020)
Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
by: Zhou, Yufan, et al.
Published: (2025)
by: Zhou, Yufan, et al.
Published: (2025)
An Efficient Memory Module for Graph Few-Shot Class-Incremental Learning
by: Li, Dong, et al.
Published: (2024)
by: Li, Dong, et al.
Published: (2024)
Exploring Micro Accidents and Driver Responses in Automated Driving: Insights from Real-world Videos
by: Xiang, Wei, et al.
Published: (2025)
by: Xiang, Wei, et al.
Published: (2025)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
On the fractional parts of certain sequences of $ξα^{n}$
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
by: Zhang, Zhuoyang, et al.
Published: (2025)
by: Zhang, Zhuoyang, et al.
Published: (2025)
LIVE: Long-horizon Interactive Video World Modeling
by: Huang, Junchao, et al.
Published: (2026)
by: Huang, Junchao, et al.
Published: (2026)
SMR: State Memory Replay for Long Sequence Modeling
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
by: Chen, Yabo, et al.
Published: (2025)
by: Chen, Yabo, et al.
Published: (2025)
NFIG: Multi-Scale Autoregressive Image Generation via Frequency Ordering
by: Huang, Zhihao, et al.
Published: (2025)
by: Huang, Zhihao, et al.
Published: (2025)
HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Evolution of Thought: Diverse and High-Quality Reasoning via Multi-Objective Optimization
by: Qi, Biqing, et al.
Published: (2024)
by: Qi, Biqing, et al.
Published: (2024)
SRAGAN: Saliency Regularized and Attended Generative Adversarial Network for Chinese Ink-wash Painting Style Transfer
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
Micro-, Meso- and Macro-Connectomics of the Brain
Published: (2018)
Published: (2018)
Robust Micro-Macro Entangled States
by: Mirkamali, Maryam Sadat, et al.
Published: (2024)
by: Mirkamali, Maryam Sadat, et al.
Published: (2024)
Micro-, Meso- and Macro-Dynamics of the Brain
Published: (2018)
Published: (2018)
Similar Items
-
Pathwise Test-Time Correction for Autoregressive Long Video Generation
by: Xiang, Xunzhi, et al.
Published: (2026) -
SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
by: Zhang, Guiyu, et al.
Published: (2026) -
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025) -
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025) -
Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
by: Zhang, Guiyu, et al.
Published: (2025)