Real-Time Motion-Controllable Autoregressive Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Kesen, Shi, Jiaxin, Zhu, Beier, Zhou, Junbao, Shen, Xiaolong, Zhou, Yuan, Sun, Qianru, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
by: Zhao, Kesen, et al.
Published: (2025)
by: Zhao, Kesen, et al.
Published: (2025)
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
by: Zhou, Junbao, et al.
Published: (2025)
by: Zhou, Junbao, et al.
Published: (2025)
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
by: Zhao, Kesen, et al.
Published: (2026)
by: Zhao, Kesen, et al.
Published: (2026)
Reducing Class-Wise Performance Disparity via Margin Regularization
by: Zhu, Beier, et al.
Published: (2026)
by: Zhu, Beier, et al.
Published: (2026)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)
by: Lyu, Hengye, et al.
Published: (2026)
Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models
by: Zhu, Beier, et al.
Published: (2023)
by: Zhu, Beier, et al.
Published: (2023)
DragNeXt: Rethinking Drag-Based Image Editing
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Few-shot Learner Parameterization by Diffusion Time-steps
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
by: Li, Zisu, et al.
Published: (2025)
by: Li, Zisu, et al.
Published: (2025)
Hierarchical Semantic Alignment for Image Clustering
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models
by: Zhu, Beier, et al.
Published: (2025)
by: Zhu, Beier, et al.
Published: (2025)
Robust Fine-tuning of Zero-shot Models via Variance Reduction
by: Zhu, Beier, et al.
Published: (2024)
by: Zhu, Beier, et al.
Published: (2024)
DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control
by: Zhao, Kaifeng, et al.
Published: (2024)
by: Zhao, Kaifeng, et al.
Published: (2024)
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
by: Zhou, Yuan, et al.
Published: (2024)
by: Zhou, Yuan, et al.
Published: (2024)
Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-Of-Distribution Generalization
by: Qi, Jiaxin, et al.
Published: (2022)
by: Qi, Jiaxin, et al.
Published: (2022)
Exploring Diffusion Time-steps for Unsupervised Representation Learning
by: Yue, Zhongqi, et al.
Published: (2024)
by: Yue, Zhongqi, et al.
Published: (2024)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
NeuSpring: Neural Spring Fields for Reconstruction and Simulation of Deformable Objects from Videos
by: Xu, Qingshan, et al.
Published: (2025)
by: Xu, Qingshan, et al.
Published: (2025)
Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness
by: Zhu, Beier, et al.
Published: (2025)
by: Zhu, Beier, et al.
Published: (2025)
MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering
by: Zhou, Yuan, et al.
Published: (2026)
by: Zhou, Yuan, et al.
Published: (2026)
Generalized Visual Relation Detection with Diffusion Models
by: Gao, Kaifeng, et al.
Published: (2025)
by: Gao, Kaifeng, et al.
Published: (2025)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
Parallel Diffusion Solver via Residual Dirichlet Policy Optimization
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
Adapting VACE for Real-Time Autoregressive Video Diffusion
by: Fosdick, Ryan
Published: (2026)
by: Fosdick, Ryan
Published: (2026)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
by: Shin, Joonghyuk, et al.
Published: (2025)
by: Shin, Joonghyuk, et al.
Published: (2025)
Prompt-aligned Gradient for Prompt Tuning
by: Zhu, Beier, et al.
Published: (2022)
by: Zhu, Beier, et al.
Published: (2022)
Debiased Fine-Tuning for Vision-language Models by Prompt Regularization
by: Zhu, Beier, et al.
Published: (2023)
by: Zhu, Beier, et al.
Published: (2023)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
by: Meng, Yihao, et al.
Published: (2026)
by: Meng, Yihao, et al.
Published: (2026)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
by: Zhou, Zhenghong, et al.
Published: (2026)
by: Zhou, Zhenghong, et al.
Published: (2026)
Non-confusing Generation of Customized Concepts in Diffusion Models
by: Lin, Wang, et al.
Published: (2024)
by: Lin, Wang, et al.
Published: (2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control
by: Xu, Youcan, et al.
Published: (2026)
by: Xu, Youcan, et al.
Published: (2026)
Similar Items
-
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
by: Zhao, Kesen, et al.
Published: (2025) -
Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
by: Zhou, Junbao, et al.
Published: (2025) -
Thinking with Images as Continuous Actions: Numerical Visual Chain-of-Thought
by: Zhao, Kesen, et al.
Published: (2026) -
Reducing Class-Wise Performance Disparity via Margin Regularization
by: Zhu, Beier, et al.
Published: (2026) -
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)