MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ming, Cui, Liyuan, Zhang, Wenyuan, Zhang, Haoxian, Zhou, Yan, Li, Xiaohan, Tang, Songlin, Liu, Jiwen, Liao, Borui, Chen, Hejia, Liu, Xiaoqiang, Wan, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PEAR: Pixel-aligned Expressive humAn mesh Recovery
by: Wu, Jiahao, et al.
Published: (2026)
by: Wu, Jiahao, et al.
Published: (2026)
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
by: Chen, Hejia, et al.
Published: (2025)
by: Chen, Hejia, et al.
Published: (2025)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
by: He, Xu, et al.
Published: (2025)
by: He, Xu, et al.
Published: (2025)
Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
by: Ding, Yikang, et al.
Published: (2025)
by: Ding, Yikang, et al.
Published: (2025)
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
by: Fang, Zhixue, et al.
Published: (2026)
by: Fang, Zhixue, et al.
Published: (2026)
Semantic-Aware Prefix Learning for Token-Efficient Image Generation
by: Li, Qingfeng, et al.
Published: (2026)
by: Li, Qingfeng, et al.
Published: (2026)
GameFactory: Creating New Games with Generative Interactive Videos
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
by: Cui, Liyuan, et al.
Published: (2026)
by: Cui, Liyuan, et al.
Published: (2026)
A Survey of Interactive Generative Video
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Position: Interactive Generative Video as Next-Generation Game Engine
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Demand for catastrophe insurance under the path-dependent effects
by: Cui, Liyuan, et al.
Published: (2025)
by: Cui, Liyuan, et al.
Published: (2025)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Astra: General Interactive World Model with Autoregressive Denoising
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
by: Guo, Ying, et al.
Published: (2025)
by: Guo, Ying, et al.
Published: (2025)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
by: Zhang, Yidan, et al.
Published: (2025)
by: Zhang, Yidan, et al.
Published: (2025)
The Application of Digital Life Stories in Elderly Care: Methodological Limitations and Future Directions
by: Zilin Zhao, et al.
Published: (2025)
by: Zilin Zhao, et al.
Published: (2025)
IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation
by: Xu, Zhufeng, et al.
Published: (2026)
by: Xu, Zhufeng, et al.
Published: (2026)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
by: Li, Zisu, et al.
Published: (2025)
by: Li, Zisu, et al.
Published: (2025)
Path Choice Matters for Clear Attribution in Path Methods
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
Preventing Local Pitfalls in Vector Quantization via Optimal Transport
by: Zhang, Borui, et al.
Published: (2024)
by: Zhang, Borui, et al.
Published: (2024)
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Kling-MotionControl Technical Report
by: Kling Team, et al.
Published: (2026)
by: Kling Team, et al.
Published: (2026)
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
by: Chen, Kaijin, et al.
Published: (2026)
by: Chen, Kaijin, et al.
Published: (2026)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
by: Zhang, Shiyi, et al.
Published: (2024)
by: Zhang, Shiyi, et al.
Published: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
An approach to hummed-tune and song sequences matching
by: Pham, Loc Bao, et al.
Published: (2024)
by: Pham, Loc Bao, et al.
Published: (2024)
Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
by: Xiao, Steven, et al.
Published: (2025)
by: Xiao, Steven, et al.
Published: (2025)
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
by: Liao, Borui, et al.
Published: (2025)
by: Liao, Borui, et al.
Published: (2025)
SFTok: Bridging the Performance Gap in Discrete Tokenizers
by: Rao, Qihang, et al.
Published: (2025)
by: Rao, Qihang, et al.
Published: (2025)
Quantize-then-Rectify: Efficient VQ-VAE Training
by: Zhang, Borui, et al.
Published: (2025)
by: Zhang, Borui, et al.
Published: (2025)
Fast Shapley Value Estimation: A Unified Approach
by: Zhang, Borui, et al.
Published: (2023)
by: Zhang, Borui, et al.
Published: (2023)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
by: Liu, Yilian, et al.
Published: (2026)
by: Liu, Yilian, et al.
Published: (2026)
KlingAvatar 2.0 Technical Report
by: Kling Team, et al.
Published: (2025)
by: Kling Team, et al.
Published: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
by: Wei, Hongyang, et al.
Published: (2025)
by: Wei, Hongyang, et al.
Published: (2025)
Threshold MIDAS Forecasting of Canadian Inflation Rate
by: Chaoyi Chen, et al.
Published: (2025)
by: Chaoyi Chen, et al.
Published: (2025)
Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
by: Zhang, Juan, et al.
Published: (2024)
by: Zhang, Juan, et al.
Published: (2024)
Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
Similar Items
-
PEAR: Pixel-aligned Expressive humAn mesh Recovery
by: Wu, Jiahao, et al.
Published: (2026) -
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
by: Chen, Hejia, et al.
Published: (2025) -
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
by: Peng, Ziqiao, et al.
Published: (2025) -
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
by: He, Xu, et al.
Published: (2025) -
Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
by: Ding, Yikang, et al.
Published: (2025)