Active Intelligence in Video Avatars via Closed-loop World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xuanhua, Yang, Tianyu, Cao, Ke, Wu, Ruiqi, Meng, Cheng, Zhang, Yong, Kang, Zhuoliang, Wei, Xiaoming, Chen, Qifeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
by: Wu, Ruiqi, et al.
Published: (2026)
by: Wu, Ruiqi, et al.
Published: (2026)
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026)
by: Guo, Qin, et al.
Published: (2026)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
LongCat-Video-Avatar 1.5 Technical Report
by: Meituan LongCat Team, et al.
Published: (2026)
by: Meituan LongCat Team, et al.
Published: (2026)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026)
by: Deng, Xiang, et al.
Published: (2026)
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
by: Yu, Haojie, et al.
Published: (2025)
by: Yu, Haojie, et al.
Published: (2025)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation
by: Rao, Zhefan, et al.
Published: (2026)
by: Rao, Zhefan, et al.
Published: (2026)
ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment
by: Chen, Yiyang, et al.
Published: (2025)
by: Chen, Yiyang, et al.
Published: (2025)
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
by: Luo, Xianrui, et al.
Published: (2025)
by: Luo, Xianrui, et al.
Published: (2025)
Hawk: Learning to Understand Open-World Video Anomalies
by: Tang, Jiaqi, et al.
Published: (2024)
by: Tang, Jiaqi, et al.
Published: (2024)
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior
by: Guo, Chen, et al.
Published: (2025)
by: Guo, Chen, et al.
Published: (2025)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
by: He, Xuanhua, et al.
Published: (2024)
by: He, Xuanhua, et al.
Published: (2024)
LongCat-Video Technical Report
by: Meituan LongCat Team, et al.
Published: (2025)
by: Meituan LongCat Team, et al.
Published: (2025)
UNIC: Unified In-Context Video Editing
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
GameGen-X: Interactive Open-world Game Video Generation
by: Che, Haoxuan, et al.
Published: (2024)
by: Che, Haoxuan, et al.
Published: (2024)
AvatarPointillist: AutoRegressive 4D Gaussian Avatarization
by: Liu, Hongyu, et al.
Published: (2026)
by: Liu, Hongyu, et al.
Published: (2026)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
by: Wang, Xinjie, et al.
Published: (2025)
by: Wang, Xinjie, et al.
Published: (2025)
Fast and Physically-based Neural Explicit Surface for Relightable Human Avatars
by: Wu, Jiacheng, et al.
Published: (2025)
by: Wu, Jiacheng, et al.
Published: (2025)
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
by: Yan, Tianyi, et al.
Published: (2024)
by: Yan, Tianyi, et al.
Published: (2024)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026)
by: Zheng, Mingzhe, et al.
Published: (2026)
HR Human: Modeling Human Avatars with Triangular Mesh and High-Resolution Textures from Videos
by: Chen, Qifeng, et al.
Published: (2024)
by: Chen, Qifeng, et al.
Published: (2024)
AvatarArtist: Open-Domain 4D Avatarization
by: Liu, Hongyu, et al.
Published: (2025)
by: Liu, Hongyu, et al.
Published: (2025)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion
by: Cao, Ke, et al.
Published: (2024)
by: Cao, Ke, et al.
Published: (2024)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
by: Lu, Feichi, et al.
Published: (2024)
by: Lu, Feichi, et al.
Published: (2024)
AvatarVTON: 4D Virtual Try-On for Animatable Avatars
by: Jiang, Zicheng, et al.
Published: (2025)
by: Jiang, Zicheng, et al.
Published: (2025)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Refined Geometry-guided Head Avatar Reconstruction from Monocular RGB Video
by: Park, Pilseo, et al.
Published: (2025)
by: Park, Pilseo, et al.
Published: (2025)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
SelfieAvatar: Real-time Head Avatar reenactment from a Selfie Video
by: Liang, Wei, et al.
Published: (2026)
by: Liang, Wei, et al.
Published: (2026)
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024)
by: Tan, Shuhan, et al.
Published: (2024)
Deterministic World Models for Verification of Closed-loop Vision-based Systems
by: Geng, Yuang, et al.
Published: (2025)
by: Geng, Yuang, et al.
Published: (2025)
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Self-supervised Multiplex Consensus Mamba for General Image Fusion
by: Wang, Yingying, et al.
Published: (2025)
by: Wang, Yingying, et al.
Published: (2025)
Distilling Textual Priors from LLM to Efficient Image Fusion
by: Zhang, Ran, et al.
Published: (2025)
by: Zhang, Ran, et al.
Published: (2025)
Similar Items
-
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
by: Wu, Ruiqi, et al.
Published: (2026) -
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026) -
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025) -
LongCat-Video-Avatar 1.5 Technical Report
by: Meituan LongCat Team, et al.
Published: (2026) -
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025)