MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Men, Yifang, Yao, Yuan, Cui, Miaomiao, Bo, Liefeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
by: Chen, Yingjie, et al.
Published: (2025)
by: Chen, Yingjie, et al.
Published: (2025)
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
by: Yang, Qiushi, et al.
Published: (2025)
by: Yang, Qiushi, et al.
Published: (2025)
Towards Fine-grained Interactive Segmentation in Images and Videos
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
by: Yang, Qiushi, et al.
Published: (2025)
by: Yang, Qiushi, et al.
Published: (2025)
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
by: Hu, Li, et al.
Published: (2023)
by: Hu, Li, et al.
Published: (2023)
En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic Data
by: Men, Yifang, et al.
Published: (2024)
by: Men, Yifang, et al.
Published: (2024)
GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior
by: Tang, Zichen, et al.
Published: (2025)
by: Tang, Zichen, et al.
Published: (2025)
ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
by: Guo, Xiefan, et al.
Published: (2025)
by: Guo, Xiefan, et al.
Published: (2025)
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation
by: Guo, Xiefan, et al.
Published: (2024)
by: Guo, Xiefan, et al.
Published: (2024)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
by: Yao, Yuan, et al.
Published: (2026)
by: Yao, Yuan, et al.
Published: (2026)
Textoon: Generating Vivid 2D Cartoon Characters from Text Descriptions
by: He, Chao, et al.
Published: (2025)
by: He, Chao, et al.
Published: (2025)
DiffuEraser: A Diffusion Model for Video Inpainting
by: Li, Xiaowen, et al.
Published: (2025)
by: Li, Xiaowen, et al.
Published: (2025)
Controllable and Expressive One-Shot Video Head Swapping
by: Ji, Chaonan, et al.
Published: (2025)
by: Ji, Chaonan, et al.
Published: (2025)
EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
by: Tian, Linrui, et al.
Published: (2024)
by: Tian, Linrui, et al.
Published: (2024)
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
by: Qiu, Di, et al.
Published: (2024)
by: Qiu, Di, et al.
Published: (2024)
Make-A-Character 2: Animatable 3D Character Generation From a Single Image
by: Liu, Lin, et al.
Published: (2025)
by: Liu, Lin, et al.
Published: (2025)
MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling
by: Yuan, Weihao, et al.
Published: (2024)
by: Yuan, Weihao, et al.
Published: (2024)
Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting
by: Liu, Hanxi, et al.
Published: (2025)
by: Liu, Hanxi, et al.
Published: (2025)
ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
by: Qi, Jinwei, et al.
Published: (2025)
by: Qi, Jinwei, et al.
Published: (2025)
Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance
by: Hu, Li, et al.
Published: (2025)
by: Hu, Li, et al.
Published: (2025)
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
by: Tian, Linrui, et al.
Published: (2025)
by: Tian, Linrui, et al.
Published: (2025)
AnyText2: Visual Text Generation and Editing With Customizable Attributes
by: Tuo, Yuxiang, et al.
Published: (2024)
by: Tuo, Yuxiang, et al.
Published: (2024)
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
by: He, Junjie, et al.
Published: (2024)
by: He, Junjie, et al.
Published: (2024)
Generative Omnimatte: Learning to Decompose Video into Layers
by: Lee, Yao-Chih, et al.
Published: (2024)
by: Lee, Yao-Chih, et al.
Published: (2024)
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
by: He, Tianyao, et al.
Published: (2025)
by: He, Tianyao, et al.
Published: (2025)
Motion In-Betweening for Densely Interacting Characters
by: Zhang, Xiaotang, et al.
Published: (2025)
by: Zhang, Xiaotang, et al.
Published: (2025)
Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
by: Zhu, Lingyu, et al.
Published: (2024)
by: Zhu, Lingyu, et al.
Published: (2024)
StdGEN: Semantic-Decomposed 3D Character Generation from Single Images
by: He, Yuze, et al.
Published: (2024)
by: He, Yuze, et al.
Published: (2024)
StdGEN++: A Comprehensive System for Semantic-Decomposed 3D Character Generation
by: He, Yuze, et al.
Published: (2026)
by: He, Yuze, et al.
Published: (2026)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
CFSynthesis: Controllable and Free-view 3D Human Video Synthesis
by: Cui, Liyuan, et al.
Published: (2024)
by: Cui, Liyuan, et al.
Published: (2024)
Exploring Timeline Control for Facial Motion Generation
by: Ma, Yifeng, et al.
Published: (2025)
by: Ma, Yifeng, et al.
Published: (2025)
MotionCharacter: Fine-Grained Motion Controllable Human Video Generation
by: Fang, Haopeng, et al.
Published: (2024)
by: Fang, Haopeng, et al.
Published: (2024)
Character Mixing for Video Generation
by: Liao, Tingting, et al.
Published: (2025)
by: Liao, Tingting, et al.
Published: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Freditor: High-Fidelity and Transferable NeRF Editing by Frequency Decomposition
by: He, Yisheng, et al.
Published: (2024)
by: He, Yisheng, et al.
Published: (2024)
The Character Error Vector: Decomposable errors for page-level OCR evaluation
by: Bourne, Jonathan, et al.
Published: (2026)
by: Bourne, Jonathan, et al.
Published: (2026)
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
AdaptiveDrag: Semantic-Driven Dragging on Diffusion-Based Image Editing
by: Chen, DuoSheng, et al.
Published: (2024)
by: Chen, DuoSheng, et al.
Published: (2024)
Similar Items
-
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
by: Chen, Yingjie, et al.
Published: (2025) -
MoSAM: Motion-Guided Segment Anything Model with Spatial-Temporal Memory Selection
by: Yang, Qiushi, et al.
Published: (2025) -
Towards Fine-grained Interactive Segmentation in Images and Videos
by: Yao, Yuan, et al.
Published: (2025) -
McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
by: Yang, Qiushi, et al.
Published: (2025) -
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
by: Hu, Li, et al.
Published: (2023)