Versatile Multimodal Controls for Expressive Talking Human Animation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Zheng, Zheng, Ruobing, Wang, Yabing, Li, Tianqi, Zhu, Zixin, Zhou, Sanping, Yang, Ming, Wang, Le |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
Robust Noisy Label Learning via Two-Stream Sample Distillation
von: Bai, Sihan, et al.
Veröffentlicht: (2024)
von: Bai, Sihan, et al.
Veröffentlicht: (2024)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
von: Wang, Baiqin, et al.
Veröffentlicht: (2025)
von: Wang, Baiqin, et al.
Veröffentlicht: (2025)
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
von: Chen, Hejia, et al.
Veröffentlicht: (2025)
von: Chen, Hejia, et al.
Veröffentlicht: (2025)
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
von: Wang, Yabing, et al.
Veröffentlicht: (2025)
von: Wang, Yabing, et al.
Veröffentlicht: (2025)
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
von: Shi, Hanlei, et al.
Veröffentlicht: (2025)
von: Shi, Hanlei, et al.
Veröffentlicht: (2025)
Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
von: Liu, Kangwei, et al.
Veröffentlicht: (2025)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
von: Xie, You, et al.
Veröffentlicht: (2024)
von: Xie, You, et al.
Veröffentlicht: (2024)
Implicit Preference Alignment for Human Image Animation
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2026)
Recurrent Aligned Network for Generalized Pedestrian Trajectory Prediction
von: Dong, Yonghao, et al.
Veröffentlicht: (2024)
von: Dong, Yonghao, et al.
Veröffentlicht: (2024)
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
von: Song, Guoxian, et al.
Veröffentlicht: (2025)
von: Song, Guoxian, et al.
Veröffentlicht: (2025)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
von: Wang, Haotian, et al.
Veröffentlicht: (2024)
von: Wang, Haotian, et al.
Veröffentlicht: (2024)
Enabling Versatile Controls for Video Diffusion Models
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
von: Qu, Qiang, et al.
Veröffentlicht: (2025)
DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
von: Tian, Jingyi, et al.
Veröffentlicht: (2025)
von: Tian, Jingyi, et al.
Veröffentlicht: (2025)
LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head Synthesis
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
von: Li, Wuyang, et al.
Veröffentlicht: (2026)
von: Li, Wuyang, et al.
Veröffentlicht: (2026)
PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning
von: Li, Yizhe, et al.
Veröffentlicht: (2025)
von: Li, Yizhe, et al.
Veröffentlicht: (2025)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
von: Niu, Muyao, et al.
Veröffentlicht: (2024)
von: Niu, Muyao, et al.
Veröffentlicht: (2024)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
Embedded Representation Learning Network for Animating Styled Video Portrait
von: Wang, Tianyong, et al.
Veröffentlicht: (2024)
von: Wang, Tianyong, et al.
Veröffentlicht: (2024)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
von: Bigata, Antoni, et al.
Veröffentlicht: (2025)
von: Bigata, Antoni, et al.
Veröffentlicht: (2025)
XHand: Real-time Expressive Hand Avatar
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
von: Gan, Qijun, et al.
Veröffentlicht: (2024)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
First Multi-Dimensional Evaluation of Flowchart Comprehension for Multimodal Large Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
Multi-identity Human Image Animation with Structural Video Diffusion
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAE
von: Wu, Sichun, et al.
Veröffentlicht: (2024)
von: Wu, Sichun, et al.
Veröffentlicht: (2024)
HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
von: Zhang, Shiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiyi, et al.
Veröffentlicht: (2025)
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
von: Liu, Tao, et al.
Veröffentlicht: (2024)
von: Liu, Tao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
von: Qin, Zheng, et al.
Veröffentlicht: (2025) -
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025) -
Robust Noisy Label Learning via Two-Stream Sample Distillation
von: Bai, Sihan, et al.
Veröffentlicht: (2024) -
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
von: Wang, Baiqin, et al.
Veröffentlicht: (2025) -
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
von: Chen, Hejia, et al.
Veröffentlicht: (2025)