ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Mughal, Muhammad Hamza, Dabral, Rishabh, Habibie, Ikhsanul, Donatelli, Lucia, Habermann, Marc, Theobalt, Christian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MIBURI: Towards Expressive Interactive Gesture Synthesis
di: Mughal, M. Hamza, et al.
Pubblicazione: (2026)
di: Mughal, M. Hamza, et al.
Pubblicazione: (2026)
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
di: Mughal, M. Hamza, et al.
Pubblicazione: (2024)
di: Mughal, M. Hamza, et al.
Pubblicazione: (2024)
MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering
di: Sun, Guoxing, et al.
Pubblicazione: (2024)
di: Sun, Guoxing, et al.
Pubblicazione: (2024)
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
di: Sun, Guoxing, et al.
Pubblicazione: (2024)
di: Sun, Guoxing, et al.
Pubblicazione: (2024)
ROAM: Robust and Object-Aware Motion Generation Using Neural Pose Descriptors
di: Zhang, Wanyue, et al.
Pubblicazione: (2023)
di: Zhang, Wanyue, et al.
Pubblicazione: (2023)
PocoLoco: A Point Cloud Diffusion Model of Human Shape in Loose Clothing
di: Seth, Siddharth, et al.
Pubblicazione: (2024)
di: Seth, Siddharth, et al.
Pubblicazione: (2024)
Betsu-Betsu: Multi-View Separable 3D Reconstruction of Two Interacting Objects
di: Gopal, Suhas, et al.
Pubblicazione: (2025)
di: Gopal, Suhas, et al.
Pubblicazione: (2025)
BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects
di: Zhang, Wanyue, et al.
Pubblicazione: (2024)
di: Zhang, Wanyue, et al.
Pubblicazione: (2024)
ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions
di: Ghosh, Anindita, et al.
Pubblicazione: (2023)
di: Ghosh, Anindita, et al.
Pubblicazione: (2023)
Physics-based Human Pose Estimation from a Single Moving RGB Camera
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2025)
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2025)
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
di: Camiletto, Andrea Boscolo, et al.
Pubblicazione: (2025)
di: Camiletto, Andrea Boscolo, et al.
Pubblicazione: (2025)
PractiLight: Practical Light Control Using Foundational Diffusion Models
di: Erel, Yotam, et al.
Pubblicazione: (2025)
di: Erel, Yotam, et al.
Pubblicazione: (2025)
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
di: Wang, Jian, et al.
Pubblicazione: (2025)
di: Wang, Jian, et al.
Pubblicazione: (2025)
Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2025)
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2025)
SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens
di: Ghosh, Anindita, et al.
Pubblicazione: (2026)
di: Ghosh, Anindita, et al.
Pubblicazione: (2026)
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
di: Meric, Adil, et al.
Pubblicazione: (2026)
di: Meric, Adil, et al.
Pubblicazione: (2026)
Modeling Turn-Taking with Semantically Informed Gestures
di: Suresh, Varsha, et al.
Pubblicazione: (2025)
di: Suresh, Varsha, et al.
Pubblicazione: (2025)
Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views
di: Singh, Kunwar Maheep, et al.
Pubblicazione: (2025)
di: Singh, Kunwar Maheep, et al.
Pubblicazione: (2025)
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
di: Suresh, Varsha, et al.
Pubblicazione: (2025)
di: Suresh, Varsha, et al.
Pubblicazione: (2025)
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
di: Zhang, Wanyue, et al.
Pubblicazione: (2025)
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
di: Suresh, Varsha, et al.
Pubblicazione: (2026)
di: Suresh, Varsha, et al.
Pubblicazione: (2026)
UMA: Ultra-detailed Human Avatars via Multi-level Surface Alignment
di: Zhu, Heming, et al.
Pubblicazione: (2025)
di: Zhu, Heming, et al.
Pubblicazione: (2025)
Attention (as Discrete-Time Markov) Chains
di: Erel, Yotam, et al.
Pubblicazione: (2025)
di: Erel, Yotam, et al.
Pubblicazione: (2025)
Matrix-free Second-order Optimization of Gaussian Splats with Residual Sampling
di: Pehlivan, Hamza, et al.
Pubblicazione: (2025)
di: Pehlivan, Hamza, et al.
Pubblicazione: (2025)
EVA: Expressive Virtual Avatars from Multi-view Videos
di: Junkawitsch, Hendrik, et al.
Pubblicazione: (2025)
di: Junkawitsch, Hendrik, et al.
Pubblicazione: (2025)
A Latent Implicit 3D Shape Model for Multiple Levels of Detail
di: Guillard, Benoit, et al.
Pubblicazione: (2024)
di: Guillard, Benoit, et al.
Pubblicazione: (2024)
Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2026)
di: Aytekin, Ayce Idil, et al.
Pubblicazione: (2026)
LiveGesture Streamable Co-Speech Gesture Generation Model
di: Saleem, Muhammad Usama, et al.
Pubblicazione: (2026)
di: Saleem, Muhammad Usama, et al.
Pubblicazione: (2026)
ASH: Animatable Gaussian Splats for Efficient and Photoreal Human Rendering
di: Pang, Haokai, et al.
Pubblicazione: (2023)
di: Pang, Haokai, et al.
Pubblicazione: (2023)
Relightable Neural Actor with Intrinsic Decomposition and Pose Control
di: Luvizon, Diogo, et al.
Pubblicazione: (2023)
di: Luvizon, Diogo, et al.
Pubblicazione: (2023)
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
di: Liu, Binjie, et al.
Pubblicazione: (2025)
di: Liu, Binjie, et al.
Pubblicazione: (2025)
Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fields
di: Kairanda, Navami, et al.
Pubblicazione: (2025)
di: Kairanda, Navami, et al.
Pubblicazione: (2025)
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
di: Nagy, Rajmund, et al.
Pubblicazione: (2025)
di: Nagy, Rajmund, et al.
Pubblicazione: (2025)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
di: Ghosh, Anindita, et al.
Pubblicazione: (2025)
di: Ghosh, Anindita, et al.
Pubblicazione: (2025)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
di: Sun, Yasheng, et al.
Pubblicazione: (2025)
di: Sun, Yasheng, et al.
Pubblicazione: (2025)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
di: Fu, Chencan, et al.
Pubblicazione: (2024)
di: Fu, Chencan, et al.
Pubblicazione: (2024)
Audio-Driven Universal Gaussian Head Avatars
di: Teotia, Kartik, et al.
Pubblicazione: (2025)
di: Teotia, Kartik, et al.
Pubblicazione: (2025)
EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar Rendering
di: Chen, Jianchun, et al.
Pubblicazione: (2026)
di: Chen, Jianchun, et al.
Pubblicazione: (2026)
Holoported Characters: Real-time Free-viewpoint Rendering of Humans from Sparse RGB Cameras
di: Shetty, Ashwath, et al.
Pubblicazione: (2023)
di: Shetty, Ashwath, et al.
Pubblicazione: (2023)
TEDRA: Text-based Editing of Dynamic and Photoreal Actors
di: Sunagad, Basavaraj, et al.
Pubblicazione: (2024)
di: Sunagad, Basavaraj, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MIBURI: Towards Expressive Interactive Gesture Synthesis
di: Mughal, M. Hamza, et al.
Pubblicazione: (2026) -
Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
di: Mughal, M. Hamza, et al.
Pubblicazione: (2024) -
MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering
di: Sun, Guoxing, et al.
Pubblicazione: (2024) -
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
di: Sun, Guoxing, et al.
Pubblicazione: (2024) -
ROAM: Robust and Object-Aware Motion Generation Using Neural Pose Descriptors
di: Zhang, Wanyue, et al.
Pubblicazione: (2023)