MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Chencan, Wang, Yabiao, Zhang, Jiangning, Jiang, Zhengkai, Mao, Xiaofeng, Wu, Jiafu, Cao, Weijian, Wang, Chengjie, Ge, Yanhao, Liu, Yong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
por: Mao, Xiaofeng, et al.
Publicado: (2024)
por: Mao, Xiaofeng, et al.
Publicado: (2024)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
por: Zhang, Fan, et al.
Publicado: (2024)
por: Zhang, Fan, et al.
Publicado: (2024)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024)
por: He, Xu, et al.
Publicado: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
por: Wang, Sen, et al.
Publicado: (2024)
por: Wang, Sen, et al.
Publicado: (2024)
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
por: Yang, Quanwei, et al.
Publicado: (2025)
por: Yang, Quanwei, et al.
Publicado: (2025)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
por: Cheng, Hongye, et al.
Publicado: (2025)
por: Cheng, Hongye, et al.
Publicado: (2025)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
por: Kim, Yonghyun, et al.
Publicado: (2025)
por: Kim, Yonghyun, et al.
Publicado: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
por: Yang, Sicheng, et al.
Publicado: (2024)
por: Yang, Sicheng, et al.
Publicado: (2024)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
por: Qi, Xingqun, et al.
Publicado: (2023)
por: Qi, Xingqun, et al.
Publicado: (2023)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
por: Zhu, Yixin, et al.
Publicado: (2026)
por: Zhu, Yixin, et al.
Publicado: (2026)
G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition
por: Deng, Kaikai, et al.
Publicado: (2024)
por: Deng, Kaikai, et al.
Publicado: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
por: Zhu, Haodong, et al.
Publicado: (2025)
por: Zhu, Haodong, et al.
Publicado: (2025)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
por: Zhao, Baoquan, et al.
Publicado: (2025)
por: Zhao, Baoquan, et al.
Publicado: (2025)
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
por: Tang, Jinghua, et al.
Publicado: (2024)
por: Tang, Jinghua, et al.
Publicado: (2024)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
por: Wong, Kam Kwai, et al.
Publicado: (2023)
por: Wong, Kam Kwai, et al.
Publicado: (2023)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
por: Cai, Zhuojiang, et al.
Publicado: (2024)
por: Cai, Zhuojiang, et al.
Publicado: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
por: Shi, Weiyan, et al.
Publicado: (2026)
por: Shi, Weiyan, et al.
Publicado: (2026)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
por: Chen, Yuheng, et al.
Publicado: (2026)
por: Chen, Yuheng, et al.
Publicado: (2026)
Rethinking Fusion: Disentangled Learning of Shared and Modality-Specific Information for Stance Detection
por: Xie, Zhiyu, et al.
Publicado: (2026)
por: Xie, Zhiyu, et al.
Publicado: (2026)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
por: Guo, Zixin, et al.
Publicado: (2024)
por: Guo, Zixin, et al.
Publicado: (2024)
Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture Generation
por: Gao, Zhilin, et al.
Publicado: (2025)
por: Gao, Zhilin, et al.
Publicado: (2025)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
por: Ning, Zheng, et al.
Publicado: (2024)
por: Ning, Zheng, et al.
Publicado: (2024)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
por: Ghaleb, Esam, et al.
Publicado: (2025)
por: Ghaleb, Esam, et al.
Publicado: (2025)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
por: Wang, Xianghan
Publicado: (2025)
por: Wang, Xianghan
Publicado: (2025)
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
por: Kumar, Lokesh, et al.
Publicado: (2026)
por: Kumar, Lokesh, et al.
Publicado: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025)
por: Shi, Weiyan, et al.
Publicado: (2025)
Deep Mamba Multi-modal Learning
por: Zhu, Jian, et al.
Publicado: (2024)
por: Zhu, Jian, et al.
Publicado: (2024)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
por: Guo, Xin, et al.
Publicado: (2025)
por: Guo, Xin, et al.
Publicado: (2025)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
por: Song, Jiarun, et al.
Publicado: (2026)
por: Song, Jiarun, et al.
Publicado: (2026)
MetaDragonBoat: Exploring Paddling Techniques of Virtual Dragon Boating in a Metaverse Campus
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
Promisedland: An XR Narrative Attraction Integrating Diorama-to-Virtual Workflow and Elemental Storytelling
por: Wang, Xianghan, et al.
Publicado: (2025)
por: Wang, Xianghan, et al.
Publicado: (2025)
Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
por: Voss, Hendric, et al.
Publicado: (2025)
por: Voss, Hendric, et al.
Publicado: (2025)
MV-Crafter: An Intelligent System for Music-guided Video Generation
por: Chen, Chuer, et al.
Publicado: (2025)
por: Chen, Chuer, et al.
Publicado: (2025)
Unveiling the Visual Rhetoric of Persuasive Cartography: A Case Study of the Design of Octopus Maps
por: Lin, Daocheng, et al.
Publicado: (2025)
por: Lin, Daocheng, et al.
Publicado: (2025)
Ejemplares similares
-
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
por: Mao, Xiaofeng, et al.
Publicado: (2024) -
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
por: Zhang, Fan, et al.
Publicado: (2024) -
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024) -
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
por: Wang, Sen, et al.
Publicado: (2024) -
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)