EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xiangyue, Li, Jianfang, Zhang, Jiaxu, Ren, Jianqiang, Bo, Liefeng, Tu, Zhigang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
di: Cheng, Shihao, et al.
Pubblicazione: (2026)
di: Cheng, Shihao, et al.
Pubblicazione: (2026)
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis
di: Zhang, Xiangyue, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyue, et al.
Pubblicazione: (2024)
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
di: Zhang, Xiangyue, et al.
Pubblicazione: (2025)
di: Zhang, Xiangyue, et al.
Pubblicazione: (2025)
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2026)
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2026)
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
di: Liu, Binjie, et al.
Pubblicazione: (2025)
di: Liu, Binjie, et al.
Pubblicazione: (2025)
Generative Motion Stylization of Cross-structure Characters within Canonical Motion Space
di: Zhang, Jiaxu, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxu, et al.
Pubblicazione: (2024)
Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
di: Zhang, Zeyi, et al.
Pubblicazione: (2024)
di: Zhang, Zeyi, et al.
Pubblicazione: (2024)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
di: Guo, Xin, et al.
Pubblicazione: (2025)
di: Guo, Xin, et al.
Pubblicazione: (2025)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
di: Ghosh, Anindita, et al.
Pubblicazione: (2025)
di: Ghosh, Anindita, et al.
Pubblicazione: (2025)
DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
di: Chen, Junming, et al.
Pubblicazione: (2024)
di: Chen, Junming, et al.
Pubblicazione: (2024)
It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model
di: Shi, Mingyi, et al.
Pubblicazione: (2024)
di: Shi, Mingyi, et al.
Pubblicazione: (2024)
DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech
di: Cheng, Yongkang, et al.
Pubblicazione: (2025)
di: Cheng, Yongkang, et al.
Pubblicazione: (2025)
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
di: Huang, Mingyang, et al.
Pubblicazione: (2025)
di: Huang, Mingyang, et al.
Pubblicazione: (2025)
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
di: Gan, Yuan, et al.
Pubblicazione: (2025)
di: Gan, Yuan, et al.
Pubblicazione: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
di: Razlighi, AmirHossein Naghi, et al.
Pubblicazione: (2026)
di: Razlighi, AmirHossein Naghi, et al.
Pubblicazione: (2026)
DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
di: Zhang, Hengyuan, et al.
Pubblicazione: (2025)
A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
di: Airale, Louis, et al.
Pubblicazione: (2023)
di: Airale, Louis, et al.
Pubblicazione: (2023)
GaussianSpeech: Audio-Driven Gaussian Avatars
di: Aneja, Shivangi, et al.
Pubblicazione: (2024)
di: Aneja, Shivangi, et al.
Pubblicazione: (2024)
ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
di: Wang, Xuanchen, et al.
Pubblicazione: (2025)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
di: Gupta, Prerit, et al.
Pubblicazione: (2025)
di: Gupta, Prerit, et al.
Pubblicazione: (2025)
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
di: Dahan, Aviad, et al.
Pubblicazione: (2026)
di: Dahan, Aviad, et al.
Pubblicazione: (2026)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
DiM-Gesture: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2 framework
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
MaskAdapt: Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters
di: Park, Soomin, et al.
Pubblicazione: (2026)
di: Park, Soomin, et al.
Pubblicazione: (2026)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
UniMuMo: Unified Text, Music and Motion Generation
di: Yang, Han, et al.
Pubblicazione: (2024)
di: Yang, Han, et al.
Pubblicazione: (2024)
Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers
di: Anisetty, Sohan, et al.
Pubblicazione: (2024)
di: Anisetty, Sohan, et al.
Pubblicazione: (2024)
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
di: Huang, Jinhe, et al.
Pubblicazione: (2025)
A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis
di: Mascaro, Esteve Valls, et al.
Pubblicazione: (2023)
di: Mascaro, Esteve Valls, et al.
Pubblicazione: (2023)
Click2Mask: Local Editing with Dynamic Mask Generation
di: Regev, Omer, et al.
Pubblicazione: (2024)
di: Regev, Omer, et al.
Pubblicazione: (2024)
GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling
di: Liu, Pinxin, et al.
Pubblicazione: (2025)
di: Liu, Pinxin, et al.
Pubblicazione: (2025)
MATT-GS: Masked Attention-based 3DGS for Robot Perception and Object Detection
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
di: Park, Sangjune, et al.
Pubblicazione: (2026)
di: Park, Sangjune, et al.
Pubblicazione: (2026)
InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions
di: Li, Ronghui, et al.
Pubblicazione: (2024)
di: Li, Ronghui, et al.
Pubblicazione: (2024)
Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation
di: Zhou, Zhizhen, et al.
Pubblicazione: (2024)
di: Zhou, Zhizhen, et al.
Pubblicazione: (2024)
MIDGET: Music Conditioned 3D Dance Generation
di: Wang, Jinwu, et al.
Pubblicazione: (2024)
di: Wang, Jinwu, et al.
Pubblicazione: (2024)
Masked Extended Attention for Zero-Shot Virtual Try-On In The Wild
di: Orzech, Nadav, et al.
Pubblicazione: (2024)
di: Orzech, Nadav, et al.
Pubblicazione: (2024)
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
di: Petermann, Darius, et al.
Pubblicazione: (2025)
di: Petermann, Darius, et al.
Pubblicazione: (2025)
MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
di: Zhou, Hao, et al.
Pubblicazione: (2025)
di: Zhou, Hao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
di: Cheng, Shihao, et al.
Pubblicazione: (2026) -
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis
di: Zhang, Xiangyue, et al.
Pubblicazione: (2024) -
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
di: Zhang, Xiangyue, et al.
Pubblicazione: (2025) -
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2026) -
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
di: Liu, Binjie, et al.
Pubblicazione: (2025)