Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhattacharya, Uttaran, Childs, Elizabeth, Rewkowski, Nicholas, Manocha, Dinesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
por: Bhattacharya, Uttaran, et al.
Publicado: (2024)
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
por: Bhattacharya, Uttaran, et al.
Publicado: (2021)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
por: Fu, Chencan, et al.
Publicado: (2024)
por: Fu, Chencan, et al.
Publicado: (2024)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
por: Qi, Xingqun, et al.
Publicado: (2023)
por: Qi, Xingqun, et al.
Publicado: (2023)
SyMuPe: Affective and Controllable Symbolic Music Performance
por: Borovik, Ilya, et al.
Publicado: (2025)
por: Borovik, Ilya, et al.
Publicado: (2025)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
por: Cheng, Hongye, et al.
Publicado: (2025)
por: Cheng, Hongye, et al.
Publicado: (2025)
3DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action Control
por: Sha, Xuanmeng, et al.
Publicado: (2026)
por: Sha, Xuanmeng, et al.
Publicado: (2026)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024)
por: He, Xu, et al.
Publicado: (2024)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
por: Guo, Xin, et al.
Publicado: (2025)
por: Guo, Xin, et al.
Publicado: (2025)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
por: Guo, Zixin, et al.
Publicado: (2024)
por: Guo, Zixin, et al.
Publicado: (2024)
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
por: Kumar, Lokesh, et al.
Publicado: (2026)
por: Kumar, Lokesh, et al.
Publicado: (2026)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes
por: Ratnarajah, Anton, et al.
Publicado: (2023)
por: Ratnarajah, Anton, et al.
Publicado: (2023)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
por: Yang, Quanwei, et al.
Publicado: (2025)
por: Yang, Quanwei, et al.
Publicado: (2025)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
por: Bhattacharya, Aneesh, et al.
Publicado: (2023)
por: Bhattacharya, Aneesh, et al.
Publicado: (2023)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
por: Zhang, Fan, et al.
Publicado: (2024)
por: Zhang, Fan, et al.
Publicado: (2024)
Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
por: Bhattacharya, Uttaran, et al.
Publicado: (2019)
por: Bhattacharya, Uttaran, et al.
Publicado: (2019)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
por: Lin, Ronghao, et al.
Publicado: (2024)
por: Lin, Ronghao, et al.
Publicado: (2024)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
por: Jeyaraj, Rathinaraja, et al.
Publicado: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
por: Ghaleb, Esam, et al.
Publicado: (2025)
por: Ghaleb, Esam, et al.
Publicado: (2025)
Sec2Sec Co-attention for Video-Based Apparent Affective Prediction
por: Sun, Mingwei, et al.
Publicado: (2024)
por: Sun, Mingwei, et al.
Publicado: (2024)
Improving Generative Adversarial Network Generalization for Facial Expression Synthesis
por: Akram, Arbish, et al.
Publicado: (2026)
por: Akram, Arbish, et al.
Publicado: (2026)
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
por: Jain, Arnav, et al.
Publicado: (2024)
por: Jain, Arnav, et al.
Publicado: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2024)
por: Lou, Haowei, et al.
Publicado: (2024)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
por: Kim, Sungnyun, et al.
Publicado: (2025)
por: Kim, Sungnyun, et al.
Publicado: (2025)
Multimodal Speech Enhancement Using Burst Propagation
por: Raza, Mohsin, et al.
Publicado: (2022)
por: Raza, Mohsin, et al.
Publicado: (2022)
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
por: Wang, Wupeng, et al.
Publicado: (2024)
por: Wang, Wupeng, et al.
Publicado: (2024)
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
por: Zheng, Haolong, et al.
Publicado: (2025)
por: Zheng, Haolong, et al.
Publicado: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
por: Lin, Meng-Ping, et al.
Publicado: (2025)
por: Lin, Meng-Ping, et al.
Publicado: (2025)
Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture Generation
por: Gao, Zhilin, et al.
Publicado: (2025)
por: Gao, Zhilin, et al.
Publicado: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
por: Yang, Sicheng, et al.
Publicado: (2024)
por: Yang, Sicheng, et al.
Publicado: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
por: Wu, Linzhi, et al.
Publicado: (2026)
por: Wu, Linzhi, et al.
Publicado: (2026)
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
por: Song, Zeyang, et al.
Publicado: (2025)
por: Song, Zeyang, et al.
Publicado: (2025)
Stemphonic: All-at-once Flexible Multi-stem Music Generation
por: Wu, Shih-Lun, et al.
Publicado: (2026)
por: Wu, Shih-Lun, et al.
Publicado: (2026)
AttentionStitch: How Attention Solves the Speech Editing Problem
por: Alexos, Antonios, et al.
Publicado: (2024)
por: Alexos, Antonios, et al.
Publicado: (2024)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
por: Munyer, Travis, et al.
Publicado: (2023)
por: Munyer, Travis, et al.
Publicado: (2023)
Adversarially Robust Deepfake Detection via Adversarial Feature Similarity Learning
por: Khan, Sarwar
Publicado: (2024)
por: Khan, Sarwar
Publicado: (2024)
Sequence-to-Sequence Multi-Modal Speech In-Painting
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
por: Elyaderani, Mahsa Kadkhodaei, et al.
Publicado: (2024)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
por: Wu, Wenxuan, et al.
Publicado: (2025)
por: Wu, Wenxuan, et al.
Publicado: (2025)
Ejemplares similares
-
Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
por: Bhattacharya, Uttaran, et al.
Publicado: (2024) -
Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
por: Bhattacharya, Uttaran, et al.
Publicado: (2021) -
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025) -
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
por: Fu, Chencan, et al.
Publicado: (2024) -
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
por: Qi, Xingqun, et al.
Publicado: (2023)