Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Xu, Huang, Qiaochu, Zhang, Zhensong, Lin, Zhiwei, Wu, Zhiyong, Yang, Sicheng, Li, Minglei, Chen, Zhiyi, Xu, Songcen, Wu, Xiaofei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models
von: Xue, Haiwei, et al.
Veröffentlicht: (2023)
von: Xue, Haiwei, et al.
Veröffentlicht: (2023)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
von: Fu, Chencan, et al.
Veröffentlicht: (2024)
von: Fu, Chencan, et al.
Veröffentlicht: (2024)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
von: Yang, Sicheng, et al.
Veröffentlicht: (2024)
von: Yang, Sicheng, et al.
Veröffentlicht: (2024)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
von: Zhao, Baoquan, et al.
Veröffentlicht: (2025)
von: Zhao, Baoquan, et al.
Veröffentlicht: (2025)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
MV-Crafter: An Intelligent System for Music-guided Video Generation
von: Chen, Chuer, et al.
Veröffentlicht: (2025)
von: Chen, Chuer, et al.
Veröffentlicht: (2025)
MetaDragonBoat: Exploring Paddling Techniques of Virtual Dragon Boating in a Metaverse Campus
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
von: Wong, Kam Kwai, et al.
Veröffentlicht: (2023)
von: Wong, Kam Kwai, et al.
Veröffentlicht: (2023)
G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition
von: Deng, Kaikai, et al.
Veröffentlicht: (2024)
von: Deng, Kaikai, et al.
Veröffentlicht: (2024)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
DiffMesh: A Motion-aware Diffusion Framework for Human Mesh Recovery from Videos
von: Zheng, Ce, et al.
Veröffentlicht: (2023)
von: Zheng, Ce, et al.
Veröffentlicht: (2023)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
von: Nishida, Naoto, et al.
Veröffentlicht: (2025)
Explainable Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
ICE: Interactive 3D Game Character Editing via Dialogue
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
Language-Guided Multimodal Texture Authoring via Generative Models
von: Qian, Wanli, et al.
Veröffentlicht: (2026)
von: Qian, Wanli, et al.
Veröffentlicht: (2026)
User-Generated Content and Editors in Games: A Comprehensive Survey
von: Liu, Yuyue, et al.
Veröffentlicht: (2024)
von: Liu, Yuyue, et al.
Veröffentlicht: (2024)
Unveiling the Visual Rhetoric of Persuasive Cartography: A Case Study of the Design of Octopus Maps
von: Lin, Daocheng, et al.
Veröffentlicht: (2025)
von: Lin, Daocheng, et al.
Veröffentlicht: (2025)
SVFAP: Self-supervised Video Facial Affect Perceiver
von: Sun, Licai, et al.
Veröffentlicht: (2023)
von: Sun, Licai, et al.
Veröffentlicht: (2023)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
von: Song, Jiarun, et al.
Veröffentlicht: (2026)
von: Song, Jiarun, et al.
Veröffentlicht: (2026)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
von: Taniguchi, Tadahiro
Veröffentlicht: (2025)
von: Taniguchi, Tadahiro
Veröffentlicht: (2025)
Multimodal Infusion Tuning for Large Models
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
von: De La Fuente, Tessa
Veröffentlicht: (2025)
von: De La Fuente, Tessa
Veröffentlicht: (2025)
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
von: Kyaw, Alexander Htet, et al.
Veröffentlicht: (2025)
von: Kyaw, Alexander Htet, et al.
Veröffentlicht: (2025)
Video-Mediated Emotion Disclosure: Expressions of Fear, Sadness, and Joy by People with Schizophrenia on YouTube
von: Liu, Jiaying Lizzy, et al.
Veröffentlicht: (2025)
von: Liu, Jiaying Lizzy, et al.
Veröffentlicht: (2025)
The perceptual gap between video see-through displays and natural human vision
von: Wang, Jialin, et al.
Veröffentlicht: (2026)
von: Wang, Jialin, et al.
Veröffentlicht: (2026)
Human Aesthetic Preference-Based Large Text-to-Image Model Personalization: Kandinsky Generation as an Example
von: Zhou, Aven-Le, et al.
Veröffentlicht: (2024)
von: Zhou, Aven-Le, et al.
Veröffentlicht: (2024)
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
von: Wang, Bryan, et al.
Veröffentlicht: (2024)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
von: Xu, Yutong, et al.
Veröffentlicht: (2024)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
von: Cole, Adam, et al.
Veröffentlicht: (2026)
von: Cole, Adam, et al.
Veröffentlicht: (2026)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
von: Li, Changyang, et al.
Veröffentlicht: (2024)
von: Li, Changyang, et al.
Veröffentlicht: (2024)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
von: Lin, David Chuan-En, et al.
Veröffentlicht: (2022)
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models
von: Xue, Haiwei, et al.
Veröffentlicht: (2023) -
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
von: Fu, Chencan, et al.
Veröffentlicht: (2024) -
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
von: Yang, Sicheng, et al.
Veröffentlicht: (2024) -
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
von: Cai, Weitong, et al.
Veröffentlicht: (2026) -
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
von: Zhao, Baoquan, et al.
Veröffentlicht: (2025)