ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input
Fuente:
arXiv
Saved in:
| Main Authors: | Voss, Hendric, Kopp, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
by: Voss, Hendric, et al.
Published: (2025)
by: Voss, Hendric, et al.
Published: (2025)
JAX-IK: Real-Time Inverse Kinematics for Generating Multi-Constrained Movements of Virtual Human Characters
by: Voss, Hendric, et al.
Published: (2025)
by: Voss, Hendric, et al.
Published: (2025)
Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
by: Robrecht, Amelie Sophie, et al.
Published: (2024)
by: Robrecht, Amelie Sophie, et al.
Published: (2024)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
by: He, Xu, et al.
Published: (2024)
by: He, Xu, et al.
Published: (2024)
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
by: Vu, Evgeniia, et al.
Published: (2025)
by: Vu, Evgeniia, et al.
Published: (2025)
Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
by: Nagy, Rajmund, et al.
Published: (2025)
by: Nagy, Rajmund, et al.
Published: (2025)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models
by: Meinecke, Christofer, et al.
Published: (2025)
by: Meinecke, Christofer, et al.
Published: (2025)
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
by: Gao, Nan, et al.
Published: (2023)
by: Gao, Nan, et al.
Published: (2023)
Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation
by: De Simone, Zoe, et al.
Published: (2026)
by: De Simone, Zoe, et al.
Published: (2026)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
Unsupervised Domain Adaptation for RF-based Gesture Recognition
by: Zhang, Bin-Bin, et al.
Published: (2021)
by: Zhang, Bin-Bin, et al.
Published: (2021)
SynthoGestures: A Novel Framework for Synthetic Dynamic Hand Gesture Generation for Driving Scenarios
by: Gomaa, Amr, et al.
Published: (2023)
by: Gomaa, Amr, et al.
Published: (2023)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2024)
by: Garg, Mallika, et al.
Published: (2024)
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
by: Lin, Haichuan, et al.
Published: (2025)
by: Lin, Haichuan, et al.
Published: (2025)
DeepSORT-Driven Visual Tracking Approach for Gesture Recognition in Interactive Systems
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
by: Xu, Zunnan, et al.
Published: (2024)
by: Xu, Zunnan, et al.
Published: (2024)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
SASG-DA: Sparse-Aware Semantic-Guided Diffusion Augmentation For Myoelectric Gesture Recognition
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Dreamcrafter: Immersive Editing of 3D Radiance Fields Through Flexible, Generative Inputs and Outputs
by: Vachha, Cyrus, et al.
Published: (2025)
by: Vachha, Cyrus, et al.
Published: (2025)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2025)
by: Garg, Mallika, et al.
Published: (2025)
SketchPlay: Intuitive Creation of Physically Realistic VR Content with Gesture-Driven Sketching
by: Zhang, Xiangwen, et al.
Published: (2025)
by: Zhang, Xiangwen, et al.
Published: (2025)
GestFormer: Multiscale Wavelet Pooling Transformer Network for Dynamic Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2024)
by: Garg, Mallika, et al.
Published: (2024)
MIBURI: Towards Expressive Interactive Gesture Synthesis
by: Mughal, M. Hamza, et al.
Published: (2026)
by: Mughal, M. Hamza, et al.
Published: (2026)
Zero-Shot Pupil Segmentation with SAM 2: A Case Study of Over 14 Million Images
by: Maquiling, Virmarie, et al.
Published: (2024)
by: Maquiling, Virmarie, et al.
Published: (2024)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
by: Chang, Haochen, et al.
Published: (2025)
by: Chang, Haochen, et al.
Published: (2025)
Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN
by: Yusuf, Oluwaleke, et al.
Published: (2024)
by: Yusuf, Oluwaleke, et al.
Published: (2024)
Towards a GENEA Leaderboard -- an Extended, Living Benchmark for Evaluating and Advancing Conversational Motion Synthesis
by: Nagy, Rajmund, et al.
Published: (2024)
by: Nagy, Rajmund, et al.
Published: (2024)
Resource-Efficient Gesture Recognition using Low-Resolution Thermal Camera via Spiking Neural Networks and Sparse Segmentation
by: Safa, Ali, et al.
Published: (2024)
by: Safa, Ali, et al.
Published: (2024)
G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition
by: Deng, Kaikai, et al.
Published: (2024)
by: Deng, Kaikai, et al.
Published: (2024)
Steering Generative Models for Accessibility: EasyRead Image Generation
by: Dickenmann, Nicolas, et al.
Published: (2026)
by: Dickenmann, Nicolas, et al.
Published: (2026)
Zero-Shot Segmentation of Eye Features Using the Segment Anything Model (SAM)
by: Maquiling, Virmarie, et al.
Published: (2023)
by: Maquiling, Virmarie, et al.
Published: (2023)
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
by: Zhang, Xijie, et al.
Published: (2025)
by: Zhang, Xijie, et al.
Published: (2025)
Semantic and Expressive Variation in Image Captions Across Languages
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
Is Medieval Distant Viewing Possible? : Extending and Enriching Annotation of Legacy Image Collections using Visual Analytics
by: Meinecke, Christofer, et al.
Published: (2022)
by: Meinecke, Christofer, et al.
Published: (2022)
Resource-Efficient Gesture Recognition through Convexified Attention
by: Schwartz, Daniel, et al.
Published: (2026)
by: Schwartz, Daniel, et al.
Published: (2026)
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
by: Nowicki, Filip, et al.
Published: (2026)
by: Nowicki, Filip, et al.
Published: (2026)
Similar Items
-
Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
by: Voss, Hendric, et al.
Published: (2025) -
JAX-IK: Real-Time Inverse Kinematics for Generating Multi-Constrained Movements of Virtual Human Characters
by: Voss, Hendric, et al.
Published: (2025) -
Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
by: Robrecht, Amelie Sophie, et al.
Published: (2024) -
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
by: He, Xu, et al.
Published: (2024) -
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
by: Vu, Evgeniia, et al.
Published: (2025)