Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Mughal, M. Hamza, Dabral, Rishabh, Scholman, Merel C. J., Demberg, Vera, Theobalt, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIBURI: Towards Expressive Interactive Gesture Synthesis
by: Mughal, M. Hamza, et al.
Published: (2026)
by: Mughal, M. Hamza, et al.
Published: (2026)
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
by: Mughal, Muhammad Hamza, et al.
Published: (2024)
by: Mughal, Muhammad Hamza, et al.
Published: (2024)
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions
by: Ghosh, Anindita, et al.
Published: (2023)
by: Ghosh, Anindita, et al.
Published: (2023)
Betsu-Betsu: Multi-View Separable 3D Reconstruction of Two Interacting Objects
by: Gopal, Suhas, et al.
Published: (2025)
by: Gopal, Suhas, et al.
Published: (2025)
Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
by: Aytekin, Ayce Idil, et al.
Published: (2025)
by: Aytekin, Ayce Idil, et al.
Published: (2025)
MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering
by: Sun, Guoxing, et al.
Published: (2024)
by: Sun, Guoxing, et al.
Published: (2024)
SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens
by: Ghosh, Anindita, et al.
Published: (2026)
by: Ghosh, Anindita, et al.
Published: (2026)
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
PractiLight: Practical Light Control Using Foundational Diffusion Models
by: Erel, Yotam, et al.
Published: (2025)
by: Erel, Yotam, et al.
Published: (2025)
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
by: Sun, Guoxing, et al.
Published: (2024)
by: Sun, Guoxing, et al.
Published: (2024)
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
ROAM: Robust and Object-Aware Motion Generation Using Neural Pose Descriptors
by: Zhang, Wanyue, et al.
Published: (2023)
by: Zhang, Wanyue, et al.
Published: (2023)
Attention (as Discrete-Time Markov) Chains
by: Erel, Yotam, et al.
Published: (2025)
by: Erel, Yotam, et al.
Published: (2025)
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
by: Meric, Adil, et al.
Published: (2026)
by: Meric, Adil, et al.
Published: (2026)
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video
by: Camiletto, Andrea Boscolo, et al.
Published: (2025)
by: Camiletto, Andrea Boscolo, et al.
Published: (2025)
Physics-based Human Pose Estimation from a Single Moving RGB Camera
by: Aytekin, Ayce Idil, et al.
Published: (2025)
by: Aytekin, Ayce Idil, et al.
Published: (2025)
BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects
by: Zhang, Wanyue, et al.
Published: (2024)
by: Zhang, Wanyue, et al.
Published: (2024)
PocoLoco: A Point Cloud Diffusion Model of Human Shape in Loose Clothing
by: Seth, Siddharth, et al.
Published: (2024)
by: Seth, Siddharth, et al.
Published: (2024)
Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions
by: Aytekin, Ayce Idil, et al.
Published: (2026)
by: Aytekin, Ayce Idil, et al.
Published: (2026)
Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views
by: Singh, Kunwar Maheep, et al.
Published: (2025)
by: Singh, Kunwar Maheep, et al.
Published: (2025)
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
by: Delitzas, Alexandros, et al.
Published: (2026)
by: Delitzas, Alexandros, et al.
Published: (2026)
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
by: Nagy, Rajmund, et al.
Published: (2025)
by: Nagy, Rajmund, et al.
Published: (2025)
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
by: Wang, Wenjia, et al.
Published: (2026)
by: Wang, Wenjia, et al.
Published: (2026)
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
by: Ghosh, Anindita, et al.
Published: (2025)
by: Ghosh, Anindita, et al.
Published: (2025)
Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels
by: Dünkel, Olaf, et al.
Published: (2025)
by: Dünkel, Olaf, et al.
Published: (2025)
Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
by: Suresh, Varsha, et al.
Published: (2026)
by: Suresh, Varsha, et al.
Published: (2026)
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
by: Boudiaf, Abderrahmene, et al.
Published: (2026)
by: Boudiaf, Abderrahmene, et al.
Published: (2026)
Prompting Implicit Discourse Relation Annotation
by: Yung, Frances, et al.
Published: (2024)
by: Yung, Frances, et al.
Published: (2024)
Matrix-free Second-order Optimization of Gaussian Splats with Residual Sampling
by: Pehlivan, Hamza, et al.
Published: (2025)
by: Pehlivan, Hamza, et al.
Published: (2025)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
by: Dünkel, Olaf, et al.
Published: (2026)
by: Dünkel, Olaf, et al.
Published: (2026)
TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
by: Wang, Yiyao, et al.
Published: (2026)
by: Wang, Yiyao, et al.
Published: (2026)
Quantum Visual Fields with Neural Amplitude Encoding
by: Wang, Shuteng, et al.
Published: (2025)
by: Wang, Shuteng, et al.
Published: (2025)
Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
by: Lin, Pin-Jie, et al.
Published: (2024)
by: Lin, Pin-Jie, et al.
Published: (2024)
3D Human Pose Perception from Egocentric Stereo Videos
by: Akada, Hiroyasu, et al.
Published: (2023)
by: Akada, Hiroyasu, et al.
Published: (2023)
Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
by: Zhang, Zeyi, et al.
Published: (2024)
by: Zhang, Zeyi, et al.
Published: (2024)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
by: Voss, Hendric, et al.
Published: (2025)
by: Voss, Hendric, et al.
Published: (2025)
Similar Items
-
MIBURI: Towards Expressive Interactive Gesture Synthesis
by: Mughal, M. Hamza, et al.
Published: (2026) -
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
by: Mughal, Muhammad Hamza, et al.
Published: (2024) -
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025) -
ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions
by: Ghosh, Anindita, et al.
Published: (2023) -
Betsu-Betsu: Multi-View Separable 3D Reconstruction of Two Interacting Objects
by: Gopal, Suhas, et al.
Published: (2025)