Investigating the impact of 2D gesture representation on co-speech gesture generation
Fuente:
arXiv
Saved in:
| Main Authors: | Guichoux, Teo, Soulier, Laure, Obin, Nicolas, Pelachaud, Catherine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
by: Guichoux, Téo, et al.
Published: (2024)
by: Guichoux, Téo, et al.
Published: (2024)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
by: Amprimo, Gianluca, et al.
Published: (2025)
by: Amprimo, Gianluca, et al.
Published: (2025)
Online hand gesture recognition using Continual Graph Transformers
by: Slama, Rim, et al.
Published: (2025)
by: Slama, Rim, et al.
Published: (2025)
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
by: Akremi, Mohamed Sanim, et al.
Published: (2025)
by: Akremi, Mohamed Sanim, et al.
Published: (2025)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)
by: Guichoux, Téo, et al.
Published: (2025)
Prototype Learning for Micro-gesture Classification
by: Chen, Guoliang, et al.
Published: (2024)
by: Chen, Guoliang, et al.
Published: (2024)
What Makes Multimodal In-Context Learning Work?
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
by: Tang, Yiwen, et al.
Published: (2025)
by: Tang, Yiwen, et al.
Published: (2025)
Research on gesture recognition method based on SEDCNN-SVM
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
by: Nguyen, Giang Son, et al.
Published: (2026)
by: Nguyen, Giang Son, et al.
Published: (2026)
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024)
by: Faure, Gueter Josmy, et al.
Published: (2024)
Multi-object event graph representation learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
by: Teo, Rachel S. Y., et al.
Published: (2024)
by: Teo, Rachel S. Y., et al.
Published: (2024)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Micro-gesture Online Recognition using Learnable Query Points
by: Liu, Pengyu, et al.
Published: (2024)
by: Liu, Pengyu, et al.
Published: (2024)
Interpreting Hand gestures using Object Detection and Digits Classification
by: K, Sangeetha, et al.
Published: (2024)
by: K, Sangeetha, et al.
Published: (2024)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024)
by: Englebert, Alexandre, et al.
Published: (2024)
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
by: Xu, Zhiyu, et al.
Published: (2026)
by: Xu, Zhiyu, et al.
Published: (2026)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
MAIRA-1: A specialised large multimodal model for radiology report generation
by: Hyland, Stephanie L., et al.
Published: (2023)
by: Hyland, Stephanie L., et al.
Published: (2023)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
by: Zhong, Weihong, et al.
Published: (2024)
by: Zhong, Weihong, et al.
Published: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023)
by: Chen, Xinyi, et al.
Published: (2023)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
Phoenix-VL 1.5 Medium Technical Report
by: Phoenix, Team, et al.
Published: (2026)
by: Phoenix, Team, et al.
Published: (2026)
Deep self-supervised learning with visualisation for automatic gesture recognition
by: Allemand, Fabien, et al.
Published: (2024)
by: Allemand, Fabien, et al.
Published: (2024)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
by: Guo, Ziyu, et al.
Published: (2024)
by: Guo, Ziyu, et al.
Published: (2024)
S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation
by: Gao, Jiechao, et al.
Published: (2025)
by: Gao, Jiechao, et al.
Published: (2025)
CADENet: Condition-Adaptive Asynchronous Dual-Stream Enhancement Network for Adverse Weather Perception in Autonomous Driving
by: Khairy, Sherif, et al.
Published: (2026)
by: Khairy, Sherif, et al.
Published: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
by: Omri, Yasmine, et al.
Published: (2025)
by: Omri, Yasmine, et al.
Published: (2025)
Elliptical Attention
by: Nielsen, Stefan K., et al.
Published: (2024)
by: Nielsen, Stefan K., et al.
Published: (2024)
Hybrid-supervised Hypergraph-enhanced Transformer for Micro-gesture Based Emotion Recognition
by: Xia, Zhaoqiang, et al.
Published: (2025)
by: Xia, Zhaoqiang, et al.
Published: (2025)
Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention
by: Liu, Pengyu, et al.
Published: (2025)
by: Liu, Pengyu, et al.
Published: (2025)
Panoptic Diffusion Models: co-generation of images and segmentation maps
by: Long, Yinghan, et al.
Published: (2024)
by: Long, Yinghan, et al.
Published: (2024)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
by: Patock, Jake R., et al.
Published: (2025)
by: Patock, Jake R., et al.
Published: (2025)
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding
by: Wang, Austin T., et al.
Published: (2025)
by: Wang, Austin T., et al.
Published: (2025)
3D-LEX v1.0: 3D Lexicons for American Sign Language and Sign Language of the Netherlands
by: Ranum, Oline, et al.
Published: (2024)
by: Ranum, Oline, et al.
Published: (2024)
Similar Items
-
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
by: Guichoux, Téo, et al.
Published: (2024) -
EHWGesture -- A dataset for multimodal understanding of clinical gestures
by: Amprimo, Gianluca, et al.
Published: (2025) -
Online hand gesture recognition using Continual Graph Transformers
by: Slama, Rim, et al.
Published: (2025) -
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
by: Akremi, Mohamed Sanim, et al.
Published: (2025) -
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)