Investigating the impact of 2D gesture representation on co-speech gesture generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Guichoux, Teo, Soulier, Laure, Obin, Nicolas, Pelachaud, Catherine |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
di: Guichoux, Téo, et al.
Pubblicazione: (2024)
di: Guichoux, Téo, et al.
Pubblicazione: (2024)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
di: Amprimo, Gianluca, et al.
Pubblicazione: (2025)
di: Amprimo, Gianluca, et al.
Pubblicazione: (2025)
Online hand gesture recognition using Continual Graph Transformers
di: Slama, Rim, et al.
Pubblicazione: (2025)
di: Slama, Rim, et al.
Pubblicazione: (2025)
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
di: Akremi, Mohamed Sanim, et al.
Pubblicazione: (2025)
di: Akremi, Mohamed Sanim, et al.
Pubblicazione: (2025)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
di: Guichoux, Téo, et al.
Pubblicazione: (2025)
di: Guichoux, Téo, et al.
Pubblicazione: (2025)
Prototype Learning for Micro-gesture Classification
di: Chen, Guoliang, et al.
Pubblicazione: (2024)
di: Chen, Guoliang, et al.
Pubblicazione: (2024)
What Makes Multimodal In-Context Learning Work?
di: Baldassini, Folco Bertini, et al.
Pubblicazione: (2024)
di: Baldassini, Folco Bertini, et al.
Pubblicazione: (2024)
Explaining latent representations of generative models with large multimodal models
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
di: Tang, Yiwen, et al.
Pubblicazione: (2025)
di: Tang, Yiwen, et al.
Pubblicazione: (2025)
Research on gesture recognition method based on SEDCNN-SVM
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
di: Zhang, Mingjin, et al.
Pubblicazione: (2024)
An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
di: Nguyen, Giang Son, et al.
Pubblicazione: (2026)
di: Nguyen, Giang Son, et al.
Pubblicazione: (2026)
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
di: Faure, Gueter Josmy, et al.
Pubblicazione: (2024)
di: Faure, Gueter Josmy, et al.
Pubblicazione: (2024)
Multi-object event graph representation learning for Video Question Answering
di: Wang, Yanan, et al.
Pubblicazione: (2024)
di: Wang, Yanan, et al.
Pubblicazione: (2024)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2024)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
di: Teo, Rachel S. Y., et al.
Pubblicazione: (2025)
Micro-gesture Online Recognition using Learnable Query Points
di: Liu, Pengyu, et al.
Pubblicazione: (2024)
di: Liu, Pengyu, et al.
Pubblicazione: (2024)
Interpreting Hand gestures using Object Detection and Digits Classification
di: K, Sangeetha, et al.
Pubblicazione: (2024)
di: K, Sangeetha, et al.
Pubblicazione: (2024)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
di: Xu, Zhiyu, et al.
Pubblicazione: (2026)
di: Xu, Zhiyu, et al.
Pubblicazione: (2026)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
di: Madasu, Avinash, et al.
Pubblicazione: (2023)
di: Madasu, Avinash, et al.
Pubblicazione: (2023)
MAIRA-1: A specialised large multimodal model for radiology report generation
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
di: Hyland, Stephanie L., et al.
Pubblicazione: (2023)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
di: Chen, Xinyi, et al.
Pubblicazione: (2023)
A multimodal gesture recognition dataset for desktop human-computer interaction
di: Wang, Qi, et al.
Pubblicazione: (2024)
di: Wang, Qi, et al.
Pubblicazione: (2024)
Phoenix-VL 1.5 Medium Technical Report
di: Phoenix, Team, et al.
Pubblicazione: (2026)
di: Phoenix, Team, et al.
Pubblicazione: (2026)
Deep self-supervised learning with visualisation for automatic gesture recognition
di: Allemand, Fabien, et al.
Pubblicazione: (2024)
di: Allemand, Fabien, et al.
Pubblicazione: (2024)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
di: Guo, Ziyu, et al.
Pubblicazione: (2024)
di: Guo, Ziyu, et al.
Pubblicazione: (2024)
S2D-ALIGN: Shallow-to-Deep Auxiliary Learning for Anatomically-Grounded Radiology Report Generation
di: Gao, Jiechao, et al.
Pubblicazione: (2025)
di: Gao, Jiechao, et al.
Pubblicazione: (2025)
CADENet: Condition-Adaptive Asynchronous Dual-Stream Enhancement Network for Adverse Weather Perception in Autonomous Driving
di: Khairy, Sherif, et al.
Pubblicazione: (2026)
di: Khairy, Sherif, et al.
Pubblicazione: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
di: Hu, Wenbo, et al.
Pubblicazione: (2025)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
di: Omri, Yasmine, et al.
Pubblicazione: (2025)
di: Omri, Yasmine, et al.
Pubblicazione: (2025)
Elliptical Attention
di: Nielsen, Stefan K., et al.
Pubblicazione: (2024)
di: Nielsen, Stefan K., et al.
Pubblicazione: (2024)
Hybrid-supervised Hypergraph-enhanced Transformer for Micro-gesture Based Emotion Recognition
di: Xia, Zhaoqiang, et al.
Pubblicazione: (2025)
di: Xia, Zhaoqiang, et al.
Pubblicazione: (2025)
Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention
di: Liu, Pengyu, et al.
Pubblicazione: (2025)
di: Liu, Pengyu, et al.
Pubblicazione: (2025)
Panoptic Diffusion Models: co-generation of images and segmentation maps
di: Long, Yinghan, et al.
Pubblicazione: (2024)
di: Long, Yinghan, et al.
Pubblicazione: (2024)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
di: Patock, Jake R., et al.
Pubblicazione: (2025)
di: Patock, Jake R., et al.
Pubblicazione: (2025)
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
di: Kang, Weitai, et al.
Pubblicazione: (2024)
di: Kang, Weitai, et al.
Pubblicazione: (2024)
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding
di: Wang, Austin T., et al.
Pubblicazione: (2025)
di: Wang, Austin T., et al.
Pubblicazione: (2025)
3D-LEX v1.0: 3D Lexicons for American Sign Language and Sign Language of the Netherlands
di: Ranum, Oline, et al.
Pubblicazione: (2024)
di: Ranum, Oline, et al.
Pubblicazione: (2024)
Documenti analoghi
-
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
di: Guichoux, Téo, et al.
Pubblicazione: (2024) -
EHWGesture -- A dataset for multimodal understanding of clinical gestures
di: Amprimo, Gianluca, et al.
Pubblicazione: (2025) -
Online hand gesture recognition using Continual Graph Transformers
di: Slama, Rim, et al.
Pubblicazione: (2025) -
Accurate online action and gesture recognition system using detectors and Deep SPD Siamese Networks
di: Akremi, Mohamed Sanim, et al.
Pubblicazione: (2025) -
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
di: Guichoux, Téo, et al.
Pubblicazione: (2025)