I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
Fuente:
arXiv
Guardado en:
| Autores principales: | Ghaleb, Esam, Khaertdinov, Bulat, Özyürek, Aslı, Fernández, Raquel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
por: Ghaleb, Esam, et al.
Publicado: (2024)
por: Ghaleb, Esam, et al.
Publicado: (2024)
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
por: Liu, Lanmiao, et al.
Publicado: (2026)
por: Liu, Lanmiao, et al.
Publicado: (2026)
SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
por: Liu, Lanmiao, et al.
Publicado: (2025)
por: Liu, Lanmiao, et al.
Publicado: (2025)
DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
por: Paar, Ferdinand, et al.
Publicado: (2026)
por: Paar, Ferdinand, et al.
Publicado: (2026)
Co-Speech Gesture Detection through Multi-Phase Sequence Labeling
por: Ghaleb, Esam, et al.
Publicado: (2023)
por: Ghaleb, Esam, et al.
Publicado: (2023)
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
por: Keleş, Onur, et al.
Publicado: (2025)
por: Keleş, Onur, et al.
Publicado: (2025)
Leveraging Speech for Gesture Detection in Multimodal Communication
por: Ghaleb, Esam, et al.
Publicado: (2024)
por: Ghaleb, Esam, et al.
Publicado: (2024)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
por: Cheng, Hongye, et al.
Publicado: (2025)
por: Cheng, Hongye, et al.
Publicado: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
por: Jiang, Jingjing, et al.
Publicado: (2025)
por: Jiang, Jingjing, et al.
Publicado: (2025)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
por: Qi, Xingqun, et al.
Publicado: (2023)
por: Qi, Xingqun, et al.
Publicado: (2023)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024)
por: He, Xu, et al.
Publicado: (2024)
Analysing Cross-Speaker Convergence in Face-to-Face Dialogue through the Lens of Automatically Detected Shared Linguistic Constructions
por: Ghaleb, Esam, et al.
Publicado: (2024)
por: Ghaleb, Esam, et al.
Publicado: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
por: Cheng, Zebang, et al.
Publicado: (2024)
por: Cheng, Zebang, et al.
Publicado: (2024)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2024)
por: Wang, Bing, et al.
Publicado: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
por: Bin, Yi, et al.
Publicado: (2024)
por: Bin, Yi, et al.
Publicado: (2024)
MLLM can see? Dynamic Correction Decoding for Hallucination Mitigation
por: Wang, Chenxi, et al.
Publicado: (2024)
por: Wang, Chenxi, et al.
Publicado: (2024)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
por: An, Wenbin, et al.
Publicado: (2025)
por: An, Wenbin, et al.
Publicado: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
por: Wang, Kangsheng, et al.
Publicado: (2025)
por: Wang, Kangsheng, et al.
Publicado: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
por: Gan, Chengguang, et al.
Publicado: (2025)
por: Gan, Chengguang, et al.
Publicado: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
por: Zhao, Zhixian, et al.
Publicado: (2026)
por: Zhao, Zhixian, et al.
Publicado: (2026)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
por: Masumura, Ryo, et al.
Publicado: (2025)
por: Masumura, Ryo, et al.
Publicado: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
por: Wu, Jiaying, et al.
Publicado: (2025)
por: Wu, Jiaying, et al.
Publicado: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
por: Zhang, Xueqiao, et al.
Publicado: (2025)
por: Zhang, Xueqiao, et al.
Publicado: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
por: Jiang, Chaoya, et al.
Publicado: (2024)
por: Jiang, Chaoya, et al.
Publicado: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
por: Li, Jinyuan, et al.
Publicado: (2024)
por: Li, Jinyuan, et al.
Publicado: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
por: Zhang, Dongxu, et al.
Publicado: (2026)
por: Zhang, Dongxu, et al.
Publicado: (2026)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
por: Liang, Zhengyang, et al.
Publicado: (2024)
por: Liang, Zhengyang, et al.
Publicado: (2024)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
por: Fernandez-Lopez, Adriana, et al.
Publicado: (2024)
por: Fernandez-Lopez, Adriana, et al.
Publicado: (2024)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
por: Zhang, Zhi, et al.
Publicado: (2023)
por: Zhang, Zhi, et al.
Publicado: (2023)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
por: Zhang, Liyun
Publicado: (2024)
por: Zhang, Liyun
Publicado: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
por: Lai, Zhengzhao, et al.
Publicado: (2025)
por: Lai, Zhengzhao, et al.
Publicado: (2025)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
por: Guo, Xin, et al.
Publicado: (2025)
por: Guo, Xin, et al.
Publicado: (2025)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
por: Yan, Zehong, et al.
Publicado: (2025)
por: Yan, Zehong, et al.
Publicado: (2025)
A New Hybrid Intelligent Approach for Multimodal Detection of Suspected Disinformation on TikTok
por: Guerrero-Sosa, Jared D. T., et al.
Publicado: (2025)
por: Guerrero-Sosa, Jared D. T., et al.
Publicado: (2025)
Ejemplares similares
-
Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
por: Ghaleb, Esam, et al.
Publicado: (2024) -
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
por: Liu, Lanmiao, et al.
Publicado: (2026) -
SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
por: Liu, Lanmiao, et al.
Publicado: (2025) -
DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
por: Paar, Ferdinand, et al.
Publicado: (2026) -
Co-Speech Gesture Detection through Multi-Phase Sequence Labeling
por: Ghaleb, Esam, et al.
Publicado: (2023)