MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | Deichler, Anna, O'Regan, Jim, Dogan, Fethiye Irmak, Marcinek, Lubos, Klezovich, Anna, Leite, Iolanda, Beskow, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Incorporating Spatial Awareness in Data-Driven Gesture Generation for Virtual Agents
by: Deichler, Anna, et al.
Published: (2024)
by: Deichler, Anna, et al.
Published: (2024)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
by: Deichler, Anna
Published: (2026)
by: Deichler, Anna
Published: (2026)
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Gesture Evaluation in Virtual Reality
by: Werner, Axel Wiebe, et al.
Published: (2025)
by: Werner, Axel Wiebe, et al.
Published: (2025)
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
by: Deichler, Anna, et al.
Published: (2024)
by: Deichler, Anna, et al.
Published: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024)
by: Bonial, Claire, et al.
Published: (2024)
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
by: Portelance, Eva, et al.
Published: (2023)
by: Portelance, Eva, et al.
Published: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
by: Yang, Shan
Published: (2026)
by: Yang, Shan
Published: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
by: Dai, Song, et al.
Published: (2025)
by: Dai, Song, et al.
Published: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
by: Lim, Shoon Kit, et al.
Published: (2025)
by: Lim, Shoon Kit, et al.
Published: (2025)
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
by: Dong, Yi, et al.
Published: (2025)
by: Dong, Yi, et al.
Published: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
by: Parcalabescu, Letitia, et al.
Published: (2022)
by: Parcalabescu, Letitia, et al.
Published: (2022)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
by: Lukin, Stephanie M., et al.
Published: (2024)
by: Lukin, Stephanie M., et al.
Published: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
PC-SNN: Predictive Coding-based Local Hebbian Plasticity Learning in Spiking Neural Networks
by: Wang, Haidong, et al.
Published: (2022)
by: Wang, Haidong, et al.
Published: (2022)
Precision at Scale: Domain-Specific Datasets On-Demand
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
by: Portelance, Eva, et al.
Published: (2024)
by: Portelance, Eva, et al.
Published: (2024)
Habitat Classification from Ground-Level Imagery Using Deep Neural Networks
by: Shi, Hongrui, et al.
Published: (2025)
by: Shi, Hongrui, et al.
Published: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
by: Mansour, Jad, et al.
Published: (2024)
by: Mansour, Jad, et al.
Published: (2024)
Conditional Compatibility Learning for Context-Dependent Anomaly Detection
by: Mishra, Shashank, et al.
Published: (2026)
by: Mishra, Shashank, et al.
Published: (2026)
Perceptual Flow Network for Visually Grounded Reasoning
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
by: Liu, Bingnan, et al.
Published: (2026)
by: Liu, Bingnan, et al.
Published: (2026)
Learning Association via Track-Detection Matching for Multi-Object Tracking
by: Adžemović, Momir
Published: (2025)
by: Adžemović, Momir
Published: (2025)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
by: Dehghani, Mahshid, et al.
Published: (2024)
by: Dehghani, Mahshid, et al.
Published: (2024)
Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
by: Adhikari, Sayanta, et al.
Published: (2024)
by: Adhikari, Sayanta, et al.
Published: (2024)
Growing Perspectives: Modelling Embodied Perspective Taking and Inner Narrative Development Using Large Language Models
by: Patania, Sabrina, et al.
Published: (2025)
by: Patania, Sabrina, et al.
Published: (2025)
Similar Items
-
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
by: Deichler, Anna, et al.
Published: (2025) -
Incorporating Spatial Awareness in Data-Driven Gesture Generation for Virtual Agents
by: Deichler, Anna, et al.
Published: (2024) -
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
by: Deichler, Anna
Published: (2026) -
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025) -
Gesture Evaluation in Virtual Reality
by: Werner, Axel Wiebe, et al.
Published: (2025)