Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
Fuente:
arXiv
Saved in:
| Main Authors: | Deichler, Anna, Beskow, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
by: Deichler, Anna, et al.
Published: (2026)
by: Deichler, Anna, et al.
Published: (2026)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024)
by: Bonial, Claire, et al.
Published: (2024)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
by: Lukin, Stephanie M., et al.
Published: (2024)
by: Lukin, Stephanie M., et al.
Published: (2024)
A Replicable Robotics Awareness Method Using LLM-Enabled Robotics Interaction: Evidence from a Corporate Challenge
by: Prieto, S. A., et al.
Published: (2026)
by: Prieto, S. A., et al.
Published: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Exploring the Effect of Robotic Embodiment and Empathetic Tone of LLMs on Empathy Elicitation
by: Darwesh, Liza, et al.
Published: (2025)
by: Darwesh, Liza, et al.
Published: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
by: Lim, Shoon Kit, et al.
Published: (2025)
by: Lim, Shoon Kit, et al.
Published: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
3HANDS Dataset: Learning from Humans for Generating Naturalistic Handovers with Supernumerary Robotic Limbs
by: Abadian, Artin Saberpour, et al.
Published: (2025)
by: Abadian, Artin Saberpour, et al.
Published: (2025)
PoseRefer: Pathway-Local Parameters for Semantically Grounded Reference Resolution
by: Deichler, Anna
Published: (2026)
by: Deichler, Anna
Published: (2026)
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
by: Menon, Anjali R., et al.
Published: (2025)
by: Menon, Anjali R., et al.
Published: (2025)
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
by: Li, Wenbo, et al.
Published: (2025)
by: Li, Wenbo, et al.
Published: (2025)
GIST: Multimodal Knowledge Extraction and Spatial Grounding via Intelligent Semantic Topology
by: Agrawal, Shivendra, et al.
Published: (2026)
by: Agrawal, Shivendra, et al.
Published: (2026)
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
by: Wachowiak, Lennart, et al.
Published: (2025)
by: Wachowiak, Lennart, et al.
Published: (2025)
Grounded Gesture Generation: Language, Motion, and Space
by: Deichler, Anna, et al.
Published: (2025)
by: Deichler, Anna, et al.
Published: (2025)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
by: Chen, Yiteng, et al.
Published: (2025)
by: Chen, Yiteng, et al.
Published: (2025)
Speech to Reality: On-Demand Production using Natural Language, 3D Generative AI, and Discrete Robotic Assembly
by: Kyaw, Alexander Htet, et al.
Published: (2024)
by: Kyaw, Alexander Htet, et al.
Published: (2024)
Immersive Robot Programming Interface for Human-Guided Automation and Randomized Path Planning
by: Malek, Kaveh, et al.
Published: (2024)
by: Malek, Kaveh, et al.
Published: (2024)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
by: Sharma, Aditya, et al.
Published: (2025)
by: Sharma, Aditya, et al.
Published: (2025)
M2HRI: An LLM-Driven Multimodal Multi-Agent Framework for Personalized Human-Robot Interaction
by: Hasan, Shaid, et al.
Published: (2026)
by: Hasan, Shaid, et al.
Published: (2026)
A Surveillance Based Interactive Robot
by: Kavimandan, Kshitij, et al.
Published: (2025)
by: Kavimandan, Kshitij, et al.
Published: (2025)
Emotions in the Loop: A Survey of Affective Computing for Emotional Support
by: Hegde, Karishma, et al.
Published: (2025)
by: Hegde, Karishma, et al.
Published: (2025)
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
by: Shaw, Ankit Kumar, et al.
Published: (2025)
by: Shaw, Ankit Kumar, et al.
Published: (2025)
Diffusion-SAFE: Diffusion-Native Human-to-Robot Driving Handover for Shared Autonomy
by: Fan, Yunxin, et al.
Published: (2025)
by: Fan, Yunxin, et al.
Published: (2025)
Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages
by: Murtadak, Divendar, et al.
Published: (2025)
by: Murtadak, Divendar, et al.
Published: (2025)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
by: Agia, Christopher, et al.
Published: (2024)
by: Agia, Christopher, et al.
Published: (2024)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
Conversations with Andrea: Visitors' Opinions on Android Robots in a Museum
by: Heisler, Marcel, et al.
Published: (2025)
by: Heisler, Marcel, et al.
Published: (2025)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
by: Hu, X., et al.
Published: (2025)
by: Hu, X., et al.
Published: (2025)
PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
by: Patania, Sabrina, et al.
Published: (2025)
by: Patania, Sabrina, et al.
Published: (2025)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
Industrial Robot Motion Planning with GPUs: Integration of cuRobo for Extended DOF Systems
by: Abuelsamen, Luai, et al.
Published: (2025)
by: Abuelsamen, Luai, et al.
Published: (2025)
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
by: de Carvalho, Marcus Vinicius Leal, et al.
Published: (2024)
by: de Carvalho, Marcus Vinicius Leal, et al.
Published: (2024)
Unconscious and Intentional Human Motion Cues for Expressive Robot-Arm Motion Design
by: Tashiro, Taito, et al.
Published: (2025)
by: Tashiro, Taito, et al.
Published: (2025)
Otherness as a Quality in Designing Expressive Robotic Touch
by: Zhou, Ran, et al.
Published: (2026)
by: Zhou, Ran, et al.
Published: (2026)
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
by: Zhang, Chengwen, et al.
Published: (2026)
by: Zhang, Chengwen, et al.
Published: (2026)
HoloSpot: Intuitive Object Manipulation via Mixed Reality Drag-and-Drop
by: Garcia, Pablo Soler, et al.
Published: (2024)
by: Garcia, Pablo Soler, et al.
Published: (2024)
Yanyun-3: Enabling Cross-Platform Strategy Game Operation with Vision-Language Models
by: Wang, Guoyan, et al.
Published: (2025)
by: Wang, Guoyan, et al.
Published: (2025)
Similar Items
-
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
by: Deichler, Anna, et al.
Published: (2026) -
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
by: Bonial, Claire, et al.
Published: (2024) -
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
by: Lukin, Stephanie M., et al.
Published: (2024) -
A Replicable Robotics Awareness Method Using LLM-Enabled Robotics Interaction: Evidence from a Corporate Challenge
by: Prieto, S. A., et al.
Published: (2026) -
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)