Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
Fuente:
arXiv
Salvato in:
| Autori principali: | Jha, Saurav, Ehrlich, Stefan K. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Backbone for Long-Horizon Robot Task Understanding
di: Chen, Xiaoshuai, et al.
Pubblicazione: (2024)
di: Chen, Xiaoshuai, et al.
Pubblicazione: (2024)
Acoustic Field Video for Multimodal Scene Understanding
di: Kim, Daehwa, et al.
Pubblicazione: (2026)
di: Kim, Daehwa, et al.
Pubblicazione: (2026)
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
di: Han, Kyungtae, et al.
Pubblicazione: (2025)
di: Han, Kyungtae, et al.
Pubblicazione: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
di: Wang, Kohou, et al.
Pubblicazione: (2025)
di: Wang, Kohou, et al.
Pubblicazione: (2025)
Multi-face emotion detection for effective Human-Robot Interaction
di: Yahyaoui, Mohamed Ala, et al.
Pubblicazione: (2025)
di: Yahyaoui, Mohamed Ala, et al.
Pubblicazione: (2025)
DEXOP: A Device for Robotic Transfer of Dexterous Human Manipulation
di: Fang, Hao-Shu, et al.
Pubblicazione: (2025)
di: Fang, Hao-Shu, et al.
Pubblicazione: (2025)
A Survey on Improving Human Robot Collaboration through Vision-and-Language Navigation
di: Yakolli, Nivedan, et al.
Pubblicazione: (2025)
di: Yakolli, Nivedan, et al.
Pubblicazione: (2025)
Gaze Detection and Analysis for Initiating Joint Activity in Industrial Human-Robot Collaboration
di: Prajod, Pooja, et al.
Pubblicazione: (2023)
di: Prajod, Pooja, et al.
Pubblicazione: (2023)
Creativity and Visual Communication from Machine to Musician: Sharing a Score through a Robotic Camera
di: Greer, Ross, et al.
Pubblicazione: (2024)
di: Greer, Ross, et al.
Pubblicazione: (2024)
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
di: Li, Xun, et al.
Pubblicazione: (2025)
di: Li, Xun, et al.
Pubblicazione: (2025)
ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement
di: Su, Chenyu, et al.
Pubblicazione: (2025)
di: Su, Chenyu, et al.
Pubblicazione: (2025)
Generating Robot Constitutions & Benchmarks for Semantic Safety
di: Sermanet, Pierre, et al.
Pubblicazione: (2025)
di: Sermanet, Pierre, et al.
Pubblicazione: (2025)
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
di: Chen, Zhen, et al.
Pubblicazione: (2024)
di: Chen, Zhen, et al.
Pubblicazione: (2024)
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
di: Wu, Yuekun, et al.
Pubblicazione: (2025)
di: Wu, Yuekun, et al.
Pubblicazione: (2025)
VLM-driven Behavior Tree for Context-aware Task Planning
di: Wake, Naoki, et al.
Pubblicazione: (2025)
di: Wake, Naoki, et al.
Pubblicazione: (2025)
Skeleton-Based Transformer for Classification of Errors and Better Feedback in Low Back Pain Physical Rehabilitation Exercises
di: Marusic, Aleksa, et al.
Pubblicazione: (2025)
di: Marusic, Aleksa, et al.
Pubblicazione: (2025)
ICPR 2024 Competition on Rider Intention Prediction
di: Gangisetty, Shankar, et al.
Pubblicazione: (2025)
di: Gangisetty, Shankar, et al.
Pubblicazione: (2025)
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
di: Chavan, Vivek, et al.
Pubblicazione: (2025)
di: Chavan, Vivek, et al.
Pubblicazione: (2025)
CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
di: Zeng, Tianle, et al.
Pubblicazione: (2026)
di: Zeng, Tianle, et al.
Pubblicazione: (2026)
Logic-Free Building Automation: Learning the Control of Room Facilities with Wall Switches and Ceiling Camera
di: Ochiai, Hideya, et al.
Pubblicazione: (2024)
di: Ochiai, Hideya, et al.
Pubblicazione: (2024)
SurgBox: Agent-Driven Operating Room Sandbox with Surgery Copilot
di: Wu, Jinlin, et al.
Pubblicazione: (2024)
di: Wu, Jinlin, et al.
Pubblicazione: (2024)
GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation
di: Lu, Haoran, et al.
Pubblicazione: (2024)
di: Lu, Haoran, et al.
Pubblicazione: (2024)
Magma: A Foundation Model for Multimodal AI Agents
di: Yang, Jianwei, et al.
Pubblicazione: (2025)
di: Yang, Jianwei, et al.
Pubblicazione: (2025)
A Multimodal Depth-Aware Method For Embodied Reference Understanding
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
di: Eyiokur, Fevziye Irem, et al.
Pubblicazione: (2025)
Human-Agent Joint Learning for Efficient Robot Manipulation Skill Acquisition
di: Luo, Shengcheng, et al.
Pubblicazione: (2024)
di: Luo, Shengcheng, et al.
Pubblicazione: (2024)
HAPI: A Model for Learning Robot Facial Expressions from Human Preferences
di: Yang, Dongsheng, et al.
Pubblicazione: (2025)
di: Yang, Dongsheng, et al.
Pubblicazione: (2025)
Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction
di: Wahl, Ari, et al.
Pubblicazione: (2026)
di: Wahl, Ari, et al.
Pubblicazione: (2026)
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
CD-TWINSAFE: A ROS-enabled Digital Twin for Scene Understanding and Safety Emerging V2I Technology
di: Khaled, Amro, et al.
Pubblicazione: (2026)
di: Khaled, Amro, et al.
Pubblicazione: (2026)
ShelfHelp: Empowering Humans to Perform Vision-Independent Manipulation Tasks with a Socially Assistive Robotic Cane
di: Agrawal, Shivendra, et al.
Pubblicazione: (2024)
di: Agrawal, Shivendra, et al.
Pubblicazione: (2024)
Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
di: Yano, Yuga, et al.
Pubblicazione: (2024)
di: Yano, Yuga, et al.
Pubblicazione: (2024)
ObjectGS: Object-aware Scene Reconstruction and Scene Understanding via Gaussian Splatting
di: Zhu, Ruijie, et al.
Pubblicazione: (2025)
di: Zhu, Ruijie, et al.
Pubblicazione: (2025)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
di: Duan, Lin, et al.
Pubblicazione: (2025)
di: Duan, Lin, et al.
Pubblicazione: (2025)
Real-Time Multimodal Signal Processing for HRI in RoboCup: Understanding a Human Referee
di: Ansalone, Filippo, et al.
Pubblicazione: (2024)
di: Ansalone, Filippo, et al.
Pubblicazione: (2024)
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces
di: Payandeh, Amirreza, et al.
Pubblicazione: (2024)
di: Payandeh, Amirreza, et al.
Pubblicazione: (2024)
User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals
di: Miyoshi, Ryo, et al.
Pubblicazione: (2025)
di: Miyoshi, Ryo, et al.
Pubblicazione: (2025)
SasMamba: A Lightweight Structure-Aware Stride State Space Model for 3D Human Pose Estimation
di: Cui, Hu, et al.
Pubblicazione: (2025)
di: Cui, Hu, et al.
Pubblicazione: (2025)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
di: Zhao, Haoyu, et al.
Pubblicazione: (2025)
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
di: Barakat, Nadim, et al.
Pubblicazione: (2025)
di: Barakat, Nadim, et al.
Pubblicazione: (2025)
Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
di: Chiatti, Agnese, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Backbone for Long-Horizon Robot Task Understanding
di: Chen, Xiaoshuai, et al.
Pubblicazione: (2024) -
Acoustic Field Video for Multimodal Scene Understanding
di: Kim, Daehwa, et al.
Pubblicazione: (2026) -
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
di: Han, Kyungtae, et al.
Pubblicazione: (2025) -
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
di: Wang, Kohou, et al.
Pubblicazione: (2025) -
Multi-face emotion detection for effective Human-Robot Interaction
di: Yahyaoui, Mohamed Ala, et al.
Pubblicazione: (2025)