A Multimodal Depth-Aware Method For Embodied Reference Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Eyiokur, Fevziye Irem, Yaman, Dogucan, Ekenel, Hazım Kemal, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
by: Eyiokur, Fevziye Irem, et al.
Published: (2025)
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)
by: Yaman, Dogucan, et al.
Published: (2024)
Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
by: Yaman, Dogucan, et al.
Published: (2025)
by: Yaman, Dogucan, et al.
Published: (2025)
Acoustic Field Video for Multimodal Scene Understanding
by: Kim, Daehwa, et al.
Published: (2026)
by: Kim, Daehwa, et al.
Published: (2026)
Real-Time Multimodal Signal Processing for HRI in RoboCup: Understanding a Human Referee
by: Ansalone, Filippo, et al.
Published: (2024)
by: Ansalone, Filippo, et al.
Published: (2024)
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
by: Jha, Saurav, et al.
Published: (2025)
by: Jha, Saurav, et al.
Published: (2025)
CD-TWINSAFE: A ROS-enabled Digital Twin for Scene Understanding and Safety Emerging V2I Technology
by: Khaled, Amro, et al.
Published: (2026)
by: Khaled, Amro, et al.
Published: (2026)
Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
by: Yano, Yuga, et al.
Published: (2024)
by: Yano, Yuga, et al.
Published: (2024)
User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals
by: Miyoshi, Ryo, et al.
Published: (2025)
by: Miyoshi, Ryo, et al.
Published: (2025)
CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
by: Zeng, Tianle, et al.
Published: (2026)
by: Zeng, Tianle, et al.
Published: (2026)
Assessing the Use of Face Swapping Methods as Face Anonymizers in Videos
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
by: Muştu, Mustafa İzzet, et al.
Published: (2025)
A Backbone for Long-Horizon Robot Task Understanding
by: Chen, Xiaoshuai, et al.
Published: (2024)
by: Chen, Xiaoshuai, et al.
Published: (2024)
MILE: A Mechanically Isomorphic Exoskeleton Data Collection System with Fingertip Visuotactile Sensing for Dexterous Manipulation
by: Du, Jinda, et al.
Published: (2025)
by: Du, Jinda, et al.
Published: (2025)
Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
by: Li, Xun, et al.
Published: (2025)
by: Li, Xun, et al.
Published: (2025)
Inclusive STEAM Education: A Framework for Teaching Cod-2 ing and Robotics to Students with Visually Impairment Using 3 Advanced Computer Vision
by: Hamash, Mahmoud, et al.
Published: (2025)
by: Hamash, Mahmoud, et al.
Published: (2025)
Analyzing the Effect of Combined Degradations on Face Recognition
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
Analyzing the Feature Extractor Networks for Face Image Synthesis
by: Sarıtaş, Erdi, et al.
Published: (2024)
by: Sarıtaş, Erdi, et al.
Published: (2024)
LocoVR: Multiuser Indoor Locomotion Dataset in Virtual Reality
by: Takeyama, Kojiro, et al.
Published: (2024)
by: Takeyama, Kojiro, et al.
Published: (2024)
EgoTouch: On-Body Touch Input Using AR/VR Headset Cameras
by: Mollyn, Vimal, et al.
Published: (2025)
by: Mollyn, Vimal, et al.
Published: (2025)
SpiritSight Agent: Advanced GUI Agent with One Look
by: Huang, Zhiyuan, et al.
Published: (2025)
by: Huang, Zhiyuan, et al.
Published: (2025)
Extending 3D body pose estimation for robotic-assistive therapies of autistic children
by: Santos, Laura, et al.
Published: (2024)
by: Santos, Laura, et al.
Published: (2024)
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
by: Wu, Yuekun, et al.
Published: (2025)
by: Wu, Yuekun, et al.
Published: (2025)
Robot Interaction Behavior Generation based on Social Motion Forecasting for Human-Robot Interaction
by: Mascaro, Esteve Valls, et al.
Published: (2024)
by: Mascaro, Esteve Valls, et al.
Published: (2024)
Toward Reliable Human Pose Forecasting with Uncertainty
by: Saadatnejad, Saeed, et al.
Published: (2023)
by: Saadatnejad, Saeed, et al.
Published: (2023)
AnyUser: Translating Sketched User Intent into Domestic Robots
by: Yang, Songyuan, et al.
Published: (2026)
by: Yang, Songyuan, et al.
Published: (2026)
TBD Pedestrian Data Collection: Towards Rich, Portable, and Large-Scale Natural Pedestrian Data
by: Wang, Allan, et al.
Published: (2023)
by: Wang, Allan, et al.
Published: (2023)
Next-Best-Trajectory Planning of Robot Manipulators for Effective Observation and Exploration
by: Renz, Heiko, et al.
Published: (2025)
by: Renz, Heiko, et al.
Published: (2025)
Low-Back Pain Physical Rehabilitation by Movement Analysis in Clinical Trial
by: Nguyen, Sao Mai
Published: (2026)
by: Nguyen, Sao Mai
Published: (2026)
GentleHumanoid: Learning Upper-body Compliance for Contact-rich Human and Object Interaction
by: Lu, Qingzhou, et al.
Published: (2025)
by: Lu, Qingzhou, et al.
Published: (2025)
Stable Tracking of Eye Gaze Direction During Ophthalmic Surgery
by: Hong, Tinghe, et al.
Published: (2025)
by: Hong, Tinghe, et al.
Published: (2025)
GuideNav: User-Informed Development of a Vision-Only Robotic Navigation Assistant For Blind Travelers
by: Hwang, Hochul, et al.
Published: (2025)
by: Hwang, Hochul, et al.
Published: (2025)
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces
by: Payandeh, Amirreza, et al.
Published: (2024)
by: Payandeh, Amirreza, et al.
Published: (2024)
ConceptFactory: Facilitate 3D Object Knowledge Annotation with Object Conceptualization
by: Sun, Jianhua, et al.
Published: (2024)
by: Sun, Jianhua, et al.
Published: (2024)
Probabilistic Human Intent Prediction for Mobile Manipulation: An Evaluation with Human-Inspired Constraints
by: Contreras, Cesar Alan, et al.
Published: (2025)
by: Contreras, Cesar Alan, et al.
Published: (2025)
Benchmarking Tesla's Traffic Light and Stop Sign Control: Field Dataset and Behavior Insights
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
by: Han, Kyungtae, et al.
Published: (2025)
by: Han, Kyungtae, et al.
Published: (2025)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
by: Guo, Hao, et al.
Published: (2025)
by: Guo, Hao, et al.
Published: (2025)
Similar Items
-
CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
by: Eyiokur, Fevziye Irem, et al.
Published: (2025) -
Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework
by: Yaman, Dogucan, et al.
Published: (2025) -
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023) -
Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation
by: Yaman, Dogucan, et al.
Published: (2025) -
Audio-Visual Speech Representation Expert for Enhanced Talking Face Video Generation and Evaluation
by: Yaman, Dogucan, et al.
Published: (2024)