CLUE: Crossmodal disambiguation via Language-vision Understanding with attEntion
Fuente:
arXiv
Saved in:
| Main Authors: | Abrini, Mouad, Chetouani, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Legibot: Generating Legible Motions for Service Robots Using Cost-Based Local Planners
by: Amirian, Javad, et al.
Published: (2024)
by: Amirian, Javad, et al.
Published: (2024)
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Reasoning LLMs for User-Aware Multimodal Conversational Agents
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs
by: Malécot, Jeanne, et al.
Published: (2026)
by: Malécot, Jeanne, et al.
Published: (2026)
Demographic User Modeling for Social Robotics with Multimodal Pre-trained Models
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Iterative On-Policy Refinement of Hierarchical Diffusion Policies for Language-Conditioned Manipulation
by: Grislain, Clemence, et al.
Published: (2026)
by: Grislain, Clemence, et al.
Published: (2026)
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025)
by: Grislain, Clemence, et al.
Published: (2025)
Upgrading Pepper Robot s Social Interaction with Advanced Hardware and Perception Enhancements
by: Magri, Paolo, et al.
Published: (2024)
by: Magri, Paolo, et al.
Published: (2024)
Controlling Intent Expressiveness in Robot Motion with Diffusion Models
by: Shi, Wenli, et al.
Published: (2025)
by: Shi, Wenli, et al.
Published: (2025)
Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
CLUE: Adaptively Prioritized Contextual Cues by Leveraging a Unified Semantic Map for Effective Zero-Shot Object-Goal Navigation
by: Kim, Taeyun, et al.
Published: (2026)
by: Kim, Taeyun, et al.
Published: (2026)
Inferring Implicit Goals Across Differing Task Models
by: Tulli, Silvia, et al.
Published: (2025)
by: Tulli, Silvia, et al.
Published: (2025)
Task-Aware Robotic Grasping by evaluating Quality Diversity Solutions through Foundation Models
by: Appius, Aurel X., et al.
Published: (2024)
by: Appius, Aurel X., et al.
Published: (2024)
VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making
by: Aissi, Mohamed Salim, et al.
Published: (2025)
by: Aissi, Mohamed Salim, et al.
Published: (2025)
Tid att städa
by: Ambjörnsson, Fanny
Published: (2018)
by: Ambjörnsson, Fanny
Published: (2018)
Konsten att kontextualisera
Published: (2022)
Published: (2022)
Sign Language: Towards Sign Understanding for Robot Autonomy
by: Agrawal, Ayush, et al.
Published: (2025)
by: Agrawal, Ayush, et al.
Published: (2025)
Incremental Language Understanding for Online Motion Planning of Robot Manipulators
by: Abrams, Mitchell, et al.
Published: (2025)
by: Abrams, Mitchell, et al.
Published: (2025)
Object Depth and Size Estimation using Stereo-vision and Integration with SLAM
by: Hamad, Layth, et al.
Published: (2024)
by: Hamad, Layth, et al.
Published: (2024)
MUVLA: Learning to Explore Object Navigation via Map Understanding
by: Han, Peilong, et al.
Published: (2025)
by: Han, Peilong, et al.
Published: (2025)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
by: Zhang, Hanxin, et al.
Published: (2026)
by: Zhang, Hanxin, et al.
Published: (2026)
Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments
by: Elnoor, Mohamed, et al.
Published: (2024)
by: Elnoor, Mohamed, et al.
Published: (2024)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
by: Ma, Wenxuan, et al.
Published: (2025)
by: Ma, Wenxuan, et al.
Published: (2025)
Robust Multi-Agent Target Tracking in Intermittent Communication Environments via Analytical Belief Merging
by: Abdelnaby, Mohamed, et al.
Published: (2026)
by: Abdelnaby, Mohamed, et al.
Published: (2026)
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
by: Jie, Haoxiang, et al.
Published: (2026)
by: Jie, Haoxiang, et al.
Published: (2026)
NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
by: Chen, Jiahong, et al.
Published: (2025)
by: Chen, Jiahong, et al.
Published: (2025)
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
by: Sun, Xuefei, et al.
Published: (2025)
by: Sun, Xuefei, et al.
Published: (2025)
Purely vision-based collective movement of robots
by: Mezey, David, et al.
Published: (2024)
by: Mezey, David, et al.
Published: (2024)
CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
by: Wei, Yi-Lin, et al.
Published: (2025)
by: Wei, Yi-Lin, et al.
Published: (2025)
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024)
by: Ryu, Chanhoe, et al.
Published: (2024)
Learning Longitudinal Stress Dynamics from Irregular Self-Reports via Time Embeddings
by: Simon, Louis, et al.
Published: (2025)
by: Simon, Louis, et al.
Published: (2025)
A vision-based robotic system for precision pollination of apples
by: Bhattarai, Uddhav, et al.
Published: (2024)
by: Bhattarai, Uddhav, et al.
Published: (2024)
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
by: Zhang, Yuhang, et al.
Published: (2025)
by: Zhang, Yuhang, et al.
Published: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
by: Xia, Zhongyu, et al.
Published: (2026)
by: Xia, Zhongyu, et al.
Published: (2026)
Sparsh: Self-supervised touch representations for vision-based tactile sensing
by: Higuera, Carolina, et al.
Published: (2024)
by: Higuera, Carolina, et al.
Published: (2024)
Collision avoidance from monocular vision trained with novel view synthesis
by: Tordjman--Levavasseur, Valentin, et al.
Published: (2025)
by: Tordjman--Levavasseur, Valentin, et al.
Published: (2025)
Traversability analysis with vision and terrain probing for safe legged robot navigation
by: Haddeler, Garen, et al.
Published: (2022)
by: Haddeler, Garen, et al.
Published: (2022)
Similar Items
-
Legibot: Generating Legible Motions for Service Robots Using Cost-Based Local Planners
by: Amirian, Javad, et al.
Published: (2024) -
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026) -
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
by: Rahimi, Hamed, et al.
Published: (2025) -
Reasoning LLMs for User-Aware Multimodal Conversational Agents
by: Rahimi, Hamed, et al.
Published: (2025) -
HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs
by: Malécot, Jeanne, et al.
Published: (2026)