MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fung, Angus, Tan, Aaron Hao, Wang, Haitong, Benhabib, Beno, Nejat, Goldie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
by: Fung, Angus, et al.
Published: (2024)
by: Fung, Angus, et al.
Published: (2024)
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
by: Fung, Angus, et al.
Published: (2022)
by: Fung, Angus, et al.
Published: (2022)
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
by: Lisondra, Matthew, et al.
Published: (2025)
by: Lisondra, Matthew, et al.
Published: (2025)
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
by: Tan, Aaron Hao, et al.
Published: (2025)
by: Tan, Aaron Hao, et al.
Published: (2025)
X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
by: Wang, Haitong, et al.
Published: (2025)
by: Wang, Haitong, et al.
Published: (2025)
Find Everything: A General Vision Language Model Approach to Multi-Object Search
by: Choi, Daniel, et al.
Published: (2024)
by: Choi, Daniel, et al.
Published: (2024)
NavFormer: A Transformer Architecture for Robot Target-Driven Navigation in Unknown and Dynamic Environments
by: Wang, Haitong, et al.
Published: (2024)
by: Wang, Haitong, et al.
Published: (2024)
The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
by: Pashangpour, Souren, et al.
Published: (2024)
by: Pashangpour, Souren, et al.
Published: (2024)
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
by: Narasimhan, Siddarth, et al.
Published: (2025)
by: Narasimhan, Siddarth, et al.
Published: (2025)
4CNet: A Diffusion Approach to Map Prediction for Decentralized Multi-Robot Exploration
by: Tan, Aaron Hao, et al.
Published: (2024)
by: Tan, Aaron Hao, et al.
Published: (2024)
OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation
by: Narasimhan, Siddarth, et al.
Published: (2024)
by: Narasimhan, Siddarth, et al.
Published: (2024)
MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
by: Wang, Liman, et al.
Published: (2025)
by: Wang, Liman, et al.
Published: (2025)
ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions
by: Pashangpour, Souren, et al.
Published: (2026)
by: Pashangpour, Souren, et al.
Published: (2026)
MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation
by: Singh, Harsh, et al.
Published: (2024)
by: Singh, Harsh, et al.
Published: (2024)
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
by: Wen, Congcong, et al.
Published: (2025)
by: Wen, Congcong, et al.
Published: (2025)
ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition
by: Li, Samuel, et al.
Published: (2024)
by: Li, Samuel, et al.
Published: (2024)
Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation
by: Subedi, Nitesh, et al.
Published: (2025)
by: Subedi, Nitesh, et al.
Published: (2025)
NVP-HRI: Zero Shot Natural Voice and Posture-based Human-Robot Interaction via Large Language Model
by: Lai, Yuzhi, et al.
Published: (2025)
by: Lai, Yuzhi, et al.
Published: (2025)
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment Planning
by: Yang, Dongrong, et al.
Published: (2025)
by: Yang, Dongrong, et al.
Published: (2025)
History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation
by: Habibpour, Mobin, et al.
Published: (2025)
by: Habibpour, Mobin, et al.
Published: (2025)
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
Conformal Temporal Logic Planning using Large Language Models
by: Wang, Jun, et al.
Published: (2023)
by: Wang, Jun, et al.
Published: (2023)
SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
by: As, Yarden, et al.
Published: (2025)
by: As, Yarden, et al.
Published: (2025)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
Neural Algorithmic Reasoners informed Large Language Model for Multi-Agent Path Finding
by: Feng, Pu, et al.
Published: (2025)
by: Feng, Pu, et al.
Published: (2025)
Deep Reinforcement Learning for Decentralized Multi-Robot Exploration With Macro Actions
by: Tan, Aaron Hao, et al.
Published: (2021)
by: Tan, Aaron Hao, et al.
Published: (2021)
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots
by: Varela, Iñaki Dellibarda, et al.
Published: (2025)
by: Varela, Iñaki Dellibarda, et al.
Published: (2025)
OpenNav: Open-World Navigation with Multimodal Large Language Models
by: Yuan, Mingfeng, et al.
Published: (2025)
by: Yuan, Mingfeng, et al.
Published: (2025)
Trajectory Adaptation using Large Language Models
by: Maurya, Anurag, et al.
Published: (2025)
by: Maurya, Anurag, et al.
Published: (2025)
Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model
by: Lu, Qiujing, et al.
Published: (2024)
by: Lu, Qiujing, et al.
Published: (2024)
Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
by: Dorbala, Vishnu Sashank, et al.
Published: (2023)
by: Dorbala, Vishnu Sashank, et al.
Published: (2023)
DexGrasp-Zero: A Morphology-Aligned Policy for Zero-Shot Cross-Embodiment Dexterous Grasping
by: Wu, Yuliang, et al.
Published: (2026)
by: Wu, Yuliang, et al.
Published: (2026)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
by: Liu, Jinyi, et al.
Published: (2024)
by: Liu, Jinyi, et al.
Published: (2024)
Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs
by: Wang, Zidan, et al.
Published: (2024)
by: Wang, Zidan, et al.
Published: (2024)
Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition
by: Quartey, Benedict, et al.
Published: (2026)
by: Quartey, Benedict, et al.
Published: (2026)
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
by: Debnath, Arnab, et al.
Published: (2025)
by: Debnath, Arnab, et al.
Published: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
by: Xiong, Chuyan, et al.
Published: (2024)
by: Xiong, Chuyan, et al.
Published: (2024)
COMRES-VLM: Coordinated Multi-Robot Exploration and Search using Vision Language Models
by: Wang, Ruiyang, et al.
Published: (2025)
by: Wang, Ruiyang, et al.
Published: (2025)
Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
by: Liu, Yuchen, et al.
Published: (2025)
by: Liu, Yuchen, et al.
Published: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
by: Sasabuchi, Kazuhiro, et al.
Published: (2025)
Similar Items
-
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
by: Fung, Angus, et al.
Published: (2024) -
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
by: Fung, Angus, et al.
Published: (2022) -
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
by: Lisondra, Matthew, et al.
Published: (2025) -
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
by: Tan, Aaron Hao, et al.
Published: (2025) -
X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
by: Wang, Haitong, et al.
Published: (2025)