MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fung, Angus, Tan, Aaron Hao, Wang, Haitong, Benhabib, Beno, Nejat, Goldie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
von: Fung, Angus, et al.
Veröffentlicht: (2024)
von: Fung, Angus, et al.
Veröffentlicht: (2024)
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
von: Fung, Angus, et al.
Veröffentlicht: (2022)
von: Fung, Angus, et al.
Veröffentlicht: (2022)
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
von: Lisondra, Matthew, et al.
Veröffentlicht: (2025)
von: Lisondra, Matthew, et al.
Veröffentlicht: (2025)
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2025)
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2025)
X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
von: Wang, Haitong, et al.
Veröffentlicht: (2025)
von: Wang, Haitong, et al.
Veröffentlicht: (2025)
Find Everything: A General Vision Language Model Approach to Multi-Object Search
von: Choi, Daniel, et al.
Veröffentlicht: (2024)
von: Choi, Daniel, et al.
Veröffentlicht: (2024)
NavFormer: A Transformer Architecture for Robot Target-Driven Navigation in Unknown and Dynamic Environments
von: Wang, Haitong, et al.
Veröffentlicht: (2024)
von: Wang, Haitong, et al.
Veröffentlicht: (2024)
The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
von: Pashangpour, Souren, et al.
Veröffentlicht: (2024)
von: Pashangpour, Souren, et al.
Veröffentlicht: (2024)
SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models
von: Narasimhan, Siddarth, et al.
Veröffentlicht: (2025)
von: Narasimhan, Siddarth, et al.
Veröffentlicht: (2025)
4CNet: A Diffusion Approach to Map Prediction for Decentralized Multi-Robot Exploration
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2024)
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2024)
OLiVia-Nav: An Online Lifelong Vision Language Approach for Mobile Robot Social Navigation
von: Narasimhan, Siddarth, et al.
Veröffentlicht: (2024)
von: Narasimhan, Siddarth, et al.
Veröffentlicht: (2024)
MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection
von: Wang, Liman, et al.
Veröffentlicht: (2025)
von: Wang, Liman, et al.
Veröffentlicht: (2025)
ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions
von: Pashangpour, Souren, et al.
Veröffentlicht: (2026)
von: Pashangpour, Souren, et al.
Veröffentlicht: (2026)
MALMM: Multi-Agent Large Language Models for Zero-Shot Robotics Manipulation
von: Singh, Harsh, et al.
Veröffentlicht: (2024)
von: Singh, Harsh, et al.
Veröffentlicht: (2024)
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
von: Wen, Congcong, et al.
Veröffentlicht: (2025)
von: Wen, Congcong, et al.
Veröffentlicht: (2025)
ShapeGrasp: Zero-Shot Task-Oriented Grasping with Large Language Models through Geometric Decomposition
von: Li, Samuel, et al.
Veröffentlicht: (2024)
von: Li, Samuel, et al.
Veröffentlicht: (2024)
Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation
von: Subedi, Nitesh, et al.
Veröffentlicht: (2025)
von: Subedi, Nitesh, et al.
Veröffentlicht: (2025)
NVP-HRI: Zero Shot Natural Voice and Posture-based Human-Robot Interaction via Large Language Model
von: Lai, Yuzhi, et al.
Veröffentlicht: (2025)
von: Lai, Yuzhi, et al.
Veröffentlicht: (2025)
Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment Planning
von: Yang, Dongrong, et al.
Veröffentlicht: (2025)
von: Yang, Dongrong, et al.
Veröffentlicht: (2025)
History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation
von: Habibpour, Mobin, et al.
Veröffentlicht: (2025)
von: Habibpour, Mobin, et al.
Veröffentlicht: (2025)
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
von: Shi, Junyao, et al.
Veröffentlicht: (2025)
Conformal Temporal Logic Planning using Large Language Models
von: Wang, Jun, et al.
Veröffentlicht: (2023)
von: Wang, Jun, et al.
Veröffentlicht: (2023)
SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
von: As, Yarden, et al.
Veröffentlicht: (2025)
von: As, Yarden, et al.
Veröffentlicht: (2025)
Zero-shot Object Navigation with Vision-Language Models Reasoning
von: Wen, Congcong, et al.
Veröffentlicht: (2024)
von: Wen, Congcong, et al.
Veröffentlicht: (2024)
Neural Algorithmic Reasoners informed Large Language Model for Multi-Agent Path Finding
von: Feng, Pu, et al.
Veröffentlicht: (2025)
von: Feng, Pu, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Decentralized Multi-Robot Exploration With Macro Actions
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2021)
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2021)
Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots
von: Varela, Iñaki Dellibarda, et al.
Veröffentlicht: (2025)
von: Varela, Iñaki Dellibarda, et al.
Veröffentlicht: (2025)
OpenNav: Open-World Navigation with Multimodal Large Language Models
von: Yuan, Mingfeng, et al.
Veröffentlicht: (2025)
von: Yuan, Mingfeng, et al.
Veröffentlicht: (2025)
Trajectory Adaptation using Large Language Models
von: Maurya, Anurag, et al.
Veröffentlicht: (2025)
von: Maurya, Anurag, et al.
Veröffentlicht: (2025)
Realistic Corner Case Generation for Autonomous Vehicles with Multimodal Large Language Model
von: Lu, Qiujing, et al.
Veröffentlicht: (2024)
von: Lu, Qiujing, et al.
Veröffentlicht: (2024)
Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2023)
von: Dorbala, Vishnu Sashank, et al.
Veröffentlicht: (2023)
DexGrasp-Zero: A Morphology-Aligned Policy for Zero-Shot Cross-Embodiment Dexterous Grasping
von: Wu, Yuliang, et al.
Veröffentlicht: (2026)
von: Wu, Yuliang, et al.
Veröffentlicht: (2026)
Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
von: Liu, Jinyi, et al.
Veröffentlicht: (2024)
von: Liu, Jinyi, et al.
Veröffentlicht: (2024)
Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition
von: Quartey, Benedict, et al.
Veröffentlicht: (2026)
von: Quartey, Benedict, et al.
Veröffentlicht: (2026)
SemNav: A Model-Based Planner for Zero-Shot Object Goal Navigation Using Vision-Foundation Models
von: Debnath, Arnab, et al.
Veröffentlicht: (2025)
von: Debnath, Arnab, et al.
Veröffentlicht: (2025)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
COMRES-VLM: Coordinated Multi-Robot Exploration and Search using Vision Language Models
von: Wang, Ruiyang, et al.
Veröffentlicht: (2025)
von: Wang, Ruiyang, et al.
Veröffentlicht: (2025)
Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
von: Liu, Yuchen, et al.
Veröffentlicht: (2025)
von: Liu, Yuchen, et al.
Veröffentlicht: (2025)
Plan-and-Act using Large Language Models for Interactive Agreement
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
von: Sasabuchi, Kazuhiro, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
von: Fung, Angus, et al.
Veröffentlicht: (2024) -
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
von: Fung, Angus, et al.
Veröffentlicht: (2022) -
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
von: Lisondra, Matthew, et al.
Veröffentlicht: (2025) -
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
von: Tan, Aaron Hao, et al.
Veröffentlicht: (2025) -
X-Nav: Learning End-to-End Cross-Embodiment Navigation for Mobile Robots
von: Wang, Haitong, et al.
Veröffentlicht: (2025)