Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nwankwo, Linus, Rueckert, Elmar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Conversation is the Command: Interacting with Real-World Autonomous Robot Through Natural Language
by: Nwankwo, Linus, et al.
Published: (2024)
by: Nwankwo, Linus, et al.
Published: (2024)
SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
by: Nwankwo, Linus, et al.
Published: (2025)
by: Nwankwo, Linus, et al.
Published: (2025)
Understanding why SLAM algorithms fail in modern indoor environments
by: Linus, Nwankwo, et al.
Published: (2023)
by: Linus, Nwankwo, et al.
Published: (2023)
ReLI: A Language-Agnostic Approach to Human-Robot Interaction
by: Nwankwo, Linus, et al.
Published: (2025)
by: Nwankwo, Linus, et al.
Published: (2025)
Real-Time 3D Vision-Language Embedding Mapping
by: Rauch, Christian, et al.
Published: (2025)
by: Rauch, Christian, et al.
Published: (2025)
ROMR: A ROS-based Open-source Mobile Robot
by: Linus, Nwankwo, et al.
Published: (2022)
by: Linus, Nwankwo, et al.
Published: (2022)
Multimodal Visual-Tactile Representation Learning through Self-Supervised Contrastive Pre-Training
by: Dave, Vedant, et al.
Published: (2024)
by: Dave, Vedant, et al.
Published: (2024)
EnvoDat: A Large-Scale Multisensory Dataset for Robotic Spatial Awareness and Semantic Reasoning in Heterogeneous Environments
by: Nwankwo, Linus, et al.
Published: (2024)
by: Nwankwo, Linus, et al.
Published: (2024)
Integrating Human Expertise in Continuous Spaces: A Novel Interactive Bayesian Optimization Framework with Preference Expected Improvement
by: Feith, Nikolaus, et al.
Published: (2024)
by: Feith, Nikolaus, et al.
Published: (2024)
M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
by: Lygerakis, Fotios, et al.
Published: (2024)
by: Lygerakis, Fotios, et al.
Published: (2024)
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024)
by: Ryu, Chanhoe, et al.
Published: (2024)
GPT-Fabric: Smoothing and Folding Fabric by Leveraging Pre-Trained Foundation Models
by: Raval, Vedant, et al.
Published: (2024)
by: Raval, Vedant, et al.
Published: (2024)
An Interactive Agent Foundation Model
by: Durante, Zane, et al.
Published: (2024)
by: Durante, Zane, et al.
Published: (2024)
SInViG: A Self-Evolving Interactive Visual Agent for Human-Robot Interaction
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
Multi-Agent Planning Using Visual Language Models
by: Brienza, Michele, et al.
Published: (2024)
by: Brienza, Michele, et al.
Published: (2024)
Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction
by: Bu, Fanjun, et al.
Published: (2025)
by: Bu, Fanjun, et al.
Published: (2025)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
A Multimodal Framework for Human-Multi-Agent Interaction
by: Hasan, Shaid, et al.
Published: (2026)
by: Hasan, Shaid, et al.
Published: (2026)
Autonomous Algorithm for Training Autonomous Vehicles with Minimal Human Intervention
by: Lee, Sang-Hyun, et al.
Published: (2024)
by: Lee, Sang-Hyun, et al.
Published: (2024)
Autonomous Surface Selection For Manipulator-Based UV Disinfection In Hospitals Using Foundation Models
by: Oh, Xueyan, et al.
Published: (2025)
by: Oh, Xueyan, et al.
Published: (2025)
MAGIC-VFM: Meta-learning Adaptation for Ground Interaction Control with Visual Foundation Models
by: Lupu, Elena Sorina, et al.
Published: (2024)
by: Lupu, Elena Sorina, et al.
Published: (2024)
Large Language Model based Interactive Decision-Making for Autonomous Driving
by: Dong, Xinwei, et al.
Published: (2026)
by: Dong, Xinwei, et al.
Published: (2026)
Multimodal Safe Control for Human-Robot Interaction
by: Pandya, Ravi, et al.
Published: (2023)
by: Pandya, Ravi, et al.
Published: (2023)
Unidirectional Human-Robot-Human Physical Interaction for Gait Training
by: Amato, Lorenzo, et al.
Published: (2024)
by: Amato, Lorenzo, et al.
Published: (2024)
Human-Robot Mutual Learning through Affective-Linguistic Interaction and Differential Outcomes Training [Pre-Print]
by: Heikkinen, Emilia, et al.
Published: (2024)
by: Heikkinen, Emilia, et al.
Published: (2024)
DVRP-MHSI: Dynamic Visualization Research Platform for Multimodal Human-Swarm Interaction
by: Zhu, Pengming, et al.
Published: (2024)
by: Zhu, Pengming, et al.
Published: (2024)
Robotic Applications of Pre-Trained Vision-Language Models to Various Recognition Behaviors
by: Kawaharazuka, Kento, et al.
Published: (2023)
by: Kawaharazuka, Kento, et al.
Published: (2023)
Reasoning Multi-Agent Behavioral Topology for Interactive Autonomous Driving
by: Liu, Haochen, et al.
Published: (2024)
by: Liu, Haochen, et al.
Published: (2024)
Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model
by: Lai, Yuzhi, et al.
Published: (2025)
by: Lai, Yuzhi, et al.
Published: (2025)
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
Autonomously Unweaving Multiple Cables Using Visual Feedback
by: Tian, Tina, et al.
Published: (2025)
by: Tian, Tina, et al.
Published: (2025)
VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving
by: Long, Keke, et al.
Published: (2024)
by: Long, Keke, et al.
Published: (2024)
Unifying Deep Predicate Invention with Pre-trained Foundation Models
by: Wang, Qianwei, et al.
Published: (2025)
by: Wang, Qianwei, et al.
Published: (2025)
Where to Fetch: Extracting Visual Scene Representation from Large Pre-Trained Models for Robotic Goal Navigation
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Learning-Based Modeling of Human-Autonomous Vehicle Interaction for Improved Safety in Mixed-Vehicle Platooning Control
by: Wang, Jie, et al.
Published: (2023)
by: Wang, Jie, et al.
Published: (2023)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
by: He, Linfeng, et al.
Published: (2024)
by: He, Linfeng, et al.
Published: (2024)
Autonomous Human-Robot Interaction via Operator Imitation
by: Christen, Sammy, et al.
Published: (2025)
by: Christen, Sammy, et al.
Published: (2025)
A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving
by: Zhang, Liangdong, et al.
Published: (2026)
by: Zhang, Liangdong, et al.
Published: (2026)
SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation
by: Neubauer, Melanie, et al.
Published: (2026)
by: Neubauer, Melanie, et al.
Published: (2026)
Benchmarking Autonomous Vehicles: A Driver Foundation Model Framework
by: Zhang, Yuxin, et al.
Published: (2026)
by: Zhang, Yuxin, et al.
Published: (2026)
Similar Items
-
The Conversation is the Command: Interacting with Real-World Autonomous Robot Through Natural Language
by: Nwankwo, Linus, et al.
Published: (2024) -
SIL: Symbiotic Interactive Learning for Language-Conditioned Human-Agent Co-Adaptation
by: Nwankwo, Linus, et al.
Published: (2025) -
Understanding why SLAM algorithms fail in modern indoor environments
by: Linus, Nwankwo, et al.
Published: (2023) -
ReLI: A Language-Agnostic Approach to Human-Robot Interaction
by: Nwankwo, Linus, et al.
Published: (2025) -
Real-Time 3D Vision-Language Embedding Mapping
by: Rauch, Christian, et al.
Published: (2025)