MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Junpeng, Xu, Xinrun, Karlsson, Börje F., Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025)
by: Yuan, Haoqi, et al.
Published: (2025)
SELU: Self-Learning Embodied MLLMs in Unknown Environments
by: Li, Boyu, et al.
Published: (2024)
by: Li, Boyu, et al.
Published: (2024)
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025)
by: Liu, Jiazheng, et al.
Published: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026)
by: Li, Boyu, et al.
Published: (2026)
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
ICPRL: Acquiring Physical Intuition from Interactive Control
by: Xu, Xinrun, et al.
Published: (2026)
by: Xu, Xinrun, et al.
Published: (2026)
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
by: Yue, Junpeng, et al.
Published: (2025)
by: Yue, Junpeng, et al.
Published: (2025)
Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
by: Li, Boyu, et al.
Published: (2025)
by: Li, Boyu, et al.
Published: (2025)
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation
by: Xie, Quanting, et al.
Published: (2024)
by: Xie, Quanting, et al.
Published: (2024)
Fully Decentralized Cooperative Multi-Agent Reinforcement Learning: A Survey
by: Jiang, Jiechuan, et al.
Published: (2024)
by: Jiang, Jiechuan, et al.
Published: (2024)
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
by: Chen, Zhanpeng, et al.
Published: (2024)
by: Chen, Zhanpeng, et al.
Published: (2024)
From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Settling Decentralized Multi-Agent Coordinated Exploration by Novelty Sharing
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning
by: Liu, Jiajin, et al.
Published: (2025)
by: Liu, Jiajin, et al.
Published: (2025)
Creative Agents: Empowering Agents with Imagination for Creative Tasks
by: Cai, Penglin, et al.
Published: (2023)
by: Cai, Penglin, et al.
Published: (2023)
Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation
by: Löwenmark, Karl, et al.
Published: (2025)
by: Löwenmark, Karl, et al.
Published: (2025)
From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
by: Szot, Andrew, et al.
Published: (2024)
by: Szot, Andrew, et al.
Published: (2024)
Tackling Non-Stationarity in Reinforcement Learning via Causal-Origin Representation
by: Zhang, Wanpeng, et al.
Published: (2023)
by: Zhang, Wanpeng, et al.
Published: (2023)
Generalized Contrastive Learning for Universal Multimodal Retrieval
by: Lee, Jungsoo, et al.
Published: (2025)
by: Lee, Jungsoo, et al.
Published: (2025)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
by: Xu, Binqian, et al.
Published: (2024)
by: Xu, Binqian, et al.
Published: (2024)
Cross-Embodiment Dexterous Grasping with Reinforcement Learning
by: Yuan, Haoqi, et al.
Published: (2024)
by: Yuan, Haoqi, et al.
Published: (2024)
Mildly Conservative Q-Learning for Offline Reinforcement Learning
by: Lyu, Jiafei, et al.
Published: (2022)
by: Lyu, Jiafei, et al.
Published: (2022)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
by: Tan, Weihao, et al.
Published: (2024)
by: Tan, Weihao, et al.
Published: (2024)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping
by: Huang, Ziye, et al.
Published: (2024)
by: Huang, Ziye, et al.
Published: (2024)
Learning Diverse Bimanual Dexterous Manipulation Skills from Human Demonstrations
by: Zhou, Bohan, et al.
Published: (2024)
by: Zhou, Bohan, et al.
Published: (2024)
ChemMLLM: Chemical Multimodal Large Language Model
by: Tan, Qian, et al.
Published: (2025)
by: Tan, Qian, et al.
Published: (2025)
RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
by: Kagaya, Tomoyuki, et al.
Published: (2024)
by: Kagaya, Tomoyuki, et al.
Published: (2024)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
by: Yang, Zeyu, et al.
Published: (2026)
by: Yang, Zeyu, et al.
Published: (2026)
Bottleneck Tokens for Unified Multimodal Retrieval
by: Sun, Siyu, et al.
Published: (2026)
by: Sun, Siyu, et al.
Published: (2026)
Retrieval-Augmented Multimodal Depression Detection
by: Hou, Ruibo, et al.
Published: (2025)
by: Hou, Ruibo, et al.
Published: (2025)
Representation Retrieval Learning for Heterogeneous Data Integration
by: Xu, Qi, et al.
Published: (2025)
by: Xu, Qi, et al.
Published: (2025)
Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
by: Xu, Zerui, et al.
Published: (2024)
by: Xu, Zerui, et al.
Published: (2024)
MINER: Mining Multimodal Internal Representation for Efficient Retrieval
by: Li, Weien, et al.
Published: (2026)
by: Li, Weien, et al.
Published: (2026)
Locality Preserving Markovian Transition for Instance Retrieval
by: Luo, Jifei, et al.
Published: (2025)
by: Luo, Jifei, et al.
Published: (2025)
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
by: Lin, Jieru, et al.
Published: (2025)
by: Lin, Jieru, et al.
Published: (2025)
Spike Hijacking in Late-Interaction Retrieval
by: Suresh, Karthik, et al.
Published: (2026)
by: Suresh, Karthik, et al.
Published: (2026)
Event-Centric World Modeling with Memory-Augmented Retrieval for Embodied Decision-Making
by: Fan, Zhaowen, et al.
Published: (2026)
by: Fan, Zhaowen, et al.
Published: (2026)
Efficient Table Retrieval and Understanding with Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2026)
by: Xu, Zhuoyan, et al.
Published: (2026)
Understanding What Affects the Generalization Gap in Visual Reinforcement Learning: Theory and Empirical Evidence
by: Lyu, Jiafei, et al.
Published: (2024)
by: Lyu, Jiafei, et al.
Published: (2024)
Similar Items
-
Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills
by: Yuan, Haoqi, et al.
Published: (2025) -
SELU: Self-Learning Embodied MLLMs in Unknown Environments
by: Li, Boyu, et al.
Published: (2024) -
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning
by: Liu, Jiazheng, et al.
Published: (2025) -
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
by: Li, Boyu, et al.
Published: (2026) -
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)