SELU: Self-Learning Embodied MLLMs in Unknown Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Boyu, Jiang, Haobin, Ding, Ziluo, Xu, Xinrun, Li, Haoran, Zhao, Dongbin, Lu, Zongqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023)
by: Jiang, Haobin, et al.
Published: (2023)
Settling Decentralized Multi-Agent Coordinated Exploration by Novelty Sharing
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
Egocentric Vision Language Planning
by: Fang, Zhirui, et al.
Published: (2024)
by: Fang, Zhirui, et al.
Published: (2024)
A Survey on Deep Clustering: From the Prior Perspective
by: Lu, Yiding, et al.
Published: (2024)
by: Lu, Yiding, et al.
Published: (2024)
From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
Dream to Drive with Predictive Individual World Model
by: Gao, Yinfeng, et al.
Published: (2025)
by: Gao, Yinfeng, et al.
Published: (2025)
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
by: Guo, Haiyang, et al.
Published: (2025)
by: Guo, Haiyang, et al.
Published: (2025)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
by: Zhao, Yanpeng, et al.
Published: (2026)
by: Zhao, Yanpeng, et al.
Published: (2026)
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
by: Xu, Xinrun, et al.
Published: (2025)
by: Xu, Xinrun, et al.
Published: (2025)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Federated Joint Learning for Domain and Class Generalization
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era
by: Hong, Qiuhe, et al.
Published: (2026)
by: Hong, Qiuhe, et al.
Published: (2026)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
by: Peng, Yibo, et al.
Published: (2026)
by: Peng, Yibo, et al.
Published: (2026)
Growing Visual Generative Capacity for Pre-Trained MLLMs
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
by: Yue, Junpeng, et al.
Published: (2024)
by: Yue, Junpeng, et al.
Published: (2024)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
by: Du, Yao, et al.
Published: (2026)
by: Du, Yao, et al.
Published: (2026)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Reliable Thinking with Images
by: Li, Haobin, et al.
Published: (2026)
by: Li, Haobin, et al.
Published: (2026)
Scaling Large Motion Models with Million-Level Human Motions
by: Wang, Ye, et al.
Published: (2024)
by: Wang, Ye, et al.
Published: (2024)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
by: Feng, Chen, et al.
Published: (2021)
by: Feng, Chen, et al.
Published: (2021)
Three Creates All: You Only Sample 3 Steps
by: Cai, Yuren, et al.
Published: (2026)
by: Cai, Yuren, et al.
Published: (2026)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Confidence Self-Calibration for Multi-Label Class-Incremental Learning
by: Du, Kaile, et al.
Published: (2024)
by: Du, Kaile, et al.
Published: (2024)
BECAME: BayEsian Continual Learning with Adaptive Model MErging
by: Li, Mei, et al.
Published: (2025)
by: Li, Mei, et al.
Published: (2025)
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
by: Yang, Jingkang, et al.
Published: (2023)
by: Yang, Jingkang, et al.
Published: (2023)
From Isolation to Integration: Building an Adaptive Expert Forest for Pre-Trained Model-based Class-Incremental Learning
by: Liu, Ruiqi, et al.
Published: (2026)
by: Liu, Ruiqi, et al.
Published: (2026)
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning
by: Jiang, Huchen, et al.
Published: (2024)
by: Jiang, Huchen, et al.
Published: (2024)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Revisiting Unknowns: Towards Effective and Efficient Open-Set Active Learning
by: Zong, Chen-Chen, et al.
Published: (2026)
by: Zong, Chen-Chen, et al.
Published: (2026)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
by: Chen, Haoran, et al.
Published: (2024)
by: Chen, Haoran, et al.
Published: (2024)
EmbodiTTA: Resource-Efficient Test-Time Adaptation for Embodied Visual Systems
by: Ma, Xiao, et al.
Published: (2025)
by: Ma, Xiao, et al.
Published: (2025)
HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Similar Items
-
Reinforcement Learning Friendly Vision-Language Model for Minecraft
by: Jiang, Haobin, et al.
Published: (2023) -
Settling Decentralized Multi-Agent Coordinated Exploration by Novelty Sharing
by: Jiang, Haobin, et al.
Published: (2024) -
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024) -
Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
by: Li, Haoran, et al.
Published: (2024) -
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)