RynnEC: Bringing MLLMs into Embodied World

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Ronghao, Yuan, Yuqian, Mao, Yunxuan, Li, Kehan, Liu, Jiangpin, Wang, Zhikai, Li, Xin, Wang, Fan, Zhao, Deli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915625197109248
author Dang, Ronghao
Yuan, Yuqian
Mao, Yunxuan
Li, Kehan
Liu, Jiangpin
Wang, Zhikai
Li, Xin
Wang, Fan
Zhao, Deli
author_facet Dang, Ronghao
Yuan, Yuqian
Mao, Yunxuan
Li, Kehan
Liu, Jiangpin
Wang, Zhikai
Li, Xin
Wang, Fan
Zhao, Deli
contents We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible region-level video interaction. Despite its compact architecture, RynnEC achieves state-of-the-art performance in object property understanding, object segmentation, and spatial reasoning. Conceptually, it offers a region-centric video paradigm for the brain of embodied agents, providing fine-grained perception of the physical world and enabling more precise interactions. To mitigate the scarcity of annotated 3D datasets, we propose an egocentric video based pipeline for generating embodied cognition data. Furthermore, we introduce RynnEC-Bench, a region-centered benchmark for evaluating embodied cognitive capabilities. We anticipate that RynnEC will advance the development of general-purpose cognitive cores for embodied agents and facilitate generalization across diverse embodied tasks. The code, model checkpoints, and benchmark are available at: https://github.com/alibaba-damo-academy/RynnEC
format Preprint
id arxiv_https___arxiv_org_abs_2508_14160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RynnEC: Bringing MLLMs into Embodied World
Dang, Ronghao
Yuan, Yuqian
Mao, Yunxuan
Li, Kehan
Liu, Jiangpin
Wang, Zhikai
Li, Xin
Wang, Fan
Zhao, Deli
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible region-level video interaction. Despite its compact architecture, RynnEC achieves state-of-the-art performance in object property understanding, object segmentation, and spatial reasoning. Conceptually, it offers a region-centric video paradigm for the brain of embodied agents, providing fine-grained perception of the physical world and enabling more precise interactions. To mitigate the scarcity of annotated 3D datasets, we propose an egocentric video based pipeline for generating embodied cognition data. Furthermore, we introduce RynnEC-Bench, a region-centered benchmark for evaluating embodied cognitive capabilities. We anticipate that RynnEC will advance the development of general-purpose cognitive cores for embodied agents and facilitate generalization across diverse embodied tasks. The code, model checkpoints, and benchmark are available at: https://github.com/alibaba-damo-academy/RynnEC
title RynnEC: Bringing MLLMs into Embodied World
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2508.14160