Static and Plugged: Make Embodied Evaluation Simple

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Jiahao, Zhang, Jianbo, Yan, BoWen, Guo, Shengyu, Ye, Tongrui, Zhang, Kaiwei, Zhang, Zicheng, Liu, Xiaohong, Cheng, Zhengxue, Fan, Lei, Li, Chuyi, Zhai, Guangtao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916889683296256
author Xiao, Jiahao
Zhang, Jianbo
Yan, BoWen
Guo, Shengyu
Ye, Tongrui
Zhang, Kaiwei
Zhang, Zicheng
Liu, Xiaohong
Cheng, Zhengxue
Fan, Lei
Li, Chuyi
Zhai, Guangtao
author_facet Xiao, Jiahao
Zhang, Jianbo
Yan, BoWen
Guo, Shengyu
Ye, Tongrui
Zhang, Kaiwei
Zhang, Zicheng
Liu, Xiaohong
Cheng, Zhengxue
Fan, Lei
Li, Chuyi
Zhai, Guangtao
contents Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Static and Plugged: Make Embodied Evaluation Simple
Xiao, Jiahao
Zhang, Jianbo
Yan, BoWen
Guo, Shengyu
Ye, Tongrui
Zhang, Kaiwei
Zhang, Zicheng
Liu, Xiaohong
Cheng, Zhengxue
Fan, Lei
Li, Chuyi
Zhai, Guangtao
Computer Vision and Pattern Recognition
Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.
title Static and Plugged: Make Embodied Evaluation Simple
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06553