Static and Plugged: Make Embodied Evaluation Simple
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916889683296256 |
|---|---|
| author | Xiao, Jiahao Zhang, Jianbo Yan, BoWen Guo, Shengyu Ye, Tongrui Zhang, Kaiwei Zhang, Zicheng Liu, Xiaohong Cheng, Zhengxue Fan, Lei Li, Chuyi Zhai, Guangtao |
| author_facet | Xiao, Jiahao Zhang, Jianbo Yan, BoWen Guo, Shengyu Ye, Tongrui Zhang, Kaiwei Zhang, Zicheng Liu, Xiaohong Cheng, Zhengxue Fan, Lei Li, Chuyi Zhai, Guangtao |
| contents | Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06553 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Static and Plugged: Make Embodied Evaluation Simple Xiao, Jiahao Zhang, Jianbo Yan, BoWen Guo, Shengyu Ye, Tongrui Zhang, Kaiwei Zhang, Zicheng Liu, Xiaohong Cheng, Zhengxue Fan, Lei Li, Chuyi Zhai, Guangtao Computer Vision and Pattern Recognition Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence. |
| title | Static and Plugged: Make Embodied Evaluation Simple |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.06553 |