BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ge, Yunhao, Tang, Yihe, Xu, Jiashu, Gokmen, Cem, Li, Chengshu, Ai, Wensi, Martinez, Benjamin Jose, Aydin, Arman, Anvari, Mona, Chakravarthy, Ayush K, Yu, Hong-Xing, Wong, Josiah, Srivastava, Sanjana, Lee, Sharon, Zha, Shengxin, Itti, Laurent, Li, Yunzhu, Martín-Martín, Roberto, Liu, Miao, Zhang, Pengchuan, Zhang, Ruohan, Fei-Fei, Li, Wu, Jiajun
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910448561946624
author Ge, Yunhao
Tang, Yihe
Xu, Jiashu
Gokmen, Cem
Li, Chengshu
Ai, Wensi
Martinez, Benjamin Jose
Aydin, Arman
Anvari, Mona
Chakravarthy, Ayush K
Yu, Hong-Xing
Wong, Josiah
Srivastava, Sanjana
Lee, Sharon
Zha, Shengxin
Itti, Laurent
Li, Yunzhu
Martín-Martín, Roberto
Liu, Miao
Zhang, Pengchuan
Zhang, Ruohan
Fei-Fei, Li
Wu, Jiajun
author_facet Ge, Yunhao
Tang, Yihe
Xu, Jiashu
Gokmen, Cem
Li, Chengshu
Ai, Wensi
Martinez, Benjamin Jose
Aydin, Arman
Anvari, Mona
Chakravarthy, Ayush K
Yu, Hong-Xing
Wong, Josiah
Srivastava, Sanjana
Lee, Sharon
Zha, Shengxin
Itti, Laurent
Li, Yunzhu
Martín-Martín, Roberto
Liu, Miao
Zhang, Pengchuan
Zhang, Ruohan
Fei-Fei, Li
Wu, Jiajun
contents The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as "filled" and "folded"), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2405_09546
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation
Ge, Yunhao
Tang, Yihe
Xu, Jiashu
Gokmen, Cem
Li, Chengshu
Ai, Wensi
Martinez, Benjamin Jose
Aydin, Arman
Anvari, Mona
Chakravarthy, Ayush K
Yu, Hong-Xing
Wong, Josiah
Srivastava, Sanjana
Lee, Sharon
Zha, Shengxin
Itti, Laurent
Li, Yunzhu
Martín-Martín, Roberto
Liu, Miao
Zhang, Pengchuan
Zhang, Ruohan
Fei-Fei, Li
Wu, Jiajun
Computer Vision and Pattern Recognition
The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as "filled" and "folded"), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/
title BEHAVIOR Vision Suite: Customizable Dataset Generation via Simulation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.09546