Nimbus: A Unified Embodied Synthetic Data Generation Framework
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910016212041728 |
|---|---|
| author | He, Zeyu Zhang, Yuchang Zhou, Yuanzhen Tao, Miao Li, Hengjie Wang, Hui Tian, Yang Zeng, Jia Wang, Tai Cai, Wenzhe Chen, Yilun Gao, Ning Pang, Jiangmiao |
| author_facet | He, Zeyu Zhang, Yuchang Zhou, Yuanzhen Tao, Miao Li, Hengjie Wang, Hui Tian, Yang Zeng, Jia Wang, Tai Cai, Wenzhe Chen, Yilun Gao, Ning Pang, Jiangmiao |
| contents | Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, existing pipelines remain fragmented and task-specific. This isolation leads to significant engineering inefficiency and system instability, failing to support the sustained, high-throughput data generation required for foundation model training. To address these challenges, we present Nimbus, a unified synthetic data generation framework designed to integrate heterogeneous navigation and manipulation pipelines. Nimbus introduces a modular four-layer architecture featuring a decoupled execution model that separates trajectory planning, rendering, and storage into asynchronous stages. By implementing dynamic pipeline scheduling, global load balancing, distributed fault tolerance, and backend-specific rendering optimizations, the system maximizes resource utilization across CPU, GPU, and I/O resources. Our evaluation demonstrates that Nimbus achieves a 2-3X improvement in end-to-end throughput compared to unoptimized baselines and ensuring robust, long-term operation in large-scale distributed environments. This framework serves as the production backbone for the InternData suite, enabling seamless cross-domain data synthesis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_21449 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Nimbus: A Unified Embodied Synthetic Data Generation Framework He, Zeyu Zhang, Yuchang Zhou, Yuanzhen Tao, Miao Li, Hengjie Wang, Hui Tian, Yang Zeng, Jia Wang, Tai Cai, Wenzhe Chen, Yilun Gao, Ning Pang, Jiangmiao Robotics Distributed, Parallel, and Cluster Computing Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, existing pipelines remain fragmented and task-specific. This isolation leads to significant engineering inefficiency and system instability, failing to support the sustained, high-throughput data generation required for foundation model training. To address these challenges, we present Nimbus, a unified synthetic data generation framework designed to integrate heterogeneous navigation and manipulation pipelines. Nimbus introduces a modular four-layer architecture featuring a decoupled execution model that separates trajectory planning, rendering, and storage into asynchronous stages. By implementing dynamic pipeline scheduling, global load balancing, distributed fault tolerance, and backend-specific rendering optimizations, the system maximizes resource utilization across CPU, GPU, and I/O resources. Our evaluation demonstrates that Nimbus achieves a 2-3X improvement in end-to-end throughput compared to unoptimized baselines and ensuring robust, long-term operation in large-scale distributed environments. This framework serves as the production backbone for the InternData suite, enabling seamless cross-domain data synthesis. |
| title | Nimbus: A Unified Embodied Synthetic Data Generation Framework |
| topic | Robotics Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2601.21449 |