Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915945946021888 |
|---|---|
| author | Cao, Tianshi Ren, Jiawei Zhang, Yuxuan Seo, Jaewoo Huang, Jiahui Solanki, Shikhar Zhang, Haotian Guo, Mingfei Turki, Haithem Li, Muxingzi Zhu, Yue Zhang, Sipeng Gojcic, Zan Fidler, Sanja Yin, Kangxue |
| author_facet | Cao, Tianshi Ren, Jiawei Zhang, Yuxuan Seo, Jaewoo Huang, Jiahui Solanki, Shikhar Zhang, Haotian Guo, Mingfei Turki, Haithem Li, Muxingzi Zhu, Yue Zhang, Sipeng Gojcic, Zan Fidler, Sanja Yin, Kangxue |
| contents | Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for simulation, but it does not produce complete 3D object assets required for agent manipulation and large-viewpoint novel-view synthesis. To address this challenge, we present Asset Harvester, an image-to-3D model and end-to-end pipeline that converts sparse, in-the-wild object observations from real driving logs into complete, simulation-ready assets. Rather than relying on a single model component, we developed a system-level design for real-world AV data that combines large-scale curation of object-centric training tuples, geometry-aware preprocessing across heterogeneous sensors, and a robust training recipe that couples sparse-view-conditioned multiview generation with 3D Gaussian lifting. Within this system, SparseViewDiT is explicitly designed to address limited-angle views and other real-world data challenges. Together with hybrid data curation, augmentation, and self-distillation, this system enables scalable conversion of sparse AV object observations into reusable 3D assets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_18468 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation Cao, Tianshi Ren, Jiawei Zhang, Yuxuan Seo, Jaewoo Huang, Jiahui Solanki, Shikhar Zhang, Haotian Guo, Mingfei Turki, Haithem Li, Muxingzi Zhu, Yue Zhang, Sipeng Gojcic, Zan Fidler, Sanja Yin, Kangxue Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before real-world deployment. Neural scene reconstruction converts driving logs into interactive 3D environments for simulation, but it does not produce complete 3D object assets required for agent manipulation and large-viewpoint novel-view synthesis. To address this challenge, we present Asset Harvester, an image-to-3D model and end-to-end pipeline that converts sparse, in-the-wild object observations from real driving logs into complete, simulation-ready assets. Rather than relying on a single model component, we developed a system-level design for real-world AV data that combines large-scale curation of object-centric training tuples, geometry-aware preprocessing across heterogeneous sensors, and a robust training recipe that couples sparse-view-conditioned multiview generation with 3D Gaussian lifting. Within this system, SparseViewDiT is explicitly designed to address limited-angle views and other real-world data challenges. Together with hybrid data curation, augmentation, and self-distillation, this system enables scalable conversion of sparse AV object observations into reusable 3D assets. |
| title | Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics Machine Learning |
| url | https://arxiv.org/abs/2604.18468 |