Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Yuxi, Wang, Zi, Lu, Yifan, Xu, Chenxin, Liu, Changxing, Zhao, Hao, Chen, Siheng, Wang, Yanfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914848991870976
author Wei, Yuxi
Wang, Zi
Lu, Yifan
Xu, Chenxin
Liu, Changxing
Zhao, Hao
Chen, Siheng
Wang, Yanfeng
author_facet Wei, Yuxi
Wang, Zi
Lu, Yifan
Xu, Chenxin
Liu, Changxing
Zhao, Hao
Chen, Siheng
Wang, Yanfeng
contents Scene simulation in autonomous driving has gained significant attention because of its huge potential for generating customized data. However, existing editable scene simulation approaches face limitations in terms of user interaction efficiency, multi-camera photo-realistic rendering and external digital assets integration. To address these challenges, this paper introduces ChatSim, the first system that enables editable photo-realistic 3D driving scene simulations via natural language commands with external digital assets. To enable editing with high command flexibility,~ChatSim leverages a large language model (LLM) agent collaboration framework. To generate photo-realistic outcomes, ChatSim employs a novel multi-camera neural radiance field method. Furthermore, to unleash the potential of extensive high-quality digital assets, ChatSim employs a novel multi-camera lighting estimation method to achieve scene-consistent assets' rendering. Our experiments on Waymo Open Dataset demonstrate that ChatSim can handle complex language commands and generate corresponding photo-realistic scene videos.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05746
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
Wei, Yuxi
Wang, Zi
Lu, Yifan
Xu, Chenxin
Liu, Changxing
Zhao, Hao
Chen, Siheng
Wang, Yanfeng
Computer Vision and Pattern Recognition
Scene simulation in autonomous driving has gained significant attention because of its huge potential for generating customized data. However, existing editable scene simulation approaches face limitations in terms of user interaction efficiency, multi-camera photo-realistic rendering and external digital assets integration. To address these challenges, this paper introduces ChatSim, the first system that enables editable photo-realistic 3D driving scene simulations via natural language commands with external digital assets. To enable editing with high command flexibility,~ChatSim leverages a large language model (LLM) agent collaboration framework. To generate photo-realistic outcomes, ChatSim employs a novel multi-camera neural radiance field method. Furthermore, to unleash the potential of extensive high-quality digital assets, ChatSim employs a novel multi-camera lighting estimation method to achieve scene-consistent assets' rendering. Our experiments on Waymo Open Dataset demonstrate that ChatSim can handle complex language commands and generate corresponding photo-realistic scene videos.
title Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.05746