NuPlanQA: A Large-Scale Dataset and Benchmark for Multi-View Driving Scene Understanding in Multi-Modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Sung-Yeon, Cui, Can, Ma, Yunsheng, Moradipari, Ahmadreza, Gupta, Rohit, Han, Kyungtae, Wang, Ziran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
by: Park, Sung-Yeon, et al.
Published: (2025)
by: Park, Sung-Yeon, et al.
Published: (2025)
On Learning Closed-Loop Probabilistic Multi-Agent Simulator
by: Lu, Juanwu, et al.
Published: (2025)
by: Lu, Juanwu, et al.
Published: (2025)
Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
by: Ma, Yunsheng, et al.
Published: (2023)
by: Ma, Yunsheng, et al.
Published: (2023)
LLM4AD: Large Language Models for Autonomous Driving -- Concept, Review, Benchmark, Experiments, and Future Trends
by: Cui, Can, et al.
Published: (2024)
by: Cui, Can, et al.
Published: (2024)
Agentic AI for Trip Planning Optimization Application
by: Chen, Tiejin, et al.
Published: (2026)
by: Chen, Tiejin, et al.
Published: (2026)
Multi-Agent Stage-wise Conservative Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2025)
by: Afsharrad, Amirhossein, et al.
Published: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
by: Qian, Tianwen, et al.
Published: (2023)
by: Qian, Tianwen, et al.
Published: (2023)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
Cooperative Multi-Agent Constrained Stochastic Linear Bandits
by: Afsharrad, Amirhossein, et al.
Published: (2024)
by: Afsharrad, Amirhossein, et al.
Published: (2024)
ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
by: Cui, Can, et al.
Published: (2025)
by: Cui, Can, et al.
Published: (2025)
On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
by: Afsharrad, Amirhossein, et al.
Published: (2026)
by: Afsharrad, Amirhossein, et al.
Published: (2026)
Personalized Autonomous Driving with Large Language Models: Field Experiments
by: Cui, Can, et al.
Published: (2023)
by: Cui, Can, et al.
Published: (2023)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
by: Ma, Yunsheng, et al.
Published: (2022)
by: Ma, Yunsheng, et al.
Published: (2022)
Network-Efficient World Model Token Streaming
by: Mishra, Shatadal, et al.
Published: (2026)
by: Mishra, Shatadal, et al.
Published: (2026)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024)
by: Ding, Xinpeng, et al.
Published: (2024)
PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
by: Han, Kyungtae, et al.
Published: (2025)
by: Han, Kyungtae, et al.
Published: (2025)
Formation and Investigation of Cooperative Platooning at the Early Stage of Connected and Automated Vehicles Deployment
by: Mu, Zeyu, et al.
Published: (2025)
by: Mu, Zeyu, et al.
Published: (2025)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
by: Cui, Erfei, et al.
Published: (2023)
by: Cui, Erfei, et al.
Published: (2023)
PDB: Not All Drivers Are the Same -- A Personalized Dataset for Understanding Driving Behavior
by: Wei, Chuheng, et al.
Published: (2025)
by: Wei, Chuheng, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
OT-MeanFlow3D: Bridging Optimal Transport and Meanflow for Efficient 3D Point Cloud Generation
by: Akbari, Elaheh, et al.
Published: (2025)
by: Akbari, Elaheh, et al.
Published: (2025)
Quantifying Uncertainty in Motion Prediction with Variational Bayesian Mixture
by: Lu, Juanwu, et al.
Published: (2024)
by: Lu, Juanwu, et al.
Published: (2024)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
by: Mo, Wentao, et al.
Published: (2025)
by: Mo, Wentao, et al.
Published: (2025)
ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
by: Choi, Tae-Min, et al.
Published: (2025)
by: Choi, Tae-Min, et al.
Published: (2025)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
$\mathtt{M^3VIR}$: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation
by: Li, Yuanzhi, et al.
Published: (2025)
by: Li, Yuanzhi, et al.
Published: (2025)
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
by: Moradipari, Ahmadreza, et al.
Published: (2023)
by: Moradipari, Ahmadreza, et al.
Published: (2023)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
by: Li, Chuhan, et al.
Published: (2024)
by: Li, Chuhan, et al.
Published: (2024)
V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views
by: You, Junwei, et al.
Published: (2026)
by: You, Junwei, et al.
Published: (2026)
WayveScenes101: A Dataset and Benchmark for Novel View Synthesis in Autonomous Driving
by: Zürn, Jannik, et al.
Published: (2024)
by: Zürn, Jannik, et al.
Published: (2024)
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2025)
by: Ma, Yunsheng, et al.
Published: (2025)
Similar Items
-
SIMSplat: Predictive Driving Scene Editing with Language-aligned 4D Gaussian Splatting
by: Park, Sung-Yeon, et al.
Published: (2025) -
On Learning Closed-Loop Probabilistic Multi-Agent Simulator
by: Lu, Juanwu, et al.
Published: (2025) -
Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2024) -
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
by: Ma, Yunsheng, et al.
Published: (2023) -
LLM4AD: Large Language Models for Autonomous Driving -- Concept, Review, Benchmark, Experiments, and Future Trends
by: Cui, Can, et al.
Published: (2024)