Learning Traffic Crashes as Language: Datasets, Benchmarks, and What-if Causal Analyses
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Zhiwen, Wang, Pu, Zhao, Yang, Zhao, Yibo, Ivanovic, Boris, Wang, Zhangyang, Pavone, Marco, Yang, Hao Frank |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
InstantSplat: Sparse-view Gaussian Splatting in Seconds
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024)
by: Tan, Shuhan, et al.
Published: (2024)
Towards Reliable and Interpretable Traffic Crash Pattern Prediction and Safety Interventions Using Customized Large Language Models
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Parallelized Spatiotemporal Binding
by: Singh, Gautam, et al.
Published: (2024)
by: Singh, Gautam, et al.
Published: (2024)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
by: Ivanovic, Boris, et al.
Published: (2025)
by: Ivanovic, Boris, et al.
Published: (2025)
Distributed NeRF Learning for Collaborative Multi-Robot Perception
by: Zhao, Hongrui, et al.
Published: (2024)
by: Zhao, Hongrui, et al.
Published: (2024)
FSGS: Real-Time Few-shot View Synthesis using Gaussian Splatting
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
by: Ma, Yingzi, et al.
Published: (2025)
by: Ma, Yingzi, et al.
Published: (2025)
CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning
by: Gan, Rui, et al.
Published: (2026)
by: Gan, Rui, et al.
Published: (2026)
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
by: T, Mukund Varma, et al.
Published: (2024)
by: T, Mukund Varma, et al.
Published: (2024)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
by: Li, Longfei, et al.
Published: (2025)
by: Li, Longfei, et al.
Published: (2025)
Extrapolated Urban View Synthesis Benchmark
by: Han, Xiangyu, et al.
Published: (2024)
by: Han, Xiangyu, et al.
Published: (2024)
CrashChat: A Multimodal Large Language Model for Multitask Traffic Crash Video Analysis
by: Liang, Kaidi, et al.
Published: (2025)
by: Liang, Kaidi, et al.
Published: (2025)
SCB-Dataset3: A Benchmark for Detecting Student Classroom Behavior
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS
by: Fan, Zhiwen, et al.
Published: (2023)
by: Fan, Zhiwen, et al.
Published: (2023)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
by: Yang, Junjie, et al.
Published: (2023)
by: Yang, Junjie, et al.
Published: (2023)
SAVeD: A First-Person Social Media Video Dataset for ADAS-equipped vehicle Near-Miss and Crash Event Analyses
by: Zhai, Shaoyan, et al.
Published: (2025)
by: Zhai, Shaoyan, et al.
Published: (2025)
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
by: Zhao, Kaikai, et al.
Published: (2025)
by: Zhao, Kaikai, et al.
Published: (2025)
MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
by: Li, Renjie, et al.
Published: (2025)
by: Li, Renjie, et al.
Published: (2025)
CausalFSFG: Rethinking Few-Shot Fine-Grained Visual Categorization from Causal Perspective
by: Yang, Zhiwen, et al.
Published: (2025)
by: Yang, Zhiwen, et al.
Published: (2025)
DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
by: Wang, Letian, et al.
Published: (2024)
by: Wang, Letian, et al.
Published: (2024)
NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
by: Dauner, Daniel, et al.
Published: (2024)
by: Dauner, Daniel, et al.
Published: (2024)
Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
by: Wu, Guanlin, et al.
Published: (2025)
by: Wu, Guanlin, et al.
Published: (2025)
Traffic Sign Recognition in Autonomous Driving: Dataset, Benchmark, and Field Experiment
by: Zhao, Guoyang, et al.
Published: (2026)
by: Zhao, Guoyang, et al.
Published: (2026)
OmniRe: Omni Urban Scene Reconstruction
by: Chen, Ziyu, et al.
Published: (2024)
by: Chen, Ziyu, et al.
Published: (2024)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
by: Hu, Hezhen, et al.
Published: (2026)
by: Hu, Hezhen, et al.
Published: (2026)
Language-Image Models with 3D Understanding
by: Cho, Jang Hyun, et al.
Published: (2024)
by: Cho, Jang Hyun, et al.
Published: (2024)
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Expressive Gaussian Human Avatars from Monocular RGB Video
by: Hu, Hezhen, et al.
Published: (2024)
by: Hu, Hezhen, et al.
Published: (2024)
LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild
by: Wang, Zhiqiang, et al.
Published: (2024)
by: Wang, Zhiqiang, et al.
Published: (2024)
CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning
by: Ma, Wenxin, et al.
Published: (2026)
by: Ma, Wenxin, et al.
Published: (2026)
Similar Items
-
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
by: Cong, Wenyan, et al.
Published: (2025) -
InstantSplat: Sparse-view Gaussian Splatting in Seconds
by: Fan, Zhiwen, et al.
Published: (2024) -
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
by: Fan, Zhiwen, et al.
Published: (2024) -
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024) -
Towards Reliable and Interpretable Traffic Crash Pattern Prediction and Safety Interventions Using Customized Large Language Models
by: Zhao, Yang, et al.
Published: (2025)