STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Fruhwirth-Reisinger, Christian, Malić, Dušan, Lin, Wei, Schinagl, David, Schulter, Samuel, Possegger, Horst |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detection
by: Malić, Dušan, et al.
Published: (2025)
by: Malić, Dušan, et al.
Published: (2025)
LiSu: A Dataset and Method for LiDAR Surface Normal Estimation
by: Malić, Dušan, et al.
Published: (2025)
by: Malić, Dušan, et al.
Published: (2025)
GBlobs: Local LiDAR Geometry for Improved Sensor Placement Generalization
by: Malić, Dušan, et al.
Published: (2025)
by: Malić, Dušan, et al.
Published: (2025)
Learn to Rank: Visual Attribution by Learning Importance Ranking
by: Schinagl, David, et al.
Published: (2026)
by: Schinagl, David, et al.
Published: (2026)
Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection
by: Fruhwirth-Reisinger, Christian, et al.
Published: (2024)
by: Fruhwirth-Reisinger, Christian, et al.
Published: (2024)
SHARP: Short-Window Streaming for Accurate and Robust Prediction in Motion Forecasting
by: Prutsch, Alexander, et al.
Published: (2026)
by: Prutsch, Alexander, et al.
Published: (2026)
Streaming Real-Time Trajectory Prediction Using Endpoint-Aware Modeling
by: Prutsch, Alexander, et al.
Published: (2026)
by: Prutsch, Alexander, et al.
Published: (2026)
ASCENT: Transformer-Based Aircraft Trajectory Prediction in Non-Towered Terminal Airspace
by: Prutsch, Alexander, et al.
Published: (2026)
by: Prutsch, Alexander, et al.
Published: (2026)
An Investigation of Beam Density on LiDAR Object Detection Performance
by: Griesbacher, Christoph, et al.
Published: (2025)
by: Griesbacher, Christoph, et al.
Published: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
by: Qian, Tianwen, et al.
Published: (2023)
by: Qian, Tianwen, et al.
Published: (2023)
Distilling Multi-modal Large Language Models for Autonomous Driving
by: Hegde, Deepti, et al.
Published: (2025)
by: Hegde, Deepti, et al.
Published: (2025)
Efficient Motion Prediction: A Lightweight & Accurate Trajectory Prediction Model With Fast Training and Inference Speed
by: Prutsch, Alexander, et al.
Published: (2024)
by: Prutsch, Alexander, et al.
Published: (2024)
Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving
by: Yamazaki, Shota, et al.
Published: (2024)
by: Yamazaki, Shota, et al.
Published: (2024)
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
by: Liang, Mingfu, et al.
Published: (2024)
by: Liang, Mingfu, et al.
Published: (2024)
BEVPredFormer: Spatio-temporal Attention for BEV Instance Prediction in Autonomous Driving
by: Antunes-García, Miguel, et al.
Published: (2026)
by: Antunes-García, Miguel, et al.
Published: (2026)
One Model, Many Behaviors: Training-Induced Effects on Out-of-Distribution Detection
by: Krumpl, Gerhard, et al.
Published: (2026)
by: Krumpl, Gerhard, et al.
Published: (2026)
ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
by: Krumpl, Gerhard, et al.
Published: (2026)
by: Krumpl, Gerhard, et al.
Published: (2026)
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
ADGaussian: Generalizable Gaussian Splatting for Autonomous Driving via Multi-modal Joint Learning
by: Song, Qi, et al.
Published: (2025)
by: Song, Qi, et al.
Published: (2025)
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
Into the Fog: Evaluating Robustness of Multiple Object Tracking
by: Kirillova, Nadezda, et al.
Published: (2024)
by: Kirillova, Nadezda, et al.
Published: (2024)
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
by: Hannan, Tanveer, et al.
Published: (2025)
by: Hannan, Tanveer, et al.
Published: (2025)
MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly Detection
by: Micorek, Jakub, et al.
Published: (2024)
by: Micorek, Jakub, et al.
Published: (2024)
LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
by: Sun, Qihao, et al.
Published: (2026)
by: Sun, Qihao, et al.
Published: (2026)
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
by: Tu, Zhi, et al.
Published: (2025)
by: Tu, Zhi, et al.
Published: (2025)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
by: Spoecklberger, Johannes, et al.
Published: (2025)
by: Spoecklberger, Johannes, et al.
Published: (2025)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
by: Zhang, Zhenguo, et al.
Published: (2025)
by: Zhang, Zhenguo, et al.
Published: (2025)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
by: Sha, Lin, et al.
Published: (2026)
by: Sha, Lin, et al.
Published: (2026)
LLM-attacker: Enhancing Closed-loop Adversarial Scenario Generation for Autonomous Driving with Large Language Models
by: Mei, Yuewen, et al.
Published: (2025)
by: Mei, Yuewen, et al.
Published: (2025)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
Generative Scenario Rollouts for End-to-End Autonomous Driving
by: Yasarla, Rajeev, et al.
Published: (2026)
by: Yasarla, Rajeev, et al.
Published: (2026)
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024)
by: Aich, Abhishek, et al.
Published: (2024)
Hyperspectral Imaging-Based Perception in Autonomous Driving Scenarios: Benchmarking Baseline Semantic Segmentation Models
by: Shah, Imad Ali, et al.
Published: (2024)
by: Shah, Imad Ali, et al.
Published: (2024)
OccGen: Generative Multi-modal 3D Occupancy Prediction for Autonomous Driving
by: Wang, Guoqing, et al.
Published: (2024)
by: Wang, Guoqing, et al.
Published: (2024)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
by: Cui, Erfei, et al.
Published: (2023)
by: Cui, Erfei, et al.
Published: (2023)
Enhancing Autonomous Driving Safety with Collision Scenario Integration
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
by: Guo, Xianda, et al.
Published: (2024)
by: Guo, Xianda, et al.
Published: (2024)
Similar Items
-
GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detection
by: Malić, Dušan, et al.
Published: (2025) -
LiSu: A Dataset and Method for LiDAR Surface Normal Estimation
by: Malić, Dušan, et al.
Published: (2025) -
GBlobs: Local LiDAR Geometry for Improved Sensor Placement Generalization
by: Malić, Dušan, et al.
Published: (2025) -
Learn to Rank: Visual Attribution by Learning Importance Ranking
by: Schinagl, David, et al.
Published: (2026) -
Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection
by: Fruhwirth-Reisinger, Christian, et al.
Published: (2024)