Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Chang, Chun-Peng, Wang, Chen-Yu, Caesar, Holger, Pagani, Alain |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
por: Chang, Chun-Peng, et al.
Publicado: (2025)
por: Chang, Chun-Peng, et al.
Publicado: (2025)
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
por: Chang, Chun-Peng, et al.
Publicado: (2026)
por: Chang, Chun-Peng, et al.
Publicado: (2026)
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
por: Chang, Chun-Peng, et al.
Publicado: (2024)
por: Chang, Chun-Peng, et al.
Publicado: (2024)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
por: Chang, Chun-Peng, et al.
Publicado: (2024)
por: Chang, Chun-Peng, et al.
Publicado: (2024)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
por: Li, Weiming, et al.
Publicado: (2025)
por: Li, Weiming, et al.
Publicado: (2025)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
por: Guo, Chaohong, et al.
Publicado: (2025)
por: Guo, Chaohong, et al.
Publicado: (2025)
ICP-Flow: LiDAR Scene Flow Estimation with ICP
por: Lin, Yancong, et al.
Publicado: (2024)
por: Lin, Yancong, et al.
Publicado: (2024)
Offline Tracking with Object Permanence
por: Liu, Xianzhong, et al.
Publicado: (2023)
por: Liu, Xianzhong, et al.
Publicado: (2023)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
por: Jiang, Bo, et al.
Publicado: (2025)
por: Jiang, Bo, et al.
Publicado: (2025)
Uni-SLAM: Uncertainty-Aware Neural Implicit SLAM for Real-Time Dense Indoor Scene Reconstruction
por: Wang, Shaoxiang, et al.
Publicado: (2024)
por: Wang, Shaoxiang, et al.
Publicado: (2024)
Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs
por: Lim, Jeongkee, et al.
Publicado: (2024)
por: Lim, Jeongkee, et al.
Publicado: (2024)
BikeScenes: Online LiDAR Semantic Segmentation for Bicycles
por: Goren, Denniz, et al.
Publicado: (2025)
por: Goren, Denniz, et al.
Publicado: (2025)
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
por: Li, Jiangyang, et al.
Publicado: (2026)
por: Li, Jiangyang, et al.
Publicado: (2026)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
por: Zheng, Zelin, et al.
Publicado: (2026)
por: Zheng, Zelin, et al.
Publicado: (2026)
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
por: Xie, Shaoyuan, et al.
Publicado: (2025)
por: Xie, Shaoyuan, et al.
Publicado: (2025)
nuScenes Revisited: Progress and Challenges in Autonomous Driving
por: Fong, Whye Kit, et al.
Publicado: (2025)
por: Fong, Whye Kit, et al.
Publicado: (2025)
Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
por: Chen, Xin, et al.
Publicado: (2025)
por: Chen, Xin, et al.
Publicado: (2025)
4DRC-OCC: Robust Semantic Occupancy Prediction Through Fusion of 4D Radar and Camera
por: Ninfa, David, et al.
Publicado: (2026)
por: Ninfa, David, et al.
Publicado: (2026)
4D-RaDiff: Latent Diffusion for 4D Radar Point Cloud Generation
por: Kwok, Jimmie, et al.
Publicado: (2025)
por: Kwok, Jimmie, et al.
Publicado: (2025)
BaSAL: Size-Balanced Warm Start Active Learning for LiDAR Semantic Segmentation
por: Wei, Jiarong, et al.
Publicado: (2023)
por: Wei, Jiarong, et al.
Publicado: (2023)
DPFT: Dual Perspective Fusion Transformer for Camera-Radar-based Object Detection
por: Fent, Felix, et al.
Publicado: (2024)
por: Fent, Felix, et al.
Publicado: (2024)
VLPrompt: Vision-Language Prompting for Panoptic Scene Graph Generation
por: Zhou, Zijian, et al.
Publicado: (2023)
por: Zhou, Zijian, et al.
Publicado: (2023)
Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs
por: Ma, Wen, et al.
Publicado: (2026)
por: Ma, Wen, et al.
Publicado: (2026)
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
NeuroNCAP: Photorealistic Closed-loop Safety Testing for Autonomous Driving
por: Ljungbergh, William, et al.
Publicado: (2024)
por: Ljungbergh, William, et al.
Publicado: (2024)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
por: Lian, Weitong, et al.
Publicado: (2026)
por: Lian, Weitong, et al.
Publicado: (2026)
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
por: Zong, Yongshuo, et al.
Publicado: (2025)
por: Zong, Yongshuo, et al.
Publicado: (2025)
LeAP: Consistent multi-domain 3D labeling using Foundation Models
por: Gebraad, Simon, et al.
Publicado: (2025)
por: Gebraad, Simon, et al.
Publicado: (2025)
A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs
por: Vaishnav, Mohit, et al.
Publicado: (2025)
por: Vaishnav, Mohit, et al.
Publicado: (2025)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
por: Zhang, Ben, et al.
Publicado: (2025)
por: Zhang, Ben, et al.
Publicado: (2025)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
por: Wang, Fanyi, et al.
Publicado: (2025)
por: Wang, Fanyi, et al.
Publicado: (2025)
G3FA: Geometry-guided GAN for Face Animation
por: Javanmardi, Alireza, et al.
Publicado: (2024)
por: Javanmardi, Alireza, et al.
Publicado: (2024)
OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
por: Zhou, Zijian, et al.
Publicado: (2024)
por: Zhou, Zijian, et al.
Publicado: (2024)
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
por: Lin, Jiaping, et al.
Publicado: (2026)
por: Lin, Jiaping, et al.
Publicado: (2026)
CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving
por: Song, Rui, et al.
Publicado: (2025)
por: Song, Rui, et al.
Publicado: (2025)
VLMs Guided Interpretable Decision Making for Autonomous Driving
por: Hu, Xin, et al.
Publicado: (2025)
por: Hu, Xin, et al.
Publicado: (2025)
AsyncBEV: Cross-modal Flow Alignment in Asynchronous 3D Object Detection
por: Wang, Shiming, et al.
Publicado: (2026)
por: Wang, Shiming, et al.
Publicado: (2026)
NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics
por: Lan, Jian, et al.
Publicado: (2026)
por: Lan, Jian, et al.
Publicado: (2026)
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
por: Vo, Hao, et al.
Publicado: (2026)
por: Vo, Hao, et al.
Publicado: (2026)
Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers
por: Wang, Jiancheng, et al.
Publicado: (2026)
por: Wang, Jiancheng, et al.
Publicado: (2026)
Ejemplares similares
-
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
por: Chang, Chun-Peng, et al.
Publicado: (2025) -
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
por: Chang, Chun-Peng, et al.
Publicado: (2026) -
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
por: Chang, Chun-Peng, et al.
Publicado: (2024) -
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
por: Chang, Chun-Peng, et al.
Publicado: (2024) -
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
por: Li, Weiming, et al.
Publicado: (2025)