EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schäfer, Finn Rasmus, Gao, Yuan, Wang, Dingrui, Stauner, Thomas, Günnemann, Stephan, Piccinini, Mattia, Schmidt, Sebastian, Betz, Johannes
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915955942096896
author Schäfer, Finn Rasmus
Gao, Yuan
Wang, Dingrui
Stauner, Thomas
Günnemann, Stephan
Piccinini, Mattia
Schmidt, Sebastian
Betz, Johannes
author_facet Schäfer, Finn Rasmus
Gao, Yuan
Wang, Dingrui
Stauner, Thomas
Günnemann, Stephan
Piccinini, Mattia
Schmidt, Sebastian
Betz, Johannes
contents While Vision-Language Models (VLMs) have advanced highlevel reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench, a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20 + models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they consistently fail to accurately align them with visual observations, frequently underperforming classical non-learned geometric baselines. This failure persists across model scales and domain-specific training, indicating a structural deficit in how current architectures couple visual perception with physical reasoning. We demonstrate that providing explicit trajectory encodings substantially restores physical consistency across all evaluated models, revealing a functional disentanglement between vision and language: egomotion logic is derived almost exclusively from the language modality, while visual observations contribute negligible additional signal. This structural finding provides a standardized diagnostic framework and a practical pathway toward physically aligned embodied AI. Keywords: Ego-motion - Physical Reasoning - Foundation Models
format Preprint
id arxiv_https___arxiv_org_abs_2604_22851
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
Schäfer, Finn Rasmus
Gao, Yuan
Wang, Dingrui
Stauner, Thomas
Günnemann, Stephan
Piccinini, Mattia
Schmidt, Sebastian
Betz, Johannes
Computer Vision and Pattern Recognition
Computation and Language
Robotics
While Vision-Language Models (VLMs) have advanced highlevel reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench, a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20 + models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they consistently fail to accurately align them with visual observations, frequently underperforming classical non-learned geometric baselines. This failure persists across model scales and domain-specific training, indicating a structural deficit in how current architectures couple visual perception with physical reasoning. We demonstrate that providing explicit trajectory encodings substantially restores physical consistency across all evaluated models, revealing a functional disentanglement between vision and language: egomotion logic is derived almost exclusively from the language modality, while visual observations contribute negligible additional signal. This structural finding provides a standardized diagnostic framework and a practical pathway toward physically aligned embodied AI. Keywords: Ego-motion - Physical Reasoning - Foundation Models
title EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
topic Computer Vision and Pattern Recognition
Computation and Language
Robotics
url https://arxiv.org/abs/2604.22851