From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gado, Ahmed Y., Goba, Omar Y., Hassanein, Alaa, Elias, Catherine M., Hussein, Ahmed
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918511636381696
author Gado, Ahmed Y.
Goba, Omar Y.
Hassanein, Alaa
Elias, Catherine M.
Hussein, Ahmed
author_facet Gado, Ahmed Y.
Goba, Omar Y.
Hassanein, Alaa
Elias, Catherine M.
Hussein, Ahmed
contents Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondary property. This lack of temporal grounding leads to inconsistencies in reasoning about continuous actions, undermining both safety and interpretability. This work explores whether temporal conditioning within inter-agent communication can preserve or enhance coherence without introducing degradation in semantic or logical consistency. To investigate this, we introduce three planner architectures with progressively increasing temporal integration and evaluate them on curated subsets of the BDD-X dataset using semantic, syntactic, and logical metrics. Results show that while temporal conditioning reshapes reasoning style, it yields no statistically significant improvements in standard NLP-based correctness metrics. However, qualitative analysis reveals predictive hazard reasoning, stable corrective behavior, and strategic divergence in the Sentinel. These findings clarify the limits of prompt-based temporal grounding and establish the first empirical benchmark for temporal scene-to-plan reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19824
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
Gado, Ahmed Y.
Goba, Omar Y.
Hassanein, Alaa
Elias, Catherine M.
Hussein, Ahmed
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and Large Multimodal Models (LMMs) continue to treat time as a secondary property. This lack of temporal grounding leads to inconsistencies in reasoning about continuous actions, undermining both safety and interpretability. This work explores whether temporal conditioning within inter-agent communication can preserve or enhance coherence without introducing degradation in semantic or logical consistency. To investigate this, we introduce three planner architectures with progressively increasing temporal integration and evaluate them on curated subsets of the BDD-X dataset using semantic, syntactic, and logical metrics. Results show that while temporal conditioning reshapes reasoning style, it yields no statistically significant improvements in standard NLP-based correctness metrics. However, qualitative analysis reveals predictive hazard reasoning, stable corrective behavior, and strategic divergence in the Sentinel. These findings clarify the limits of prompt-based temporal grounding and establish the first empirical benchmark for temporal scene-to-plan reasoning.
title From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2605.19824