SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lai, Yuzhi, Yuan, Shenghai, Li, Peizheng, Lou, Jun, Zell, Andreas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912551705509888
author Lai, Yuzhi
Yuan, Shenghai
Li, Peizheng
Lou, Jun
Zell, Andreas
author_facet Lai, Yuzhi
Yuan, Shenghai
Li, Peizheng
Lou, Jun
Zell, Andreas
contents We present SEER-VAR, a novel framework for egocentric vehicle-based augmented reality (AR) that unifies semantic decomposition, Context-Aware SLAM Branches (CASB), and LLM-driven recommendation. Unlike existing systems that assume static or single-view settings, SEER-VAR dynamically separates cabin and road scenes via depth-guided vision-language grounding. Two SLAM branches track egocentric motion in each context, while a GPT-based module generates context-aware overlays such as dashboard cues and hazard alerts. To support evaluation, we introduce EgoSLAM-Drive, a real-world dataset featuring synchronized egocentric views, 6DoF ground-truth poses, and AR annotations across diverse driving scenarios. Experiments demonstrate that SEER-VAR achieves robust spatial alignment and perceptually coherent AR rendering across varied environments. As one of the first to explore LLM-based AR recommendation in egocentric driving, we address the lack of comparable systems through structured prompting and detailed user studies. Results show that SEER-VAR enhances perceived scene understanding, overlay relevance, and driver ease, providing an effective foundation for future research in this direction. Code and dataset will be made open source.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17255
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality
Lai, Yuzhi
Yuan, Shenghai
Li, Peizheng
Lou, Jun
Zell, Andreas
Computer Vision and Pattern Recognition
Robotics
We present SEER-VAR, a novel framework for egocentric vehicle-based augmented reality (AR) that unifies semantic decomposition, Context-Aware SLAM Branches (CASB), and LLM-driven recommendation. Unlike existing systems that assume static or single-view settings, SEER-VAR dynamically separates cabin and road scenes via depth-guided vision-language grounding. Two SLAM branches track egocentric motion in each context, while a GPT-based module generates context-aware overlays such as dashboard cues and hazard alerts. To support evaluation, we introduce EgoSLAM-Drive, a real-world dataset featuring synchronized egocentric views, 6DoF ground-truth poses, and AR annotations across diverse driving scenarios. Experiments demonstrate that SEER-VAR achieves robust spatial alignment and perceptually coherent AR rendering across varied environments. As one of the first to explore LLM-based AR recommendation in egocentric driving, we address the lack of comparable systems through structured prompting and detailed user studies. Results show that SEER-VAR enhances perceived scene understanding, overlay relevance, and driver ease, providing an effective foundation for future research in this direction. Code and dataset will be made open source.
title SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2508.17255