Where, What, Why: Towards Explainable Driver Attention Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuchen, Tang, Jiayu, Xiao, Xiaoyan, Lin, Yueyao, Liu, Linkai, Guo, Zipeng, Fei, Hao, Xia, Xiaobo, Gou, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916815207137280
author Zhou, Yuchen
Tang, Jiayu
Xiao, Xiaoyan
Lin, Yueyao
Liu, Linkai
Guo, Zipeng
Fei, Hao
Xia, Xiaobo
Gou, Chao
author_facet Zhou, Yuchen
Tang, Jiayu
Xiao, Xiaoyan
Lin, Yueyao
Liu, Linkai
Guo, Zipeng
Fei, Hao
Xia, Xiaobo
Gou, Chao
contents Modeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to capture the cognitive motivations behind attention allocation in specific contexts, which limits deeper understanding of attention mechanisms. To bridge this gap, we introduce Explainable Driver Attention Prediction, a novel task paradigm that jointly predicts spatial attention regions (where), parses attended semantics (what), and provides cognitive reasoning for attention allocation (why). To support this, we present W3DA, the first large-scale explainable driver attention dataset. It enriches existing benchmarks with detailed semantic and causal annotations across diverse driving scenarios, including normal conditions, safety-critical situations, and traffic accidents. We further propose LLada, a Large Language model-driven framework for driver attention prediction, which unifies pixel modeling, semantic parsing, and cognitive reasoning within an end-to-end architecture. Extensive experiments demonstrate the effectiveness of LLada, exhibiting robust generalization across datasets and driving conditions. This work serves as a key step toward a deeper understanding of driver attention mechanisms, with significant implications for autonomous driving, intelligent driver training, and human-computer interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Where, What, Why: Towards Explainable Driver Attention Prediction
Zhou, Yuchen
Tang, Jiayu
Xiao, Xiaoyan
Lin, Yueyao
Liu, Linkai
Guo, Zipeng
Fei, Hao
Xia, Xiaobo
Gou, Chao
Computer Vision and Pattern Recognition
Modeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to capture the cognitive motivations behind attention allocation in specific contexts, which limits deeper understanding of attention mechanisms. To bridge this gap, we introduce Explainable Driver Attention Prediction, a novel task paradigm that jointly predicts spatial attention regions (where), parses attended semantics (what), and provides cognitive reasoning for attention allocation (why). To support this, we present W3DA, the first large-scale explainable driver attention dataset. It enriches existing benchmarks with detailed semantic and causal annotations across diverse driving scenarios, including normal conditions, safety-critical situations, and traffic accidents. We further propose LLada, a Large Language model-driven framework for driver attention prediction, which unifies pixel modeling, semantic parsing, and cognitive reasoning within an end-to-end architecture. Extensive experiments demonstrate the effectiveness of LLada, exhibiting robust generalization across datasets and driving conditions. This work serves as a key step toward a deeper understanding of driver attention mechanisms, with significant implications for autonomous driving, intelligent driver training, and human-computer interaction.
title Where, What, Why: Towards Explainable Driver Attention Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23088