Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cartella, Giuseppe, Cuculo, Vittorio, D'Amelio, Alessandro, Cornia, Marcella, Boccignone, Giuseppe, Cucchiara, Rita
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909712722690048
author Cartella, Giuseppe
Cuculo, Vittorio
D'Amelio, Alessandro
Cornia, Marcella
Boccignone, Giuseppe
Cucchiara, Rita
author_facet Cartella, Giuseppe
Cuculo, Vittorio
D'Amelio, Alessandro
Cornia, Marcella
Boccignone, Giuseppe
Cucchiara, Rita
contents Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to capture the variability of human visual exploration. In this work, we present ScanDiff, a novel architecture that combines diffusion models with Vision Transformers to generate diverse and realistic scanpaths. Our method explicitly models scanpath variability by leveraging the stochastic nature of diffusion models, producing a wide range of plausible gaze trajectories. Additionally, we introduce textual conditioning to enable task-driven scanpath generation, allowing the model to adapt to different visual search objectives. Experiments on benchmark datasets show that ScanDiff surpasses state-of-the-art methods in both free-viewing and task-driven scenarios, producing more diverse and accurate scanpaths. These results highlight its ability to better capture the complexity of human visual behavior, pushing forward gaze prediction research. Source code and models are publicly available at https://aimagelab.github.io/ScanDiff.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
Cartella, Giuseppe
Cuculo, Vittorio
D'Amelio, Alessandro
Cornia, Marcella
Boccignone, Giuseppe
Cucchiara, Rita
Computer Vision and Pattern Recognition
Artificial Intelligence
Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to capture the variability of human visual exploration. In this work, we present ScanDiff, a novel architecture that combines diffusion models with Vision Transformers to generate diverse and realistic scanpaths. Our method explicitly models scanpath variability by leveraging the stochastic nature of diffusion models, producing a wide range of plausible gaze trajectories. Additionally, we introduce textual conditioning to enable task-driven scanpath generation, allowing the model to adapt to different visual search objectives. Experiments on benchmark datasets show that ScanDiff surpasses state-of-the-art methods in both free-viewing and task-driven scenarios, producing more diverse and accurate scanpaths. These results highlight its ability to better capture the complexity of human visual behavior, pushing forward gaze prediction research. Source code and models are publicly available at https://aimagelab.github.io/ScanDiff.
title Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.23021