OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Yini, Yu, Jingling, Zhang, Haozheng, van der Lans, Ralf, Shi, Bertram
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929426326880256
author Fang, Yini
Yu, Jingling
Zhang, Haozheng
van der Lans, Ralf
Shi, Bertram
author_facet Fang, Yini
Yu, Jingling
Zhang, Haozheng
van der Lans, Ralf
Shi, Bertram
contents Visual search is important in our daily life. The efficient allocation of visual attention is critical to effectively complete visual search tasks. Prior research has predominantly modelled the spatial allocation of visual attention in images at the pixel level, e.g. using a saliency map. However, emerging evidence shows that visual attention is guided by objects rather than pixel intensities. This paper introduces the Object-level Attention Transformer (OAT), which predicts human scanpaths as they search for a target object within a cluttered scene of distractors. OAT uses an encoder-decoder architecture. The encoder captures information about the position and appearance of the objects within an image and about the target. The decoder predicts the gaze scanpath as a sequence of object fixations, by integrating output features from both the encoder and decoder. We also propose a new positional encoding that better reflects spatial relationships between objects. We evaluated OAT on the Amazon book cover dataset and a new dataset for visual search that we collected. OAT's predicted gaze scanpaths align more closely with human gaze patterns, compared to predictions by algorithms based on spatial attention on both established metrics and a novel behavioural-based metric. Our results demonstrate the generalization ability of OAT, as it accurately predicts human scanpaths for unseen layouts and target objects.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13335
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
Fang, Yini
Yu, Jingling
Zhang, Haozheng
van der Lans, Ralf
Shi, Bertram
Computer Vision and Pattern Recognition
Visual search is important in our daily life. The efficient allocation of visual attention is critical to effectively complete visual search tasks. Prior research has predominantly modelled the spatial allocation of visual attention in images at the pixel level, e.g. using a saliency map. However, emerging evidence shows that visual attention is guided by objects rather than pixel intensities. This paper introduces the Object-level Attention Transformer (OAT), which predicts human scanpaths as they search for a target object within a cluttered scene of distractors. OAT uses an encoder-decoder architecture. The encoder captures information about the position and appearance of the objects within an image and about the target. The decoder predicts the gaze scanpath as a sequence of object fixations, by integrating output features from both the encoder and decoder. We also propose a new positional encoding that better reflects spatial relationships between objects. We evaluated OAT on the Amazon book cover dataset and a new dataset for visual search that we collected. OAT's predicted gaze scanpaths align more closely with human gaze patterns, compared to predictions by algorithms based on spatial attention on both established metrics and a novel behavioural-based metric. Our results demonstrate the generalization ability of OAT, as it accurately predicts human scanpaths for unseen layouts and target objects.
title OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.13335