Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kolner, Oleh, Ortner, Thomas, Woźniak, Stanisław, Pantazi, Angeliki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909560544952320
author Kolner, Oleh
Ortner, Thomas
Woźniak, Stanisław
Pantazi, Angeliki
author_facet Kolner, Oleh
Ortner, Thomas
Woźniak, Stanisław
Pantazi, Angeliki
contents Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to out-of-distribution visual inputs than prior models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
Kolner, Oleh
Ortner, Thomas
Woźniak, Stanisław
Pantazi, Angeliki
Computer Vision and Pattern Recognition
Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to out-of-distribution visual inputs than prior models.
title Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.20213