Context-Infused Visual Grounding for Art

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Khan, Selina, van Noord, Nanne
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916441862701056
author Khan, Selina
van Noord, Nanne
author_facet Khan, Selina
van Noord, Nanne
contents Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks. Visual grounding offers the potential for localising subjects within these descriptions on images, however, existing approaches are trained on natural images and generalise poorly to art. In this paper, we present CIGAr (Context-Infused GroundingDINO for Art), a visual grounding approach which utilises the artwork descriptions during training as context, thereby enabling visual grounding on art. In addition, we present a new dataset, Ukiyo-eVG, with manually annotated phrase-grounding annotations, and we set a new state-of-the-art for object detection on two artwork datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12369
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Context-Infused Visual Grounding for Art
Khan, Selina
van Noord, Nanne
Computer Vision and Pattern Recognition
Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks. Visual grounding offers the potential for localising subjects within these descriptions on images, however, existing approaches are trained on natural images and generalise poorly to art. In this paper, we present CIGAr (Context-Infused GroundingDINO for Art), a visual grounding approach which utilises the artwork descriptions during training as context, thereby enabling visual grounding on art. In addition, we present a new dataset, Ukiyo-eVG, with manually annotated phrase-grounding annotations, and we set a new state-of-the-art for object detection on two artwork datasets.
title Context-Infused Visual Grounding for Art
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.12369