GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Ruijie, Jin, Sheng, Xu, Lumin, Zeng, Wang, Liu, Wentao, Qian, Chen, Luo, Ping, Wu, Ji
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914876698394624
author Yao, Ruijie
Jin, Sheng
Xu, Lumin
Zeng, Wang
Liu, Wentao
Qian, Chen
Luo, Ping
Wu, Ji
author_facet Yao, Ruijie
Jin, Sheng
Xu, Lumin
Zeng, Wang
Liu, Wentao
Qian, Chen
Luo, Ping
Wu, Ji
contents Multi-Label Image Recognition (MLIR) is a challenging task that aims to predict multiple object labels in a single image while modeling the complex relationships between labels and image regions. Although convolutional neural networks and vision transformers have succeeded in processing images as regular grids of pixels or patches, these representations are sub-optimal for capturing irregular and discontinuous regions of interest. In this work, we present the first fully graph convolutional model, Group K-nearest neighbor based Graph convolutional Network (GKGNet), which models the connections between semantic label embeddings and image patches in a flexible and unified graph structure. To address the scale variance of different objects and to capture information from multiple perspectives, we propose the Group KGCN module for dynamic graph construction and message passing. Our experiments demonstrate that GKGNet achieves state-of-the-art performance with significantly lower computational costs on the challenging multi-label datasets, i.e., MS-COCO and VOC2007 datasets. Codes are available at https://github.com/jin-s13/GKGNet.
format Preprint
id arxiv_https___arxiv_org_abs_2308_14378
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
Yao, Ruijie
Jin, Sheng
Xu, Lumin
Zeng, Wang
Liu, Wentao
Qian, Chen
Luo, Ping
Wu, Ji
Computer Vision and Pattern Recognition
Multi-Label Image Recognition (MLIR) is a challenging task that aims to predict multiple object labels in a single image while modeling the complex relationships between labels and image regions. Although convolutional neural networks and vision transformers have succeeded in processing images as regular grids of pixels or patches, these representations are sub-optimal for capturing irregular and discontinuous regions of interest. In this work, we present the first fully graph convolutional model, Group K-nearest neighbor based Graph convolutional Network (GKGNet), which models the connections between semantic label embeddings and image patches in a flexible and unified graph structure. To address the scale variance of different objects and to capture information from multiple perspectives, we propose the Group KGCN module for dynamic graph construction and message passing. Our experiments demonstrate that GKGNet achieves state-of-the-art performance with significantly lower computational costs on the challenging multi-label datasets, i.e., MS-COCO and VOC2007 datasets. Codes are available at https://github.com/jin-s13/GKGNet.
title GKGNet: Group K-Nearest Neighbor based Graph Convolutional Network for Multi-Label Image Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.14378