Improving Multi-label Recognition using Class Co-Occurrence Probabilities

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Rawlekar, Samyak, Bhatnagar, Shubhang, Srinivasulu, Vishnuvardhan Pogunulu, Ahuja, Narendra
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916402009473024
author Rawlekar, Samyak
Bhatnagar, Shubhang
Srinivasulu, Vishnuvardhan Pogunulu
Ahuja, Narendra
author_facet Rawlekar, Samyak
Bhatnagar, Shubhang
Srinivasulu, Vishnuvardhan Pogunulu
Ahuja, Narendra
contents Multi-label Recognition (MLR) involves the identification of multiple objects within an image. To address the additional complexity of this problem, recent works have leveraged information from vision-language models (VLMs) trained on large text-images datasets for the task. These methods learn an independent classifier for each object (class), overlooking correlations in their occurrences. Such co-occurrences can be captured from the training data as conditional probabilities between a pair of classes. We propose a framework to extend the independent classifiers by incorporating the co-occurrence information for object pairs to improve the performance of independent classifiers. We use a Graph Convolutional Network (GCN) to enforce the conditional probabilities between classes, by refining the initial estimates derived from image and text sources obtained using VLMs. We validate our method on four MLR datasets, where our approach outperforms all state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16193
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Multi-label Recognition using Class Co-Occurrence Probabilities
Rawlekar, Samyak
Bhatnagar, Shubhang
Srinivasulu, Vishnuvardhan Pogunulu
Ahuja, Narendra
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Image and Video Processing
Multi-label Recognition (MLR) involves the identification of multiple objects within an image. To address the additional complexity of this problem, recent works have leveraged information from vision-language models (VLMs) trained on large text-images datasets for the task. These methods learn an independent classifier for each object (class), overlooking correlations in their occurrences. Such co-occurrences can be captured from the training data as conditional probabilities between a pair of classes. We propose a framework to extend the independent classifiers by incorporating the co-occurrence information for object pairs to improve the performance of independent classifiers. We use a Graph Convolutional Network (GCN) to enforce the conditional probabilities between classes, by refining the initial estimates derived from image and text sources obtained using VLMs. We validate our method on four MLR datasets, where our approach outperforms all state-of-the-art methods.
title Improving Multi-label Recognition using Class Co-Occurrence Probabilities
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2404.16193