MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qu, Xiangyan, Yu, Jing, Zhuang, Jiamin, Gou, Gaopeng, Xiong, Gang, Wu, Qi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915189582987264
author Qu, Xiangyan
Yu, Jing
Zhuang, Jiamin
Gou, Gaopeng
Xiong, Gang
Wu, Qi
author_facet Qu, Xiangyan
Yu, Jing
Zhuang, Jiamin
Gou, Gaopeng
Xiong, Gang
Wu, Qi
contents Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that documents from encyclopedias provide helpful auxiliary information. However, existing methods align noisy documents, entangled in visual and non-visual descriptions, with image regions, yet solely depend on implicit learning. These models fail to filter non-visual noise reliably and incorrectly align non-visual words to image regions, which is harmful to knowledge transfer. In this work, we propose a novel multi-attribute document supervision framework to remove noises at both document collection and model learning stages. With the help of large language models, we introduce a novel prompt algorithm that automatically removes non-visual descriptions and enriches less-described documents in multiple attribute views. Our proposed model, MADS, extracts multi-view transferable knowledge with information decoupling and semantic interactions for semantic alignment at local and global levels. Besides, we introduce a model-agnostic focus loss to explicitly enhance attention to visually discriminative information during training, also improving existing methods without additional parameters. With comparable computation costs, MADS consistently outperforms the SOTA by 7.2% and 8.2% on average in three benchmarks for document-based ZSL and GZSL settings, respectively. Moreover, we qualitatively offer interpretable predictions from multiple attribute views.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
Qu, Xiangyan
Yu, Jing
Zhuang, Jiamin
Gou, Gaopeng
Xiong, Gang
Wu, Qi
Computer Vision and Pattern Recognition
Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that documents from encyclopedias provide helpful auxiliary information. However, existing methods align noisy documents, entangled in visual and non-visual descriptions, with image regions, yet solely depend on implicit learning. These models fail to filter non-visual noise reliably and incorrectly align non-visual words to image regions, which is harmful to knowledge transfer. In this work, we propose a novel multi-attribute document supervision framework to remove noises at both document collection and model learning stages. With the help of large language models, we introduce a novel prompt algorithm that automatically removes non-visual descriptions and enriches less-described documents in multiple attribute views. Our proposed model, MADS, extracts multi-view transferable knowledge with information decoupling and semantic interactions for semantic alignment at local and global levels. Besides, we introduce a model-agnostic focus loss to explicitly enhance attention to visually discriminative information during training, also improving existing methods without additional parameters. With comparable computation costs, MADS consistently outperforms the SOTA by 7.2% and 8.2% on average in three benchmarks for document-based ZSL and GZSL settings, respectively. Moreover, we qualitatively offer interpretable predictions from multiple attribute views.
title MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06847