Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Xiangyan, Yu, Jing, Gai, Keke, Zhuang, Jiamin, Tang, Yuanmin, Xiong, Gang, Gou, Gaopeng, Wu, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916332533972992
author Qu, Xiangyan
Yu, Jing
Gai, Keke
Zhuang, Jiamin
Tang, Yuanmin
Xiong, Gang
Gou, Gaopeng
Wu, Qi
author_facet Qu, Xiangyan
Yu, Jing
Gai, Keke
Zhuang, Jiamin
Tang, Yuanmin
Xiong, Gang
Gou, Gaopeng
Wu, Qi
contents Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with corresponding images to transfer knowledge. However, they disregard that semantic information is not equivalent between them, resulting in a suboptimal alignment. In this work, we propose a novel network to extract multi-view semantic concepts from documents and images and align the matching rather than entire concepts. Specifically, we propose a semantic decomposition module to generate multi-view semantic embeddings from visual and textual sides, providing the basic concepts for partial alignment. To alleviate the issue of information redundancy among embeddings, we propose the local-to-semantic variance loss to capture distinct local details and multiple semantic diversity loss to enforce orthogonality among embeddings. Subsequently, two losses are introduced to partially align visual-semantic embedding pairs according to their semantic relevance at the view and word-to-patch levels. Consequently, we consistently outperform state-of-the-art methods under two document sources in three standard benchmarks for document-based zero-shot learning. Qualitatively, we show that our model learns the interpretable partial association.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
Qu, Xiangyan
Yu, Jing
Gai, Keke
Zhuang, Jiamin
Tang, Yuanmin
Xiong, Gang
Gou, Gaopeng
Wu, Qi
Computer Vision and Pattern Recognition
Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with corresponding images to transfer knowledge. However, they disregard that semantic information is not equivalent between them, resulting in a suboptimal alignment. In this work, we propose a novel network to extract multi-view semantic concepts from documents and images and align the matching rather than entire concepts. Specifically, we propose a semantic decomposition module to generate multi-view semantic embeddings from visual and textual sides, providing the basic concepts for partial alignment. To alleviate the issue of information redundancy among embeddings, we propose the local-to-semantic variance loss to capture distinct local details and multiple semantic diversity loss to enforce orthogonality among embeddings. Subsequently, two losses are introduced to partially align visual-semantic embedding pairs according to their semantic relevance at the view and word-to-patch levels. Consequently, we consistently outperform state-of-the-art methods under two document sources in three standard benchmarks for document-based zero-shot learning. Qualitatively, we show that our model learns the interpretable partial association.
title Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.15613