Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Zhong, Xu, Yishi, Wang, Gerong, Chen, Wenchao, Chen, Bo, Zhang, Jing, Liu, Hongwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908803521314816
author Peng, Zhong
Xu, Yishi
Wang, Gerong
Chen, Wenchao
Chen, Bo
Zhang, Jing
Liu, Hongwei
author_facet Peng, Zhong
Xu, Yishi
Wang, Gerong
Chen, Wenchao
Chen, Bo
Zhang, Jing
Liu, Hongwei
contents Compositional Zero-Shot Learning (CZSL) seeks to recognize unseen state-object pairs by recombining primitives learned from seen compositions. Despite recent progress with vision-language models (VLMs), two limitations remain: (i) text-driven semantic prototypes are weakly discriminative in the visual feature space; and (ii) unseen pairs are optimized passively, thereby inducing seen bias. To address these limitations, we present Duplex, a framework that couples dual-prototype learning with dynamic local-graph refinement of visual prototypes. For each composition, Duplex maintains a semantic prototype via prompt learning and a visual prototype for unseen pairs constructed by recombining disentangled state and object primitives from seen images. The visual prototypes are updated dynamically through lightweight aggregation on mini-batch local graphs, which incorporates unseen compositions during training without labels. This design introduces fine-grained visual evidence while preserving semantic structure. It enriches class prototypes, better disambiguates semantically similar yet visually distinct pairs, and mitigates seen bias. Experiments on MIT-States, UT-Zappos, and CGQA in closed-world and open-world settings achieve competitive performance and consistent compositional generalization. Our source code is available at https://github.com/ISPZ/Duplex-CZSL.
format Preprint
id arxiv_https___arxiv_org_abs_2501_07114
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
Peng, Zhong
Xu, Yishi
Wang, Gerong
Chen, Wenchao
Chen, Bo
Zhang, Jing
Liu, Hongwei
Computer Vision and Pattern Recognition
Compositional Zero-Shot Learning (CZSL) seeks to recognize unseen state-object pairs by recombining primitives learned from seen compositions. Despite recent progress with vision-language models (VLMs), two limitations remain: (i) text-driven semantic prototypes are weakly discriminative in the visual feature space; and (ii) unseen pairs are optimized passively, thereby inducing seen bias. To address these limitations, we present Duplex, a framework that couples dual-prototype learning with dynamic local-graph refinement of visual prototypes. For each composition, Duplex maintains a semantic prototype via prompt learning and a visual prototype for unseen pairs constructed by recombining disentangled state and object primitives from seen images. The visual prototypes are updated dynamically through lightweight aggregation on mini-batch local graphs, which incorporates unseen compositions during training without labels. This design introduces fine-grained visual evidence while preserving semantic structure. It enriches class prototypes, better disambiguates semantically similar yet visually distinct pairs, and mitigates seen bias. Experiments on MIT-States, UT-Zappos, and CGQA in closed-world and open-world settings achieve competitive performance and consistent compositional generalization. Our source code is available at https://github.com/ISPZ/Duplex-CZSL.
title Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.07114