EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiao, Ma, Yongqiang, Jing, Haodong, Zheng, Nanning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912450660532224
author Zhang, Xiao
Ma, Yongqiang
Jing, Haodong
Zheng, Nanning
author_facet Zhang, Xiao
Ma, Yongqiang
Jing, Haodong
Zheng, Nanning
contents Compositional Zero-Shot Learning (CZSL) investigates compositional generalization capacity to recognize unknown state-object pairs based on learned primitive concepts. Existing CZSL methods typically derive primitives features through a simple composition-prototype mapping, which is suboptimal for a set of individuals that can be divided into distinct semantic subsets. Moreover, the all-to-one cross-modal primitives matching neglects compositional divergence within identical states or objects, limiting fine-grained image-composition alignment. In this study, we propose EVA, a Mixture-of-Experts Semantic Variant Alignment framework for CZSL. Specifically, we introduce domain-expert adaption, leveraging multiple experts to achieve token-aware learning and model high-quality primitive representations. To enable accurate compositional generalization, we further present semantic variant alignment to select semantically relevant representation for image-primitives matching. Our method significantly outperforms other state-of-the-art CZSL methods on three popular benchmarks in both closed- and open-world settings, demonstrating the efficacy of the proposed insight.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20986
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
Zhang, Xiao
Ma, Yongqiang
Jing, Haodong
Zheng, Nanning
Computer Vision and Pattern Recognition
Compositional Zero-Shot Learning (CZSL) investigates compositional generalization capacity to recognize unknown state-object pairs based on learned primitive concepts. Existing CZSL methods typically derive primitives features through a simple composition-prototype mapping, which is suboptimal for a set of individuals that can be divided into distinct semantic subsets. Moreover, the all-to-one cross-modal primitives matching neglects compositional divergence within identical states or objects, limiting fine-grained image-composition alignment. In this study, we propose EVA, a Mixture-of-Experts Semantic Variant Alignment framework for CZSL. Specifically, we introduce domain-expert adaption, leveraging multiple experts to achieve token-aware learning and model high-quality primitive representations. To enable accurate compositional generalization, we further present semantic variant alignment to select semantically relevant representation for image-primitives matching. Our method significantly outperforms other state-of-the-art CZSL methods on three popular benchmarks in both closed- and open-world settings, demonstrating the efficacy of the proposed insight.
title EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.20986