Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yiming, Li, Hangfei, Wang, Fangfang, Zhang, Yilong, Liang, Ronghua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917674269802496
author Wu, Yiming
Li, Hangfei
Wang, Fangfang
Zhang, Yilong
Liang, Ronghua
author_facet Wu, Yiming
Li, Hangfei
Wang, Fangfang
Zhang, Yilong
Liang, Ronghua
contents In the domain of language-based fashion image retrieval, pinpointing the desired fashion item using both a reference image and its accompanying textual description is an intriguing challenge. Existing approaches lean heavily on static fusion techniques, intertwining image and text. Despite their commendable advancements, these approaches are still limited by a deficiency in flexibility. In response, we propose a Self-distilled Dynamic Fusion Network to compose the multi-granularity features dynamically by considering the consistency of routing path and modality-specific information simultaneously. Two new modules are included in our proposed method: (1) Dynamic Fusion Network with Modality Specific Routers. The dynamic network enables a flexible determination of the routing for each reference image and modification text, taking into account their distinct semantics and distributions. (2) Self Path Distillation Loss. A stable path decision for queries benefits the optimization of feature extraction as well as routing, and we approach this by progressively refine the path decision with previous path information. Extensive experiments demonstrate the effectiveness of our proposed model compared to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15451
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
Wu, Yiming
Li, Hangfei
Wang, Fangfang
Zhang, Yilong
Liang, Ronghua
Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
In the domain of language-based fashion image retrieval, pinpointing the desired fashion item using both a reference image and its accompanying textual description is an intriguing challenge. Existing approaches lean heavily on static fusion techniques, intertwining image and text. Despite their commendable advancements, these approaches are still limited by a deficiency in flexibility. In response, we propose a Self-distilled Dynamic Fusion Network to compose the multi-granularity features dynamically by considering the consistency of routing path and modality-specific information simultaneously. Two new modules are included in our proposed method: (1) Dynamic Fusion Network with Modality Specific Routers. The dynamic network enables a flexible determination of the routing for each reference image and modification text, taking into account their distinct semantics and distributions. (2) Self Path Distillation Loss. A stable path decision for queries benefits the optimization of feature extraction as well as routing, and we approach this by progressively refine the path decision with previous path information. Extensive experiments demonstrate the effectiveness of our proposed model compared to existing methods.
title Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
topic Computer Vision and Pattern Recognition
Information Retrieval
Multimedia
url https://arxiv.org/abs/2405.15451