CLIP Multi-modal Hashing for Multimedia Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Jian, Sheng, Mingkai, Huang, Zhangmin, Chang, Jingfei, Jiang, Jinling, Long, Jian, Luo, Cheng, Liu, Lei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929535606325248
author Zhu, Jian
Sheng, Mingkai
Huang, Zhangmin
Chang, Jingfei
Jiang, Jinling
Long, Jian
Luo, Cheng
Liu, Lei
author_facet Zhu, Jian
Sheng, Mingkai
Huang, Zhangmin
Chang, Jingfei
Jiang, Jinling
Long, Jian
Luo, Cheng
Liu, Lei
contents Multi-modal hashing methods are widely used in multimedia retrieval, which can fuse multi-source data to generate binary hash code. However, the individual backbone networks have limited feature expression capabilities and are not jointly pre-trained on large-scale unsupervised multi-modal data, resulting in low retrieval accuracy. To address this issue, we propose a novel CLIP Multi-modal Hashing (CLIPMH) method. Our method employs the CLIP framework to extract both text and vision features and then fuses them to generate hash code. Due to enhancement on each modal feature, our method has great improvement in the retrieval performance of multi-modal hashing methods. Compared with state-of-the-art unsupervised and supervised multi-modal hashing methods, experiments reveal that the proposed CLIPMH can significantly improve performance (a maximum increase of 8.38% in mAP).
format Preprint
id arxiv_https___arxiv_org_abs_2410_07783
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIP Multi-modal Hashing for Multimedia Retrieval
Zhu, Jian
Sheng, Mingkai
Huang, Zhangmin
Chang, Jingfei
Jiang, Jinling
Long, Jian
Luo, Cheng
Liu, Lei
Computer Vision and Pattern Recognition
Multi-modal hashing methods are widely used in multimedia retrieval, which can fuse multi-source data to generate binary hash code. However, the individual backbone networks have limited feature expression capabilities and are not jointly pre-trained on large-scale unsupervised multi-modal data, resulting in low retrieval accuracy. To address this issue, we propose a novel CLIP Multi-modal Hashing (CLIPMH) method. Our method employs the CLIP framework to extract both text and vision features and then fuses them to generate hash code. Due to enhancement on each modal feature, our method has great improvement in the retrieval performance of multi-modal hashing methods. Compared with state-of-the-art unsupervised and supervised multi-modal hashing methods, experiments reveal that the proposed CLIPMH can significantly improve performance (a maximum increase of 8.38% in mAP).
title CLIP Multi-modal Hashing for Multimedia Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.07783