Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kubota, Yugo, Haraguchi, Daichi, Uchida, Seiichi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910343709589504
author Kubota, Yugo
Haraguchi, Daichi
Uchida, Seiichi
author_facet Kubota, Yugo
Haraguchi, Daichi
Uchida, Seiichi
contents Fonts convey different impressions to readers. These impressions often come from the font shapes. However, the correlation between fonts and their impression is weak and unstable because impressions are subjective. To capture such weak and unstable cross-modal correlation between font shapes and their impressions, we propose Impression-CLIP, which is a novel machine-learning model based on CLIP (Contrastive Language-Image Pre-training). By using the CLIP-based model, font image features and their impression features are pulled closer, and font image features and unrelated impression features are pushed apart. This procedure realizes co-embedding between font image and their impressions. In our experiment, we perform cross-modal retrieval between fonts and impressions through co-embedding. The results indicate that Impression-CLIP achieves better retrieval accuracy than the state-of-the-art method. Additionally, our model shows the robustness to noise and missing tags.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16350
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts
Kubota, Yugo
Haraguchi, Daichi
Uchida, Seiichi
Computer Vision and Pattern Recognition
Fonts convey different impressions to readers. These impressions often come from the font shapes. However, the correlation between fonts and their impression is weak and unstable because impressions are subjective. To capture such weak and unstable cross-modal correlation between font shapes and their impressions, we propose Impression-CLIP, which is a novel machine-learning model based on CLIP (Contrastive Language-Image Pre-training). By using the CLIP-based model, font image features and their impression features are pulled closer, and font image features and unrelated impression features are pushed apart. This procedure realizes co-embedding between font image and their impressions. In our experiment, we perform cross-modal retrieval between fonts and impressions through co-embedding. The results indicate that Impression-CLIP achieves better retrieval accuracy than the state-of-the-art method. Additionally, our model shows the robustness to noise and missing tags.
title Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.16350