GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915604079837184 |
|---|---|
| author | Chee, Heng Er Metilda Wang, Jiayin Guo, Zhiqiang Ma, Weizhi Zhang, Min |
| author_facet | Chee, Heng Er Metilda Wang, Jiayin Guo, Zhiqiang Ma, Weizhi Zhang, Min |
| contents | Stickers have become a popular form of visual communication, yet understanding their semantic relationships remains challenging due to their highly diverse and symbolic content. In this work, we formally {define the Sticker Semantic Similarity task} and introduce {Triple-S}, the first benchmark for this task, consisting of 905 human-annotated positive and negative sticker pairs. Through extensive evaluation, we show that existing pretrained vision and multimodal models struggle to capture nuanced sticker semantics. To address this, we propose the {General Sticker Encoder (GSE)}, a lightweight and versatile model that learns robust sticker embeddings using both Triple-S and additional datasets. GSE achieves superior performance on unseen stickers, and demonstrates strong results on downstream tasks such as emotion classification and sticker-to-sticker retrieval. By releasing both Triple-S and GSE, we provide standardized evaluation tools and robust embeddings, enabling future research in sticker understanding, retrieval, and multimodal content generation. The Triple-S benchmark and GSE have been publicly released and are available here. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_04977 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder Chee, Heng Er Metilda Wang, Jiayin Guo, Zhiqiang Ma, Weizhi Zhang, Min Computer Vision and Pattern Recognition Multimedia Stickers have become a popular form of visual communication, yet understanding their semantic relationships remains challenging due to their highly diverse and symbolic content. In this work, we formally {define the Sticker Semantic Similarity task} and introduce {Triple-S}, the first benchmark for this task, consisting of 905 human-annotated positive and negative sticker pairs. Through extensive evaluation, we show that existing pretrained vision and multimodal models struggle to capture nuanced sticker semantics. To address this, we propose the {General Sticker Encoder (GSE)}, a lightweight and versatile model that learns robust sticker embeddings using both Triple-S and additional datasets. GSE achieves superior performance on unseen stickers, and demonstrates strong results on downstream tasks such as emotion classification and sticker-to-sticker retrieval. By releasing both Triple-S and GSE, we provide standardized evaluation tools and robust embeddings, enabling future research in sticker understanding, retrieval, and multimodal content generation. The Triple-S benchmark and GSE have been publicly released and are available here. |
| title | GSE: Evaluating Sticker Visual Semantic Similarity via a General Sticker Encoder |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2511.04977 |