Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bhattacharyya, Sree, Singla, Yaman Kumar, Yarram, Sudhir, Singh, Somesh Kumar, S I, Harini, Wang, James Z.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912729732743168
author Bhattacharyya, Sree
Singla, Yaman Kumar
Yarram, Sudhir
Singh, Somesh Kumar
S I, Harini
Wang, James Z.
author_facet Bhattacharyya, Sree
Singla, Yaman Kumar
Yarram, Sudhir
Singh, Somesh Kumar
S I, Harini
Wang, James Z.
contents Visual content memorability has intrigued the scientific community for decades, with applications ranging widely, from understanding nuanced aspects of human memory to enhancing content design. A significant challenge in progressing the field lies in the expensive process of collecting memorability annotations from humans. This limits the diversity and scalability of datasets for modeling visual content memorability. Most existing datasets are limited to collecting aggregate memorability scores for visual content, not capturing the nuanced memorability signals present in natural, open-ended recall descriptions. In this work, we introduce the first large-scale unsupervised dataset designed explicitly for modeling visual memorability signals, containing over 82,000 videos, accompanied by descriptive recall data. We leverage tip-of-the-tongue (ToT) retrieval queries from online platforms such as Reddit. We demonstrate that our unsupervised dataset provides rich signals for two memorability-related tasks: recall generation and ToT retrieval. Large vision-language models fine-tuned on our dataset outperform state-of-the-art models such as GPT-4o in generating open-ended memorability descriptions for visual content. We also employ a contrastive training strategy to create the first model capable of performing multimodal ToT retrieval. Our dataset and models present a novel direction, facilitating progress in visual content memorability research.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
Bhattacharyya, Sree
Singla, Yaman Kumar
Yarram, Sudhir
Singh, Somesh Kumar
S I, Harini
Wang, James Z.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Visual content memorability has intrigued the scientific community for decades, with applications ranging widely, from understanding nuanced aspects of human memory to enhancing content design. A significant challenge in progressing the field lies in the expensive process of collecting memorability annotations from humans. This limits the diversity and scalability of datasets for modeling visual content memorability. Most existing datasets are limited to collecting aggregate memorability scores for visual content, not capturing the nuanced memorability signals present in natural, open-ended recall descriptions. In this work, we introduce the first large-scale unsupervised dataset designed explicitly for modeling visual memorability signals, containing over 82,000 videos, accompanied by descriptive recall data. We leverage tip-of-the-tongue (ToT) retrieval queries from online platforms such as Reddit. We demonstrate that our unsupervised dataset provides rich signals for two memorability-related tasks: recall generation and ToT retrieval. Large vision-language models fine-tuned on our dataset outperform state-of-the-art models such as GPT-4o in generating open-ended memorability descriptions for visual content. We also employ a contrastive training strategy to create the first model capable of performing multimodal ToT retrieval. Our dataset and models present a novel direction, facilitating progress in visual content memorability research.
title Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.20854