Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jing, Jie, Huang, Yongjian, Wang, Serena J. -W., Han, Shuangpeng, Schiatti, Lucia, Kuo, Yen-Ling, Lin, Qing, Zhang, Mengmi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911177144008704
author Jing, Jie
Huang, Yongjian
Wang, Serena J. -W.
Han, Shuangpeng
Schiatti, Lucia
Kuo, Yen-Ling
Lin, Qing
Zhang, Mengmi
author_facet Jing, Jie
Huang, Yongjian
Wang, Serena J. -W.
Han, Shuangpeng
Schiatti, Lucia
Kuo, Yen-Ling
Lin, Qing
Zhang, Mengmi
contents We introduce intra-class memorability, where certain images within the same class are more memorable than others despite shared category characteristics. To investigate what features make one object instance more memorable than others, we design and conduct human behavior experiments, where participants are shown a series of images, and they must identify when the current image matches the image presented a few steps back in the sequence. To quantify memorability, we propose the Intra-Class Memorability score (ICMscore), a novel metric that incorporates the temporal intervals between repeated image presentations into its calculation. Furthermore, we curate the Intra-Class Memorability Dataset (ICMD), comprising over 5,000 images across ten object classes with their ICMscores derived from 2,000 participants' responses. Subsequently, we demonstrate the usefulness of ICMD by training AI models on this dataset for various downstream tasks: memorability prediction, image recognition, continual learning, and memorability-controlled image editing. Surprisingly, high-ICMscore images impair AI performance in image recognition and continual learning tasks, while low-ICMscore images improve outcomes in these tasks. Additionally, we fine-tune a state-of-the-art image diffusion model on ICMD image pairs with and without masked semantic objects. The diffusion model can successfully manipulate image elements to enhance or reduce memorability. Our contributions open new pathways in understanding intra-class memorability by scrutinizing fine-grained visual features behind the most and least memorable images and laying the groundwork for real-world applications in computer vision. We will release all code, data, and models publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20761
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
Jing, Jie
Huang, Yongjian
Wang, Serena J. -W.
Han, Shuangpeng
Schiatti, Lucia
Kuo, Yen-Ling
Lin, Qing
Zhang, Mengmi
Computer Vision and Pattern Recognition
We introduce intra-class memorability, where certain images within the same class are more memorable than others despite shared category characteristics. To investigate what features make one object instance more memorable than others, we design and conduct human behavior experiments, where participants are shown a series of images, and they must identify when the current image matches the image presented a few steps back in the sequence. To quantify memorability, we propose the Intra-Class Memorability score (ICMscore), a novel metric that incorporates the temporal intervals between repeated image presentations into its calculation. Furthermore, we curate the Intra-Class Memorability Dataset (ICMD), comprising over 5,000 images across ten object classes with their ICMscores derived from 2,000 participants' responses. Subsequently, we demonstrate the usefulness of ICMD by training AI models on this dataset for various downstream tasks: memorability prediction, image recognition, continual learning, and memorability-controlled image editing. Surprisingly, high-ICMscore images impair AI performance in image recognition and continual learning tasks, while low-ICMscore images improve outcomes in these tasks. Additionally, we fine-tune a state-of-the-art image diffusion model on ICMD image pairs with and without masked semantic objects. The diffusion model can successfully manipulate image elements to enhance or reduce memorability. Our contributions open new pathways in understanding intra-class memorability by scrutinizing fine-grained visual features behind the most and least memorable images and laying the groundwork for real-world applications in computer vision. We will release all code, data, and models publicly.
title Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.20761