Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nareti, Utsav Kumar, Adak, Chandranath, Chattopadhyay, Soumi, Wang, Pichao
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912138406133760
author Nareti, Utsav Kumar
Adak, Chandranath
Chattopadhyay, Soumi
Wang, Pichao
author_facet Nareti, Utsav Kumar
Adak, Chandranath
Chattopadhyay, Soumi
Wang, Pichao
contents Movie posters are not just decorative; they are meticulously designed to capture the essence of a movie, such as its genre, storyline, and tone/vibe. For decades, movie posters have graced cinema walls, billboards, and now our digital screens as a form of digital posters. Movie genre classification plays a pivotal role in film marketing, audience engagement, and recommendation systems. Previous explorations into movie genre classification have been mostly examined in plot summaries, subtitles, trailers and movie scenes. Movie posters provide a pre-release tantalizing glimpse into a film's key aspects, which can ignite public interest. In this paper, we presented the framework that exploits movie posters from a visual and textual perspective to address the multilabel movie genre classification problem. Firstly, we extracted text from movie posters using an OCR and retrieved the relevant embedding. Next, we introduce a cross-attention-based fusion module to allocate attention weights to visual and textual embedding. In validating our framework, we utilized 13882 posters sourced from the Internet Movie Database (IMDb). The outcomes of the experiments indicate that our model exhibited promising performance and outperformed even some prominent contemporary architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
Nareti, Utsav Kumar
Adak, Chandranath
Chattopadhyay, Soumi
Wang, Pichao
Information Retrieval
Artificial Intelligence
Multimedia
Movie posters are not just decorative; they are meticulously designed to capture the essence of a movie, such as its genre, storyline, and tone/vibe. For decades, movie posters have graced cinema walls, billboards, and now our digital screens as a form of digital posters. Movie genre classification plays a pivotal role in film marketing, audience engagement, and recommendation systems. Previous explorations into movie genre classification have been mostly examined in plot summaries, subtitles, trailers and movie scenes. Movie posters provide a pre-release tantalizing glimpse into a film's key aspects, which can ignite public interest. In this paper, we presented the framework that exploits movie posters from a visual and textual perspective to address the multilabel movie genre classification problem. Firstly, we extracted text from movie posters using an OCR and retrieved the relevant embedding. Next, we introduce a cross-attention-based fusion module to allocate attention weights to visual and textual embedding. In validating our framework, we utilized 13882 posters sourced from the Internet Movie Database (IMDb). The outcomes of the experiments indicate that our model exhibited promising performance and outperformed even some prominent contemporary architectures.
title Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
topic Information Retrieval
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2410.19764