Multimodal Transformer for Comics Text-Cloze

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vivoli, Emanuele, Baeza, Joan Lafuente, Llobet, Ernest Valveny, Karatzas, Dimosthenis
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910355419037696
author Vivoli, Emanuele
Baeza, Joan Lafuente
Llobet, Ernest Valveny
Karatzas, Dimosthenis
author_facet Vivoli, Emanuele
Baeza, Joan Lafuente
Llobet, Ernest Valveny
Karatzas, Dimosthenis
contents This work explores a closure task in comics, a medium where visual and textual elements are intricately intertwined. Specifically, Text-cloze refers to the task of selecting the correct text to use in a comic panel, given its neighboring panels. Traditional methods based on recurrent neural networks have struggled with this task due to limited OCR accuracy and inherent model limitations. We introduce a novel Multimodal Large Language Model (Multimodal-LLM) architecture, specifically designed for Text-cloze, achieving a 10% improvement over existing state-of-the-art models in both its easy and hard variants. Central to our approach is a Domain-Adapted ResNet-50 based visual encoder, fine-tuned to the comics domain in a self-supervised manner using SimCLR. This encoder delivers comparable results to more complex models with just one-fifth of the parameters. Additionally, we release new OCR annotations for this dataset, enhancing model input quality and resulting in another 1% improvement. Finally, we extend the task to a generative format, establishing new baselines and expanding the research possibilities in the field of comics analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03719
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Transformer for Comics Text-Cloze
Vivoli, Emanuele
Baeza, Joan Lafuente
Llobet, Ernest Valveny
Karatzas, Dimosthenis
Computer Vision and Pattern Recognition
This work explores a closure task in comics, a medium where visual and textual elements are intricately intertwined. Specifically, Text-cloze refers to the task of selecting the correct text to use in a comic panel, given its neighboring panels. Traditional methods based on recurrent neural networks have struggled with this task due to limited OCR accuracy and inherent model limitations. We introduce a novel Multimodal Large Language Model (Multimodal-LLM) architecture, specifically designed for Text-cloze, achieving a 10% improvement over existing state-of-the-art models in both its easy and hard variants. Central to our approach is a Domain-Adapted ResNet-50 based visual encoder, fine-tuned to the comics domain in a self-supervised manner using SimCLR. This encoder delivers comparable results to more complex models with just one-fifth of the parameters. Additionally, we release new OCR annotations for this dataset, enhancing model input quality and resulting in another 1% improvement. Finally, we extend the task to a generative format, establishing new baselines and expanding the research possibilities in the field of comics analysis.
title Multimodal Transformer for Comics Text-Cloze
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.03719