Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Villegas, Danae Sánchez, Preoţiuc-Pietro, Daniel, Aletras, Nikolaos
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916113853448192
author Villegas, Danae Sánchez
Preoţiuc-Pietro, Daniel
Aletras, Nikolaos
author_facet Villegas, Danae Sánchez
Preoţiuc-Pietro, Daniel
Aletras, Nikolaos
contents Effectively leveraging multimodal information from social media posts is essential to various downstream tasks such as sentiment analysis, sarcasm detection or hate speech classification. Jointly modeling text and images is challenging because cross-modal semantics might be hidden or the relation between image and text is weak. However, prior work on multimodal classification of social media posts has not yet addressed these challenges. In this work, we present an extensive study on the effectiveness of using two auxiliary losses jointly with the main task during fine-tuning multimodal models. First, Image-Text Contrastive (ITC) is designed to minimize the distance between image-text representations within a post, thereby effectively bridging the gap between posts where the image plays an important role in conveying the post's meaning. Second, Image-Text Matching (ITM) enhances the model's ability to understand the semantic relationship between images and text, thus improving its capacity to handle ambiguous or loosely related modalities. We combine these objectives with five multimodal models across five diverse social media datasets, demonstrating consistent improvements of up to 2.6 points F1. Our comprehensive analysis shows the specific scenarios where each auxiliary task is most effective.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07794
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
Villegas, Danae Sánchez
Preoţiuc-Pietro, Daniel
Aletras, Nikolaos
Computation and Language
Machine Learning
Social and Information Networks
Effectively leveraging multimodal information from social media posts is essential to various downstream tasks such as sentiment analysis, sarcasm detection or hate speech classification. Jointly modeling text and images is challenging because cross-modal semantics might be hidden or the relation between image and text is weak. However, prior work on multimodal classification of social media posts has not yet addressed these challenges. In this work, we present an extensive study on the effectiveness of using two auxiliary losses jointly with the main task during fine-tuning multimodal models. First, Image-Text Contrastive (ITC) is designed to minimize the distance between image-text representations within a post, thereby effectively bridging the gap between posts where the image plays an important role in conveying the post's meaning. Second, Image-Text Matching (ITM) enhances the model's ability to understand the semantic relationship between images and text, thus improving its capacity to handle ambiguous or loosely related modalities. We combine these objectives with five multimodal models across five diverse social media datasets, demonstrating consistent improvements of up to 2.6 points F1. Our comprehensive analysis shows the specific scenarios where each auxiliary task is most effective.
title Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks
topic Computation and Language
Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2309.07794