Exploring Vision Language Models for Multimodal and Multilingual Stance Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vasilakes, Jake, Scarton, Carolina, Zhao, Zhixue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915127947689984
author Vasilakes, Jake
Scarton, Carolina
Zhao, Zhixue
author_facet Vasilakes, Jake
Scarton, Carolina
Zhao, Zhixue
contents Social media's global reach amplifies the spread of information, highlighting the need for robust Natural Language Processing tasks like stance detection across languages and modalities. Prior research predominantly focuses on text-only inputs, leaving multimodal scenarios, such as those involving both images and text, relatively underexplored. Meanwhile, the prevalence of multimodal posts has increased significantly in recent years. Although state-of-the-art Vision-Language Models (VLMs) show promise, their performance on multimodal and multilingual stance detection tasks remains largely unexamined. This paper evaluates state-of-the-art VLMs on a newly extended dataset covering seven languages and multimodal inputs, investigating their use of visual cues, language-specific performance, and cross-modality interactions. Our results show that VLMs generally rely more on text than images for stance detection and this trend persists across languages. Additionally, VLMs rely significantly more on text contained within the images than other visual content. Regarding multilinguality, the models studied tend to generate consistent predictions across languages whether they are explicitly multilingual or not, although there are outliers that are incongruous with macro F1, language support, and model size.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
Vasilakes, Jake
Scarton, Carolina
Zhao, Zhixue
Computation and Language
Artificial Intelligence
Social media's global reach amplifies the spread of information, highlighting the need for robust Natural Language Processing tasks like stance detection across languages and modalities. Prior research predominantly focuses on text-only inputs, leaving multimodal scenarios, such as those involving both images and text, relatively underexplored. Meanwhile, the prevalence of multimodal posts has increased significantly in recent years. Although state-of-the-art Vision-Language Models (VLMs) show promise, their performance on multimodal and multilingual stance detection tasks remains largely unexamined. This paper evaluates state-of-the-art VLMs on a newly extended dataset covering seven languages and multimodal inputs, investigating their use of visual cues, language-specific performance, and cross-modality interactions. Our results show that VLMs generally rely more on text than images for stance detection and this trend persists across languages. Additionally, VLMs rely significantly more on text contained within the images than other visual content. Regarding multilinguality, the models studied tend to generate consistent predictions across languages whether they are explicitly multilingual or not, although there are outliers that are incongruous with macro F1, language support, and model size.
title Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.17654