Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Achmann-Denkler, Michael, Haim, Mario, Wolff, Christian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908983920427008
author Achmann-Denkler, Michael
Haim, Mario
Wolff, Christian
author_facet Achmann-Denkler, Michael
Haim, Mario
Wolff, Christian
contents This paper presents a computational case study that evaluates the capabilities of specialized machine learning models and emerging multimodal large language models for Visual Political Communication (VPC) analysis. Focusing on concentrated visibility in Instagram stories and posts during the 2021 German federal election campaign, we compare the performance of traditional computer vision models (FaceNet512, RetinaFace, Google Cloud Vision) with a multimodal large language model (GPT-4o) in identifying front-runner politicians and counting individuals in images. GPT-4o outperformed the other models, achieving a macro F1-score of 0.89 for face recognition and 0.86 for person counting in stories. These findings demonstrate the potential of advanced AI systems to scale and refine visual content analysis in political communication while highlighting methodological considerations for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19489
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram
Achmann-Denkler, Michael
Haim, Mario
Wolff, Christian
Computer Vision and Pattern Recognition
Computers and Society
This paper presents a computational case study that evaluates the capabilities of specialized machine learning models and emerging multimodal large language models for Visual Political Communication (VPC) analysis. Focusing on concentrated visibility in Instagram stories and posts during the 2021 German federal election campaign, we compare the performance of traditional computer vision models (FaceNet512, RetinaFace, Google Cloud Vision) with a multimodal large language model (GPT-4o) in identifying front-runner politicians and counting individuals in images. GPT-4o outperformed the other models, achieving a macro F1-score of 0.89 for face recognition and 0.86 for person counting in stories. These findings demonstrate the potential of advanced AI systems to scale and refine visual content analysis in political communication while highlighting methodological considerations for future research.
title Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram
topic Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2604.19489