An Image is Worth $K$ Topics: A Visual Structural Topic Model with Pretrained Image Embeddings

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Piqueras, Matías, Segerberg, Alexandra, Magnani, Matteo, Magnusson, Måns, Sladoje, Nataša
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916688496164864
author Piqueras, Matías
Segerberg, Alexandra
Magnani, Matteo
Magnusson, Måns
Sladoje, Nataša
author_facet Piqueras, Matías
Segerberg, Alexandra
Magnani, Matteo
Magnusson, Måns
Sladoje, Nataša
contents Political scientists are increasingly interested in analyzing visual content at scale. However, the existing computational toolbox is still in need of methods and models attuned to the specific challenges and goals of social and political inquiry. In this article, we introduce a visual Structural Topic Model (vSTM) that combines pretrained image embeddings with a structural topic model. This has important advantages compared to existing approaches. First, pretrained embeddings allow the model to capture the semantic complexity of images relevant to political contexts. Second, the structural topic model provides the ability to analyze how topics and covariates are related, while maintaining a nuanced representation of images as a mixture of multiple topics. In our empirical application, we show that the vSTM is able to identify topics that are interpretable, coherent, and substantively relevant to the study of online political communication.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Image is Worth $K$ Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
Piqueras, Matías
Segerberg, Alexandra
Magnani, Matteo
Magnusson, Måns
Sladoje, Nataša
Computer Vision and Pattern Recognition
Computers and Society
Applications
Methodology
Political scientists are increasingly interested in analyzing visual content at scale. However, the existing computational toolbox is still in need of methods and models attuned to the specific challenges and goals of social and political inquiry. In this article, we introduce a visual Structural Topic Model (vSTM) that combines pretrained image embeddings with a structural topic model. This has important advantages compared to existing approaches. First, pretrained embeddings allow the model to capture the semantic complexity of images relevant to political contexts. Second, the structural topic model provides the ability to analyze how topics and covariates are related, while maintaining a nuanced representation of images as a mixture of multiple topics. In our empirical application, we show that the vSTM is able to identify topics that are interpretable, coherent, and substantively relevant to the study of online political communication.
title An Image is Worth $K$ Topics: A Visual Structural Topic Model with Pretrained Image Embeddings
topic Computer Vision and Pattern Recognition
Computers and Society
Applications
Methodology
url https://arxiv.org/abs/2504.10004