A Study on Scaling Up Multilingual News Framing Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akter, Syeda Sabrina, Anastasopoulos, Antonios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917628822421504
author Akter, Syeda Sabrina
Anastasopoulos, Antonios
author_facet Akter, Syeda Sabrina
Anastasopoulos, Antonios
contents Media framing is the study of strategically selecting and presenting specific aspects of political issues to shape public opinion. Despite its relevance to almost all societies around the world, research has been limited due to the lack of available datasets and other resources. This study explores the possibility of dataset creation through crowdsourcing, utilizing non-expert annotators to develop training corpora. We first extend framing analysis beyond English news to a multilingual context (12 typologically diverse languages) through automatic translation. We also present a novel benchmark in Bengali and Portuguese on the immigration and same-sex marriage domains. Additionally, we show that a system trained on our crowd-sourced dataset, combined with other existing ones, leads to a 5.32 percentage point increase from the baseline, showing that crowdsourcing is a viable option. Last, we study the performance of large language models (LLMs) for this task, finding that task-specific fine-tuning is a better approach than employing bigger non-specialized models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01481
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Study on Scaling Up Multilingual News Framing Analysis
Akter, Syeda Sabrina
Anastasopoulos, Antonios
Computation and Language
Media framing is the study of strategically selecting and presenting specific aspects of political issues to shape public opinion. Despite its relevance to almost all societies around the world, research has been limited due to the lack of available datasets and other resources. This study explores the possibility of dataset creation through crowdsourcing, utilizing non-expert annotators to develop training corpora. We first extend framing analysis beyond English news to a multilingual context (12 typologically diverse languages) through automatic translation. We also present a novel benchmark in Bengali and Portuguese on the immigration and same-sex marriage domains. Additionally, we show that a system trained on our crowd-sourced dataset, combined with other existing ones, leads to a 5.32 percentage point increase from the baseline, showing that crowdsourcing is a viable option. Last, we study the performance of large language models (LLMs) for this task, finding that task-specific fine-tuning is a better approach than employing bigger non-specialized models.
title A Study on Scaling Up Multilingual News Framing Analysis
topic Computation and Language
url https://arxiv.org/abs/2404.01481