taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Urchs, Stefanie, Thurner, Veronika, Aßenmacher, Matthias, Heumann, Christian, Thiemichen, Stephanie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915328727973888
author Urchs, Stefanie
Thurner, Veronika
Aßenmacher, Matthias
Heumann, Christian
Thiemichen, Stephanie
author_facet Urchs, Stefanie
Thurner, Veronika
Aßenmacher, Matthias
Heumann, Christian
Thiemichen, Stephanie
contents Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal issues such as gender bias. We present taz2024full, the largest publicly available corpus of German newspaper articles to date, comprising over 1.8 million texts from taz, spanning 1980 to 2024. As a demonstration of the corpus's utility for bias and discrimination research, we analyse gender representation across four decades of reporting. We find a consistent overrepresentation of men, but also a gradual shift toward more balanced coverage in recent years. Using a scalable, structured analysis pipeline, we provide a foundation for studying actor mentions, sentiment, and linguistic framing in German journalistic texts. The corpus supports a wide range of applications, from diachronic language analysis to critical media studies, and is freely available to foster inclusive and reproducible research in German-language NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades
Urchs, Stefanie
Thurner, Veronika
Aßenmacher, Matthias
Heumann, Christian
Thiemichen, Stephanie
Computation and Language
Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal issues such as gender bias. We present taz2024full, the largest publicly available corpus of German newspaper articles to date, comprising over 1.8 million texts from taz, spanning 1980 to 2024. As a demonstration of the corpus's utility for bias and discrimination research, we analyse gender representation across four decades of reporting. We find a consistent overrepresentation of men, but also a gradual shift toward more balanced coverage in recent years. Using a scalable, structured analysis pipeline, we provide a foundation for studying actor mentions, sentiment, and linguistic framing in German journalistic texts. The corpus supports a wide range of applications, from diachronic language analysis to critical media studies, and is freely available to foster inclusive and reproducible research in German-language NLP.
title taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades
topic Computation and Language
url https://arxiv.org/abs/2506.05388