WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Santos, Jônatas H. dos, Reis, Julio C. S., Melo, Philipe, Olivetti, João F. H., Silva, Thales H., Guimaraes, Matheus Gontijo, de Souza, Glaucio, Gonçalves, Marcos A., Benevenuto, Fabricio, Zanovello, Filipe B. B., Rodrigues, Marco A. G., Lima, Cristiano X.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917488087793664
author Santos, Jônatas H. dos
Reis, Julio C. S.
Melo, Philipe
Olivetti, João F. H.
Silva, Thales H.
Guimaraes, Matheus Gontijo
de Souza, Glaucio
Gonçalves, Marcos A.
Benevenuto, Fabricio
Zanovello, Filipe B. B.
Rodrigues, Marco A. G.
Lima, Cristiano X.
author_facet Santos, Jônatas H. dos
Reis, Julio C. S.
Melo, Philipe
Olivetti, João F. H.
Silva, Thales H.
Guimaraes, Matheus Gontijo
de Souza, Glaucio
Gonçalves, Marcos A.
Benevenuto, Fabricio
Zanovello, Filipe B. B.
Rodrigues, Marco A. G.
Lima, Cristiano X.
contents We introduce WhaVax, a new expert-annotated dataset of vaccine-related WhatsApp messages collected from large Brazilian public groups spanning multiple pandemic years. The dataset was constructed through a rigorous, carefully designed pipeline that integrates keyword-based data collection, semantic deduplication to remove near-duplicate content, and a multi-stage annotation protocol conducted by medical specialists. This process produced a high-quality gold-standard corpus, characterized by substantial inter-annotator agreement and strong reliability for downstream analysis. Additionally, we provide a detailed characterization of WhatsApp misinformation, revealing distinctive linguistic, structural, lexical, temporal, and group-level patterns, as well as a meaningful layer of ambiguous cases that reflect the complexity of health discourse in private messaging. We also benchmark classical models, fine-tuned Small Language Models, and zero- or few-shot Large Language Models under realistic data-scarcity constraints, demonstrating that strong embeddings and LLM approaches perform competitively, while domain alignment and data availability remain critical factors. This study provides a rare, high-quality resource to support misinformation research and computational modeling in encrypted communication environments.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12510
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection
Santos, Jônatas H. dos
Reis, Julio C. S.
Melo, Philipe
Olivetti, João F. H.
Silva, Thales H.
Guimaraes, Matheus Gontijo
de Souza, Glaucio
Gonçalves, Marcos A.
Benevenuto, Fabricio
Zanovello, Filipe B. B.
Rodrigues, Marco A. G.
Lima, Cristiano X.
Social and Information Networks
Computation and Language
Computers and Society
We introduce WhaVax, a new expert-annotated dataset of vaccine-related WhatsApp messages collected from large Brazilian public groups spanning multiple pandemic years. The dataset was constructed through a rigorous, carefully designed pipeline that integrates keyword-based data collection, semantic deduplication to remove near-duplicate content, and a multi-stage annotation protocol conducted by medical specialists. This process produced a high-quality gold-standard corpus, characterized by substantial inter-annotator agreement and strong reliability for downstream analysis. Additionally, we provide a detailed characterization of WhatsApp misinformation, revealing distinctive linguistic, structural, lexical, temporal, and group-level patterns, as well as a meaningful layer of ambiguous cases that reflect the complexity of health discourse in private messaging. We also benchmark classical models, fine-tuned Small Language Models, and zero- or few-shot Large Language Models under realistic data-scarcity constraints, demonstrating that strong embeddings and LLM approaches perform competitively, while domain alignment and data availability remain critical factors. This study provides a rare, high-quality resource to support misinformation research and computational modeling in encrypted communication environments.
title WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection
topic Social and Information Networks
Computation and Language
Computers and Society
url https://arxiv.org/abs/2605.12510