WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917488087793664 |
|---|---|
| author | Santos, Jônatas H. dos Reis, Julio C. S. Melo, Philipe Olivetti, João F. H. Silva, Thales H. Guimaraes, Matheus Gontijo de Souza, Glaucio Gonçalves, Marcos A. Benevenuto, Fabricio Zanovello, Filipe B. B. Rodrigues, Marco A. G. Lima, Cristiano X. |
| author_facet | Santos, Jônatas H. dos Reis, Julio C. S. Melo, Philipe Olivetti, João F. H. Silva, Thales H. Guimaraes, Matheus Gontijo de Souza, Glaucio Gonçalves, Marcos A. Benevenuto, Fabricio Zanovello, Filipe B. B. Rodrigues, Marco A. G. Lima, Cristiano X. |
| contents | We introduce WhaVax, a new expert-annotated dataset of vaccine-related WhatsApp messages collected from large Brazilian public groups spanning multiple pandemic years. The dataset was constructed through a rigorous, carefully designed pipeline that integrates keyword-based data collection, semantic deduplication to remove near-duplicate content, and a multi-stage annotation protocol conducted by medical specialists. This process produced a high-quality gold-standard corpus, characterized by substantial inter-annotator agreement and strong reliability for downstream analysis. Additionally, we provide a detailed characterization of WhatsApp misinformation, revealing distinctive linguistic, structural, lexical, temporal, and group-level patterns, as well as a meaningful layer of ambiguous cases that reflect the complexity of health discourse in private messaging. We also benchmark classical models, fine-tuned Small Language Models, and zero- or few-shot Large Language Models under realistic data-scarcity constraints, demonstrating that strong embeddings and LLM approaches perform competitively, while domain alignment and data availability remain critical factors. This study provides a rare, high-quality resource to support misinformation research and computational modeling in encrypted communication environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_12510 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection Santos, Jônatas H. dos Reis, Julio C. S. Melo, Philipe Olivetti, João F. H. Silva, Thales H. Guimaraes, Matheus Gontijo de Souza, Glaucio Gonçalves, Marcos A. Benevenuto, Fabricio Zanovello, Filipe B. B. Rodrigues, Marco A. G. Lima, Cristiano X. Social and Information Networks Computation and Language Computers and Society We introduce WhaVax, a new expert-annotated dataset of vaccine-related WhatsApp messages collected from large Brazilian public groups spanning multiple pandemic years. The dataset was constructed through a rigorous, carefully designed pipeline that integrates keyword-based data collection, semantic deduplication to remove near-duplicate content, and a multi-stage annotation protocol conducted by medical specialists. This process produced a high-quality gold-standard corpus, characterized by substantial inter-annotator agreement and strong reliability for downstream analysis. Additionally, we provide a detailed characterization of WhatsApp misinformation, revealing distinctive linguistic, structural, lexical, temporal, and group-level patterns, as well as a meaningful layer of ambiguous cases that reflect the complexity of health discourse in private messaging. We also benchmark classical models, fine-tuned Small Language Models, and zero- or few-shot Large Language Models under realistic data-scarcity constraints, demonstrating that strong embeddings and LLM approaches perform competitively, while domain alignment and data availability remain critical factors. This study provides a rare, high-quality resource to support misinformation research and computational modeling in encrypted communication environments. |
| title | WhatsApp Vaccine Discourse (WhaVax): An Expert-Annotated Dataset and Benchmark for Health Misinformation Detection |
| topic | Social and Information Networks Computation and Language Computers and Society |
| url | https://arxiv.org/abs/2605.12510 |