Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912069533564928 |
|---|---|
| author | Johnson, Isaac Kaffee, Lucie-Aimée Redi, Miriam |
| author_facet | Johnson, Isaac Kaffee, Lucie-Aimée Redi, Miriam |
| contents | Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across pre-training, post-training, and model evaluations. We point to opportunities for greater use of Wikimedia content but also identify ways in which the language modeling community could better center the needs of Wikimedia editors. In particular, we call for incorporating additional sources of Wikimedia data, a greater focus on benchmarks for LLMs that encode Wikimedia principles, and greater multilingualism in Wikimedia-derived datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_08918 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing Johnson, Isaac Kaffee, Lucie-Aimée Redi, Miriam Computers and Society Wikimedia content is used extensively by the AI community and within the language modeling community in particular. In this paper, we provide a review of the different ways in which Wikimedia data is curated to use in NLP tasks across pre-training, post-training, and model evaluations. We point to opportunities for greater use of Wikimedia content but also identify ways in which the language modeling community could better center the needs of Wikimedia editors. In particular, we call for incorporating additional sources of Wikimedia data, a greater focus on benchmarks for LLMs that encode Wikimedia principles, and greater multilingualism in Wikimedia-derived datasets. |
| title | Wikimedia data for AI: a review of Wikimedia datasets for NLP tasks and AI-assisted editing |
| topic | Computers and Society |
| url | https://arxiv.org/abs/2410.08918 |