Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Atuhurra, Jesse, Kamigaito, Hidetaka
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916324034215936
author Atuhurra, Jesse
Kamigaito, Hidetaka
author_facet Atuhurra, Jesse
Kamigaito, Hidetaka
contents Natural language processing (NLP) has grown significantly since the advent of the Transformer architecture. Transformers have given birth to pre-trained large language models (PLMs). There has been tremendous improvement in the performance of NLP systems across several tasks. NLP systems are on par or, in some cases, better than humans at accomplishing specific tasks. However, it remains the norm that \emph{better quality datasets at the time of pretraining enable PLMs to achieve better performance, regardless of the task.} The need to have quality datasets has prompted NLP researchers to continue creating new datasets to satisfy particular needs. For example, the two top NLP conferences, ACL and EMNLP, accepted ninety-two papers in 2022, introducing new datasets. This work aims to uncover the trends and insights mined within these datasets. Moreover, we provide valuable suggestions to researchers interested in curating datasets in the future.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08666
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences
Atuhurra, Jesse
Kamigaito, Hidetaka
Computation and Language
Machine Learning
Natural language processing (NLP) has grown significantly since the advent of the Transformer architecture. Transformers have given birth to pre-trained large language models (PLMs). There has been tremendous improvement in the performance of NLP systems across several tasks. NLP systems are on par or, in some cases, better than humans at accomplishing specific tasks. However, it remains the norm that \emph{better quality datasets at the time of pretraining enable PLMs to achieve better performance, regardless of the task.} The need to have quality datasets has prompted NLP researchers to continue creating new datasets to satisfy particular needs. For example, the two top NLP conferences, ACL and EMNLP, accepted ninety-two papers in 2022, introducing new datasets. This work aims to uncover the trends and insights mined within these datasets. Moreover, we provide valuable suggestions to researchers interested in curating datasets in the future.
title Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2404.08666