Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Minzhi, Shi, Weiyan, Ziems, Caleb, Yang, Diyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911807601377280
author Li, Minzhi
Shi, Weiyan
Ziems, Caleb
Yang, Diyi
author_facet Li, Minzhi
Shi, Weiyan
Ziems, Caleb
Yang, Diyi
contents As Natural Language Processing (NLP) systems become increasingly integrated into human social life, these technologies will need to increasingly rely on social intelligence. Although there are many valuable datasets that benchmark isolated dimensions of social intelligence, there does not yet exist any body of work to join these threads into a cohesive subfield in which researchers can quickly identify research gaps and future directions. Towards this goal, we build a Social AI Data Infrastructure, which consists of a comprehensive social AI taxonomy and a data library of 480 NLP datasets. Our infrastructure allows us to analyze existing dataset efforts, and also evaluate language models' performance in different social intelligence aspects. Our analyses demonstrate its utility in enabling a thorough understanding of current data landscape and providing a holistic perspective on potential directions for future dataset development. We show there is a need for multifaceted datasets, increased diversity in language and culture, more long-tailed social situations, and more interactive data in future social intelligence data efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14659
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future
Li, Minzhi
Shi, Weiyan
Ziems, Caleb
Yang, Diyi
Computers and Society
Artificial Intelligence
Computation and Language
As Natural Language Processing (NLP) systems become increasingly integrated into human social life, these technologies will need to increasingly rely on social intelligence. Although there are many valuable datasets that benchmark isolated dimensions of social intelligence, there does not yet exist any body of work to join these threads into a cohesive subfield in which researchers can quickly identify research gaps and future directions. Towards this goal, we build a Social AI Data Infrastructure, which consists of a comprehensive social AI taxonomy and a data library of 480 NLP datasets. Our infrastructure allows us to analyze existing dataset efforts, and also evaluate language models' performance in different social intelligence aspects. Our analyses demonstrate its utility in enabling a thorough understanding of current data landscape and providing a holistic perspective on potential directions for future dataset development. We show there is a need for multifaceted datasets, increased diversity in language and culture, more long-tailed social situations, and more interactive data in future social intelligence data efforts.
title Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2403.14659