MMT: A Multilingual and Multi-Topic Indian Social Media Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dalal, Dwip, Srivastava, Vivek, Singh, Mayank
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908773556158464
author Dalal, Dwip
Srivastava, Vivek
Singh, Mayank
author_facet Dalal, Dwip
Srivastava, Vivek
Singh, Mayank
contents Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for processing such information, like language identification, topic modeling, and named-entity recognition. To address this, we introduce a large-scale multilingual, and multi-topic dataset (MMT) collected from Twitter (1.7 million Tweets), encompassing 13 coarse-grained and 63 fine-grained topics in the Indian context. We further annotate a subset of 5,346 tweets from the MMT dataset with various Indian languages and their code-mixed counterparts. Also, we demonstrate that the currently existing tools fail to capture the linguistic diversity in MMT on two downstream tasks, i.e., topic modeling and language identification. To facilitate future research, we have make the anonymized and annotated dataset available at https://huggingface.co/datasets/LingoIITGN/MMT.
format Preprint
id arxiv_https___arxiv_org_abs_2304_00634
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MMT: A Multilingual and Multi-Topic Indian Social Media Dataset
Dalal, Dwip
Srivastava, Vivek
Singh, Mayank
Computation and Language
Machine Learning
Social and Information Networks
Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for processing such information, like language identification, topic modeling, and named-entity recognition. To address this, we introduce a large-scale multilingual, and multi-topic dataset (MMT) collected from Twitter (1.7 million Tweets), encompassing 13 coarse-grained and 63 fine-grained topics in the Indian context. We further annotate a subset of 5,346 tweets from the MMT dataset with various Indian languages and their code-mixed counterparts. Also, we demonstrate that the currently existing tools fail to capture the linguistic diversity in MMT on two downstream tasks, i.e., topic modeling and language identification. To facilitate future research, we have make the anonymized and annotated dataset available at https://huggingface.co/datasets/LingoIITGN/MMT.
title MMT: A Multilingual and Multi-Topic Indian Social Media Dataset
topic Computation and Language
Machine Learning
Social and Information Networks
url https://arxiv.org/abs/2304.00634