YT-30M: A multi-lingual multi-category dataset of YouTube comments
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916507549696000 |
|---|---|
| author | Dutta, Hridoy Sankar |
| author_facet | Dutta, Hridoy Sankar |
| contents | This paper introduces two large-scale multilingual comment datasets, YT-30M (and YT-100K) from YouTube. The analysis in this paper is performed on a smaller sample (YT-100K) of YT-30M. Both the datasets: YT-30M (full) and YT-100K (randomly selected 100K sample from YT-30M) are publicly released for further research. YT-30M (YT-100K) contains 32236173 (108694) comments posted by YouTube channel that belong to YouTube categories. Each comment is associated with a video ID, comment ID, commentor name, commentor channel ID, comment text, upvotes, original channel ID and category of the YouTube channel (e.g., 'News & Politics', 'Science & Technology', etc.). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_03465 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | YT-30M: A multi-lingual multi-category dataset of YouTube comments Dutta, Hridoy Sankar Social and Information Networks Artificial Intelligence Computation and Language Information Retrieval Machine Learning This paper introduces two large-scale multilingual comment datasets, YT-30M (and YT-100K) from YouTube. The analysis in this paper is performed on a smaller sample (YT-100K) of YT-30M. Both the datasets: YT-30M (full) and YT-100K (randomly selected 100K sample from YT-30M) are publicly released for further research. YT-30M (YT-100K) contains 32236173 (108694) comments posted by YouTube channel that belong to YouTube categories. Each comment is associated with a video ID, comment ID, commentor name, commentor channel ID, comment text, upvotes, original channel ID and category of the YouTube channel (e.g., 'News & Politics', 'Science & Technology', etc.). |
| title | YT-30M: A multi-lingual multi-category dataset of YouTube comments |
| topic | Social and Information Networks Artificial Intelligence Computation and Language Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2412.03465 |