Sentiment Analysis of COVID-19 Scientific Publication Dissemination on Social Media X: A Dataset Analyzed with ChatGPT 3.5 and Gemini 1.5 Flash
Fuente:
Zenodo
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2025
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901968988930048 |
|---|---|
| author | Pontes, Danielle Maricato, João de Melo |
| author_facet | Pontes, Danielle Maricato, João de Melo |
| contents | <p>The dataset provided includes a sample of posts on X that mentioned the editorial published in the journal “Dying in a Leadership Vacuum” on October 7, 2020, with the title “Dying in a Leadership Vacuum” (DOI: 10.1056/NEJMe2029812).</p> <p>A sample of posts on X that referenced the publication was collected. The posts were extracted from the Altmetric platform using a Python 3.12 algorithm with the Beautiful Soup 4.12 library and the Google Colab development environment. As a result, a dataset was generated containing 9,792 posts on X that specifically commented on the aforementioned editorial. Among these posts, 5,601 unique profiles were identified and cross-referenced with the profiles classified and made available in the dataset created by Pontes and Maricato (2023a). From the accounts that had an existing classification (bot or human), 41 accounts that had made more than four posts were selected.</p> <p>According to the dataset provided by Pontes and Maricato (2023), 10 accounts were classified as bots by Botometer, while 31 were classified as human. Considering that Pontes and Maricato (2023) highlighted the limitations of using Botometer for classifying accounts in the altmetric attention network, a manual classification of the 41 selected accounts was conducted. The manual classification was based on criteria such as the number of posts, posting times, time intervals between posts, account creation dates, and profile pictures. Through this manual classification, it was determined that 20 accounts were bots and 21 were human.</p> <p>The classified accounts posted a total of 3,493 posts, which are included in this dataset and were used in the analyses presented in the article.</p> <p>The metadata structure of the dataset is presented below.</p> <div> <p>Variable Name: ACCOUNT <br>Data Type: String (Text) <br>Description: Anonymized account code to preserve user identity. <br>Possible Values: ACCOUNT + sequential number </p> <p>Variable Name: ACCOUNT CLASS (BTM) <br>Data Type: Categorical <br>Description: Automatic account classification using a tool like Botometer. <br>Possible Values: human, bot </p> <p>Variable Name: ACCOUNT CLASS (MANUAL) <br>Data Type: Categorical <br>Description: Manual account classification based on researcher analysis. <br>Possible Values: human, bot </p> <p>Variable Name: POST CONTENT <br>Data Type: Text (String) <br>Description: Full text of the collected post. </p> <p>Variable Name: SENTIMENT CLASS (GPT) <br>Data Type: Categorical <br>Description: Sentiment classification assigned by ChatGPT. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: SENTIMENT CLASS (GEMINI) <br>Data Type: Categorical <br>Description: Sentiment classification assigned by Gemini. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: GPT X GEM (MATCH/DIFFERENCE) <br>Data Type: Binary <br>Description: Indicates whether the sentiment classification was the same or different between ChatGPT and Gemini. <br>Possible Values: match, different </p> <p>Variable Name: POST CLASS (MANUAL) <br>Data Type: Categorical <br>Description: Manual sentiment classification of the post. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: GPT RESULT <br>Data Type: Categorical <br>Description: Evaluation of ChatGPT's classification against the manual standard. <br>Possible Values: correct, incorrect </p> <p>Variable Name: GEMINI RESULT <br>Data Type: Categorical <br>Description: Evaluation of Gemini's classification against the manual standard. <br>Possible Values: correct, incorrect </p> </div> <p><span> </span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_14919674 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Sentiment Analysis of COVID-19 Scientific Publication Dissemination on Social Media X: A Dataset Analyzed with ChatGPT 3.5 and Gemini 1.5 Flash Pontes, Danielle Maricato, João de Melo <p>The dataset provided includes a sample of posts on X that mentioned the editorial published in the journal “Dying in a Leadership Vacuum” on October 7, 2020, with the title “Dying in a Leadership Vacuum” (DOI: 10.1056/NEJMe2029812).</p> <p>A sample of posts on X that referenced the publication was collected. The posts were extracted from the Altmetric platform using a Python 3.12 algorithm with the Beautiful Soup 4.12 library and the Google Colab development environment. As a result, a dataset was generated containing 9,792 posts on X that specifically commented on the aforementioned editorial. Among these posts, 5,601 unique profiles were identified and cross-referenced with the profiles classified and made available in the dataset created by Pontes and Maricato (2023a). From the accounts that had an existing classification (bot or human), 41 accounts that had made more than four posts were selected.</p> <p>According to the dataset provided by Pontes and Maricato (2023), 10 accounts were classified as bots by Botometer, while 31 were classified as human. Considering that Pontes and Maricato (2023) highlighted the limitations of using Botometer for classifying accounts in the altmetric attention network, a manual classification of the 41 selected accounts was conducted. The manual classification was based on criteria such as the number of posts, posting times, time intervals between posts, account creation dates, and profile pictures. Through this manual classification, it was determined that 20 accounts were bots and 21 were human.</p> <p>The classified accounts posted a total of 3,493 posts, which are included in this dataset and were used in the analyses presented in the article.</p> <p>The metadata structure of the dataset is presented below.</p> <div> <p>Variable Name: ACCOUNT <br>Data Type: String (Text) <br>Description: Anonymized account code to preserve user identity. <br>Possible Values: ACCOUNT + sequential number </p> <p>Variable Name: ACCOUNT CLASS (BTM) <br>Data Type: Categorical <br>Description: Automatic account classification using a tool like Botometer. <br>Possible Values: human, bot </p> <p>Variable Name: ACCOUNT CLASS (MANUAL) <br>Data Type: Categorical <br>Description: Manual account classification based on researcher analysis. <br>Possible Values: human, bot </p> <p>Variable Name: POST CONTENT <br>Data Type: Text (String) <br>Description: Full text of the collected post. </p> <p>Variable Name: SENTIMENT CLASS (GPT) <br>Data Type: Categorical <br>Description: Sentiment classification assigned by ChatGPT. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: SENTIMENT CLASS (GEMINI) <br>Data Type: Categorical <br>Description: Sentiment classification assigned by Gemini. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: GPT X GEM (MATCH/DIFFERENCE) <br>Data Type: Binary <br>Description: Indicates whether the sentiment classification was the same or different between ChatGPT and Gemini. <br>Possible Values: match, different </p> <p>Variable Name: POST CLASS (MANUAL) <br>Data Type: Categorical <br>Description: Manual sentiment classification of the post. <br>Possible Values: positive, neutral, negative </p> <p>Variable Name: GPT RESULT <br>Data Type: Categorical <br>Description: Evaluation of ChatGPT's classification against the manual standard. <br>Possible Values: correct, incorrect </p> <p>Variable Name: GEMINI RESULT <br>Data Type: Categorical <br>Description: Evaluation of Gemini's classification against the manual standard. <br>Possible Values: correct, incorrect </p> </div> <p><span> </span></p> |
| title | Sentiment Analysis of COVID-19 Scientific Publication Dissemination on Social Media X: A Dataset Analyzed with ChatGPT 3.5 and Gemini 1.5 Flash |
| url | https://doi.org/10.5281/zenodo.14919674 |