Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gatto, Joseph, Basak, Madhusudan, Srivastava, Yash, Bohlman, Philip, Preum, Sarah M.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916148583333888
author Gatto, Joseph
Basak, Madhusudan
Srivastava, Yash
Bohlman, Philip
Preum, Sarah M.
author_facet Gatto, Joseph
Basak, Madhusudan
Srivastava, Yash
Bohlman, Philip
Preum, Sarah M.
contents In this paper, we develop an LLM-powered framework for the curation and evaluation of emerging opinion mining in online health communities. We formulate emerging opinion mining as a pairwise stance detection problem between (title, comment) pairs sourced from Reddit, where post titles contain emerging health-related claims on a topic that is not predefined. The claims are either explicitly or implicitly expressed by the user. We detail (i) a method of claim identification -- the task of identifying if a post title contains a claim and (ii) an opinion mining-driven evaluation framework for stance detection using LLMs. We facilitate our exploration by releasing a novel test dataset, Long COVID-Stance, or LC-stance, which can be used to evaluate LLMs on the tasks of claim identification and stance detection in online health communities. Long Covid is an emerging post-COVID disorder with uncertain and complex treatment guidelines, thus making it a suitable use case for our task. LC-Stance contains long COVID treatment related discourse sourced from a Reddit community. Our evaluation shows that GPT-4 significantly outperforms prior works on zero-shot stance detection. We then perform thorough LLM model diagnostics, identifying the role of claim type (i.e. implicit vs explicit claims) and comment length as sources of model error.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03336
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse
Gatto, Joseph
Basak, Madhusudan
Srivastava, Yash
Bohlman, Philip
Preum, Sarah M.
Computation and Language
Social and Information Networks
In this paper, we develop an LLM-powered framework for the curation and evaluation of emerging opinion mining in online health communities. We formulate emerging opinion mining as a pairwise stance detection problem between (title, comment) pairs sourced from Reddit, where post titles contain emerging health-related claims on a topic that is not predefined. The claims are either explicitly or implicitly expressed by the user. We detail (i) a method of claim identification -- the task of identifying if a post title contains a claim and (ii) an opinion mining-driven evaluation framework for stance detection using LLMs. We facilitate our exploration by releasing a novel test dataset, Long COVID-Stance, or LC-stance, which can be used to evaluate LLMs on the tasks of claim identification and stance detection in online health communities. Long Covid is an emerging post-COVID disorder with uncertain and complex treatment guidelines, thus making it a suitable use case for our task. LC-Stance contains long COVID treatment related discourse sourced from a Reddit community. Our evaluation shows that GPT-4 significantly outperforms prior works on zero-shot stance detection. We then perform thorough LLM model diagnostics, identifying the role of claim type (i.e. implicit vs explicit claims) and comment length as sources of model error.
title Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2403.03336