PVminer: A Domain-Specific Tool to Detect the Patient Voice in Patient Generated Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fodeh, Samah, Ma, Linhai, Wang, Yan, Talakokkul, Srivani, Puthiaraju, Ganesh, Khan, Afshan, Hagaman, Ashley, Lowe, Sarah, Roundtree, Aimee
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914348218187776
author Fodeh, Samah
Ma, Linhai
Wang, Yan
Talakokkul, Srivani
Puthiaraju, Ganesh
Khan, Afshan
Hagaman, Ashley
Lowe, Sarah
Roundtree, Aimee
author_facet Fodeh, Samah
Ma, Linhai
Wang, Yan
Talakokkul, Srivani
Puthiaraju, Ganesh
Khan, Afshan
Hagaman, Ashley
Lowe, Sarah
Roundtree, Aimee
contents Patient-generated text such as secure messages, surveys, and interviews contains rich expressions of the patient voice (PV), reflecting communicative behaviors and social determinants of health (SDoH). Traditional qualitative coding frameworks are labor intensive and do not scale to large volumes of patient-authored messages across health systems. Existing machine learning (ML) and natural language processing (NLP) approaches provide partial solutions but often treat patient-centered communication (PCC) and SDoH as separate tasks or rely on models not well suited to patient-facing language. We introduce PVminer, a domain-adapted NLP framework for structuring patient voice in secure patient-provider communication. PVminer formulates PV detection as a multi-label, multi-class prediction task integrating patient-specific BERT encoders (PV-BERT-base and PV-BERT-large), unsupervised topic modeling for thematic augmentation (PV-Topic-BERT), and fine-tuned classifiers for Code, Subcode, and Combo-level labels. Topic representations are incorporated during fine-tuning and inference to enrich semantic inputs. PVminer achieves strong performance across hierarchical tasks and outperforms biomedical and clinical pre-trained baselines, achieving F1 scores of 82.25% (Code), 80.14% (Subcode), and up to 77.87% (Combo). An ablation study further shows that author identity and topic-based augmentation each contribute meaningful gains. Pre-trained models, source code, and documentation will be publicly released, with annotated datasets available upon request for research use.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21165
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PVminer: A Domain-Specific Tool to Detect the Patient Voice in Patient Generated Data
Fodeh, Samah
Ma, Linhai
Wang, Yan
Talakokkul, Srivani
Puthiaraju, Ganesh
Khan, Afshan
Hagaman, Ashley
Lowe, Sarah
Roundtree, Aimee
Computation and Language
Artificial Intelligence
Patient-generated text such as secure messages, surveys, and interviews contains rich expressions of the patient voice (PV), reflecting communicative behaviors and social determinants of health (SDoH). Traditional qualitative coding frameworks are labor intensive and do not scale to large volumes of patient-authored messages across health systems. Existing machine learning (ML) and natural language processing (NLP) approaches provide partial solutions but often treat patient-centered communication (PCC) and SDoH as separate tasks or rely on models not well suited to patient-facing language. We introduce PVminer, a domain-adapted NLP framework for structuring patient voice in secure patient-provider communication. PVminer formulates PV detection as a multi-label, multi-class prediction task integrating patient-specific BERT encoders (PV-BERT-base and PV-BERT-large), unsupervised topic modeling for thematic augmentation (PV-Topic-BERT), and fine-tuned classifiers for Code, Subcode, and Combo-level labels. Topic representations are incorporated during fine-tuning and inference to enrich semantic inputs. PVminer achieves strong performance across hierarchical tasks and outperforms biomedical and clinical pre-trained baselines, achieving F1 scores of 82.25% (Code), 80.14% (Subcode), and up to 77.87% (Combo). An ablation study further shows that author identity and topic-based augmentation each contribute meaningful gains. Pre-trained models, source code, and documentation will be publicly released, with annotated datasets available upon request for research use.
title PVminer: A Domain-Specific Tool to Detect the Patient Voice in Patient Generated Data
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.21165