Saved in:
Bibliographic Details
Main Authors: Ren, Libo, Belkadi, Samuel, Han, Lifeng, Del-Pinto, Warren, Nenadic, Goran
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.09501
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914949403508736
author Ren, Libo
Belkadi, Samuel
Han, Lifeng
Del-Pinto, Warren
Nenadic, Goran
author_facet Ren, Libo
Belkadi, Samuel
Han, Lifeng
Del-Pinto, Warren
Nenadic, Goran
contents Since clinical letters contain sensitive information, clinical-related datasets can not be widely applied in model training, medical research, and teaching. This work aims to generate reliable, various, and de-identified synthetic clinical letters. To achieve this goal, we explored different pre-trained language models (PLMs) for masking and generating text. After that, we worked on Bio\_ClinicalBERT, a high-performing model, and experimented with different masking strategies. Both qualitative and quantitative methods were used for evaluation. Additionally, a downstream task, Named Entity Recognition (NER), was also implemented to assess the usability of these synthetic letters. The results indicate that 1) encoder-only models outperform encoder-decoder models. 2) Among encoder-only models, those trained on general corpora perform comparably to those trained on clinical data when clinical information is preserved. 3) Additionally, preserving clinical entities and document structure better aligns with our objectives than simply fine-tuning the model. 4) Furthermore, different masking strategies can impact the quality of synthetic clinical letters. Masking stopwords has a positive impact, while masking nouns or verbs has a negative effect. 5) For evaluation, BERTScore should be the primary quantitative evaluation metric, with other metrics serving as supplementary references. 6) Contextual information does not significantly impact the models' understanding, so the synthetic clinical letters have the potential to replace the original ones in downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09501
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Synthetic4Health: Generating Annotated Synthetic Clinical Letters
Ren, Libo
Belkadi, Samuel
Han, Lifeng
Del-Pinto, Warren
Nenadic, Goran
Computation and Language
Artificial Intelligence
Since clinical letters contain sensitive information, clinical-related datasets can not be widely applied in model training, medical research, and teaching. This work aims to generate reliable, various, and de-identified synthetic clinical letters. To achieve this goal, we explored different pre-trained language models (PLMs) for masking and generating text. After that, we worked on Bio\_ClinicalBERT, a high-performing model, and experimented with different masking strategies. Both qualitative and quantitative methods were used for evaluation. Additionally, a downstream task, Named Entity Recognition (NER), was also implemented to assess the usability of these synthetic letters. The results indicate that 1) encoder-only models outperform encoder-decoder models. 2) Among encoder-only models, those trained on general corpora perform comparably to those trained on clinical data when clinical information is preserved. 3) Additionally, preserving clinical entities and document structure better aligns with our objectives than simply fine-tuning the model. 4) Furthermore, different masking strategies can impact the quality of synthetic clinical letters. Masking stopwords has a positive impact, while masking nouns or verbs has a negative effect. 5) For evaluation, BERTScore should be the primary quantitative evaluation metric, with other metrics serving as supplementary references. 6) Contextual information does not significantly impact the models' understanding, so the synthetic clinical letters have the potential to replace the original ones in downstream tasks.
title Synthetic4Health: Generating Annotated Synthetic Clinical Letters
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.09501