Automated extraction of information from free text of Spanish oncology pathology reports

Fuente: Redalyc
Saved in:
Bibliographic Details
Main Author: Diana Marcela Mendoza-Urbano
Format: Artículo científico
Language:en
Published: Universidad del Valle 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1876456614296289280
author Diana Marcela Mendoza-Urbano
author_facet Diana Marcela Mendoza-Urbano
contents Automated extraction of information from free text of Spanish oncology pathology reports Diana Marcela Mendoza-Urbano Johan Felipe Garcia Juan Sebastian Moreno Juan Carlos Bravo-Ocaña Alvaro José Riascos Angela Zambrano Harvey Sergio I Prada Medicina algorithm data science ontology learning regular expressions artificial intelligence Background: Pathology reports are stored as unstructured, ungrammatical, fragmented, and abbreviated free text with linguistic variability among pathologists. For this reason, tumor information extraction requires a significant human effort. Recording data in an efficient and high-quality format is essential in implementing and establishing a hospital-based-cancer registry Objective: This study aimed to describe implementing a natural language processing algorithm for oncology pathology reports. Methods: An algorithm was developed to process oncology pathology reports in Spanish to extract 20 medical descriptors. The approach is based on the successive coincidence of regular expressions. Results: The validation was performed with 140 pathological reports. The topography identification was performed manually by humans and the algorithm in all reports. The human identified morphology in 138 reports and by the algorithm in 137. The average fuzzy matching score was 68.3 for Topography and 89.5 for Morphology. Conclusions: A preliminary algorithm validation against human extraction was performed over a small set of reports with satisfactory results. This shows that a regular-expression approach can accurately and precisely extract multiple specimen attributes from free-text Spanish pathology reports. Additionally, we developed a website to facilitate collaborative validation at a larger scale which may be helpful for future research on the subject. 2023 artículo científico 0120-8322 https://www.redalyc.org/articulo.oa?id=28375672001 https://www.redalyc.org/journal/283/28375672001/ https://www.redalyc.org/journal/283/28375672001/html/ https://www.redalyc.org/journal/283/28375672001/28375672001.epub https://www.redalyc.org/journal/283/28375672001/movil https://doi.org/10.25100/cm.v54i1.5300 en http://www.redalyc.org/revista.oa?id=283 Colombia Médica application/pdf Universidad del Valle Colombia Médica (Colombia) Num.1 Vol.54
format Artículo científico
id redalyc_28375672001
institution Redalyc
language en
publishDate 2023
publisher Universidad del Valle
spellingShingle Automated extraction of information from free text of Spanish oncology pathology reports
Diana Marcela Mendoza-Urbano
Medicina
algorithm
data science
ontology learning
regular expressions
artificial intelligence
Automated extraction of information from free text of Spanish oncology pathology reports Diana Marcela Mendoza-Urbano Johan Felipe Garcia Juan Sebastian Moreno Juan Carlos Bravo-Ocaña Alvaro José Riascos Angela Zambrano Harvey Sergio I Prada Medicina algorithm data science ontology learning regular expressions artificial intelligence Background: Pathology reports are stored as unstructured, ungrammatical, fragmented, and abbreviated free text with linguistic variability among pathologists. For this reason, tumor information extraction requires a significant human effort. Recording data in an efficient and high-quality format is essential in implementing and establishing a hospital-based-cancer registry Objective: This study aimed to describe implementing a natural language processing algorithm for oncology pathology reports. Methods: An algorithm was developed to process oncology pathology reports in Spanish to extract 20 medical descriptors. The approach is based on the successive coincidence of regular expressions. Results: The validation was performed with 140 pathological reports. The topography identification was performed manually by humans and the algorithm in all reports. The human identified morphology in 138 reports and by the algorithm in 137. The average fuzzy matching score was 68.3 for Topography and 89.5 for Morphology. Conclusions: A preliminary algorithm validation against human extraction was performed over a small set of reports with satisfactory results. This shows that a regular-expression approach can accurately and precisely extract multiple specimen attributes from free-text Spanish pathology reports. Additionally, we developed a website to facilitate collaborative validation at a larger scale which may be helpful for future research on the subject. 2023 artículo científico 0120-8322 https://www.redalyc.org/articulo.oa?id=28375672001 https://www.redalyc.org/journal/283/28375672001/ https://www.redalyc.org/journal/283/28375672001/html/ https://www.redalyc.org/journal/283/28375672001/28375672001.epub https://www.redalyc.org/journal/283/28375672001/movil https://doi.org/10.25100/cm.v54i1.5300 en http://www.redalyc.org/revista.oa?id=283 Colombia Médica application/pdf Universidad del Valle Colombia Médica (Colombia) Num.1 Vol.54
title Automated extraction of information from free text of Spanish oncology pathology reports
topic Medicina
algorithm
data science
ontology learning
regular expressions
artificial intelligence
url https://www.redalyc.org/articulo.oa?id=28375672001
https://www.redalyc.org/journal/283/28375672001/
https://www.redalyc.org/journal/283/28375672001/html/
https://www.redalyc.org/journal/283/28375672001/28375672001.epub
https://www.redalyc.org/journal/283/28375672001/movil
https://doi.org/10.25100/cm.v54i1.5300