Facilitating phenotyping from clinical texts: the medkit library

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neuraz, Antoine, Vaillant, Ghislain, Arias, Camila, Birot, Olivier, Huynh, Kim-Tam, Fabacher, Thibaut, Rogier, Alice, Garcelon, Nicolas, Lerner, Ivan, Rance, Bastien, Coulet, Adrien
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912012548702208
author Neuraz, Antoine
Vaillant, Ghislain
Arias, Camila
Birot, Olivier
Huynh, Kim-Tam
Fabacher, Thibaut
Rogier, Alice
Garcelon, Nicolas
Lerner, Ivan
Rance, Bastien
Coulet, Adrien
author_facet Neuraz, Antoine
Vaillant, Ghislain
Arias, Camila
Birot, Olivier
Huynh, Kim-Tam
Fabacher, Thibaut
Rogier, Alice
Garcelon, Nicolas
Lerner, Ivan
Rance, Bastien
Coulet, Adrien
contents Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical information of EHRs are lying in texts, phenotyping from text takes an important role in studies that rely on the secondary use of EHRs. However, the heterogeneity and highly specialized aspect of both the content and form of clinical texts makes this task particularly tedious, and is the source of time and cost constraints in observational studies. To facilitate the development, evaluation and reproductibility of phenotyping pipelines, we developed an open-source Python library named medkit. It enables composing data processing pipelines made of easy-to-reuse software bricks, named medkit operations. In addition to the core of the library, we share the operations and pipelines we already developed and invite the phenotyping community for their reuse and enrichment. medkit is available at https://github.com/medkit-lib/medkit
format Preprint
id arxiv_https___arxiv_org_abs_2409_00164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Facilitating phenotyping from clinical texts: the medkit library
Neuraz, Antoine
Vaillant, Ghislain
Arias, Camila
Birot, Olivier
Huynh, Kim-Tam
Fabacher, Thibaut
Rogier, Alice
Garcelon, Nicolas
Lerner, Ivan
Rance, Bastien
Coulet, Adrien
Computation and Language
Information Retrieval
Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical information of EHRs are lying in texts, phenotyping from text takes an important role in studies that rely on the secondary use of EHRs. However, the heterogeneity and highly specialized aspect of both the content and form of clinical texts makes this task particularly tedious, and is the source of time and cost constraints in observational studies. To facilitate the development, evaluation and reproductibility of phenotyping pipelines, we developed an open-source Python library named medkit. It enables composing data processing pipelines made of easy-to-reuse software bricks, named medkit operations. In addition to the core of the library, we share the operations and pipelines we already developed and invite the phenotyping community for their reuse and enrichment. medkit is available at https://github.com/medkit-lib/medkit
title Facilitating phenotyping from clinical texts: the medkit library
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2409.00164