ACR: A Benchmark for Automatic Cohort Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thai, Dung Ngoc, Ardulov, Victor, Mena, Jose Ulises, Tiwari, Simran, Erofeev, Gleb, Eskander, Ramy, Tarabishy, Karim, Parikh, Ravi B, Salloum, Wael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911939785916416
author Thai, Dung Ngoc
Ardulov, Victor
Mena, Jose Ulises
Tiwari, Simran
Erofeev, Gleb
Eskander, Ramy
Tarabishy, Karim
Parikh, Ravi B
Salloum, Wael
author_facet Thai, Dung Ngoc
Ardulov, Victor
Mena, Jose Ulises
Tiwari, Simran
Erofeev, Gleb
Eskander, Ramy
Tarabishy, Karim
Parikh, Ravi B
Salloum, Wael
contents Identifying patient cohorts is fundamental to numerous healthcare tasks, including clinical trial recruitment and retrospective studies. Current cohort retrieval methods in healthcare organizations rely on automated queries of structured data combined with manual curation, which are time-consuming, labor-intensive, and often yield low-quality results. Recent advancements in large language models (LLMs) and information retrieval (IR) offer promising avenues to revolutionize these systems. Major challenges include managing extensive eligibility criteria and handling the longitudinal nature of unstructured Electronic Medical Records (EMRs) while ensuring that the solution remains cost-effective for real-world application. This paper introduces a new task, Automatic Cohort Retrieval (ACR), and evaluates the performance of LLMs and commercial, domain-specific neuro-symbolic approaches. We provide a benchmark task, a query dataset, an EMR dataset, and an evaluation framework. Our findings underscore the necessity for efficient, high-quality ACR systems capable of longitudinal reasoning across extensive patient databases.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14780
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ACR: A Benchmark for Automatic Cohort Retrieval
Thai, Dung Ngoc
Ardulov, Victor
Mena, Jose Ulises
Tiwari, Simran
Erofeev, Gleb
Eskander, Ramy
Tarabishy, Karim
Parikh, Ravi B
Salloum, Wael
Artificial Intelligence
Identifying patient cohorts is fundamental to numerous healthcare tasks, including clinical trial recruitment and retrospective studies. Current cohort retrieval methods in healthcare organizations rely on automated queries of structured data combined with manual curation, which are time-consuming, labor-intensive, and often yield low-quality results. Recent advancements in large language models (LLMs) and information retrieval (IR) offer promising avenues to revolutionize these systems. Major challenges include managing extensive eligibility criteria and handling the longitudinal nature of unstructured Electronic Medical Records (EMRs) while ensuring that the solution remains cost-effective for real-world application. This paper introduces a new task, Automatic Cohort Retrieval (ACR), and evaluates the performance of LLMs and commercial, domain-specific neuro-symbolic approaches. We provide a benchmark task, a query dataset, an EMR dataset, and an evaluation framework. Our findings underscore the necessity for efficient, high-quality ACR systems capable of longitudinal reasoning across extensive patient databases.
title ACR: A Benchmark for Automatic Cohort Retrieval
topic Artificial Intelligence
url https://arxiv.org/abs/2406.14780