Saved in:
Bibliographic Details
Main Authors: Gondara, Lovedeep, Arbour, Gregory, Ng, Raymond, Simkin, Jonathan, Devji, Shebnum
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.09991
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912538973700096
author Gondara, Lovedeep
Arbour, Gregory
Ng, Raymond
Simkin, Jonathan
Devji, Shebnum
author_facet Gondara, Lovedeep
Arbour, Gregory
Ng, Raymond
Simkin, Jonathan
Devji, Shebnum
contents Automating data extraction from clinical documents offers significant potential to improve efficiency in healthcare settings, yet deploying Natural Language Processing (NLP) solutions presents practical challenges. Drawing upon our experience implementing various NLP models for information extraction and classification tasks at the British Columbia Cancer Registry (BCCR), this paper shares key lessons learned throughout the project lifecycle. We emphasize the critical importance of defining problems based on clear business objectives rather than solely technical accuracy, adopting an iterative approach to development, and fostering deep interdisciplinary collaboration and co-design involving domain experts, end-users, and ML specialists from inception. Further insights highlight the need for pragmatic model selection (including hybrid approaches and simpler methods where appropriate), rigorous attention to data quality (representativeness, drift, annotation), robust error mitigation strategies involving human-in-the-loop validation and ongoing audits, and building organizational AI literacy. These practical considerations, generalizable beyond cancer registries, provide guidance for healthcare organizations seeking to successfully implement AI/NLP solutions to enhance data management processes and ultimately improve patient care and public health outcomes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
Gondara, Lovedeep
Arbour, Gregory
Ng, Raymond
Simkin, Jonathan
Devji, Shebnum
Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
Automating data extraction from clinical documents offers significant potential to improve efficiency in healthcare settings, yet deploying Natural Language Processing (NLP) solutions presents practical challenges. Drawing upon our experience implementing various NLP models for information extraction and classification tasks at the British Columbia Cancer Registry (BCCR), this paper shares key lessons learned throughout the project lifecycle. We emphasize the critical importance of defining problems based on clear business objectives rather than solely technical accuracy, adopting an iterative approach to development, and fostering deep interdisciplinary collaboration and co-design involving domain experts, end-users, and ML specialists from inception. Further insights highlight the need for pragmatic model selection (including hybrid approaches and simpler methods where appropriate), rigorous attention to data quality (representativeness, drift, annotation), robust error mitigation strategies involving human-in-the-loop validation and ongoing audits, and building organizational AI literacy. These practical considerations, generalizable beyond cancer registries, provide guidance for healthcare organizations seeking to successfully implement AI/NLP solutions to enhance data management processes and ultimately improve patient care and public health outcomes.
title Bridging AI Innovation and Healthcare Needs: Lessons Learned from Incorporating Modern NLP at The BC Cancer Registry
topic Computation and Language
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2508.09991