NOWJ @BioCreative IX ToxHabits: An Ensemble Deep Learning Approach for Detecting Substance Use and Contextual Information in Clinical Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tran, Huu-Huy-Hoang, Duong, Gia-Bao, Tran, Quoc-Viet-Anh, Vuong, Thi-Hai-Yen, Le, Hoang-Quynh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917263420948480
author Tran, Huu-Huy-Hoang
Duong, Gia-Bao
Tran, Quoc-Viet-Anh
Vuong, Thi-Hai-Yen
Le, Hoang-Quynh
author_facet Tran, Huu-Huy-Hoang
Duong, Gia-Bao
Tran, Quoc-Viet-Anh
Vuong, Thi-Hai-Yen
Le, Hoang-Quynh
contents Extracting drug use information from unstructured Electronic Health Records remains a major challenge in clinical Natural Language Processing. While Large Language Models demonstrate advancements, their use in clinical NLP is limited by concerns over trust, control, and efficiency. To address this, we present NOWJ submission to the ToxHabits Shared Task at BioCreative IX. This task targets the detection of toxic substance use and contextual attributes in Spanish clinical texts, a domain-specific, low-resource setting. We propose a multi-output ensemble system tackling both Subtask 1 - ToxNER and Subtask 2 - ToxUse. Our system integrates BETO with a CRF layer for sequence labeling, employs diverse training strategies, and uses sentence filtering to boost precision. Our top run achieved 0.94 F1 and 0.97 precision for Trigger Detection, and 0.91 F1 for Argument Detection.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09469
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle NOWJ @BioCreative IX ToxHabits: An Ensemble Deep Learning Approach for Detecting Substance Use and Contextual Information in Clinical Texts
Tran, Huu-Huy-Hoang
Duong, Gia-Bao
Tran, Quoc-Viet-Anh
Vuong, Thi-Hai-Yen
Le, Hoang-Quynh
Computation and Language
Artificial Intelligence
Extracting drug use information from unstructured Electronic Health Records remains a major challenge in clinical Natural Language Processing. While Large Language Models demonstrate advancements, their use in clinical NLP is limited by concerns over trust, control, and efficiency. To address this, we present NOWJ submission to the ToxHabits Shared Task at BioCreative IX. This task targets the detection of toxic substance use and contextual attributes in Spanish clinical texts, a domain-specific, low-resource setting. We propose a multi-output ensemble system tackling both Subtask 1 - ToxNER and Subtask 2 - ToxUse. Our system integrates BETO with a CRF layer for sequence labeling, employs diverse training strategies, and uses sentence filtering to boost precision. Our top run achieved 0.94 F1 and 0.97 precision for Trigger Detection, and 0.91 F1 for Argument Detection.
title NOWJ @BioCreative IX ToxHabits: An Ensemble Deep Learning Approach for Detecting Substance Use and Contextual Information in Clinical Texts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.09469