Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: de Almeida, Vinicius Anjos, da Silva, Sandro Saorin, Chire, Josimar, Vicenzi, Leonardo, Borges, Nícolas Henrique, Kociolek, Helena, Rocha, Sarah Miriã de Castro, Gomes, Frederico Nassif, Ferreira, Júlia Cristina, Marques, Oge, Oliveira, Lucas Emanuel Silva e
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911548345155584
author de Almeida, Vinicius Anjos
da Silva, Sandro Saorin
Chire, Josimar
Vicenzi, Leonardo
Borges, Nícolas Henrique
Kociolek, Helena
Rocha, Sarah Miriã de Castro
Gomes, Frederico Nassif
Ferreira, Júlia Cristina
Marques, Oge
Oliveira, Lucas Emanuel Silva e
author_facet de Almeida, Vinicius Anjos
da Silva, Sandro Saorin
Chire, Josimar
Vicenzi, Leonardo
Borges, Nícolas Henrique
Kociolek, Helena
Rocha, Sarah Miriã de Castro
Gomes, Frederico Nassif
Ferreira, Júlia Cristina
Marques, Oge
Oliveira, Lucas Emanuel Silva e
contents Clinical notes contain valuable unstructured information. Named entity recognition (NER) enables the automatic extraction of medical concepts; however, benchmarks for Portuguese remain scarce. In this study, we aimed to evaluate BERT-based models and large language models (LLMs) for clinical NER in Portuguese and to test strategies for addressing multilabel imbalance. We compared BioBERTpt, BERTimbau, ModernBERT, and mmBERT with LLMs such as GPT-5 and Gemini-2.5, using the public SemClinBr corpus and a private breast cancer dataset. Models were trained under identical conditions and evaluated using precision, recall, and F1-score. Iterative stratification, weighted loss, and oversampling were explored to mitigate class imbalance. The mmBERT-base model achieved the best performance (micro F1 = 0.76), outperforming all other models. Iterative stratification improved class balance and overall performance. Multilingual BERT models, particularly mmBERT, perform strongly for Portuguese clinical NER and can run locally with limited computational resources. Balanced data-splitting strategies further enhance performance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26510
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
de Almeida, Vinicius Anjos
da Silva, Sandro Saorin
Chire, Josimar
Vicenzi, Leonardo
Borges, Nícolas Henrique
Kociolek, Helena
Rocha, Sarah Miriã de Castro
Gomes, Frederico Nassif
Ferreira, Júlia Cristina
Marques, Oge
Oliveira, Lucas Emanuel Silva e
Computation and Language
Clinical notes contain valuable unstructured information. Named entity recognition (NER) enables the automatic extraction of medical concepts; however, benchmarks for Portuguese remain scarce. In this study, we aimed to evaluate BERT-based models and large language models (LLMs) for clinical NER in Portuguese and to test strategies for addressing multilabel imbalance. We compared BioBERTpt, BERTimbau, ModernBERT, and mmBERT with LLMs such as GPT-5 and Gemini-2.5, using the public SemClinBr corpus and a private breast cancer dataset. Models were trained under identical conditions and evaluated using precision, recall, and F1-score. Iterative stratification, weighted loss, and oversampling were explored to mitigate class imbalance. The mmBERT-base model achieved the best performance (micro F1 = 0.76), outperforming all other models. Iterative stratification improved class balance and overall performance. Multilingual BERT models, particularly mmBERT, perform strongly for Portuguese clinical NER and can run locally with limited computational resources. Balanced data-splitting strategies further enhance performance.
title Clinical named entity recognition in the Portuguese language: a benchmark of modern BERT models and LLMs
topic Computation and Language
url https://arxiv.org/abs/2603.26510