MOSAIC: A Multilingual, Taxonomy-Agnostic, and Computationally Efficient Approach for Radiological Report Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schiavone, Alice, Fraccaro, Marco, Pehrson, Lea Marie, Ingala, Silvia, Bonnevie, Rasmus, Nielsen, Michael Bachmann, Beliveau, Vincent, Ganz, Melanie, Elliott, Desmond
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911689611411456
author Schiavone, Alice
Fraccaro, Marco
Pehrson, Lea Marie
Ingala, Silvia
Bonnevie, Rasmus
Nielsen, Michael Bachmann
Beliveau, Vincent
Ganz, Melanie
Elliott, Desmond
author_facet Schiavone, Alice
Fraccaro, Marco
Pehrson, Lea Marie
Ingala, Silvia
Bonnevie, Rasmus
Nielsen, Michael Bachmann
Beliveau, Vincent
Ganz, Melanie
Elliott, Desmond
contents Radiology reports contain rich clinical information that can be used to train imaging models without relying on costly manual annotation. However, existing approaches face critical limitations: rule-based methods struggle with linguistic variability, supervised models require large annotated datasets, and recent LLM-based systems depend on closed-source or resource-intensive models that are unsuitable for clinical use. Moreover, current solutions are largely restricted to English and single-modality, single-taxonomy datasets. We introduce MOSAIC, a multilingual, taxonomy-agnostic, and computationally efficient approach for radiological report classification. Built on a compact open-access language model (MedGemma-4B), MOSAIC supports both zero-/few-shot prompting and lightweight fine-tuning, enabling deployment on consumer-grade GPUs. We evaluate MOSAIC across seven datasets in English, Spanish, French, and Danish, spanning multiple imaging modalities and label taxonomies. The model achieves a mean macro F1 score of 88 across five chest X-ray datasets, approaching or exceeding expert-level performance, while requiring only 24 GB of GPU memory. With data augmentation, as few as 80 annotated samples are sufficient to reach a weighted F1 score of 82 on Danish reports, compared to 86 with the full 1600-sample training set. MOSAIC offers a practical alternative to large or proprietary LLMs in clinical settings. Code and models are open-source. We invite the community to evaluate and extend MOSAIC on new languages, taxonomies, and modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOSAIC: A Multilingual, Taxonomy-Agnostic, and Computationally Efficient Approach for Radiological Report Classification
Schiavone, Alice
Fraccaro, Marco
Pehrson, Lea Marie
Ingala, Silvia
Bonnevie, Rasmus
Nielsen, Michael Bachmann
Beliveau, Vincent
Ganz, Melanie
Elliott, Desmond
Computation and Language
Artificial Intelligence
Radiology reports contain rich clinical information that can be used to train imaging models without relying on costly manual annotation. However, existing approaches face critical limitations: rule-based methods struggle with linguistic variability, supervised models require large annotated datasets, and recent LLM-based systems depend on closed-source or resource-intensive models that are unsuitable for clinical use. Moreover, current solutions are largely restricted to English and single-modality, single-taxonomy datasets. We introduce MOSAIC, a multilingual, taxonomy-agnostic, and computationally efficient approach for radiological report classification. Built on a compact open-access language model (MedGemma-4B), MOSAIC supports both zero-/few-shot prompting and lightweight fine-tuning, enabling deployment on consumer-grade GPUs. We evaluate MOSAIC across seven datasets in English, Spanish, French, and Danish, spanning multiple imaging modalities and label taxonomies. The model achieves a mean macro F1 score of 88 across five chest X-ray datasets, approaching or exceeding expert-level performance, while requiring only 24 GB of GPU memory. With data augmentation, as few as 80 annotated samples are sufficient to reach a weighted F1 score of 82 on Danish reports, compared to 86 with the full 1600-sample training set. MOSAIC offers a practical alternative to large or proprietary LLMs in clinical settings. Code and models are open-source. We invite the community to evaluate and extend MOSAIC on new languages, taxonomies, and modalities.
title MOSAIC: A Multilingual, Taxonomy-Agnostic, and Computationally Efficient Approach for Radiological Report Classification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.04471