Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: García-Ferrero, Iker, Agerri, Rodrigo, Salazar, Aitziber Atutxa, Cabrio, Elena, de la Iglesia, Iker, Lavelli, Alberto, Magnini, Bernardo, Molinet, Benjamin, Ramirez-Romero, Johana, Rigau, German, Villa-Gonzalez, Jose Maria, Villata, Serena, Zaninello, Andrea
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913310169890816
author García-Ferrero, Iker
Agerri, Rodrigo
Salazar, Aitziber Atutxa
Cabrio, Elena
de la Iglesia, Iker
Lavelli, Alberto
Magnini, Bernardo
Molinet, Benjamin
Ramirez-Romero, Johana
Rigau, German
Villa-Gonzalez, Jose Maria
Villata, Serena
Zaninello, Andrea
author_facet García-Ferrero, Iker
Agerri, Rodrigo
Salazar, Aitziber Atutxa
Cabrio, Elena
de la Iglesia, Iker
Lavelli, Alberto
Magnini, Bernardo
Molinet, Benjamin
Ramirez-Romero, Johana
Rigau, German
Villa-Gonzalez, Jose Maria
Villata, Serena
Zaninello, Andrea
contents Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large language models (LLMs) have recently been adapted to the medical domain, so that they can be used as a tool for mediating in human-AI interaction. While these LLMs display competitive performance on automated medical texts benchmarks, they have been pre-trained and evaluated with a focus on a single language (English mostly). This is particularly true of text-to-text models, which typically require large amounts of domain-specific pre-training data, often not easily accessible for many languages. In this paper, we address these shortcomings by compiling, to the best of our knowledge, the largest multilingual corpus for the medical domain in four languages, namely English, French, Italian and Spanish. This new corpus has been used to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. Additionally, we present two new evaluation benchmarks for all four languages with the aim of facilitating multilingual research in this domain. A comprehensive evaluation shows that Medical mT5 outperforms both encoders and similarly sized text-to-text models for the Spanish, French, and Italian benchmarks, while being competitive with current state-of-the-art LLMs in English.
format Preprint
id arxiv_https___arxiv_org_abs_2404_07613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
García-Ferrero, Iker
Agerri, Rodrigo
Salazar, Aitziber Atutxa
Cabrio, Elena
de la Iglesia, Iker
Lavelli, Alberto
Magnini, Bernardo
Molinet, Benjamin
Ramirez-Romero, Johana
Rigau, German
Villa-Gonzalez, Jose Maria
Villata, Serena
Zaninello, Andrea
Computation and Language
Artificial Intelligence
Machine Learning
Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large language models (LLMs) have recently been adapted to the medical domain, so that they can be used as a tool for mediating in human-AI interaction. While these LLMs display competitive performance on automated medical texts benchmarks, they have been pre-trained and evaluated with a focus on a single language (English mostly). This is particularly true of text-to-text models, which typically require large amounts of domain-specific pre-training data, often not easily accessible for many languages. In this paper, we address these shortcomings by compiling, to the best of our knowledge, the largest multilingual corpus for the medical domain in four languages, namely English, French, Italian and Spanish. This new corpus has been used to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. Additionally, we present two new evaluation benchmarks for all four languages with the aim of facilitating multilingual research in this domain. A comprehensive evaluation shows that Medical mT5 outperforms both encoders and similarly sized text-to-text models for the Spanish, French, and Italian benchmarks, while being competitive with current state-of-the-art LLMs in English.
title Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.07613