nach0: Multimodal Natural and Chemical Languages Foundation Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Livne, Micha, Miftahutdinov, Zulfat, Tutubalina, Elena, Kuznetsov, Maksim, Polykovskiy, Daniil, Brundyn, Annika, Jhunjhunwala, Aastha, Costa, Anthony, Aliper, Alex, Aspuru-Guzik, Alán, Zhavoronkov, Alex
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914850750332928
author Livne, Micha
Miftahutdinov, Zulfat
Tutubalina, Elena
Kuznetsov, Maksim
Polykovskiy, Daniil
Brundyn, Annika
Jhunjhunwala, Aastha
Costa, Anthony
Aliper, Alex
Aspuru-Guzik, Alán
Zhavoronkov, Alex
author_facet Livne, Micha
Miftahutdinov, Zulfat
Tutubalina, Elena
Kuznetsov, Maksim
Polykovskiy, Daniil
Brundyn, Annika
Jhunjhunwala, Aastha
Costa, Anthony
Aliper, Alex
Aspuru-Guzik, Alán
Zhavoronkov, Alex
contents Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a new foundation model, nach0, capable of solving various chemical and biological tasks: biomedical question answering, named entity recognition, molecular generation, molecular synthesis, attributes prediction, and others. nach0 is a multi-domain and multi-task encoder-decoder LLM pre-trained on unlabeled text from scientific literature, patents, and molecule strings to incorporate a range of chemical and linguistic knowledge. We employed instruction tuning, where specific task-related instructions are utilized to fine-tune nach0 for the final set of tasks. To train nach0 effectively, we leverage the NeMo framework, enabling efficient parallel optimization of both base and large model versions. Extensive experiments demonstrate that our model outperforms state-of-the-art baselines on single-domain and cross-domain tasks. Furthermore, it can generate high-quality outputs in molecular and textual formats, showcasing its effectiveness in multi-domain setups.
format Preprint
id arxiv_https___arxiv_org_abs_2311_12410
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle nach0: Multimodal Natural and Chemical Languages Foundation Model
Livne, Micha
Miftahutdinov, Zulfat
Tutubalina, Elena
Kuznetsov, Maksim
Polykovskiy, Daniil
Brundyn, Annika
Jhunjhunwala, Aastha
Costa, Anthony
Aliper, Alex
Aspuru-Guzik, Alán
Zhavoronkov, Alex
Computation and Language
Artificial Intelligence
Machine Learning
Quantitative Methods
Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a new foundation model, nach0, capable of solving various chemical and biological tasks: biomedical question answering, named entity recognition, molecular generation, molecular synthesis, attributes prediction, and others. nach0 is a multi-domain and multi-task encoder-decoder LLM pre-trained on unlabeled text from scientific literature, patents, and molecule strings to incorporate a range of chemical and linguistic knowledge. We employed instruction tuning, where specific task-related instructions are utilized to fine-tune nach0 for the final set of tasks. To train nach0 effectively, we leverage the NeMo framework, enabling efficient parallel optimization of both base and large model versions. Extensive experiments demonstrate that our model outperforms state-of-the-art baselines on single-domain and cross-domain tasks. Furthermore, it can generate high-quality outputs in molecular and textual formats, showcasing its effectiveness in multi-domain setups.
title nach0: Multimodal Natural and Chemical Languages Foundation Model
topic Computation and Language
Artificial Intelligence
Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2311.12410