Guardado en:
Detalles Bibliográficos
Autores principales: Lange, Christoph, Thiele, Isabel, Santolin, Lara, Riedel, Sebastian L., Borisyak, Maxim, Neubauer, Peter, Bournazou, M. Nicolas Cruz
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2402.00851
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914338992816128
author Lange, Christoph
Thiele, Isabel
Santolin, Lara
Riedel, Sebastian L.
Borisyak, Maxim
Neubauer, Peter
Bournazou, M. Nicolas Cruz
author_facet Lange, Christoph
Thiele, Isabel
Santolin, Lara
Riedel, Sebastian L.
Borisyak, Maxim
Neubauer, Peter
Bournazou, M. Nicolas Cruz
contents In biotechnology Raman Spectroscopy is rapidly gaining popularity as a process analytical technology (PAT) that measures cell densities, substrate- and product concentrations. As it records vibrational modes of molecules it provides that information non-invasively in a single spectrum. Typically, partial least squares (PLS) is the model of choice to infer information about variables of interest from the spectra. However, biological processes are known for their complexity where convolutional neural networks (CNN) present a powerful alternative. They can handle non-Gaussian noise and account for beam misalignment, pixel malfunctions or the presence of additional substances. However, they require a lot of data during model training, and they pick up non-linear dependencies in the process variables. In this work, we exploit the additive nature of spectra in order to generate additional data points from a given dataset that have statistically independent labels so that a network trained on such data exhibits low correlations between the model predictions. We show that training a CNN on these generated data points improves the performance on datasets where the annotations do not bear the same correlation as the dataset that was used for model training. This data augmentation technique enables us to reuse spectra as training data for new contexts that exhibit different correlations. The additional data allows for building a better and more robust model. This is of interest in scenarios where large amounts of historical data are available but are currently not used for model training. We demonstrate the capabilities of the proposed method using synthetic spectra of Ralstonia eutropha batch cultivations to monitor substrate, biomass and polyhydroxyalkanoate (PHA) biopolymer concentrations during of the experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2402_00851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Augmentation Scheme for Raman Spectra with Highly Correlated Annotations
Lange, Christoph
Thiele, Isabel
Santolin, Lara
Riedel, Sebastian L.
Borisyak, Maxim
Neubauer, Peter
Bournazou, M. Nicolas Cruz
Machine Learning
Quantitative Methods
In biotechnology Raman Spectroscopy is rapidly gaining popularity as a process analytical technology (PAT) that measures cell densities, substrate- and product concentrations. As it records vibrational modes of molecules it provides that information non-invasively in a single spectrum. Typically, partial least squares (PLS) is the model of choice to infer information about variables of interest from the spectra. However, biological processes are known for their complexity where convolutional neural networks (CNN) present a powerful alternative. They can handle non-Gaussian noise and account for beam misalignment, pixel malfunctions or the presence of additional substances. However, they require a lot of data during model training, and they pick up non-linear dependencies in the process variables. In this work, we exploit the additive nature of spectra in order to generate additional data points from a given dataset that have statistically independent labels so that a network trained on such data exhibits low correlations between the model predictions. We show that training a CNN on these generated data points improves the performance on datasets where the annotations do not bear the same correlation as the dataset that was used for model training. This data augmentation technique enables us to reuse spectra as training data for new contexts that exhibit different correlations. The additional data allows for building a better and more robust model. This is of interest in scenarios where large amounts of historical data are available but are currently not used for model training. We demonstrate the capabilities of the proposed method using synthetic spectra of Ralstonia eutropha batch cultivations to monitor substrate, biomass and polyhydroxyalkanoate (PHA) biopolymer concentrations during of the experiments.
title Data Augmentation Scheme for Raman Spectra with Highly Correlated Annotations
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2402.00851