Data augmentation for machine learning of chemical process flowsheets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Balhorn, Lukas Schulze, Hirtreiter, Edwin, Luderer, Lynn, Schweidtmann, Artur M.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916090538360832
author Balhorn, Lukas Schulze
Hirtreiter, Edwin
Luderer, Lynn
Schweidtmann, Artur M.
author_facet Balhorn, Lukas Schulze
Hirtreiter, Edwin
Luderer, Lynn
Schweidtmann, Artur M.
contents Artificial intelligence has great potential for accelerating the design and engineering of chemical processes. Recently, we have shown that transformer-based language models can learn to auto-complete chemical process flowsheets using the SFILES 2.0 string notation. Also, we showed that language translation models can be used to translate Process Flow Diagrams (PFDs) into Process and Instrumentation Diagrams (P&IDs). However, artificial intelligence methods require big data and flowsheet data is currently limited. To mitigate this challenge of limited data, we propose a new data augmentation methodology for flowsheet data that is represented in the SFILES 2.0 notation. We show that the proposed data augmentation improves the performance of artificial intelligence-based process design models. In our case study flowsheet data augmentation improved the prediction uncertainty of the flowsheet autocompletion model by 14.7%. In the future, our flowsheet data augmentation can be used for other machine learning algorithms on chemical process flowsheets that are based on SFILES notation.
format Preprint
id arxiv_https___arxiv_org_abs_2302_03379
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Data augmentation for machine learning of chemical process flowsheets
Balhorn, Lukas Schulze
Hirtreiter, Edwin
Luderer, Lynn
Schweidtmann, Artur M.
Machine Learning
Optimization and Control
Artificial intelligence has great potential for accelerating the design and engineering of chemical processes. Recently, we have shown that transformer-based language models can learn to auto-complete chemical process flowsheets using the SFILES 2.0 string notation. Also, we showed that language translation models can be used to translate Process Flow Diagrams (PFDs) into Process and Instrumentation Diagrams (P&IDs). However, artificial intelligence methods require big data and flowsheet data is currently limited. To mitigate this challenge of limited data, we propose a new data augmentation methodology for flowsheet data that is represented in the SFILES 2.0 notation. We show that the proposed data augmentation improves the performance of artificial intelligence-based process design models. In our case study flowsheet data augmentation improved the prediction uncertainty of the flowsheet autocompletion model by 14.7%. In the future, our flowsheet data augmentation can be used for other machine learning algorithms on chemical process flowsheets that are based on SFILES notation.
title Data augmentation for machine learning of chemical process flowsheets
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2302.03379