JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sousa, Joao, Darabi, Roya, Sousa, Armando, Brueckner, Frank, Reis, Luís Paulo, Reis, Ana
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916463331246080
author Sousa, Joao
Darabi, Roya
Sousa, Armando
Brueckner, Frank
Reis, Luís Paulo
Reis, Ana
author_facet Sousa, Joao
Darabi, Roya
Sousa, Armando
Brueckner, Frank
Reis, Luís Paulo
Reis, Ana
contents This work introduces JEMA (Joint Embedding with Multimodal Alignment), a novel co-learning framework tailored for laser metal deposition (LMD), a pivotal process in metal additive manufacturing. As Industry 5.0 gains traction in industrial applications, efficient process monitoring becomes increasingly crucial. However, limited data and the opaque nature of AI present challenges for its application in an industrial setting. JEMA addresses this challenges by leveraging multimodal data, including multi-view images and metadata such as process parameters, to learn transferable semantic representations. By applying a supervised contrastive loss function, JEMA enables robust learning and subsequent process monitoring using only the primary modality, simplifying hardware requirements and computational overhead. We investigate the effectiveness of JEMA in LMD process monitoring, focusing specifically on its generalization to downstream tasks such as melt pool geometry prediction, achieved without extensive fine-tuning. Our empirical evaluation demonstrates the high scalability and performance of JEMA, particularly when combined with Vision Transformer models. We report an 8% increase in performance in multimodal settings and a 1% improvement in unimodal settings compared to supervised contrastive learning. Additionally, the learned embedding representation enables the prediction of metadata, enhancing interpretability and making possible the assessment of the added metadata's contributions. Our framework lays the foundation for integrating multisensor data with metadata, enabling diverse downstream tasks within the LMD domain and beyond.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23988
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment
Sousa, Joao
Darabi, Roya
Sousa, Armando
Brueckner, Frank
Reis, Luís Paulo
Reis, Ana
Computer Vision and Pattern Recognition
This work introduces JEMA (Joint Embedding with Multimodal Alignment), a novel co-learning framework tailored for laser metal deposition (LMD), a pivotal process in metal additive manufacturing. As Industry 5.0 gains traction in industrial applications, efficient process monitoring becomes increasingly crucial. However, limited data and the opaque nature of AI present challenges for its application in an industrial setting. JEMA addresses this challenges by leveraging multimodal data, including multi-view images and metadata such as process parameters, to learn transferable semantic representations. By applying a supervised contrastive loss function, JEMA enables robust learning and subsequent process monitoring using only the primary modality, simplifying hardware requirements and computational overhead. We investigate the effectiveness of JEMA in LMD process monitoring, focusing specifically on its generalization to downstream tasks such as melt pool geometry prediction, achieved without extensive fine-tuning. Our empirical evaluation demonstrates the high scalability and performance of JEMA, particularly when combined with Vision Transformer models. We report an 8% increase in performance in multimodal settings and a 1% improvement in unimodal settings compared to supervised contrastive learning. Additionally, the learned embedding representation enables the prediction of metadata, enhancing interpretability and making possible the assessment of the added metadata's contributions. Our framework lays the foundation for integrating multisensor data with metadata, enabling diverse downstream tasks within the LMD domain and beyond.
title JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23988