Industrial Energy Disaggregation with Digital Twin-generated Dataset and Efficient Data Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Internò, Christian, Castellani, Andrea, Schmitt, Sebastian, Stella, Fabio, Hammer, Barbara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915494612697088
author Internò, Christian
Castellani, Andrea
Schmitt, Sebastian
Stella, Fabio
Hammer, Barbara
author_facet Internò, Christian
Castellani, Andrea
Schmitt, Sebastian
Stella, Fabio
Hammer, Barbara
contents Industrial Non-Intrusive Load Monitoring (NILM) is limited by the scarcity of high-quality datasets and the complex variability of industrial energy consumption patterns. To address data scarcity and privacy issues, we introduce the Synthetic Industrial Dataset for Energy Disaggregation (SIDED), an open-source dataset generated using Digital Twin simulations. SIDED includes three types of industrial facilities across three different geographic locations, capturing diverse appliance behaviors, weather conditions, and load profiles. We also propose the Appliance-Modulated Data Augmentation (AMDA) method, a computationally efficient technique that enhances NILM model generalization by intelligently scaling appliance power contributions based on their relative impact. We show in experiments that NILM models trained with AMDA-augmented data significantly improve the disaggregation of energy consumption of complex industrial appliances like combined heat and power systems. Specifically, in our out-of-sample scenarios, models trained with AMDA achieved a Normalized Disaggregation Error of 0.093, outperforming models trained without data augmentation (0.451) and those trained with random data augmentation (0.290). Data distribution analyses confirm that AMDA effectively aligns training and test data distributions, enhancing model generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Industrial Energy Disaggregation with Digital Twin-generated Dataset and Efficient Data Augmentation
Internò, Christian
Castellani, Andrea
Schmitt, Sebastian
Stella, Fabio
Hammer, Barbara
Machine Learning
Artificial Intelligence
Systems and Control
Industrial Non-Intrusive Load Monitoring (NILM) is limited by the scarcity of high-quality datasets and the complex variability of industrial energy consumption patterns. To address data scarcity and privacy issues, we introduce the Synthetic Industrial Dataset for Energy Disaggregation (SIDED), an open-source dataset generated using Digital Twin simulations. SIDED includes three types of industrial facilities across three different geographic locations, capturing diverse appliance behaviors, weather conditions, and load profiles. We also propose the Appliance-Modulated Data Augmentation (AMDA) method, a computationally efficient technique that enhances NILM model generalization by intelligently scaling appliance power contributions based on their relative impact. We show in experiments that NILM models trained with AMDA-augmented data significantly improve the disaggregation of energy consumption of complex industrial appliances like combined heat and power systems. Specifically, in our out-of-sample scenarios, models trained with AMDA achieved a Normalized Disaggregation Error of 0.093, outperforming models trained without data augmentation (0.451) and those trained with random data augmentation (0.290). Data distribution analyses confirm that AMDA effectively aligns training and test data distributions, enhancing model generalization.
title Industrial Energy Disaggregation with Digital Twin-generated Dataset and Efficient Data Augmentation
topic Machine Learning
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2506.20525