Enhancing Solar Power Forecasting with Data Imputation and Machine Learning: A Comparative Study

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Usha.N, P.S. Manoharan, D.Sneha Charis, P.M. Devie
Format: Recurso digital
Language:English
Published: Zenodo 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901255953055744
author Usha.N
P.S. Manoharan
D.Sneha Charis
P.M. Devie
author_facet Usha.N
P.S. Manoharan
D.Sneha Charis
P.M. Devie
contents <p><span><span>Abstract</span></span><span>—The surge in energy demand necessitates an integration of renewable sources into power grids, particularly solar energy. It focuses on handling solar yield datasets, which are subject to problems of intermittent data due to sensor failure and variability in weather conditions, using advanced imputation techniques—K-Nearest Neighbors, Linear Interpolation, and Multivariate Imputation by Chained Equations. Feature selection and dimensionality reduction methods such as the Pearson Correlation Coefficient, Principal Component Analysis, and Mutual Information add predictive capability to the optimal datasets using machine learning models like XGBoost, LSTM, CatBoost, Random Forest, and Decision Trees. CatBoost was shown to be the best at training, achieving a very high accuracy of 85.6%. This work forms a sound methodology for tackling issues in data quality and dimension reduction as far as renewable energy forecasting is concerned, taking it a step further in establishing standards in prediction for solar power.</span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15056491
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle Enhancing Solar Power Forecasting with Data Imputation and Machine Learning: A Comparative Study
Usha.N
P.S. Manoharan
D.Sneha Charis
P.M. Devie
Feature Selection
Imputation
Machine Learning
Performance metrics
Solar Power Prediction
<p><span><span>Abstract</span></span><span>—The surge in energy demand necessitates an integration of renewable sources into power grids, particularly solar energy. It focuses on handling solar yield datasets, which are subject to problems of intermittent data due to sensor failure and variability in weather conditions, using advanced imputation techniques—K-Nearest Neighbors, Linear Interpolation, and Multivariate Imputation by Chained Equations. Feature selection and dimensionality reduction methods such as the Pearson Correlation Coefficient, Principal Component Analysis, and Mutual Information add predictive capability to the optimal datasets using machine learning models like XGBoost, LSTM, CatBoost, Random Forest, and Decision Trees. CatBoost was shown to be the best at training, achieving a very high accuracy of 85.6%. This work forms a sound methodology for tackling issues in data quality and dimension reduction as far as renewable energy forecasting is concerned, taking it a step further in establishing standards in prediction for solar power.</span></p>
title Enhancing Solar Power Forecasting with Data Imputation and Machine Learning: A Comparative Study
topic Feature Selection
Imputation
Machine Learning
Performance metrics
Solar Power Prediction
url https://doi.org/10.5281/zenodo.15056491