IMPROVING THE RELIABILITY OF MACHINE LEARNING MODELS BY FILLING IN MISSING NAN VALUES IN MEDICAL DATASETS USING A GENETIC ALGORITHM
Fuente:
Zenodo
Saved in:
| Main Author: | |
|---|---|
| Format: | Recurso digital |
| Language: | English |
| Published: |
Zenodo
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901897845145600 |
|---|---|
| author | Shokhrukh Sariyev |
| author_facet | Shokhrukh Sariyev |
| contents | This article proposes a genetic algorithm-based approach to optimize the filling of missing NaN values in a dataset. The focus is on selecting NaN values in the dataset directly corresponding to the results of the classification task. In the proposed method, each individual is represented as a chromosome in the form of a vector of all missing values. The search space is bounded by the given intervals for numerical attributes, and by the set of appropriate categories for categorical attributes. The accuracy indicator of the Random Forest ensemble model was used as the fitness function in the genetic algorithm. |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17994364 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | IMPROVING THE RELIABILITY OF MACHINE LEARNING MODELS BY FILLING IN MISSING NAN VALUES IN MEDICAL DATASETS USING A GENETIC ALGORITHM Shokhrukh Sariyev Genetic algorithm ML AI Random Forest KNN This article proposes a genetic algorithm-based approach to optimize the filling of missing NaN values in a dataset. The focus is on selecting NaN values in the dataset directly corresponding to the results of the classification task. In the proposed method, each individual is represented as a chromosome in the form of a vector of all missing values. The search space is bounded by the given intervals for numerical attributes, and by the set of appropriate categories for categorical attributes. The accuracy indicator of the Random Forest ensemble model was used as the fitness function in the genetic algorithm. |
| title | IMPROVING THE RELIABILITY OF MACHINE LEARNING MODELS BY FILLING IN MISSING NAN VALUES IN MEDICAL DATASETS USING A GENETIC ALGORITHM |
| topic | Genetic algorithm ML AI Random Forest KNN |
| url | https://doi.org/10.5281/zenodo.17994364 |