Comparative analysis of K-Means, SVM, Decision Tree and Naive Bayes in Predicting Diabetes Presence

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Rohan Almeida, Rajeev Dessai, Ayush Noorani, Amogh Pai Raiturkar, Samuel Godinho
Natura: Recurso digital
Pubblicazione: Zenodo 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866901250106195968
author Rohan Almeida
Rajeev Dessai
Ayush Noorani
Amogh Pai Raiturkar
Samuel Godinho
author_facet Rohan Almeida
Rajeev Dessai
Ayush Noorani
Amogh Pai Raiturkar
Samuel Godinho
contents In the context of rapidly growing amounts of data present today it is imperative to quickly dig out information from this data as information determines a large portion of the decision making process. Data mining allows us to do exactly that. Data mining is the process of turning raw data into useful information as it is an analytical process that allows us to find patterns, anomalies or correlations within large data sets that helps us acquire knowledge to predict outcomes or to validate findings. In order to find these patterns there are several algorithms. This paper discusses the following algorithms. 1) Decision Tree 2) Support Vector Machines 3) KNN 4) Naive Bayes The aforementioned algorithms are described as classification algorithms, it provides an interesting starting point for the analysis of them. The purpose of this paper is to analyze each of these algorithms to compare their prediction accuracy and features. We will be using diabetes prediction as the basis of this analysis and the Pima Indians Diabetes data set from kaggle.com as our input to determine whether a patient is diabetic or not. This paper will also highlight various other shortcomings or benefits of each algorithm and in which situations each of these algorithms perform the best.
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18398040
institution Zenodo
language
publishDate 2023
publisher Zenodo
record_format zenodo
spellingShingle Comparative analysis of K-Means, SVM, Decision Tree and Naive Bayes in Predicting Diabetes Presence
Rohan Almeida
Rajeev Dessai
Ayush Noorani
Amogh Pai Raiturkar
Samuel Godinho
Data Mining
Comparative analysis
Classification algorithms
Decision Tree
Support Vector Machines
KNN
Naive Bayes
In the context of rapidly growing amounts of data present today it is imperative to quickly dig out information from this data as information determines a large portion of the decision making process. Data mining allows us to do exactly that. Data mining is the process of turning raw data into useful information as it is an analytical process that allows us to find patterns, anomalies or correlations within large data sets that helps us acquire knowledge to predict outcomes or to validate findings. In order to find these patterns there are several algorithms. This paper discusses the following algorithms. 1) Decision Tree 2) Support Vector Machines 3) KNN 4) Naive Bayes The aforementioned algorithms are described as classification algorithms, it provides an interesting starting point for the analysis of them. The purpose of this paper is to analyze each of these algorithms to compare their prediction accuracy and features. We will be using diabetes prediction as the basis of this analysis and the Pima Indians Diabetes data set from kaggle.com as our input to determine whether a patient is diabetic or not. This paper will also highlight various other shortcomings or benefits of each algorithm and in which situations each of these algorithms perform the best.
title Comparative analysis of K-Means, SVM, Decision Tree and Naive Bayes in Predicting Diabetes Presence
topic Data Mining
Comparative analysis
Classification algorithms
Decision Tree
Support Vector Machines
KNN
Naive Bayes
url https://doi.org/10.5281/zenodo.18398040