The Implementation of Support Vector Machine and Naïve Bayes Algorithm to Predict Diabetes
DOI:
https://doi.org/10.51601/ijse.v6i2.744Abstract
The significant increase in diabetes mellitus cases within the community demands a technology-based solution that can provide accurate, efficient, and reliable predictions. This study aims to evaluate the impact of various data preprocessing schemes on the performance of the Gaussian Naive Bayes (GNB) and Support Vector Machine (SVM) algorithms in classifying diabetes risk. The dataset used in this research was sourced from the UCI Machine Learning Repository and consists of 520 records with 16 symptom features and 1 target label. The preprocessing stages include handling missing values, encoding categorical features, normalizing numerical data using StandardScaler, balancing the dataset with the Synthetic Minority Over-sampling Technique (SMOTE), and feature selection using the SelectKBest method. A total of nine preprocessing scheme combinations were tested for each algorithm. The experimental results show that for the GNB model, the best performance was achieved using the combination of StandardScaler, SMOTE, and SelectKBest (k=5), reaching an accuracy of 94.53%, precision 98.36%, recall 90.91%, and f1-score 94.49%. Meanwhile, for the SVM model, the highest performance was obtained through the combination of StandardScaler and RBF kernel hyperparameter tuning, achieving an accuracy of 99.04%, precision 99.05%, recall 99.04%, and f1-score 99.03%. The evaluation was conducted using metrics such as accuracy, precision, recall, F1-score, confusion matrix, and learning curve visualization. These findings highlight the critical role of proper preprocessing in enhancing predictive model performance. This study is expected to serve as a reference for developing early detection systems for diabetes based on machine learning.
Downloads
References
[1] International Diabetes Federation, IDF Diabetes Atlas, 10th ed. Brussels, Belgium: International Diabetes Federation, 2021.
[2] E. Schleicher et al., “Definition, Classification and Diagnosis of Diabetes Mellitus,” Experimental and Clinical Endocrinology & Diabetes, vol. 130, Apr. 2022, doi: https://doi.org/10.1055/a-1624-2897.
[3] Yohanes Denggos, “Penyakit Diabetes Mellitus Umur 40-60 Tahun di Desa Bara Batu Kecamatan Pangkep,” Healthcaring, vol. 2, no. 1, pp. 55–61, Mar. 2023, doi: https://doi.org/10.47709/healthcaring.v2i1.2177.
[4] Z. Punthakee, R. Goldenberg, and P. Katz, “Definition, Classification and Diagnosis of Diabetes, Prediabetes and Metabolic Syndrome,” Canadian Journal of Diabetes, vol. 42, no. 1, pp. S10–S15, Apr. 2018, doi: https://doi.org/10.1016/j.jcjd.2017.10.003.
[5] K.-J. He, H. Wang, J. Xu, G. Gong, X. Liu, and H. Guan, “Global burden of type 2 diabetes mellitus from 1990 to 2021, with projections of prevalence to 2044: a systematic analysis across SDI levels for the global burden of disease study 2021,” Frontiers in Endocrinology, vol. 15, Nov. 2024, doi: https://doi.org/10.3389/fendo.2024.1501690.
[6] E. Oikonomou and R. Khera, “Machine learning in precision diabetes care and cardiovascular risk prediction,” Cardiovascular Diabetology, vol. 22, no. 1, Sep. 2023, doi: https://doi.org/10.1186/s12933-023-01985-3.
[7] R. G. Wardhana, G. Wang, and F. Sibuea, “Penerapan Machine Learning Dalam Prediksi Tingkat Kasus Penyakit Di Indonesia,” Journal of Information System Management (JOISM), vol. 5, no. 1, pp. 40–45, Jul. 2023, doi: https://doi.org/10.24076/joism.2023v5i1.1136.
[8] R. Hidayat, Y. S. Sy, T. Sujana, M. Husnah, H. T. Saputra, and F. Okmayura, “Implementasi Machine Learning Untuk Prediksi Penyakit Jantung Menggunakan Algoritma Support Vector Machine,” BIOS : Jurnal Teknologi Informasi dan Rekayasa Komputer, vol. 5, no. 2, pp. 161–168, Sep. 2024, doi: https://doi.org/10.37148/bios.v5i2.152.
[9] A. Rahman, Agung Nugroho, and A. H. Anshor, “Prediksi Penyakit Kanker Paru-Paru Dengan Algoritma Regresi Linier,” Bulletin of Information Technology (BIT), vol. 4, no. 1, pp. 63–74, Mar. 2023, doi: https://doi.org/10.47065/bit.v4i1.501.
[10] F. Saptawan, David, T. Wijaya, S. Kosasi, and S. Margaretha Kuway, “Prediksi Epidemiologi Penyakit Tidak Menular Menggunakan Algoritma Random Forest Pada Puskesmas,” Jurnal TIMES, vol. 13, no. 2, pp. 192–201, Dec. 2024, doi: https://doi.org/10.51351/jtm.13.2.2024788.
[11] S. Basu, K. T. Johnson, and S. A. Berkowitz, “Use of Machine Learning Approaches in Clinical Epidemiological Research of Diabetes,” Current Diabetes Reports, vol. 20, no. 12, Dec. 2020, doi: https://doi.org/10.1007/s11892-020-01353-5.
[12] U. N. Dulhare, K. Ahmad, Ahmad, and J. Wiley, Machine learning and big data : concepts, algorithms, tools and applications. Hoboken, Nj: Wiley-Scrivener, 2020.
[13] R. Sarno, S. Kom., Malikhah, S.Kom., M.Kom, D. Putra, and S. Hanif, Machine Learning dan Deep Learning-Konsep dan Pemrograman Python. Penerbit Andi.
[14] D. Effrosynidis and A. Arampatzis, “An evaluation of feature selection methods for environmental data,” Ecological Informatics, vol. 61, p. 101224, Mar. 2021, doi: https://doi.org/10.1016/j.ecoinf.2021.101224.
[15] I. Javid, R. Ghazali, I. Syed, M. Zulqarnain, and N. A. Husaini, “Study on the Pakistan stock market using a new stock crisis prediction method,” PLoS ONE, vol. 17, no. 10, p. e0275022, Oct. 2022, doi: https://doi.org/10.1371/journal.pone.0275022.
[16] Kumar Abhishek and Dr. Mounir Abdelaziz, Machine Learning for Imbalanced Data. Packt Publishing Ltd, 2023.
[17] F. R. Adi Pratama and S. I. Oktora, “Synthetic Minority Over-sampling Technique (SMOTE) for handling imbalanced data in poverty classification,” Statistical Journal of the IAOS, pp. 1–7, Dec. 2022, doi: https://doi.org/10.3233/sji-220080.
[18] R. Doshi, MACHINE LEARNING : master supervised and unsupervised learning algorithms with real... examples. S.L.: Bpb Publications, 2021.
[19] Aman Kharwal, Machine Learning Algorithms: Handbook. 2023.
[20] A. Ouldammar, A. Moulay Lakhdar, A. Bouida, and K. Merit, “Optimization machine learning models for selecting transmit antennas in 5G/6G systems,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 37, no. 2, p. 819, Feb. 2025, doi: https://doi.org/10.11591/ijeecs.v37.i2.pp819-828.
[21] Yessy Asri, S.T., MMSI and Dr. Dra. Dwina Kuswardani, M.Kom, dkk, MACHINE LEARNING & DEEP LEARNING: Analisis Sentimen Menggunakan Ulasan Pengguna Aplikasi. Uwais Inspirasi indonesia, 2024.
[22] G. Vivar, R. Strobl, E. Grill, Nassir Navab, A. Zwergal, and S.-A. Ahmadi, “Using Base-ml to Learn Classification of Common Vestibular Disorders on DizzyReg Registry Data,” Frontiers in Neurology, vol. 12, Aug. 2021, doi: https://doi.org/10.3389/fneur.2021.681140.
[23] E. Elgeldawi, A. Sayed, A. R. Galal, and A. M. Zaki, “Hyperparameter Tuning for Machine Learning Algorithms Used for Arabic Sentiment Analysis,” Informatics, vol. 8, no. 4, p. 79, Nov. 2021, doi: https://doi.org/10.3390/informatics8040079.
[24] D. U. Ozsahin, M. Taiwo Mustapha, A. S. Mubarak, Z. Said Ameen, and B. Uzun, “Impact of feature scaling on machine learning models for the diagnosis of diabetes,” 2022 International Conference on Artificial Intelligence in Everything (AIE), Aug. 2022, doi: https://doi.org/10.1109/aie57029.2022.00024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Joshua Roy Danna Lacanlale, Vitri Tundjungsari

This work is licensed under a Creative Commons Attribution 4.0 International License.


















