AUTOMATION, ІОТ, ROBOTICS AND INFORMATION-MEASUREMENT SYSTEMS

Use of Machine Learning Algorithms to Determine the Presence of Diabetes Mellitus

Authors

Zhytomyr Polytechnic State University ROR
O. M. Svintsytska ORCID 0000-0002-2613-2437
Zhytomyr Polytechnic State University ROR
Zhytomyr Polytechnic State University ROR
Zhytomyr Polytechnic State University ROR

Keywords

data analysis machine learning forecasting logistic regression k-nearest neighbors random fores

Abstract

The article is devoted to the research and practical application of modern machine learning methods for solving the urgent problem of detecting diabetes mellitus in patients. At the initial stage, the work focuses on preliminary data processing: detecting empty and abnormal values, studying data set variables, etc. Descriptive statistics are used to summarize, organize, and present the main characteristics of the dataset. Key indicators are obtained for numerical variables: mean, median, standard deviation, minimum and maximum values, as well as quartiles. This made it possible to assess the central tendency of the data, their dispersion, range, and identify irregularities in the distribution. For categorical variables, the number of unique values, the most common category, and its frequency were determined. The use of machine learning algorithms to diagnose diabetes in patients was described. The task was to predict the presence (result = 1) or absence (result = 0) of diabetes in patients based on a number of diagnostic indicators, including glucose level, blood pressure, skinfold thickness, insulin level, body mass index, age, etc. The study analyzed input data to identify patterns and correlations between diagnostic parameters, analyze the distribution of characteristics, and identify potential emissions. Several classification algorithms were considered and applied to build models, including logistic regression, the k-nearest neighbors method, and random forest. The final stage of the study was to compare the efficiency of different algorithms. The best results were demonstrated by the logistic regression model with an adapted classification threshold. The paper justifies shifting the decision threshold from the standard value of 0.5 to the range of 0.25...0.3, which significantly minimizes type II errors (false negatives). This is important for the early diagnosis of diabetes, as it allows for the detection of a larger number of patients who might have been missed with the standard model settings.

3 1

How to Cite

[1]
“Use of Machine Learning Algorithms to Determine the Presence of Diabetes Mellitus”, Вісник ВПІ, no. 4, pp. 178–187, Oct. 2026, doi: 10.31649/1997-9266-2026-187-4-178-187.

Author Biographies

O. V. Korotun, Zhytomyr Polytechnic State University

Cand. Sc. (Pedagog.), Associate Professor, Associate Professor of the Chair of Computer Science

O. M. Svintsytska, Zhytomyr Polytechnic State University

Cand. Sc. (Econom.), Associate Professor, Associate Professor of the Chair of Computer Science

D. K. Marchuk, Zhytomyr Polytechnic State University

Senior Lecturer of the Chair of Computer Scienc

M. R. Shtyl, Zhytomyr Polytechnic State University

Researcher of the Chair of Computer Science

References

[1] Л. Ліщинська, і Н. Добровольська, «Перспективні програмні інструменти для аналізу даних у бізнесі,» Вісник Хмельницького національного університету, № 1 (305), с. 78-83, 2022. http://ir.lib.vntu.edu.ua//handle/123456789/35278 .
[2] В. Нестеров, А. Шиш, і Т. Музиченко, «Ефективний економічний розвиток підприємства через інтелектуальний аналіз даних: використання AI для прогнозування та оптимізації стратегій бізнесу,» Економіка та суспільство, № 59, 2024. https://doi.org/10.32782/2524-0072/2024-59-87 .
[3] М. Ціж, «Огляд застосування сучасних методів топологічного аналізу даних,» TSynergy, № 2, с. 112-122, 2024, https://doi.org/10.53920/ITS-2024-2-7 .
[4] О. Дацок, «Етапи обробки медичних даних для розв’язання задач алгоритмами машинного навчання,» Вісник Національного технічного університету, т. 1. № 2(14), с. 144-156, 2025. https://doi.org/10.20998/2411-0558.2025.02.10 .
[5] Д. Панаскін, і Є. Білоконь, «Машинне навчання в діагностиці захворювань легеневої системи,» Технічні науки та технології, № 2 (28), с. 76-87, 2022. https://doi.org/10.25140/2411-5363-2022-2(28)-76-87 .
[6] Н. Бугаєць, і І. Лисенко, «Задача прогнозування в машинному навчанні,» Автоматизація технологічних та бізнес-процесів, т. 17, вип. 2, 2025. https://doi.org/10.15673/atbp.v17i2.3150 .
[7] М. Лисенко, і О. Пронькін, «Застосування технологій машинного навчання для встановлення медичного діагнозу», Зв’язок, № 5 (171), с. 75-82, 2025. https://doi.org/10.31673/2412-9070.2024.051397 .
[8] І. Калініна, і О. Гожий, «Дослідження ефективності методів класифікації при прогнозуванні в задачах машинного навчання,» Управління розвитком складних систем, вип. 46, с. 173-180, 2021. https://doi.org/10.32347/2412-9933.2021.46.173-180 .