Use of Machine Learning Algorithms to Determine the Presence of Diabetes Mellitus
Keywords
Abstract
The article is devoted to the research and practical application of modern machine learning methods for solving the urgent problem of detecting diabetes mellitus in patients. At the initial stage, the work focuses on preliminary data processing: detecting empty and abnormal values, studying data set variables, etc. Descriptive statistics are used to summarize, organize, and present the main characteristics of the dataset. Key indicators are obtained for numerical variables: mean, median, standard deviation, minimum and maximum values, as well as quartiles. This made it possible to assess the central tendency of the data, their dispersion, range, and identify irregularities in the distribution. For categorical variables, the number of unique values, the most common category, and its frequency were determined. The use of machine learning algorithms to diagnose diabetes in patients was described. The task was to predict the presence (result = 1) or absence (result = 0) of diabetes in patients based on a number of diagnostic indicators, including glucose level, blood pressure, skinfold thickness, insulin level, body mass index, age, etc. The study analyzed input data to identify patterns and correlations between diagnostic parameters, analyze the distribution of characteristics, and identify potential emissions. Several classification algorithms were considered and applied to build models, including logistic regression, the k-nearest neighbors method, and random forest. The final stage of the study was to compare the efficiency of different algorithms. The best results were demonstrated by the logistic regression model with an adapted classification threshold. The paper justifies shifting the decision threshold from the standard value of 0.5 to the range of 0.25...0.3, which significantly minimizes type II errors (false negatives). This is important for the early diagnosis of diabetes, as it allows for the detection of a larger number of patients who might have been missed with the standard model settings.
