COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE AND SYSTEMS ANALYSIS

Method and Tool for Analyzing the Performance of Large Language Models for Complex Tasks Solution

Authors

Ye. M. Kryzhanovskyi ORCID 0000-0003-1311-1175
Vinnytsia National Technical University ROR
Vinnytsia National Technical University ROR
Vinnytsia National Technical University ROR
O. O. Voitsekhovska ORCID 0000-0001-8504-1204
Vinnytsia National Technical University ROR

Keywords

large language models complex tasks performance evaluation multi-criteria approach data integration automation

Abstract

This paper analyzes modern approaches to evaluating the efficiency of large language models for complex tasks solution that involve the simultaneous analysis of regulatory documents and numerical data. The study examines challenges related to the integration of heterogeneous information, ensuring logical consistency of results, and comparing models across multiple criteria, including correctness, stability, processing time, and adherence to instructions. The relevance of the research is determined by the rapid development of large language models and the need to select optimal solutions for applied tasks where errors may have significant consequences.

The proposed method is based on a multi-criteria approach that involves the simultaneous evaluation of model outputs on a complex query, including the processing of regulatory documents and CSV data containing average daily air pollutant concentrations. The evaluation is carried out using metrics such as correctness, stability, execution time, cost, token usage, and adherence to instructions.

A mathematical model of results aggregation is proposed, which includes averaging indicators over a set of test tasks, selecting a set of Pareto-optimal models, and constructing a scalar utility function to determine the optimal configuration of the model depending on the specified weight coefficients.

The methodology was validated using a case study of the city of Vinnytsia, demonstrating that different models exhibit significant differences in their ability to produce consistent and coherent conclusions. Based on the analysis of regulatory thresholds by the models, the optimal number of monitoring stations for different pollutants was determined, ensuring representative air quality monitoring. The proposed approach accounts for the specific features of complex applied tasks, enables the integration of heterogeneous data, and improves the validity of model selection for practical applications.

It is established that the models of the GPT-4o and GPT-5.1 class provide a better balance between correctness and stability at an acceptable level of resource costs. At the applied level, justified recommendations on the number of atmospheric air observation points for the city of Vinnytsia were obtained, which confirms the ability of the models to integrate regulatory requirements and numerical data. It is shown that the correctness of such solutions significantly depends not only on the quality of the model, but also on its stability and ability to follow instructions.

2 0

How to Cite

[1]
“Method and Tool for Analyzing the Performance of Large Language Models for Complex Tasks Solution”, Вісник ВПІ, no. 4, pp. 97–107, Oct. 2026, doi: 10.31649/1997-9266-2026-187-4-97-107.

Author Biographies

Ye. M. Kryzhanovskyi, Vinnytsia National Technical University

Cand. Sc. (Eng.), Associate Professor, Associate Professor of the Chair of System Analysis and Information Technologies

I. M. Shtelmakh, Vinnytsia National Technical University

Cand. Sc. (Eng.), Assistant of the Chair of System Analysis and Information Technologies

M. V. Shvets, Vinnytsia National Technical University

Post-Graduate Student, of the Chair of System Analysis and Information Technologies

O. O. Voitsekhovska, Vinnytsia National Technical University

PhD, Associate Professor of the Chair of System Analysis and Information Technologies

References

[1] S. M. Sajjadi Mohammadabadi, et al., “A Survey of Large Language Models: Evolution, Architectures, Adaptation, Benchmarking, Applications, Challenges, and Societal Implications,” Electronics, vol. 14, no. 18: 3580, 2025. https://doi.org/10.3390/electronics14183580 .
[2] А. В. Лосенко, Є. М. Крижановський, І. М. Штельмах, і І. В. Варчук, «Технологія LLM-видобування ознак тестування пацієнтів з текстових звітів для удосконалення прогнозування кількості хворих на коронавірус,» Вісник Вінницького політехнічного інституту, № 6, с. 135-144. 2024. https://doi.org/10.31649/1997-9266-2024-177-6-135-144 .
[3] Є. М. Крижановський, В. О. Караваєв, І. М. Штельмах, і О. О. Войцеховська, «Автоматизація використання природномовних запитів для комплексного аналізу стану поверхневих вод басейну Південного Бугу,» Вісник Вінницького політехнічного інституту, № 4, с. 118-125. 2025. https://doi.org/10.31649/1997-9266-2025-181-4-118-125 .
[4] X. Han et al., “Pre-Trained Models: Past, Present and Future,” AI Open, vol. 2, pp. 225-250, 2021. https://doi.org/10.1016/j.aiopen.2021.08.002 .
[5] Q. He et al., “Can Large Language Models Understand Real-World Complex Instructions?” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, pp. 18188-18196, 2024. https://doi.org/10.1609/aaai.v38i16.29777 .
[6] K. Papineni, S. Roukos, T. Ward, and W-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (ACL ‘02), Philadelphia, Pennsylvania, USA, July 7—12, 2002, pp. 311-318. https://doi.org/10.3115/1073083.1073135 .
[7] C-Y. Lin, and F. J. Och, “Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics,” in Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics (ACL ‘04), Barcelona, Spain, July 21-26, 2004, p. 605. https://doi.org/10.3115/1218955.1219032 .
[8] T. Sellam, D. Das, and A. Parikh, “BLEURT: Learning Robust Metrics for Text Generation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA, July, 2020, pp. 7881-7892. https://aclanthology.org/2020.acl-main.704 .
[9] R. Bommasani, P. Liang, and T. Lee, “Holistic Evaluation of Language Models,” Annals of the New York Academy of Sciences, vol. 1525, pp. 140-146, 2023. https://doi.org/10.1111/nyas.15007 .
[10] Y. Liu, et al., “Datasets for large language models: a comprehensive survey,” Artificial Intelligence Review, vol. 58, article 403, 2025. https://doi.org/10.1007/s10462-025-11403-7 .
[11] W. Cao, et al., “Benchmarking large language models against human experts in rehabilitation medicine: a multidimensional evaluation,” Journal of NeuroEngineering and Rehabilitation, vol. 23, article 84, 2026. https://doi.org/10.1186/s12984-026-01903-0 .
[12] H. V. Pandhare, “Evaluating Large Language Models: Frameworks and Methodologies for AI/ML System Testing,” International Journal of Scientific Research and Management (IJSRM), vol. 12, issue 09, pp. 1467-1486. https://doi.org/10.18535/ijsrm/v12i09.ec08 .
[13] Є. М. Крижановський, і М. В. Швець, «Підхід до збору та візуалізації даних державного моніторингу стану атмосферного повітря міста Вінниці,» Матеріали LV Всеукраїнської науково-технічної конференції підрозділів ВНТУ, Вінниця, 24-27 березня 2026. [Електронний ресурс]. Режим доступу: https://conferences.vntu.edu.ua/index.php/all-fksa/all-fksa-2026/paper/view/27366
[14] LLM Comparator. [Online]. Available: https://llmcomparator.flitsolutions.com/ . Accessed: April 06, 2026.
[15] T. L. Saaty, “A scaling method for priorities in hierarchical structures,” Journal of Mathematical Psychology, vol. 15, no. 3, pp. 234-281, 1977. https://doi.org/10.1016/0022-2496(77)90033-5 .
[16] C.-L. Hwang, and K. Yoon, Multiple Attribute Decision Making: Methods and Applications. Berlin, Heidelberg: Springer, 1981. https://doi.org/10.1007/978-3-642-48318-9 .
[17] S. Opricovic, and G.-H. Tzeng, “Compromise solution by MCDM methods: A comparative analysis of VIKOR and TOPSIS,” European Journal of Operational Research, vol. 156, no. 2, pp. 445-455, 2004. https://doi.org/10.1016/S0377-2217(03)00020-1 .
[18] B. Roy, “Classement et choix en présence de points de vue multiples,” RAIRO – Revue d’automatique, d’informatique et de recherche opérationnelle, vol. 2, no. 1, pp. 57-75, 1968. https://doi.org/10.1051/ro/196802V100571 .
[19] Постанова Кабінету Міністрів України «Деякі питання здійснення державного моніторингу в галузі охорони атмосферного повітря» від 14.08.2019 р. № 827 зі змінами від 7.05. 2024 р. [Електронний ресурс]. Режим доступу: https://zakon.rada.gov.ua/laws/show/827-2019-%D0%BF#Text . Дата звернення: 06.04.2026.
[20] Наказ Міністерства внутрішніх справ України від 21.04.2021 № 300 «Про затвердження Порядку розміщення пунктів спостережень за забрудненням атмосферного повітря в зонах та агломераціях» [Електронний ресурс]. Режим доступу: https://zakon.rada.gov.ua/laws/show/z0635-21#Text . Дата звернення: 06.04.2026.

Most read articles by the same author(s)

1 2 3 > >>