Method and Tool for Analyzing the Performance of Large Language Models for Complex Tasks Solution
Keywords
Abstract
This paper analyzes modern approaches to evaluating the efficiency of large language models for complex tasks solution that involve the simultaneous analysis of regulatory documents and numerical data. The study examines challenges related to the integration of heterogeneous information, ensuring logical consistency of results, and comparing models across multiple criteria, including correctness, stability, processing time, and adherence to instructions. The relevance of the research is determined by the rapid development of large language models and the need to select optimal solutions for applied tasks where errors may have significant consequences.
The proposed method is based on a multi-criteria approach that involves the simultaneous evaluation of model outputs on a complex query, including the processing of regulatory documents and CSV data containing average daily air pollutant concentrations. The evaluation is carried out using metrics such as correctness, stability, execution time, cost, token usage, and adherence to instructions.
A mathematical model of results aggregation is proposed, which includes averaging indicators over a set of test tasks, selecting a set of Pareto-optimal models, and constructing a scalar utility function to determine the optimal configuration of the model depending on the specified weight coefficients.
The methodology was validated using a case study of the city of Vinnytsia, demonstrating that different models exhibit significant differences in their ability to produce consistent and coherent conclusions. Based on the analysis of regulatory thresholds by the models, the optimal number of monitoring stations for different pollutants was determined, ensuring representative air quality monitoring. The proposed approach accounts for the specific features of complex applied tasks, enables the integration of heterogeneous data, and improves the validity of model selection for practical applications.
It is established that the models of the GPT-4o and GPT-5.1 class provide a better balance between correctness and stability at an acceptable level of resource costs. At the applied level, justified recommendations on the number of atmospheric air observation points for the city of Vinnytsia were obtained, which confirms the ability of the models to integrate regulatory requirements and numerical data. It is shown that the correctness of such solutions significantly depends not only on the quality of the model, but also on its stability and ability to follow instructions.
