COMPUTER ENGINEERING, INFORMATION SYSTEMS AND TECHNOLOGIES

Intelligent System of Semantical Consistency Detection in Distributed Information Resources

Authors

Lviv Polytechnic National University ROR
Lviv Polytechnic National University ROR

Keywords

distributed information resources data quality semantic consistency intelligent system hybrid approach vector representations data validation data integration

Abstract

This article addresses the problem of data quality assessment in distributed information resources under conditions of rapidly growing data volume, its heterogeneity, and dynamism in modern information systems. It demonstrates that conventional data quality approaches, which focus mainly on accuracy, completeness, and consistency, are no longer sufficient for distributed environments because they do not fully capture semantic heterogeneity, usage context, replication effects, entity representation variability, and differences in data interpretation rules. Particular attention is paid to semantic consistency as a key property that determines the correctness of data integration, the reliability of analytical conclusions, and the quality of automated decisions in heterogeneous environments.

The paper implements an intelligent system for detecting semantic consistency in distributed information resources based on a hybrid approach that combines fast syntactic filtering with deep semantic analysis using vector representations. This architecture makes it possible to process both obviously consistent data pairs and more complex cases where the same meaning is expressed through different lexical forms, abbreviations, or domain-specific synonyms. In addition, the proposed solution implements adaptive threshold-based decision making, allowing the strictness of verification to be adjusted according to the attribute type and application domain, while also producing an integrated semantic consistency index for further monitoring of data quality.

Experimental evaluation confirmed the feasibility of the proposed approach for improving the detection of semantically equivalent but formally different records. The results showed that combining syntactic and semantic methods significantly reduces both false rejections and false matches during attribute comparison, while also providing a methodological basis for building scalable quality control systems for distributed information resources. The practical value of the study lies in the possibility of integrating the developed system into data quality monitoring, validation, and management platforms for financial, medical, scientific, and other heterogeneous information environments.

11 0

How to Cite

[1]
“Intelligent System of Semantical Consistency Detection in Distributed Information Resources”, Вісник ВПІ, no. 4, pp. 24–33, Sep. 2026, doi: 10.31649/1997-9266-2026-187-4-24-33.

Author Biographies

Yu. M. Heriak, Lviv Polytechnic National University

Post-Graduate Student of the Chair of Information Systems and Networks

A. Yu. Berko, Lviv Polytechnic National University

Dr. Sc. (Eng.), Professor of the Chair of Information Systems and Networks

References

[1] L. Cai and Y. Zhu, “The challenges of data quality and data quality assessment in the big data era,” Data Science Journal, vol. 14, p. 2, 2015. https://doi.org/10.5334/dsj-2015-002 .
[2] J. Declerck, et al., “Assessing data quality in heterogeneous healthcare integration: The AIDAVA framework (preprint),” JMIR Medical Informatics, 2025. https://doi.org/10.2196/75275 .
[3] Ю. М. Геряк, i А. Ю. Берко, «Система критеріїв оцінки якості даних в розподілених інформаційних системах,» Вісник Національного університету «Львівська політехніка». Серія: Інформаційні системи та мережі, № 16, с. 191-202, 2024. https://doi.org/10.23939/sisn2024.16.191 .
[4] S. Vujević, M. Cerjan, and K. Rabuzin, “Data quality in distributed information systems: Conceptual challenges and empirical findings,” in Proc. Central European Conf. Information and Intelligent Systems, 2025, pp. 105-112. [Online]. Available: https://archive.ceciis.foi.hr/public/conferences/2025/Proceedings/S3/3.pdf .
[5] O. Novytskyi, “Concept and evaluation of the big data quality in semantic environment,” Problemy Prohramuvannia, no. 3-4, pp. 260-270, 2022. https://doi.org/10.15407/pp2022.03-04.260 .
[6] B. Heinrich, M. Klier, A. Schiller, and G. Wagner, “Assessing data quality – A probability-based metric for semantic consistency,” Decision Support Systems, vol. 110, pp. 95-106, 2018. https://doi.org/10.1016/j.dss.2018.03.011 .
[7] Z. Tan, et al., “Semantic similarity distance: Towards better text-image consistency metric in text-to-image generation,” Pattern Recognition, vol. 144, p. 109883, 2023. https://doi.org/10.1016/j.patcog.2023.109883 .
[8] N. Wang, and S.-G. Leem, “Modernizing data quality: Evolving dimensions for unstructured text data in big data, AI and ethical contexts,” IntechOpen, 2026. https://doi.org/10.5772/intechopen.1013232 .
[9] M. Klier, A. Obermeier, C. Sparn, and T. Widmann, “Anomaly-based assessment of semantic consistency: Design and evaluation of a novel probability-based metric in cooperation with a German car manufacturer,” Journal of Data and Information Quality, vol. 17, no. 2, art. no. 24, 2025. https://doi.org/10.1145/3732783 .
[10] R. Huidrom, M. Lorandi, S. Mille, C. Thomson, and A. Belz, “Assessing semantic consistency in data-to-text generation: A meta-evaluation of textual, semantic and model-based metrics,” in Proc. 18th Int. Natural Language Generation Conf., Hanoi, Vietnam, 2025, pp. 98-107. [Online]. Available: https://aclanthology.org/2025.inlg-main.6/ .
[11] Ю. М. Геряк та А. Ю. Берко, «Інтелектуальна система виявлення плагіату в технічних текстах,» Вісник Національного університету «Львівська політехніка». Серія: Інформаційні системи та мережі, № 14, с. 235-247, 2023. https://doi.org/10.23939/sisn2023.14.235 .
[12] Ю. М. Геряк, і А. Ю. Берко, «Оцінка семантичної та часової узгодженості розподілених інформаційних ресурсів,» Model. Control Inf. Technol., № 8, с. 125-128, 2025. https://doi.org/10.31713/MCIT.2025.036 .
[13] Y. Heriak, and A. Berko, “Comparison analysis of metrics and evaluation scales of distributed information resources quality parameters,” Вісник Національного університету «Львівська політехніка». Серія: Інформаційні системи та мережі, № 17, с. 130-137, 2025. https://doi.org/10.23939/sisn2025.17.130 .
[14] M. Chen, S. Mao, and Y. Liu, “Big Data: A Survey,” Mobile Networks and Applications, vol. 19, no. 2, pp. 171-209, 2014. https://doi.org/10.1007/s11036-013-0489-0 .