Intelligent System of Semantical Consistency Detection in Distributed Information Resources
Keywords
Abstract
This article addresses the problem of data quality assessment in distributed information resources under conditions of rapidly growing data volume, its heterogeneity, and dynamism in modern information systems. It demonstrates that conventional data quality approaches, which focus mainly on accuracy, completeness, and consistency, are no longer sufficient for distributed environments because they do not fully capture semantic heterogeneity, usage context, replication effects, entity representation variability, and differences in data interpretation rules. Particular attention is paid to semantic consistency as a key property that determines the correctness of data integration, the reliability of analytical conclusions, and the quality of automated decisions in heterogeneous environments.
The paper implements an intelligent system for detecting semantic consistency in distributed information resources based on a hybrid approach that combines fast syntactic filtering with deep semantic analysis using vector representations. This architecture makes it possible to process both obviously consistent data pairs and more complex cases where the same meaning is expressed through different lexical forms, abbreviations, or domain-specific synonyms. In addition, the proposed solution implements adaptive threshold-based decision making, allowing the strictness of verification to be adjusted according to the attribute type and application domain, while also producing an integrated semantic consistency index for further monitoring of data quality.
Experimental evaluation confirmed the feasibility of the proposed approach for improving the detection of semantically equivalent but formally different records. The results showed that combining syntactic and semantic methods significantly reduces both false rejections and false matches during attribute comparison, while also providing a methodological basis for building scalable quality control systems for distributed information resources. The practical value of the study lies in the possibility of integrating the developed system into data quality monitoring, validation, and management platforms for financial, medical, scientific, and other heterogeneous information environments.
