Information technologies and computer sciences
Complex hierarchical approach for document clusterization
Keywords
text mining
data mining
dendogramms
k-means
hierarchical clusterization
vector model
cosine measure
Abstract
In this article we present integrated hieratical approach of text classification, based on dendrogramme and k-means clusterizations on computer. This approach allows us to present the computer-integrated new method of hierarchical clusterization, which can classify the amounts of classes given without a preliminary task, which allows keep structure documents on a computer. This approach is based on two methods related to the area text and data mining. The first stage is preprocessing of documents, as a result, time is reduced and a accurate result is calculated. The second stage is the use of vectorial model which allows expressly to define meaningfulness of words in a document. Then we use a hierarchical clusterization. It includes dendrogramms and k-means. Dendrogram method allows preliminary to define the amount of clusters (folders), the method of k-means attributes documents to certain clusters. The finishing stage is application of method of dendrogramms for creation of hierarchical sequence of documents into every cluster (folders).
How to Cite
[1]
“Complex hierarchical approach for document clusterization”, Вісник ВПІ, no. 1, pp. 47–50, Nov. 2010, Accessed: Oct. 05, 2026. Available: https://visnyk.vntu.edu.ua/index.php/visnyk/article/view/696
