Information technologies and computer sciences

Complex hierarchical approach for document clusterization

Authors

T. B. Shatovska
Kharkiv National University of Radio Electronics ROR
I. V. Kamenieva
Kharkiv National University of Radio Electronics ROR

Keywords

text mining data mining dendogramms k-means hierarchical clusterization vector model cosine measure

Abstract

In this article we present integrated hieratical approach of text classification, based on dendrogramme and k-means clusterizations on computer. This approach allows us to present the computer-integrated new method of hierarchical clusterization, which can classify the amounts of classes given without a preliminary task, which allows keep structure documents on a computer. This approach is based on two methods related to the area text and data mining. The first stage is preprocessing of documents, as a result, time is reduced and a accurate result is calculated. The second stage is the use of vectorial model which allows expressly to define meaningfulness of words in a document. Then we use a hierarchical clusterization. It includes dendrogramms and k-means. Dendrogram method allows preliminary to define the amount of clusters (folders), the method of k-means attributes documents to certain clusters. The finishing stage is application of method of dendrogramms for creation of hierarchical sequence of documents into every cluster (folders).
537 411

How to Cite

[1]
“Complex hierarchical approach for document clusterization”, Вісник ВПІ, no. 1, pp. 47–50, Nov. 2010, Accessed: Oct. 05, 2026. Available: https://visnyk.vntu.edu.ua/index.php/visnyk/article/view/696

Author Biographies

T. B. Shatovska, Kharkiv National University of Radio Electronics
доцент кафедри програмного забезпечення електронних обчислювальних машин
I. V. Kamenieva, Kharkiv National University of Radio Electronics
студентка