This article proposes the modified AHC (Agglomerative HierarchicalClustering) algorithm which considers the feature similarity and isapplied to the word clustering. The texts which are given asfeatures for encoding words into numerical vectors are semanticrelated entities, rather than independent ones, and the synergyeffect between the word clustering and the text clustering isexpected by combining both of them with each other. In thisresearch, we define the similarity metric between numerical vectorsconsidering the feature similarity, and modify the AHC algorithm byadopting the proposed similarity metric as the approach to the wordclustering. The proposed AHC algorithm is empirically validated asthe better approach in clustering words in news articles andopinions. The significance of this research is to improve theclustering performance by utilizing the feature similarities.
No takes yet. Share an insight, caveat, or question.
Taeho Jo (2024) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: