This article proposes the modified KNN (K Nearest Neighbor)algorithm which receives a string vector as its input data and isapplied to the index optimization. The results from applying thestring vector based algorithms to the text categorizations weresuccessful in previous works, and the index optimization is able tobe viewed into a classification task where each word is classifiedinto expansion, inclusion, and removal. In the proposed system, eachword in the given text is classified into one of the threecategories by the proposed KNN algorithm, associates words are addedto ones which are classified into expansion, and ones which areclassified into inclusion are kept by themselves without adding anyword. The proposed KNN version is empirically validated as thebetter approach in deciding the importance level of words in newsarticles and opinions. We need to define and characterizemathematically more operations on string vectors for modifying moreadvanced machine learning algorithms.
No takes yet. Share an insight, caveat, or question.
Taeho Jo (2024) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: