For text mining applications, keyword extraction is a popular problem to solve. The text mining applications as indexing, summarization and topic tracking uses keyword extraction models as baseline. There are many sequence labeling models proposed in the literature. The performance results, however, did not meet expectations. In this paper, we use many selective statistical and graph-based features to solve the keyword extraction problem. We calculated many selective features for each considerable word in input text. We used Random Forest token classification model. In addition to widely used 500N-KPCrowd dataset in the literature, we trained and tested publicly available KazakhNews and RussianNews datasets to extend language domain. Moreover, we collected two news dataset in Chinese and Arabic languages namely ChineseNews and ArabianNews. We also used a Turkish abstract dataset DergiParkTR for comparison. The performance of the graph-based and statistical features were tested with Random forest algorithm. We have obtained 0.680, 0.706 and 0.896 F1-scores for DergiParkTR, 500N-KPCrowd and RussianNews datasets respectively. We attained the best performance of 0.978 for KazakhNews using graph-based features and 0.999 for ChineseNews datasets with all statistical features.
No takes yet. Share an insight, caveat, or question.
Kılıç et al. (2024) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: