In view of the low utilization rate and the unclear value of the descriptive texts of the defects in the bridges of standard rail lines, we carry out text mining technology research. Firstly, by analyzing the characteristics of the descriptive texts of the defects in the bridges of standard rail lines, we establish a domain-specific dictionary based on the railway science and technology terms. Secondly, we clean the data and use the jieba segmentation tool to load the dictionary to segment the text data. Then, we extract key features of text information by using the TF-IDF algorithm, and split the training set and test set in a 8:2 ratio. Finally, we construct a BiLSTM_ATT classification model to train on the training set, and use the trained model and test set to validate the model. Through experimental comparison with commonly used keyword extraction algorithms, we find that for descriptive texts of defects in bridges of standard rail lines, based on the extraction of 1, 3, 5 keywords, the precision rate, recall rate and F-measure of the Chinese keyword extraction algorithm based on BiLSTM_ATT are better than other comparative classification models.
No takes yet. Share an insight, caveat, or question.
Guo et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: