There are many languages in the world, each one of them has its own characteristic. This makes languages a very interesting research, especially in text mining field. Written language need to be preprocessed first to get its normalized form. Abbreviated words are one of the problems in text mining. System cannot process the text optimally due to the different meaning the abbreviated words may have. This research objective is to develop Indonesian abbreviated words dataset, which can then be used to normalize any abbreviated words in Indonesian. Crowdsourcing was selected as a method to develop the dataset, because only human is able to translate abbreviated words into normal form of the words. From 1170 sentences that were tested, 1063 sentences were answered correctly while 107 sentences were answered incorrectly by the respondents. This research accuracy is about 90.85%. However, there is still a problem which occurred when the abbreviated word has more than one meaning/normal form, thus unique keywords are needed to determine which meaning is most accurate.
No takes yet. Share an insight, caveat, or question.
Sebastián et al. (2019) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: