In current era, we are experiencing tremendous growth in database sizes, types, users, working environments and data access speeds. This situation coined a new term Big Data which are large and complex datasets used for extracting meaningful knowledge. One of the main challenges in processing Big Data is its huge volume which is a common characteristic of huge collection of textual data also. Handling such voluminous big textual data using conventional data mining techniques such as clustering becomes impractical because of algorithmic incompetence to address the large computation time. This research work is mainly focused on big text data clustering using MapReduce based Distributed K-Means algorithm combined with corpus selection technique for a significant decrement of overall computation time. Four benchmark datasets have been used to explore the relationship between corpus size and computation time. It is found that the corpus selection technique significantly effective in reduction of overall processing time.
No takes yet. Share an insight, caveat, or question.
Ketu et al. (2015) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: