Clustering of image is one of the important steps of mining satellite images. our experiment we have simultaneously run multiple K-means algorithms with initial centroids and values of k in the same iteration of MapReduce. For initialization of initial centroids we have implemented Scalable-Means++ MapReduce (MR) job [1]. We have also run a validation algorithm of Silhouette Index [2] for multiple clustering outputs, again in the iteration of MR jobs. This paper explored the behavior of above mentioned algorithms when run on big data platforms like MapReduce and Spark. Spark has been chosen as it is popular for fast processing particularly iterations are involved.
No takes yet. Share an insight, caveat, or question.
Sharma et al. (2016) studied this question.