Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 8, 2026SymmetryOpen Access

Optimizing K-Means Clustering for Big Data: A Review

View Full Paper
Ask AI
Bookmark
Share

Authors

RMRavil MussabayevRMRavil MussabayevRMRustam Mussabayev

Discussion

Loading...

Member takes

Overview

Algorithmic review evaluates optimization techniques for big data K-means clustering, demonstrating that higher algorithmic complexity does not ensure superior practical trade-offs.

Key Points

  • To comprehensively analyze and benchmark optimization techniques designed to overcome scalability limitations in minimum sum-of-squares clustering for big data.
  • Reviewed optimization approaches including decomposition, sampling, initialization, parallel and distributed computing, data summarization, and metaheuristic hybridization.
  • Benchmarked representative algorithms using a unified protocol assessed via the 'less is more' approach (LIMA) across clustering quality, execution speed, and simplicity.
  • Demonstrated a multi-algorithm Pareto front under LIMA dominance, confirming that no individual optimization technique is universally superior across all evaluation metrics.
  • Lightweight approaches maximized processing speed at the expense of accuracy, whereas complex hybrid frameworks achieved higher clustering accuracy at substantially elevated computational costs.
  • Stochastic-sampling methods consistently occupied an intermediate position, offering an effective trade-off among accuracy, execution time, and algorithmic simplicity.

Cite This Study

Mussabayev et al. (2026) studied this question.

synapsesocial.com/papers/6a9fd80458e84d0ff5b47191https://doi.org/10.3390/sym18091489
View Full Paper
Ask AI
Bookmark
Share