Diabetes is one of the most prevalent diseases worldwide, capable of damaging various internal systems. Diabetes patients require routine check-ups, resulting in a time series of laboratory records such as hemoglobin A1c (HbA1c), which reflect each patient's health behavior over time. Clustering patients into groups based on their entire time series data assists doctors in making recommendations and choosing treatments without the need to review all records. However, clustering this type of dataset introduces some challenges; patients visit their doctors at different time points, making it difficult to match levels, trends, peaks, and patterns of their time series. To address these challenges, we introduce a novel method: Time and Trend Traveling Time Series Clustering (4TaStiC), using a base dissimilarity combined with Euclidean and Pearson correlation metrics. We evaluated this algorithm on labeled artificial datasets, comparing its performance with that of eight existing methods including a representation learning-based deep learning method. 4TaStiC outperformed the other methods based on both accuracy and Adjusted Rand Index. Finally, we applied 4TaStiC to cluster 1,989 type 2 diabetes patients at Siriraj Hospital using their HbA1c time series data. Each group of patients exhibits clear characteristics that will benefit doctors in making efficient clinical decisions.
Preedasawakul et al. (2026) studied this question.