PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 1, 2002IEEE Transactions on Knowledge and Data Engineering217 citations

Unsupervised learning with mixed numeric and nominal data

View Full Paper
CLChongchong LiGBGautam Biswas

Key Points

Key points are not available for this paper at this time.

Abstract

Presents a similarity-based agglomerative clustering (SBAC) algorithm that works well for data with mixed numeric and nominal features. A similarity measure proposed by D.W. Goodall (1966) for biological taxonomy, that gives greater weight to uncommon feature value matches in similarity computations and makes no assumptions about the underlying distributions of the feature values, is adopted to define the similarity measure between pairs of objects. An agglomerative algorithm is employed to construct a dendrogram, and a simple distinctness heuristic is used to extract a partition of the data. The performance of the SBAC algorithm has been studied on real and artificially-generated data sets. The results demonstrate the effectiveness of this algorithm in unsupervised discovery tasks. Comparisons with other clustering schemes illustrate the superior performance of this approach.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Li et al. (2002) studied this question.

synapsesocial.com/papers/6a211f71a37b8f8d9296908ahttps://doi.org/10.1109/tkde.2002.1019208
Ask AI
Helpful
Bookmark
Share
View Full Paper