PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 1, 1972IEEE Transactions on Information Theory293 citations

Considerations of sample and feature size

View Full Paper
DFDonald H. Foley

Key Points

Key points are not available for this paper at this time.

Abstract

In many practical pattern-classification problems the underlying probability distributions are not completely known. Consequently, the classification logic must be determined on the basis of vector samples gathered for each class. Although it is common knowledge that the error rate on the design set is a biased estimate of the true error rate of the classifier, the amount of bias as a function of sample size per class and feature size has been an open question. In this paper, the design-set error rate for a two-class problem with multivariate normal distributions is derived as a function of the sample size per class (N) and dimensionality (L) . The design-set error rate is compared to both the corresponding Bayes error rate and the test-set error rate. It is demonstrated that the design-set error rate is an extremely biased estimate of either the Bayes or test-set error rate if the ratio of samples per class to dimensions (N/L) is less than three. Also the variance of the design-set error rate is approximated by a function that is bounded by 1/8N .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Donald H. Foley (1972) studied this question.

synapsesocial.com/papers/6a15d70e665e751854d11948https://doi.org/10.1109/tit.1972.1054863
Ask AI
Helpful
Bookmark
Share
View Full Paper