Many chemometrics methods assume that the underlying population of experimental measurements is normally distributed. Whether using significance tests in soft independent modeling by class analogy, or confidence limits in regression, or interpreting the significance of factors using ANOVA, the assumption of normality is often buried deep inside. Often, decisions about, for example, whether a process is under control in manufacturing, or what the origins are of a food as measured by near infrared, or what the confidence limits are in the measurement of an additive in a fuel often depend on one or more underlying populations falling into a normal distribution. The normal (or Gaussian) distribution was first described by Carl Friedrich Gauss in 1809 1 in the context of measurement errors in astronomy. During the 19th century, this distribution was applied extensively in the developing area of applied probability and statistics. “Often decisions in chemometrics depend on one or more underlying populations falling into a normal distribution.” In the UK, the mean male height is 1.778 m, and the standard deviation 0.076 m. For a homogenous population, the distribution of heights should fall approximately into a normal distribution. In fact, underlying distributions are often not in themselves normal, but by the central limit theorem, the sum of symmetric independent distributions often approximates the normal distribution. In later articles, we will explore the idea of independence, but it should suffice in this article to consider such populations first. It is primarily because of the central limit theorem that the concept of a normal distribution gained great popularity in early statistics 2. We usually represent the normal distribution with the horizontal axis representing a measurement, such as men's heights, and the vertical axis representing a probability or frequency as in Figure 1(a). This type of representation is often called a probability density function (often abbreviated pdf). Sometimes, it can also be represented as a cumulative distribution function (cdf): this is a graph of the proportion of samples below a certain value against the value itself as in Figure 1(b). For a pdf, the maximum should be at the mean (in our case 1.778 m), whereas for a cdf, the mean represents the halfway point. This is illustrated in Figure 2. Measurements can consist of several different Gaussian distributions. In Figure 3, two partially overlapping normal distributions are illustrated. They may represent two groups of samples, for example, the length of adult mice from two subspecies. It is important to remember that the shape of both distributions illustrated is the same. The left-hand distribution is wider and corresponds on the whole to lower measurement values. But there is a region of overlap, so just by measuring the length, for example, of an adult mouse, we cannot perfectly distinguish between the two populations. In later articles, we will discuss how adding additional measurements may allow us to better distinguish samples from two groups even if there is overlap in the distributions for single measurements. Hence, a standard normal distribution has a mean of 0 and standard deviation of 1. Often, there is a further step of scaling the area under the distribution curve so it totals 1, so that the curve becomes a pdf rather than a frequency distribution. We will look more at tests of normality in other articles, but it is often constructive to visualise information in a simple graphical manner before moving forward. We can ask different types of questions; for example, if we measure the heights of 1000 men in the UK, how many do we expect will be more than 2 m high using the example of men's heights? The answer is 1.74; that is, typically, we would just encounter one or two people over this height (around 6 ft 7 in.). We can say that the underlying population is likely to be normally distributed, but the observed data will not fit this exactly. Unless all possible samples are recorded, we can only measure a fraction (usually a tiny one) of the underlying population; we try to use these to determine information about the population as a whole, but this depends on our sample being representative. Hence, it is possible to answer numerical questions about samples by assuming that measurements are normally distributed. In the example of men's heights, we can ask what proportion of men are likely to exceed 1.8 m in height? Or in a population of 30 million people, how many do we anticipate are between 1.6 and 1.7 m in height? Or what is the chance that a man is less than 1.4 m (4 ft 7 in.) in height? By assuming normality, this can be done just by determining the mean and standard deviation of a small but representative sample. The assumptions that a homogeneous population of samples from a single group or origin is normally distributed is often buried deeply within many chemometrics tests, and it is important to understand this as a fundamental basis of much statistical inference, as well as when the assumptions are not valid. There are many more detailed discussions of the normal distribution. The NIST Engineering and Statistics Handbook 3 is particularly recommended as an online source for further reading.
No takes yet. Share an insight, caveat, or question.
Richard G. Brereton (2014) studied this question.