[1] Statistics is a powerful tool for hydroclimatology. However, misuse of a statistical technique, often caused by ignoring unavoidable underlying assumptions that must be made to use the technique, can render the analysis meaningless and could also result in wrong conclusions. The Mann-Kendall (MK) test [Mann, 1945; Kendall, 1975] for trend has frequently been used to detect trend in hydroclimatological time series because it requires fewer assumptions than some of its parametric counterparts. However, this test does require the assumption that the observations that are to be analyzed are the realizations of a collection of independent random variables. If the test is applied to serially correlated data (which is often the case for hydroclimatological time series), the trend detection results may not be reliable because the test may then reject the null hypothesis of no trend (H0) more often than specified by the significance level [von Storch, 1995]. [2] One way to avoid this problem is to apply the test only on data that have been made serially independent [Lettenmaier et al., 1994; Gan, 1998; von Storch, 1995]. A simple remedy exists if there is a sufficient physical reason to assume that the time series to be tested is a sum of a deterministic trend and serially correlated noise generated by an AR(1) (or higher order) process. In this case, the test may be conducted after prewhitening the time series [von Storch, 1995]. The prewhitening procedure has been used in many applications for detecting trends [e.g., Douglas et al., 2000; Zhang et al., 2001; Wang and Swail, 2001]. [3] Yue and Wang [2002] (referred to as YW hereinafter) question the validity of the prewhitening serially correlated data before applying the MK test. They argue that prewhitening reduces the magnitude of any trend that may be present and that it therefore reduces the power, or sensitivity, of the test. They suggest that “when sample size and magnitude of trend are large enough, serial correlation does not significantly influence the MK test,” and in such case, “it is better to use the MK test on the original data rather than after prewhitening”. Unfortunately, this advice could lead unsuspecting users to make serious errors in the interpretation of their data because it requires the user to first judge visually whether trend is present in the data that is to be analyzed. Such subjective assessments cannot be performed reliably when data are serially correlated. For example, the upper panel of Figure 1 displays a sample of length 80 from a pure red noise process with lag 1 correlation coefficient ρ1 = 0.4. Trend is visually apparent, even though the red noise process has no trend. Once a trend of −0.02 per time step is added to the same time series, no trend is visually apparent (bottom panel of Figure 1), even though the red noise process now has a trend. Moreover, the overall performance of the combined procedure, consisting of subjective trend assessment followed by objective estimation of the trend magnitude, cannot be assessed. Thus this advice leaves the user with a trend analysis procedure with unknown characteristics. Good science should always use analytical tools that have known performance characteristics. Also, we note that this advice contracts advice given by Yue et al. [2002]: they recommend the removal of autocorrelation before assessing the statistical significance of a trend in that article. [4] In this note, we correct some misconceptions of YW, and provide methods for determining the magnitude and the statistical significance of linear trend in serially correlated data. The performance of the different methods of trend detection that we consider are assessed using Monte-Carlo simulation. We describe the methods for trend detection, and the approach used to assess these methods in the following section. Results are presented in section 3, and we conclude with comments on YW in section 4. [6] Equation (1) represents a classical linear regression problem: regression with autocorrelated errors. Standard procedures for solving this problem have been well documented in books on time series analysis [e.g., Box et al., 1994], and have been built into modern statistical packages such as SAS and S-Plus. We fit equation (1) using the S-Plus routine gls after setting options for autocorrelated AR(1) noise and to fit the trend plus noise model with the method of maximum likelihood [Venables and Ripley, 2000]. We refer this procedure as ML1. [7] As in YW, we also apply Sen's [1968] slope estimator to the simulated time series, before and after prewhitening the series, to compute the magnitude and the statistical significance of a trend. Results from the unprewhitened and prewhitened series are referred to as MK0 and MKP, respectively. It is known that trend computed from prewhitened data is smaller, since the trend in the prewhitened data has slope βw = (1 − ρ1) β [Wang and Swail, 2001]. This is a major criticism that YW have on computing trend from prewhitened data. An iterative procedure proposed by Zhang et al. [2000] and refined by Wang and Swail [2001] avoids this problem and gives nearly unbiased estimates of both β and ρ. This procedure is as follows. [8] 1. First, ρ1 is estimated from the time series and if that estimate is less than or equal to 0.05, the trend is computed from the time series directly using Sen's [1968] estimator. [9] 2. If > 0.05, the time series is prewhitened and a trend βw is estimated from the prewhitened series using Sen's method. The trend in the original data is then estimated as = w/(1 − ). [10] 3. The estimated linear trend is then removed from the original data and the detrended data is used to obtain a new estimate of ρ1. [11] 4. Repeat steps 1–3 until the differences for both parameters obtained in two iterations are close enough. [12] Details of this procedure can be found in Appendix A of Wang and Swail [2001]. It has been shown [Fleming and Clarke, 2002] that this procedure provides better estimates of trend than the prewhitening procedure of von Storch [1995]. We shall refer the results obtained with this procedure as MK1. [13] One may argue that the combination of a parametric method (to estimate ρ) and a nonparametric method (to estimate β) may be not as sound theoretically as a solely parametric method such as the method of maximum likelihood. Therefore, in addition to ML1, we also replace Sen's slope estimator with an ordinary least squares slope estimator in the above iterative procedure. This method is referred to as LS1. [14] The null hypothesis H0 that is tested by MK0 is the same as that of Mann-Kendall test; the test evaluates evidence counter to the statement: the data are a sample of n independent and identically distributed (iid) random variables. Similarly, the null hypothesis H0 tested by MK1 and MKP is that the prewhitened data are a sample of n iid random variables. In all three cases, H0 is false when β ≠ 0. In contrast, the null hypothesis H0 for LS1 and ML1 is that β = 0. Thus LS1 and ML1 test a more explicit statement about the presence or absence of a linear trend than MK0, MKP, and MK1. [15] We use three statistics to assess the performance of the different methods. Specifically, we compute the bias and root mean square errors (RMSE) of the estimates of β, defined as β − (∑)/m and , respectively, with m being the number of simulations. In addition, we also count the number of times in 500 simulations when H0 is rejected. This quantity, which will be expressed as a percentage, is referred to as the rejection rate. [16] Figure 2 shows the biases in when estimated with the different methods. In general, appears to be positively biased when ρ1 > 0.0 except for MKP. Biases increase with lag 1 autocorrelation coefficient when the sample size is small. They decrease with an increase in the sample size and become negligible when n = 200, even if the lag 1 autocorrelation is large for the same sample size. MK1 and LS1 estimates are much less biased compared with the MK0 and ML1 estimates. Results for MKP are similar to those of YW and are predictably negatively biased. [17] The RMSE of becomes larger when sample size is reduced because the variance of increases under these circumstances. The RMSE also increases with an increase in lag 1 autocorrelation because this reduces the effective (independent) sample size. Figure 3 displays RMSE for a sample size of 50 years. MK1 and MKP have largest RMSE, while ML1 usually has smaller RMSE due to much smaller variance of . [18] The rejection rates are plotted in Figure 4. When both the prescribed lag 1 autocorrelation and the trend are zero, MK0, MK1, and LS1 correctly reject H0 at the nominal level of 5% for all sample sizes examined; ML1 detects significant trend for more than 10% for sample sizes of 30 and 50 but performs at the nominal level of 5% when sample sizes are large (e.g., n ≥ 80). For large samples, ML1 also consistently outperforms other methods, as expected from asymptotic theory [e.g., Cox and Hinkley, 1974]. This is further confirmed by simulations with a sample size of 1000 (Figure 4). It should be noted however, that when the sample size is small, ML1 becomes too “liberal” and MK1 should be used. [19] When the noise is autocorrelated, MK0 detects trends much more often than the nominal level even if there is no trend. This is consistent with von Storch [1995] and with YW, indicating the need to account for the effects of serial correlation. For example, when the lag 1 autocorrelation reaches 0.8, MK0 “detects” significant trends about 50% of the time when no trend is present. When serial correlation is accounted for in the detection problem, the probability that MK1, MKP, LS1, and ML1 detect a significant trend when none is present is much closer to the nominal significance level, with smaller autocorrelation and larger sample sizes leading to better agreement between the actual and nominal (i.e., chosen) significance levels. Even so, for MK1, LS1, and ML1 the rejection rate is still much higher than the nominal level for most sample sizes considered (e.g., n ≤ 200). This indicates that for a typical climate data set that has a sample size less than 100, there is not enough information in the data for reliable trend detection with this method when the lag 1 autocorrelation is large. For MKP, the rate of rejection is never higher than the nominal level when sample sizes are small, due to the removal of autocorrelation. [20] We have compared the performance of five different methods in estimating the magnitude and the statistical significance of trend in a time series consisting a possible deterministic trend and red noise generated by an AR(1) process. Our results indicate, in agreement with theory, that the parametric maximum likelihood method is asymptotically optimal. However, its advantage diminishes when sample size is small to moderately large (say n < 80) and MK1 seems to be a better choice in those cases. Our results also confirm that the MK test and Sen's slope estimator are not reliable when applied on time series with serially correlated residuals. [21] YW conclude that “when sample size and magnitude of trend are large enough, serial correlation does not significantly influence the MK test,” and therefore recommend that “in such a case, it is better to use the MK test on the original data rather than after prewhitening”. We reiterate that this is poor advice (and that it is in contradiction with advice given by Yue et al. [2002]). [22] Investigators do not, a priori, know whether trend is present and cannot reliably distinguish between trend and persistence from serial correlation by subjective visual means, particularly in short data sets. However, investigators are often aware a priori that their data are affected by serial correlation. It is therefore simply irresponsible to recommend the use of tests that falsely “detect” trends when none are present much more frequently than the nominal false detection rate (the significance level) that is specified by the investigator. This undermines our entire basis for the statistical evaluation scientific evidence, which operates by determining whether a statistic, such as trend estimate, is unusual in the context of the null hypothesis, and a given set of assumptions that describe the statistical behavior of the observed system when the null hypothesis is true. In the case of hydroclimatology data, it is reasonable to expect that the independence assumption that is required by the MK test will be violated. Given the knowledge that our measures of what is unusual become unreliable when some basic aspect of the set of assumptions (such as the independence observations) is not satisfied, and that such departures from the assumptions are likely, it is incumbent upon us to use techniques that remain reliable when we anticipate that a key assumption will be violated. That means using techniques that produce false positives at an acceptable rate over a range of likely circumstances. It is also desirable to have techniques that are powerful (e.g., that detect trend frequently when it is present), but power is a secondary consideration that only comes to play after we are satisfied that a chosen technique will not produce an unexpectedly large number of false positives. [23] Concerns about biases in trend coefficients estimates that are raised by YW should not deter users from taking serial correlation into account when making statistical inferences about the present of trend. In fact, appropriately taking serial correlation into account also produces more accurate trend estimates. YW hint that a “true” trend might be estimated from serially correlated data without considering autocorrelation. Our results indicate that the estimated trend is positively biased when there is a positive autocorrelation in the series. The bias is much smaller when serial correlation is taken into account at an expense of larger RMSE. YW object to the use of prewhitening procedures because the trend computed from the prewhitened series tends to be smaller than the “true” trend. This can easily be corrected by using an iterative procedure [Wang and Swail, 2001] or, in the case of large samples, with the method of maximum likelihood. It should be noted that there is a clear distinction between a deterministic trend that has a physical basis, and an apparent trend that only reflects low frequency variability such as red noise: the former is likely to continue while the latter may change sign at any time in the future. The cause of a trend is usually hard to know and cannot be revealed by any of the methods considered in this study. However, providing users with carelessly computed trend estimates, and incorrect estimates of their statistical significance, will likely cause confusion and may well have serious financial and safety impacts. [24] We thank V. V. Kharin and Lucie Vincent for their comments on an earlier draft of the manuscript.
No takes yet. Share an insight, caveat, or question.
Zhang et al. (2004) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: