Criticisms of the uses of the no-observed-effect concentration (NOEC) and the lowest-observed-effect concentration (LOEC) and more generally the entire null hypothesis statistical testing scheme are hardly new or unique to the field of ecotoxicology 1-4. Among the criticisms of NOECs and LOECs is that statistically similar LOECs (in terms of p value) can represent drastically different levels of effect. For instance, my colleagues and I found that a battery of chronic toxicity tests with different species and endpoints yielded LOECs with minimum detectable differences ranging from 3% to 48% reductions from controls 5. For interpretations of field studies, recommendations for improved practices include using confidence intervals rather than hypothesis testing for group comparisons and evaluating whether apparent effects exceed predetermined "critical" effect sizes 2, 6. For interpretations of toxicity tests, recommendations for improved practices emphasize replacing NOEC and LOEC comparisons with either model-based true no-effect concentration estimates or curve fitting and from the fitted curve functions, reporting concentrations that produced x% of effects (ECx) 7-9. These developments beg the question: What levels of toxic effect are of concern or can be considered negligible? Biologically based arguments for selecting x are scarce, and instead discussions for selecting x from curve-fitting approaches have emphasized statistical or test performance considerations for selecting x rather than biological implications. Confidence limits, variability of point estimates, model dependence, and comparisons of NOECs to ECx percentages have been evaluated 10, 11. These statistical considerations are of value but are not sufficient by themselves and may be circular. If a major shortcoming of NOECs is that they may actually correspond with fairly high adverse effects 9, 12, why should it make sense to then turn about and ask what levels of effects are typically associated with NOECs to define the x in ECx values? In contrast, the reasons for the existence of toxicity testing practices relate to making some estimate of safe or unsafe concentrations for aquatic populations or communities. Thus, judgments of what constitutes a negligible level of effect in toxicity tests should consider the consequences of similar effect levels in the wild. For example, the primary adverse effects studied in an early–life stage toxicity test with fish are reduced growth and survival. From a population biology perspective, survival and reproduction are the only endpoints that directly matter for viability. Reduced growth may indirectly influence reproduction by slowing the time to reproduction or because smaller females may produce fewer offspring. However, in the wild, subtle differences in size could have disproportionately large effects on survival. For fish, growth is closely linked to survival, in part because of the importance of size in competitive interactions and predator–prey relations. For example, adult sculpin (Cottidae) may prey on juvenile salmonids in streams and vice versa, depending on relative sizes. Torrent sculpin, Cottus rhotheus (Figure 3), that only had a 15% length advantage could readily ambush, subdue, and eat coho salmon, Oncorhynchus kisutch, yet when the sculpin and salmon were similar in length, the sculpin posed no threat 13. With fish in temperate streams, contests for territory may determine profitable feeding locations, shelter from predation, and winter resting shelters. These in turn may determine whether fish have sufficient energy reserves for overwinter survival and eventual reproduction 14. A size disparity of as little as 5% in body weight may tip the outcome of such contests 15. Riverine survival of juvenile Chinook salmon (Oncorhynchus tshawytscha) in Idaho, USA, was disproportionately size-dependent, with a 10% difference in length associated with 33% to 70% reductions in migratory survival 16. Similarly, with aquatic invertebrates, the same ECx values for different effect endpoints cannot be assumed to have the same level of effect. For instance, a 10% reduction in length of mussels would predict approximately a 19% to 44% reduction in fecundity, based on length–fecundity regressions from field studies with different freshwater mussel species 17. In 28-d exposures with freshwater mussels and Cu, the maximum reductions in length in treatments in which at least some mussels survived to the end of the tests were only 13% to 29% 17. Uncritical reliance on a single, fixed ECx value such as the EC20 for a growth endpoint for which the maximum range of response may not even reach 20% would lead to reporting test results as "greater than" values, which would incorrectly discount biologically important effects as being insensitive. Relating mortality rates of aquatic organisms in toxicity tests to corresponding effects in real populations from contaminant-induced mortality is difficult. Early–life stage mortality does not directly translate to proportional reductions in recruitment, largely because the survival of juvenile cohorts in nature is often density-dependent. That is, when densities are high, food and space become limiting, growth is stunted, and survival is low. When densities are low, food and space are abundant, growth is higher, and survival is higher 18. In natural populations, different life stages have very different contributions to the population dynamics and, for example, the same percentage of mortality to young-of-year fish or sexually mature fish will have very different population implications. For instance, a bull trout, Salvelinus confluentus, population could withstand up to a 60% annual decrease in juvenile survival yet went into decline following a 5% increase in annual adult mortalities 19. A closely related problem is what level x effect is most appropriate for use with the ECx when the policy goal is to allow no adverse effects. An oxymoronic response has often been to equate some low level of adverse effect to no effect. The EC10 has been used in this manner 7, 9, 10, although in a more extreme interpretation, a 25% toxic response from controls was considered to represent nontoxicity 20. Although this may be logical if the policy goals are intended to allow some toxicity, it seems to me that if the goal is no toxicity, the only value for x in ECx that represents no toxicity is 0 (i.e., EC0). Although calculating an EC0 estimate from distribution- or regression-based curve-fitting models is an impossibility when using a distribution with infinite tails such as the normal (gaussian) distribution, an EC0 can easily be estimated from curve-fitting routines that use finite distributions 21. The triangular distribution can produce a threshold sigmoidal toxicity curve fit similar in shape to that from nonlinear logistic regression, and the rectangular distribution can produce a "broken stick" piecewise-linear fit. Whereas the logistic equation curve subtly angles downward from its start, illustrating the impossibility of an EC0, the threshold sigmoidal and piecewise-linear curves are flat until the no-effect thresholds (EC0) are reached (Figure 4). The piecewise-linear fit has the advantage of making visual explanation of the no-effect EC0 to laypeople or policy people because the break in the regression is the no-effect threshold, and its reasonableness can be interpreted relative to the underlying data points. These visual explanations may be easier than trying to explain that something (e.g., 10% effect) equates to nothing. If an objective of ecotoxicological testing and modeling is to estimate thresholds for the absence of effects, ecotoxicologists should not discount the EC0. Growth data that have shallow slopes and limited ranges of response are less than ideal for nonlinear regression models, and I tend to place little confidence in confidence limits. For instance, in the piecewise-linear EC0 for mussel growth in Figure 4, the EC0 estimate appears eminently reasonable to me, breaking at a treatment with nearly identical responses as the controls, yet the calculated confidence limits encompass both the next lower (control) and higher treatments. Rather, my confidence in the ECx estimates is based on how well the models fit the underlying data. Model selection can matter, especially when interpolation is needed because effects occurred at the lowest concentration tested. In the mussel shell length example (Figure 4A), the EC10 estimate from logistic regression was lower than the EC0 estimate from piecewise-linear regression. Finally, replicated exposure test designs such as those in Figure 4B are a legacy of null hypothesis significance testing and are inefficient for curve fitting and ECx point estimates. A gradient of 24 unreplicated exposures would be more capable of defining the thresholds and distributions of concentration responses than could 6 exposures replicated 4 times each. Acknowledging that proportional diluters or numbers of pump channels place practical limits on numbers of exposures, the point is that moving from null hypothesis significance testing to an ECx interpretation approach involves more of a mind-set change than simply also running the results of a null hypothesis significance testing–based test design through regression-fitting software. The main point of this Response is that a given ECx effect size percentage might have very different biological implications depending on the endpoint and ecological context. For instance, a 5% reduction in an endpoint with low inherent variability such as length-at-age could have comparable population-level implications to a 20% reduction in more variable and ecologically compensable endpoints such as early–life stage survival or fecundity. An underlying theme is that we should not just reduce species and biology to numbers and models and disconnect the "eco" from "toxicology." Considering the biological implications of toxicity testing in the context of species life histories and their ecological context is important, even though such considerations may be qualitative, uncertain, or even speculative. Christopher A. Mebane US Geological Survey, Boise, ID
No takes yet. Share an insight, caveat, or question.
Christopher A. Mebane (2015) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: