The important role of trained operators in the analysis and interpretation of flow cytometric data was first made clear to me when I was a PhD student. An esteemed colleague commented on the “nice light scattering peak” that was being acquired on the screen of the computer controlling the Skatron Argus 100 flow cytometer. I did not have the heart to tell him that I had left the computer acquiring data, when I had been cleaning the cover slip at the measuring point with a piece of lens tissue! This illustrated amply how easily noise can be interpreted as something of biological relevance. Over the intervening years, there have been many advances that have led to progress in the field of flow cytometry. Three parameter, single color instruments, such as, the Skatron have been replaced by multicolor instruments. Eighteen-color analysis is now available and data may thus be collected on many different parameters of each cell. An instrument capable of collecting up to 62 parameters has recently been launched (http://www.isac-net.org/index.php?option=com_content&task=view&id=795&Itemid=45), which may be expected to produce data sets with unprecedented dimensionality. Additionally procedures have been proposed that allow multiple analyses to be combined creating data files of potentially infinite dimensions (1). These approaches have found particular utility in imunophenotyping studies (2), where a range of markers exist for distinguishing between cell types but are also applicable to multiplexed bead-based assays for the detection of microorganisms (3). Other important advances have included an ever-increasing arsenal of fluorescent probes for biomolecules and determinants of physiological capability, such as, enzyme activity. Developments, such as, monoclonal antibodies and fluorescently labeled oligonucleotides that have impacted onother fields have also played a part in advancing flow cytometry. Gene reporters including, of course, green fluorescent protein and its derivatives have played an important part in extending studies of gene expression to include flow cytometric assays (4), and thus have allowed information to be obtained at the level of the single cell (5). All of these developments add to our capacity to acquire multiparametric data sets. Advances in instrumentation and fluorescent/fluorogenic reagents have been augmented by advances in electronics and computing. These have permitted analysis to be performed at higher speeds, resulting in larger data sets and the ability to detect and analyze rare subpopulations. However, the acquisition of higher-dimension data sets containing larger numbers of events is not all good news. With small number of parameters, it is possible to plot them on a single 2-D or 3-D graph, but when more than three parameters are recorded, it is impossible to visualize the raw data in such a way that all of the inter-relationships can be seen. The closest approximation is to produce a series of 2-D or 3-D dotplots [an everything vs. everything analysis (6)], but this approach has two disadvantages: first, there can be many plots to examine (the number of dual parameter plots is given by the formula where n is the number of measured parameters) and second, the true multidimensional relationships cannot be seen. Thus, for multidimensional data sets, this approach rapidly becomes impractical. There are many methods to explore such data sets and the last 20 years has witnessed the application of a variety of these to flow cytometric data. However, unlike the advances in instrumentation and fluorescent probe technology, these have generally had little impact beyond the initial publication and have not been adopted by the flow cytometric community. This article seeks to explain the reason for this and to suggest a way forward. The aim of many multivariate data analysis methods is dimensionality reduction, whereby the important differences in the data are compressed into fewer dimensions allowing the relationship between objects (in this case, analyzed particles) to be visualized. One of the best known techniques is principal components analysis (PCA), which compresses the variance of a flow cytometric data set from the familiar parameters of forward/side scatter and fluorescence at different wavelengths to a smaller number of principal components (PCs). These PCs can then be plotted against one another permitting representation of clusters in 2-D space that were previously only apparent in multidimensional space. PCA reduces the dimensionality of data sets by transforming possibly correlated measurements into a smaller number of uncorrelated PCs. This approach should be beneficial for flow cytometric data as there can be a high degree of correlation (7); for example, exponentially growing yeast cells that are about to divide will typically have increased FITC (protein) fluorescence and a higher DNA content when compared with newly divided cells. Furthermore, even with band pass filters fluorescence from a single stain may be detected by more than one detector due to wide emission peaks of the fluorophores. However, methods, such as, PCA are referred to as “unsupervised” learning methods, because they group individuals on the basis of the overall variance in the measured parameters without any attempt to incorporate prior knowledge. This may lead to clusters that are in fact more difficult to interpret than the raw data (8). In contrast, supervised learning methods take advantage of the fact that the flow cytometer operator usually has some prior knowledge of the identity (e.g., species, physiological status, cell type, etc.) of samples through analysis of controls. In this approach, a model is formed via a training process in which data patterns (channel number measured by each of the relevant detectors) plus answers (the identity of the type of particle in a control sample giving rise to that pattern) are presented in turn to the model-forming software. Through an iterative process, a suitable model is formed and this can then be used to predict the identity of samples of unknown identity. Thus, for example, one might analyze axenic cultures of a target organism as a positive control and other organisms as negative controls, and then analyze mixed samples spiked with varying concentrations of the target to determine whether the model can correctly identify the target. Supervised methods include artificial neural networks (ANNs), which are often regarded as a “black box” as the way in which the model performs its identifications can be difficult to interpret. This approach has been used with flow cytometric data many times (8-10) over the last 20 years with very effective models being produced, particularly for the identification of bacterial, fungal, and phytoplankton samples. An alternative approach is to use genetic programming (GP) to form the model. GP is a methodology based upon biological evolutionary concepts in which computer programs are evolved that are able to perform a user-defined task. The approach optimizes a population of automatically generated computer programs according to their “fitness” at performing the task in question. This approach was found to be suitable for flow cytometric detection of sporangia of the fungal pathogen Phytophthora infestans against a background of other airborne biological particulates (11). A significant drawback with these methods is that they are not embedded within flow cytometric software data acquisition packages and, therefore, cannot be used for real-time visualization and interpretation of the data. Additionally, they are methods with which the majority of the flow cytometric community are not well-versed and the software packages are unfamiliar to them. Consequently, although a number of excellent papers have been published describing successful application of the methods to a particular data set, the models or methods do not appear to have been used beyond the proof of principle studies. Many flow cytometric software packages permit the calculation and plotting of derived parameters in real-time during data acquisition, for example, the ratio of two “standard” acquired parameters has been used to estimate membrane potential in bacteria (12) and, more recently, to detect calcium signaling (13). Apart from familiarity and ease of use, there are substantial advantages to using data analysis methods that are embedded within acquisition software rather than using those that rely upon extensive off-line analysis. For example, when monitoring for a specific microbe a real-time method (where data are interpreted during acquisition) would permit reanalysis of a second aliquot of the sample to confirm the results or the performance of more specific, but more costly, tests (e.g., using fluorescently labeled antibodies). Also, in suitably equipped instruments, detection and identification during analysis would permit sorting of the target organism onto a microscope slide or (providing the samples have not been fixed) into growth medium for confirmatory analysis. Sorting for other analyses, such as PCR or even next generation sequencing would also provide confirmatory results with positive samples. Real-time interpretation can be performed using the derived parameters available in many standard acquisition packages. For example, using a ratio of signal from a fluorescent marker to the forward (or side) scatter signal from the same particle has the advantage that it reduces the effect of cell-to-cell heterogeneity, which results from the tendency for larger cells (or clumps of more than one cell) to fluoresce more brightly. Thus, the information from two parameters can essentially be encoded in a single value, the ratio, allowing their display on a 1-D plot or in combination with another parameter (or derived parameter) on a 2-D plot. For example, Figure 1A shows the range of propidium iodide (PI) fluorescence values measured from four different microorganisms acquired using previously reported protocols (8). There was considerable overlap between the organisms and a large degree of intra-species heterogeneity. When the ratio of PI to forward scatter (FS) was calculated (Fig. 1B), this provided much better resolution between the target (in this case Bacillus globigii) and the other organisms than was achieved from PI fluorescence alone. As shown in Figure 1C, a single parameter histogram could be plotted and an appropriate gate could be set on the subpopulation with the lower PI:FS ratio, which largely represents the B. globigii spores. Bacillus globigii spores and vegetative cells of Escherichia coli, Micrococcus luteus, and Saccharomyces cerevisiae were permeabilized with 70% ethanol, washed, and stained with 40 μg/mL, propidium iodide (Sigma, U.K.) for 30 min (8) before analysis on a Coulter Epics Elite. A: Broad heterogeneity of PI fluorescence channel number was observed. B: A ratio of PI:FS reduced the heterogeneity and introduced the potential to discriminate between the different microorganisms. C: The calculated parameter PI:FS can be displayed in an equivalent manner to single, measured parameters permitting gates to be drawn to analyze (or sort) B. globigii spores. [Color figure can be viewed in the online issue, which is available at www.interscience.wiley.com.] In conclusion, until and unless mainstream manufacturers provide an integrated solution to the acquisition and automated interpretation of flow cytometric, the application of techniques, such as, ANNs will remain rare, as they have been for the last 20 years (9). As such software is unlikely to become commercially available in the short term the flow cytometry community may benefit from investigating further the tools, such as, ratio calculation, that are provided in data acquisition software. Of the automated data analysis methods described earlier, the results of GP analysis are perhaps most straightforward to integrate into flow cytometry data acquisition software. At present, the evolution of a suitable GP would still have to take place externally but because the programs that are evolved are simple, interpretable mathematical rules, many of these could be represented using the mathematical operators already available for calculation of derived parameters in flow cytometric data acquisition software. This represents an important advantage of GPs over the “black box” approach of ANNs. With increased dimensionality of data sets and increasing sample throughput, it seems likely that data analysis methods will have an important role to play in the further development of flow cytometry. Hazel M. Davey*, * Institute of Biological, Environmental and Rural Sciences, Aberystwyth University, Penglais Campus, Aberystwyth, Wales SY23 3DD, United Kingdom.
No takes yet. Share an insight, caveat, or question.
Hazel M. Davey (2009) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: