Today the practising taxonomist may acknowledge a wealth of sophisticated methods, software, etc., produced to facilitate and optimise the reconstruction of phylogenies, and a large number of new studies are currently being published.However, one fundamental and still unresolved problem relates to the issue of how to code simple, "qualitative" data into a form suitable for phylogeny reconstruction.Character coding represents the link between observation and analysis and greatly influences the results, but has nevertheless received little attention.Whereas Stevens (1991) discussed qualitative versus quantitative data and delineation problems of different character states within a character, I here focus on delineation problems in the relation between characters and character states.In a more general treatment on character and character states, Pimentel and Riggins (1987) evaluated different aspects of coding procedures, with one of their conclusions being that characters should be coded as multistate variables rather then being treated independently as present or absent.Likewise arguing for multistate coding, Meier (1994) rejected absence/presence coding on the basis that it may result in selection of less parsimonious or "pseudoparsimonious" trees (see below).Other aspects of character coding which have received more attention (although no consensus) are problems related to different aspects of missing entries (Nixon and Davies, 1991;Platnick et al., 1991;Maddison, 1993), and assumptions, testability and information content in ordered versus unordered characters (Hauser and Presch, 1991;Wilkinson, 1992;Hauser, 1992;Slowinski, 1993).The need for resolution of these issues is apparent from a large number of case studies, where different coding techniques are employed without discussion and applied in mixed ways within a single analysis.The aim of this paper is to consider some aspects of transformation of character observations into a data matrix suitable for phylogenetic analysis, to formulate some commonly encountered problems, and to argue for some advantages in treating characters simply as binary present or absent.Hopefully, these issues in the future will receive more attention and be exposed to critical evaluation.For simplicity, characters are defined as columns in data matrix in which taxa constitute rows, and character states the different values assigned to the entries.With multistate characters only unordered, or nonadditive, character transformations are considered.Consider a group of organisms that may or may not have a certain feature X, and that this feature occurs in two shapes, rounded and square, and with two
No takes yet. Share an insight, caveat, or question.
Fredrik Pleijel (1995) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: