Recent years have seen a vast increase in the amount of longitudinal data that is available to social scientists. In Britain alone, there are now four major cohort studies, which are based on births in the years 1946, 1958, 1970 and (around) 2000. The British Household Panel Survey (BHPS) began in 1991 and now includes 15 waves. Both the cohort studies and the BHPS involve repeated interviews with a defined sample, the BHPS conducting its interviews annually and the cohort studies less frequently. In the case of the cohort studies, the samples were selected on the basis of date of birth, whereas with the BHPS the sample was drawn on the basis of place of residence in 1991. Longitudinal data may be constructed in other ways, however. The UK Office for National Statistics’s Longitudinal Study can be viewed as a panel of around 1% of the population of England and Wales built up by linking census, vital registration and other data through tracing information about identifiable persons in originally independent sources. A similar approach is now being suggested to construct a historical panel data set for the second half of the 19th century (Schürer, 2007). The situation in Britain reflects that in many other countries: within Europe the pioneering panel study was the German Socio-economic Panel which began in 1984, and initiatives such as the European Community Household Panel have been working towards harmonizing the data from different surveys to facilitate comparative analysis. Panel data have well-known advantages over cross-sectional data. They permit the analysis of the dynamics of social and economic processes, and, because measurements of variables can be located in time, potentially allow the testing of causal models. However, their analysis also poses greater challenges. The first five papers in this issue of Statistics in Society take up these challenges in different ways. They illustrate a range of new approaches to the analysis of longitudinal data but also suggest that the differences between these approaches might be less fundamental than appears at first sight. The papers were originally presented at the symposium on the ‘Convergence of methods for the analysis of panel data’ that was held at the University of Southampton in April 2006, with the support of the Economic and Social Research Council’s ‘Research methods programme’ and the Southampton Statistical Sciences Research Institute. An inherent feature of panel data is that they contain repeated measures of the same variable for the same individuals. This, of course, creates a problem for conventional statistical models, because the existence of unobserved individual level characteristics means that observations of the same person in different waves are not independent. Following the development of multilevel models for hierarchically clustered data in the 1980s and 1990s it was recognized that panel data could be analysed within this framework, since we typically have several (level 1) observations at different time points relating to the same (level 2) person. However, panel data have the additional feature that they are dynamic, in that the observations on each individual are ordered in time, and therefore earlier observations may influence later ones. In the first of the five papers in this issue, Fiona Steele explains how a range of models which capture the dynamics of panel data, including growth curve models and event history models, can be handled within the multilevel modelling framework. She compares and contrasts the multilevel approach with the alternative structural equation model framework, showing that multilevel models have advantages in the presence of survey attrition or where some individuals ‘miss’ waves of the panel (both of which are very common features of longitudinal data). However, structural equation models are useful where the variable of interest is not directly observed in each wave, but instead data are recorded on a set of indicators which are assumed to give information about the variable of interest. Among the event history models, she looks specifically at multiprocess models, in which previous outcomes of one process can influence the timing of events in the other process. Her exposition is illustrated with examples, including one drawn from the British National Child Development Study (a longitudinal study of a cohort of children who were born in a single week in 1958). In his paper, Stephen Pudney uses the BHPS to analyse ‘subjective wellbeing’, or how well individuals feel they are managing financially. This question was asked in several waves of the BHPS in a form which generates an ordinal response. The usual approach to modelling such data is to treat the dependent variable as discrete and to include lagged values of the dependent variable in the model together with other covariates of interest—an approach which is often termed the ‘state dependence’ model. However, as Pudney points out, wellbeing is an inherently continuous concept; moreover the question in the BHPS does not observe subjective wellbeing directly, so, whereas we observe ordinal data, the underlying ‘true’ wellbeing variable is both unobserved and continuous. Pudney develops a latent auto-regression model which is capable of overcoming these challenges; he shows that it is superior to the state dependence model. His results show the expected positive effects of the market value of the individual’s house, employment and being married or cohabiting, and negative effects of unemployment, divorce and separation. A key difference between the latent auto-regression model and the state dependence model is that the former implies greater persistence in the effects on subjective wellbeing of shocks: the effect of changes in circumstances lasts longer. Individual traits such as subjective wellbeing can be measured by a range of methods. This is frequently so in panel data where several questions can be asked which each provide information about the same underlying trait. Each question can be regarded as a different method of measuring the trait in question. Where several traits are each measured with more than one method, the data are described as multitrait, multimethod, and there is an established lineage within the psychometric literature on modelling such data. The paper by Steffi Pohl, Rolf Steyer and Katrin Kraus extends this literature by developing a new model, called the method effect model, in which the effects of using different methods to measure a trait (in their example, mood states in a German panel) are modelled as latent difference scores in structural equation models. Because statisticians who are not acquainted with the psychometric literature may not be familiar with multitrait, multimethod models, the authors provide a brief history of their development. It is clear, however, that they are potentially useful in validating the measurement of latent or underlying constructs, such as the subjective wellbeing that is studied by Pudney, the attitudes to gender roles that are analysed by Berrington and her colleagues in this issue or the ‘happiness’ that was studied by Gardner and Oswald in a recent paper in this journal (Gardner and Oswald, 2006). It is surely of interest to be able to assess the extent and nature of systematic effects arising from measuring the same underlying construct with different methods. The question of latent, or indirectly observed, variables is also central to the paper by Patrick Sturgis and Louise Sullivan. Their substantive interest is in social mobility within Britain, which can in principle be analysed by using the 1970 British Cohort Study of a sample of children who were born in a single week in April 1970, since this contains data on the position of the cohort members in childhood, and their corresponding position in early adulthood. As in Pudney’s paper, the model that they propose is based on an observed ordered outcome (social class as measured by the Registrar General’s schema) which is assumed to be related to an underlying continuous variable. They apply a latent growth curve model to analyse the latter, but then hypothesize that there are distinct groups within British society which are characterized by different typical patterns of social mobility, or different trajectories through the social hierarchy. Such groups can be identified by using latent class growth analysis, in which a range of background factors (including some attempting to measure individual ‘merit’ and others reflecting individuals’ endowment of social and economic capital) is used to predict membership of these latent trajectory groups. The results identify five distinct groups, and Sturgis and Sullivan show that social and economic capital and their chosen merit variables both influence the particular group into which a person is most likely to fall. Panel data provide opportunities to test competing hypotheses about the relationships between attitudes and behaviour. One example, which is studied by Ann Berrington, Yongjian Hu, Peter Smith and Patrick Sturgis in their contribution, concerns the effect of women’s participation in the labour force on gender role attitudes. In essence, the question is the extent to which women who participate in the labour force are ‘selected’ for positive attitudes towards working outside the home, and the extent to which attitudes can change by ‘adaptation’ as a result of experiences of paid work. Attitude data come from four waves of the BHPS, and the authors use a graphical chain model in which attitude in any wave is modelled as a function of attitudes at the previous waves, experience of parenthood and employment since the previous wave (the adaptation effect) and a set of fixed background characteristics. At the same time, change in parenthood and employment status between any two waves is modelled as a function of attitude at the initial wave and the same set of background characteristics: this captures the selection effect. The authors find that the adaptation effect is more consistent than the selection effect, though, as they remark, this might be a consequence of the particular set of women whom they study, who were predominantly young and all childless in 1991. Their approach has potential to shed light on other contexts where questions about the relative strength of selection and adaptation effects arise, e.g. the relationship between marriage and mortality. The contributions in this issue reflect a variety of approaches to the analysis of panel data. However, several of the contributions in this issue discuss the differences between alternative analytical frameworks, e.g. structural equation models, graphical chain models, the use of lagged dependent variables (which has a long tradition in econometrics) or the multilevel modelling approach. Their conclusion, overall, is that the fundamental methodological differences between the approaches are smaller than is commonly supposed, and that the choice of analytical framework is best determined by the precise questions of interest, whether the variables of interest are observed directly or indirectly and the nature of the hypotheses being tested We cannot in practice estimate completely unrestricted models which allow the simultaneous exploitation of all the richness of panel data: what we do have, as demonstrated by the papers in this issue, is a growing range of options, each of which is appropriate for the analysis of particular questions.
No takes yet. Share an insight, caveat, or question.
Andrew Hinde (2008) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: