Science depends on good data. Data are central to our understanding of the natural world, yet most data in ecology and evolution are lost to science—except perhaps in summary form—very quickly after it is collected. Once the results of a study are published (if ever), the data on which those results are based are often stored unreliably, subject to loss by hard drive failure and (even more likely) by the researcher forgetting the specific details required to use the data (Michener et al. 1997). Moreover, most data are never available to the broader community, even after publication of the results; in most cases this unavailability is permanent due to the eventual death of the researchers involved. We are losing nearly all of this important legacy. Yet these data, even after the main results for which they were collected are published, are invaluable to science, for meta-analysis, new uses, and quality control. With the increasing use of meta-analysis to summarize multiple studies, it has become clear that necessary summary statistics are often not published. In many cases, the study can only be used if the original data are available to the meta-analysts. Furthermore, data often can be used in ways beyond the questions that sparked its collection; for example, many studies contain information that can serve later as a baseline for detecting population trends, even decades later. The availability of data for published studies also allows error-checking, making science more open, and letting us more rapidly reach accurate conclusions. Finally, papers that have had data archived are more useful to—and more cited by—other scientists. A study of papers that report microarray data found that papers that archived their data were cited 69% more often than papers that did not archive (Piwowar et al. 2007). Data that are properly archived are saved for posterity, and the archives also function to preserve data in a useable form for the original authors. Moreover, if datasets are put into a readily interpretable format while the methods and structure of the data are foremost in the scientists’ minds, that data can be used later more easily by those scientists and others. The example of GenBank shows the value of the availability of data for all of these reasons. The modern synthetic use of DNA sequence data would not be possible without the near universal use of GenBank as a public archive. Moreover, GenBank would not be nearly as complete as it is without the communal decision to archive all DNA sequence data, a decision initially introduced by journals. To promote the preservation and fuller use of data, Evolution and other key journals in evolution and ecology will soon introduce a new data-archiving policy. This policy will state: Evolution requires, as a condition for publication, that data used in the paper should be archived in an appropriate public archive, such as GenBank, TreeBASE, Dryad, the NCEAS Data Repository or as supplementary online material associated with the paper published in Evolution. The data should be given with sufficient details that, together with the contents of the paper, it allows each result in the published paper to be recreated. Authors may elect to have the data publicly available at time of publication, or, if the technology of the archive allows, may opt to embargo access to the data for a period up to a year after publication. Exceptions may be granted at the discretion of the editor, especially for sensitive information such as the location of endangered species. This policy will be introduced in approximately a year, after a period during which authors are encouraged to voluntarily place their data in a public archive. Data that have an established standard repository, such as DNA sequences, should continue to be archived in the appropriate repository, such as GenBank. For other, more idiosyncratic data, the data can be placed in a more flexible database such as the NSF-sponsored Dryad archive at datadryad.org. When fully in place, the policy will require authors to archive the data required to support the conclusions in their published paper, along with sufficient details that a third party can reasonably interpret those data correctly. In most cases, this will require a short additional text document with details specifying the meaning of each column in the dataset. The preparation of such shareable datasets will be easiest if these files are prepared as part of the data analysis phase of the preparation of the paper, rather than after acceptance of a manuscript. The data-archiving policy is designed to address several concerns that some researchers may have about data sharing. To protect the ability of individual researchers to use the data that they have collected, the policy allows an embargo period after publication. Although the data will be entered into an archive at the time of publication, if the technology of the archive allows, the data may be restricted from public view for up to a year. This allows the original researcher time to publish other papers based on the dataset. The policy also allows longer embargo periods at the discretion of the editor in exceptional cases. In addition, the requirement is only for data that have already been used in the publication in question; other data from the same research project that have not yet been used in the published report need not be archived. Finally, data that are particularly sensitive, such as location information for endangered species subject to poaching, or human data subject to privacy concerns, should not be archived in a publicly accessible format. Throughout the history of ecology and evolution, enormous quantities of valuable data have been lost to future science, for a variety of technical and cultural reasons. There are no longer any meaningful technical barriers to long-term storage, and just as in the case of DNA sequence data, it is time for the culture of our shared use of data to evolve. With this general data policy, we think that our fields will reap great benefits for generations to come.
No takes yet. Share an insight, caveat, or question.
Rausher et al. (2009) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: