Modern day research -- particularly among the Scandinavian countries -- increasingly relies on data from a combination of large cohort studies and nationwide registry sources. These data are extremely valuable, with vast potential for analysis. Researchers can investigate new questions, revisit old ones using innovative methods, and seek to replicate or extend existing findings, and more. This cumulative reuse maximises the return on public investment and on the active, voluntary contributions of cohort participants. However, to fully realise this potential, researchers must have not only access to the data, but also tools and practices that facilitate efficient, robust, and transparent preparation and usage of it. Here, we present and describe two examples of such tools. The phenotools R package is an open source software package designed to facilitate efficient and reproducible use of data from the Norwegian Mother, Father and Child Cohort study sample (MoBa). The regtools R package is also an open source software designed to facilitate transparent and reproducible diagnostic trend and prevalence analysis. It utilises data from the Norwegian Patient Registry (NPR) with stratification according to linked demographic data using microdata from other Norwegian health and administrative registers, including information like income and education. The motivation for developing these packages will be presented alongside an overview of their contribution to facilitating replicable and transparent science, illustrated with real-life use cases and examples. As developers and researchers, we reflect on the process of creating these tools for use by our scientific peers and outline our understanding of how open source software can be a flexible solution to many of the reproducibility and transparency challenges currently facing research.
Sanchez et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: