Epidemiology is concerned with understanding the distribution and determinants of health in the population. A sizable fraction of epidemiological research involves secondary data analysis: statistically analysing data collected from cohorts, cross-sectional studies, or other data sources. Such research comprises a series of cognitive tasks currently conducted, or at least overseen, by humans. Historically, conducting epidemiological research was a slow, manual endeavor: scanning library shelves, reading physical papers, and manually collecting, coding, and analysing data (Fig. 1). Technological progress has now led to much of this work being done electronically, yet actual scientific progress arguably remains slow; e.g. despite the surge in large cohorts and ballooning data volumes—omics, wearables, administrative linkages, etc.—progress in identifying modifiable causes of disease has proved elusive [1, 2]. Technological progress in two key epidemiological tasks (literature reviews and data analysis): from manual work and computerization to artificial intelligence (AI)-augmented research. Note that, in some senses, the final tasks listed (e.g. plain-language analysis) are already possible with current AI systems, yet the quality of the outputs is of mixed or as-yet undetermined quality. Artificial intelligence (AI) represents the next step in the technological evolution of epidemiology (Fig. 1); it can accelerate—or even automate—cognitive tasks, boosting the efficiency of current practice and creating new opportunities for discovery. Epidemiologists often use AI-based tools—sometimes without explicitly knowing it—such as Google Scholar for paper discovery, spell-checkers for writing, and GitHub Copilot for coding. Despite previous AI “winters”, its current era of development, built around the transformer deep-learning architecture [3] that powers modern large language models (LLMs), has generated remarkable progress. LLMs shot into public consciousness in November 2022 with the release of ChatGPT, reportedly the fastest-growing consumer product of all time. Many scientists, particularly younger researchers, now use ChatGPT [4]. Other LLMs have since become publicly available and widely used, and billions of dollars are invested in their training. LLMs predict the next token (typically, a small piece of text) in a sequence and, when developed at a massive scale, have surprisingly useful properties. Productivity increases in tasks relevant to epidemiology have recently been suggested—writing [5], cognitive tasks [6], debating/reasoning [7], and coding [8]. Here, we map the landscape of epidemiological tasks that rely on existing datasets—from literature review through to idea generation, data access, analysis, write-up, and dissemination. We provide a snapshot of AI tools and a repository containing examples of AI-generated epidemiological output, along with prompt and model details (https://github.com/edlowther/automated-epidemiology). The tools were chosen to present an illustration of what, at the time of writing, frontier AI models were capable of. (We use the term AI to refer to computational systems that can perform cognitive tasks relevant to epidemiological research.) Finally, we discuss barriers to deeper AI integration and broader implications for the field, including how epidemiologists can contribute to AI development, addressing recent calls [9]. We note that the issues discussed apply to other fields, e.g. social sciences [10, 11]. Systematic reviews generally take ≥1 year to undertake [12], with much of this time spent on pain-staking and often tedious screening—manually removing the vast majority of irrelevant articles from the selected pool by comparing them against the same inclusion/exclusion criteria. At face value, this breaks the software engineering principle to “automate repetitive tasks,” but an additional motivation for automation is to reduce errors: humans do not screen without error [13]. In recent years, machine-learning tools have become available to speed up screening: authors manually screen a smaller subset of abstracts to train models, which then automatically screen the remainder [14]. Such tools appear to increase efficiency [15, 16], widening the scope to produce or update reviews more rapidly (possibly continuously) and undertake more ambitious evaluations. LLMs can help in other review-related tasks, such as creating synonym/search-term lists and extracting data. Could the entire process of reviewing be automated? Using an agentic AI system (Otto-SR), a 2025 study claimed to have reproduced and updated a Cochrane issue in 2 days—the equivalent of 12 work-years of traditional systematic review work (assuming 1 year per review) [17]. For more ad-hoc literature searches, systematic reviews are typically prohibitively costly, e.g. when informing introductions or discussion sections in original research articles. Researchers are increasingly using AI-augmented search tools such as Google Scholar to undertake literature searches—unlike PubMed, it indexes non-health publications (e.g. economics articles), as well as grey literature. Nevertheless, both tools require conversion of the search question (e.g. “What effect does X have on Y?”) into terms that are more likely effective for such databases (e.g. “associations between X and Y,” a “randomized controlled trial of X on Y,” etc.), in addition to a continued manual search of “cited by” articles. LLMs can answer such single questions directly and recent reasoning LLMs enable AI “sense checks” before a response is produced. Hallucinations—which raised considerable concern in early models—appear to have been reduced [18]. “DeepResearch” capabilities, made available in several leading LLM tools in recent months, enable more extended searches of scholarly literature; users can check the sources provided via links to the full text. The generality of LLMs means they can usefully sift, connect, and summarize evidence from far-flung disciplines—a task that has otherwise become progressively harder as scientific output has surged [19]. For epidemiologists, such sources range from mechanistic studies in cells, animal models, human autopsy studies, and the social sciences (e.g. psychology, economics, sociology). AI tools could thus make cross-disciplinary triangulation more feasible. Hallucinations remain a barrier to the trustworthiness of LLMs, but a human barrier exists in accessing research articles. Since the 1970s, the five largest for-profit publishers have steadily increased their market share, accounting for more than half of all papers by 2013 [20]. Papers—and even their abstracts—are copyrighted. This creates a particular barrier for open-source AI systems [21]. Partial access to research articles, limited performance with longer “context windows” (the amount of data the LLM uses from memory), and the capacity of LLMs to provide highly compelling but unsupported narratives mean LLMs may mislead [22]. Ongoing evaluation of such systems is required: empirical study of their sensitivity and specificity in searches, for example. This is a challenge given their rapid development—closed-source frontier LLMs can be rapidly updated or decommissioned. It is often assumed that AI systems (particularly LLMs) simply interpolate between data points contained within their training set and are thus not capable of being creative or generating novel ideas—i.e. they are “stochastic parrots” [23]. Setting aside the “incremental” nature of modern science, such claims are at least partly empirically testable: an emerging literature suggests that the creative capability of frontier models may match those of humans in discrete small-scale creative tasks [24]. Their abilities in real-life scientific creativity remain uncertain, as do the comparisons of humans alone versus human–AI collaborations in (i) forming hypotheses that advance epidemiology or (ii) selecting hypotheses that are tractable and falsifiable [25] given the existing data. In other disciplines, such as drug discovery, new scientific findings are seemingly being discovered via AI systems [26]. We prompted a recently developed AI tool (the AI Scientist [27]) to suggest novel hypotheses across two topics: (i) the links between birthweight and subsequent body mass index (BMI) and (ii) social inequalities in mental health (see github.com/edlowther/automated-epidemiology). Many hypotheses appear to have face validity, e.g. suggesting generally underutilized approaches to causal inference (sibling comparison studies and natural experiments). We note that such suggestions were created in “one shot” and are thus the equivalent of a human’s first draft. Even if only a fraction of AI-suggested hypotheses are promising, the number that can be created quickly is large and may be especially valuable with discerning humans “in the loop” to select them: an AI-augmented process akin to human brainstorming. A common approach in epidemiological research is that groups running specific epidemiological studies (e.g. cohorts or health surveys) publish research focused on using that specific dataset. In this scenario, multiple publications in the literature from different research groups address the same question; yet, subsequently synthesizing such evidence (e.g. via meta-analysis) is not always possible due to methodological differences. Consortia integrating multiple studies are one manual approach to circumventing this, but they are typically set up for specific research questions and are hard to maintain in the long run; when their funding ends, they may cease to operate. AI may enable a bolder default for epidemiological research, enabling us to ascertain, for each research question, the possible available datasets that could contribute evidence. Of these, which have harmonizable data? And what does that evidence collectively show? Current barriers to this include the high fixed costs of becoming familiar with datasets and the fragmented approaches to data discovery and access. Platforms to aid cohort discovery, e.g. the recent Atlas of Longitudinal Data (https://atlaslongitudinaldatasets.ac.uk), are a step forwards in helping to identify datasets; yet, using them highlights our barriers: 10 different cohorts may involve 10 separate access systems, with considerable overlap in the information requested. The challenge for data providers is whether a single point of entry can be provided—a cohesive streamlined data-access system with interoperable data and necessary safeguards. ORCID provides a centralized and broadly accepted system for verifying researcher identity—could existing centralized systems for data documentation and access be expanded (e.g. UK Data Service for UK cohorts; or the Gateway to Global Aging Data, for older adults) or newly created? Within such systems, AI tools can also facilitate the historically slow and manual process of harmonizing data across different datasets [28] (e.g. the Harmony tool [29]). Finally, AI tools can aid in the creation of new epidemiological data. In existing cohorts, for example, data held in historic non-electronic form (e.g. paper questionnaires or microfiche) can be digitized by using automated optical character recognition tools. Such tools can also be used to create new retrospective cohorts: many hundreds of papers have now cited the cohort profiles that arose from the discovery and digitization of records, which formed the basis for the Hertfordshire [30] and Lothian [31] cohort studies. AI tools could also improve existing metadata (e.g. annotating questionnaires with associated variable names). Much like in the literature reviews, epidemiologists are increasingly supported by AI when analysing data. Rather than manually typing out each letter when coding, AI autocompleters such as GitHub Copilot can speed up code writing. For a guide, see https://www.ncrm.ac.uk/resources/online/all/? id=20859. Frontier LLMs are now able to create a complete draft of code in response to a plain-language prompt and then execute this code. The promise is that the rapid, autonomous generation of research code will enable human researchers to spend more time at higher levels of abstraction, e.g. thinking carefully about designing research strategies. We prompted an agentic AI framework (Data Analysis Crow) to address two research questions and provided simulated data. The responses yielded an analytical plan, analytical code, execution of this code, and visualizations—see github.com/edlowther/automated-epidemiology for full workbooks and Table 1 for a summary. While the outputs were (in our view) impressive, they did contain errors and, in some cases, failed entirely, depending on the underlying LLM used. This suggests that (i) the choice of LLM is important and (ii) code review remains essential. Evaluating AI-generated analysis: illustrative results from the Data Analysis Crow ✓ Derived BMI ✓ Removed implausible values ✓ Derived BMI ⚠️ Removed some but not all implausible values ✓ Created analytical plan ⚠️ Errors (e.g. <1.2 metre cases excluded, not <1 metre) ✓ Executed results: tables/figures ✓ Created analytical plan ⚠️ Partial results (API crash) ⚠️ Identified sex, assumed value labels ✓ Log-transformed income ⚠️ Identified sex, assumed value labels ✓ Rescaled income ✓ Created analytical plan ✓ Executed results: tables/figures ✓ Created analytical plan ⚠️ Partial results (API crash) Each item was evaluated as follows: ✓: correct or plausible output; ⚠️: error or concern identified. A simulated dataset was provided, available on the accompanying repository: https://github.com/edlowther/automated-epidemiology. The Data Analysis Crow is available at https://github.com/Future-House/data-analysis-crow. API, application programming interface. Often, epidemiologists specialize in one piece of software or programming language (e.g. SAS, SPSS, Stata, R). In one sense, specialization is increasingly not needed, as the barrier to entry lowers to code in multiple languages. What will remain important is the clear articulation of the goals in plain language and code review. The fact that AI systems provide analytical syntax also aids in reproducibility: something that <2% of health researchers currently do [32]. The more mundane aspects of data analysis could also be accelerated by AI. Data cleaning, for example, is often a highly manual and time-intensive task that is required even for well-used datasets, leading to considerable duplication of work. Assuming that data cleaning involves 1 month of unnecessary work (cleaning data that should have otherwise been centrally cleaned)—a task ordinarily repeated across 1000 papers—1000 months (83 years) of scientists’ time could be saved in the future. A cursory look at the literature suggests at least six distinct AI data-cleaning tools from 2024 onwards that claim varying levels of accuracy in data cleaning [33–38]. Whether such tools are useful in epidemiological applications remains to be seen. A challenge for epidemiologists will be to make sense of the bewildering numbers of tools released in AI-related fields: the creation and curation of epidemiological benchmarks could provide objective criteria by which they can be continually evaluated. Barriers to the use of AI in data analysis include the current frequent need for uploading data to cloud providers: this is not possible for many health-related datasets held in sandboxed secure computing environments. Researchers could instead use local open-source models—such models have historically been weaker than the closed-source models, yet, in recent months, the gap has narrowed considerably [39]. Alternatively, the AI tools could be restricted to accessing metadata (e.g. variable names, labels, and result output) rather than the raw data, or data owners could release synthetic versions of their data. Epidemiologists will, as need to two of The first is a of data if AI tools are used the is a that scientific discovery is restricted if such tools are not used. The is often in our despite to the continued of epidemiological studies (e.g. funding and response and the of data which often remain given a plain-language frontier LLMs can produce entire epidemiological research github.com/edlowther/automated-epidemiology examples of this by using multiple LLMs, each to a paper on the between birthweight and BMI by using a simulated dataset that we do such papers in the current distribution of epidemiological research Despite being given the quality our outputs by at face value to the used criteria for the of epidemiology studies available on the accompanying It also of it a in by that we into the simulated data, despite the prompt the LLM to by In other the papers are of e.g. and the of result as in the LLMs can tool produce compelling and current barriers to AI-generated may partly in how the model is prompted or with tools. automation of epidemiological research generating the idea all the through to a of the capability of AI in each Such tools as AI Scientist and are recent open-source examples of A of research is the evaluation of such systems for e.g. to and Researchers are increasingly to their research with such as the public and and more generally in continued public the of are already with a of small recent on the of AI such tasks may be if public is by a and of scientists, then may LLMs can speed up the creation of and social if provided with and (e.g. the research human may be necessary to accuracy and can now be an AI-generated on this can be at AI increases then researchers may be able to a deeper with than simply that their work should creating evidence across the entire evidence for example, and carefully In our the current capability of AI suggests a to This is the whether AI is used for tasks human as a research or with human as an a or autonomous research Each may to with integration an of both and A promise of AI in the to term is that it could enable more time to be spent on tasks (e.g. designing new research questions or data rather than on tasks that are often (e.g. repetitive or (e.g. code to (see The is at our some in tasks is likely to be or even necessary to (e.g. to data, to and the for training should to this and an on which could the required to and use their outputs Could AI help to epidemiologists to on illustration of our current time versus AI-augmented the quality rather than the of our outputs is then the result could be and a of our We in an era that to produce of papers of including their this is a not an AI are rapidly the costs of cognitive tasks and may for human epidemiologists, on and analysis tasks, by a this could the training leading to epidemiologists across all levels and thus a in the are already uncertain, with of for epidemiologists generally than those in other (e.g. our to become epidemiologists in the is one for is the of on and important integration of AI with epidemiology is For example, can AI or can we use AI to improve epidemiology and a vast of AI benchmarks be created for What and can AI systems the in the AI literature be used to improve and inference in will AI AI could and aspects of epidemiology in the future. AI in its current were to take human intelligence entirely, our could be to produce new which vast use to train AI systems without our Whether such are or for scientific discovery or at large remains an question, which epidemiologists can and should contribute the of AI for epidemiology will require between epidemiologists and the first analysis: authors to generating reviewing and the as well as the accompanying and are by the and and the This paper the use of AI in a range of tools were used as in the paper and LLMs such as ChatGPT were also used to suggest means of the for some The human authors all sections of the and carefully suggestions The a for using
No takes yet. Share an insight, caveat, or question.
Bann et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: