"background": "The difference-in-differences (DiD) model is a quasi-experimental method increasingly employed to evaluate health policy impacts in low-resource settings. Its application to assess clinical outcomes within district hospital systems in Tanzania requires methodological scrutiny to ensure validity and inform future research. ", "purpose and objectives": "This systematic review aims to critically appraise the methodological application of the DiD model in studies evaluating clinical outcomes in Tanzanian district hospitals, identifying common design strengths, limitations, and reporting practices. ", "methodology": "A systematic search of multiple electronic databases was conducted following a pre-registered protocol. Eligible studies employed DiD to analyse clinical outcomes (e. g. , mortality, morbidity) following interventions in district hospital systems. Screening, data extraction, and quality assessment were performed independently by two reviewers. The core DiD model was specified as Y{it = \0 + \1 + \2 + \3 (\) +, with appraisal focusing on parallel trends testing, robustness checks, and clustering of standard errors. ", "findings": "Of the 18 included studies, only 39% presented formal statistical tests or graphical analyses to support the critical parallel trends assumption. A prominent theme was the frequent omission of discussion regarding potential contamination between intervention and control groups. The majority of studies (72%) failed to report using heteroskedasticity-robust standard errors clustered at the facility level, which may affect inference. ", "conclusion": "The methodological rigour of DiD applications in this context is inconsistent. While the design is valuable for causal inference in routine health system data, widespread shortcomings in validating key assumptions and in statistical reporting limit the reliability of many estimated effects. ", "recommendations": "Future studies must explicitly test and report on the parallel trends assumption. Researchers should employ clustered robust standard errors and conduct comprehensive sensitivity analyses,
Fatuma Mwakyembe (Sat,) studied this question.