This analysis uncovers how absence affects data science, highlighting uncertainty in AI datasets.
Absences are inescapable in data. Data collection always focuses on some elements while occluding others. Yet, how absences are considered and recorded within data infrastructures markedly transforms the inferences that can be made. Tracing a genealogy from early databases to contemporary AI datasets, this paper explores how data infrastructures have grappled with the inherent incompleteness of data. Specifically, I uncover a tension between a desire for certainty and acknowledging partiality at the foundation of data science that continues to pervade contemporary AI datasets. Drawing on archival studies and sociological perspectives, I argue that data science must embrace uncertainty by recognizing the “ghosts in the data”—the uncounted, the unrepresented, and the silenced—and how their absence shapes the outcomes of automated systems.
No takes yet. Share an insight, caveat, or question.
Will Orr (2025) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: