Key points are not available for this paper at this time.
During the last decade, Machine Learning (ML) approaches have been increasingly adapted to archaeology from related fields, supported by the expansion of open repositories, advances in computational infrastructures, and the fast development of detection, classification and segmentation algorithms. However, the multiplication of ML-based papers can give a false impression of methodological efficiency and general applicability, while masking fundamental constraints derived from the nature of archaeological data: scarcity of training samples, high variability in the features of interest, strong spatial dispersion and class imbalance, and highly heterogeneous backgrounds. In contrast to disciplines working with digital-born data, archaeological “objects” are physical phenomena that can only be accessed through indirect proxies (spectral, morphometric, textural, topographic, or contextual measures), meaning that ML systems do not learn archaeological features as such, but a mathematical definition of how those proxies were labelled, represented, and sampled. Consequently, archaeological ML workflows are intrinsically shaped by disciplinary definitions, data acquisition choices, and training/validation practices, with direct implications for bias, transferability, interpretability, and overfitting. Building on these premises, this paper provides a practice-oriented theoretical framework for the implementation of ML-based approaches without losing domain and methodological accountability. We argue that the value of archaeological ML lies not in replacing domain expertise but in its integration, which can enable scalable, explicit, and testable forms of inference and automatise analyses that are otherwise impractical, while making the chain between archaeological concepts and computational representations more transparent and critically examinable.
Orengo et al. (Thu,) studied this question.