In geotechnical engineering, machine learning success is often reported as “high accuracy”, yet models that score well under conventional metrics may fail in deployment. This article clarifies what “good performance” should mean in geotechnical machine learning and why apparently well built models can lose reliability when transferred across sites, sensing systems, and construction contexts. Performance is interpreted as an engineering claim in which predictions should remain valid under spatially structured data, interpretation dependent labels, proxy driven predictors, and decision thresholds tied to safety and serviceability. We synthesize key failure modes, including non-representative training data, validation leakage under spatial dependence, domain shift, and miscalibrated uncertainty, and show how they interact with common geotechnical data regimes such as small sample sizes, spatial autocorrelation, and heterogeneous data provenance. A case study predicting the logarithm of normalized mobilized undrained shear strength in clays shows that apparent model performance and engineering interpretation are strongly validation dependent. Errors propagate into unconservative strength classifications, while predictive intervals that appear acceptable under sample-wise validation can become less informative under deployment consistent evaluation, directly affecting engineering interpretation. Finally, we propose a minimum evidence evaluation and reporting framework that aligns model assessment with deployment consistent validation, physical admissibility, uncertainty calibration, and decision relevance. Reliable use of machine learning in geotechnical engineering depends less on algorithmic novelty than on disciplined validation strategies, transparent uncertainty reporting, and explicit definition of applicability domains for safety critical decisions.
Taherdangkoo et al. (2026) studied this question.