Advances in Artificial Intelligence (AI) and Machine Learning (ML) have significantly enhanced the practice of Earth Observation (EO), enabling complex analyses such as land cover change detection, vegetation monitoring, and disaster response. However, while model architectures have matured, the refinement of reference data remains a major challenge. Accurate and dynamic multi-temporal labelling is essential for capturing evolving ground conditions in high-dimensional EO datasets, yet key challenges persist, including spatiotemporal inconsistencies, heterogeneous data integration, and multi-resolution harmonization. Without robust preprocessing, reference labels may introduce biases, resulting in reduced model reliability and generalizability. This review tackles four core aspects of reference data preprocessing in EO: (i) essential steps for producing consistent and high-quality datasets, particularly for dynamic spatiotemporal data; (ii) best practices and guidelines that enable scalable and accurate workflows across diverse EO applications; (iii) introduction of the HELIX framework, a unified approach for standardizing, enhancing, and automating spatiotemporal label preprocessing; and (iv) a forward-looking discussion on the future of reference labels and features, including next-generation techniques for dynamic EO data integration. By synthesizing existing methodologies, highlighting emerging approaches, and addressing current gaps, this review underscores how well-engineered reference data are fundamental to advancing AI/ML-driven EO applications.
No takes yet. Share an insight, caveat, or question.
Hauser et al. (2025) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: