ABSTRACT Deep learning models, particularly long short-term memory networks (LSTMs), have set new standards in streamflow prediction but require extensive data and computational resources. This raises a practical question: which parts of the data are truly indispensable, especially when computational budgets are limited or when only sparse observations are available? To address this, we examine which type of data is most essential in model training. We systematically ablate the CAMELS-US dataset using four families of sampling strategies – hydrological extremes, event rarity/statistical representativity, temporal context, and spatial representativity – and use the ablated datasets in training. Among all tested approaches, sampling based on statistical representativity via random sampling consistently outperformed more targeted strategies, achieving strong performance (NSE 0.7) and good representativity on as little as 10% of the data. Sampling hydrological extremes is the second-most efficient strategy, particularly when high-flow and low-flow extremes are sampled jointly, but with the largest performance gains stemming from high-flow events. Concerning temporal context, surprisingly short sequence lengths (3 weeks) and training periods (2 years) were sufficient for competitive performance (NSE 0.7). These findings provide practical guidance for efficient data selection in data-driven modeling and provide groundwork for future studies on training strategies.
Heudorfer et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: