Utility-scale photovoltaic (PV) plants rely on supervisory control and data acquisition (SCADA) streams for performance verification, yet high-frequency measurements are routinely affected by sensor dropouts, intermittency, and operating-state transitions that bias regression-based capacity estimates. This study evaluates a reproducible SCADA processing workflow for capacity-style reporting and a complementary soiling–clean temperature prediction model using data from a documented October 2022 test window (5 s SCADA aggregated to 1 min). The following three filtering approaches are compared: (i) naïve thresholds (Baseline A), (ii) deterministic stability screening using ramp-rate and rolling-variability constraints (Baseline B), and (iii) an optional residual-based outlier trimming step (Method C). Capacity is estimated via a multivariate regression evaluated on a fixed-size reporting-condition subset (RC197) with day-coverage constraints. All methods achieved high fit quality on RC197 (R2≈0.99), with Baseline B improving error and uncertainty over Baseline A (RMSE 2.05 vs. 2.18 MW; U95 0.97% vs. 1.03%) while preserving day coverage; Method C yielded the lowest in-sample RMSE (1.89 MW) but reduced day coverage. For temperature prediction, a baseline-plus-residual learning formulation substantially improved leave-one-day-out performance, reducing MAE/RMSE from 2.99/3.76 °C to 1.43/1.80 °C and increasing R2 from 0.60 to 0.91. The results highlight trade-offs between fit tightness and representativeness in capacity-style filtering and demonstrate residual learning is an effective approach for SCADA-based thermal characterization.
Gao et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: