Low-cost sensors (LCSs) have rapidly expanded in urban air quality monitoring but still suffer from limited data accuracy and vulnerability to environmental interference compared with regulatory monitoring stations. To improve their reliability, we proposed a machine learning (ML)-based framework for LCS correction that integrates various meteorological factors at observation sites. Taking Tongshan District of Xuzhou City as an example, this study carried out continuous co-location data collection of hourly PM2.5 measurements by placing our LCS (American Temtop M10+ series) close to a regular fixed monitoring station. A mathematical model was developed to regress the PM2.5 deviations (PM2.5 concentrations at the fixed station—PM2.5 concentrations at the LCS) and the most important predictor variables. The data calibration was carried out based on six kinds of ML algorithms: random forest (RF), support vector regression (SVR), long short-term memory network (LSTM), decision tree regression (DTR), Gated Recurrent Unit (GRU), and Bidirectional LSTM (BiLSTM), and the final model was selected from them with the optimal performance. The performance of calibration was then evaluated by a testing dataset generated in a bootstrap fashion with ten time repetitions. The results show that RF achieved the best overall accuracy, with R2 of 0.99 (training), 0.94 (validation), and 0.94 (testing), followed by DTR, BiLSTM, and GRU, which also showed strong predictive capabilities. In contrast, LSTM and SVR produced lower accuracy with larger errors under the limited data conditions. The results demonstrate that tree-based and advanced deep learning models can effectively capture the complex nonlinear relationships influencing LCS performance. The proposed framework exhibits high scalability and transferability, allowing its application to different LCS types and regions. This study advances the development of innovative techniques that enhance air quality assessment and support environmental research.
Fan et al. (Mon,) studied this question.