PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 20260 citationsOpen Access

A Bias Correction Scheme for FY-3E/HIRAS-II Data Assimilation Based on EXtreme Gradient Boosting

View Full Paper
HCHongtao ChenNanjing University of Information Science and TechnologyLGLi GuanNanjing University of Information Science and Technology

Key Points

  • To develop a bias correction scheme for FY-3E/HIRAS-II data using XGBoost.
  • Established an XGBoost model for bias correction.
  • Selected predictors include model skin temperature and total column water vapor.
  • Conducted data assimilation experiments over two weeks for validation.
  • XGBoost bias correction outperformed static and variational bias correction methods.
  • Mean and standard deviation of biases were smallest after XGBoost correction.
  • Temperature and humidity fields matched closely with ERA5 across all levels.

Abstract

More and more spaceborne infrared hyperspectral atmospheric observations are assimilated into data assimilation systems. The key to bias correction (BC) of these instruments depends on selecting predictors. However, it is difficult to find a set of predictors that are highly correlated with the O-B biases in all FY-3E/HIRAS-II channels, due to its multi-channel characteristics. A machine learning model XGBoost (EXtreme Gradient Boosting) BC scheme for FY-3E/HIRAS-II is established in this article. The selected predictors include model skin temperature, model total column water vapor, 1000–300 hPa thickness, 200–50 hPa thickness, scan position, observed brightness temperature (BT) and simulated BT. The method is also compared with the operational static BC and the variational BC, to validate its effect. The two-week data assimilation experiments show that the XGBoost BC is the most effective among the three BC schemes. The mean and standard deviation of O-B in all channels are the smallest after BC, and the effective observations through quality control are the largest, followed by the static BC. The static BC and variational BC are performed based on linear regression, which may lead to a small loss of valid observations in some channels that are weakly correlated with the predictor, whereas machine learning algorithms can search for the nonlinear correlation between biases and predictors. Compared with ERA5, both temperature- and humidity-analysis fields based on XGBoost BC are closest to ERA5 at all levels, and the root mean square errors do not change much over time.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chen et al. (2026) studied this question.

synapsesocial.com/papers/69a67f12f353c071a6f0aeb2https://doi.org/10.3390/rs18050744
Ask AI
Helpful
Bookmark
Share
View Full Paper