National statistical offices are exploring alternative data sources for official statistics. Such data, including point-of-sale (POS) records and mobile phone GPS logs, are originally collected for operational rather than statistical purposes. In the era of big data, the private sector accumulates vast volumes of transaction data, and leveraging such data for official statistics has become an emerging priority. However, these efforts face significant challenges, primarily due to severe selection bias stemming from such data. Prior research has shown that, even if non-representative, transaction data can produce timely statistics via density ratio estimation methods from machine learning. As a proof of concept, that study demonstrated that preliminary estimates could be generated using biased data from a Japanese private employment agency, enabling the early release of a labor market indicator otherwise delayed by up to a year. Building on this, the present study incorporates deep learning into density ratio estimation to improve accuracy. While deep learning, when applied to density ratio estimation, is often considered prone to overfitting, this study demonstrates that it can improve estimation accuracy without overfitting. Moreover, although deep learning is typically regarded as requiring extensive hyperparameter tuning, we show that it can be implemented without a significant tuning burden, supporting its practical use in the production of official statistics.
Takada et al. (Sun,) studied this question.