PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 13, 2026Communication in Statistics- Theory and Methods0 citations

Leveraging big data and crowdsourced data as auxiliary variables and clustering analysis to enhance the quality of small area estimation models

View Full Paper
MSMaria A. Hasiholan SiallaganAUAzka Ubaidillah

Key Points

  • The study aims to improve the accuracy of poverty estimates at the sub-district level using small area estimation techniques.
  • Utilizes small area estimation (SAE) to measure poverty in West Java for 2022.
  • Incorporates big data and crowdsourced data as auxiliary variables in the SAE model.
  • Applies clustering analysis to group areas by similar covariate characteristics.
  • SAE model yields more precise poverty estimates compared to direct methods.
  • Clustering analysis significantly enhances the validity of the SAE estimates.
  • Big data and crowdsourced data substantially influence the formation of the SAE model.

Abstract

Measuring poverty up to the sub-district level using direct estimation from survey data often yields imprecise results due to insufficient sample size. Small area estimation (SAE) is an indirect estimation method that can enhance estimation accuracy by borrowing strength from neighboring areas and utilizing the relationship between auxiliary variables and the interest variable. Big data and crowdsourced data have the potential to provide informative auxiliary variables, available at small, up-to-date, and low-cost levels. Additionally, information from cluster analysis in SAE can reduce prediction errors by grouping areas based on the similarity of their covariate characteristics. This study aims to estimate the proportion of poor population at the sub-district level in West Java in 2022 using SAE. This study also analyses the impact of using clustering analysis as well as big data and crowdsourced data as auxiliary information in SAE modeling. The auxiliary variables are derived from administrative data, satellite imagery, Open Street Maps, and Google Maps. The research indicates that the SAE model provides more precise estimates compared to direct estimations. This research also finds that clustering analysis enhances the validity of SAE estimation. Furthermore, big data and crowdsourced data significantly influences the formation of the SAE model.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Siallagan et al. (2026) studied this question.

synapsesocial.com/papers/6a7d76172b0e0cff3f63f448https://doi.org/10.1080/03610926.2026.2712972
Ask AI
Helpful
Bookmark
Share
View Full Paper