Key points are not available for this paper at this time.
Generative data augmentation has been widely explored in Wi-Fi fingerprint-based indoor localization to reduce the cost of dense radiomap construction, with many studies reporting substantial localization improvements. However, these comparisons typically rely on default or weakly optimized baseline regressors, making it difficult to determine whether reported gains reflect genuine synthesis quality or merely compensate for suboptimal baselines. In this paper, we systematically investigate under which conditions generative augmentation is actually justified. We fine-tune four widely used localization regressors—kNN, SVR, XGBoost, and DNN—using Bayesian hyperparameter optimization and establish strong non-augmented baselines across radiomaps with controlled levels of spatial sparsity, constructed via farthest-point sampling. We then train five representative generative models—VAE, GAN, DDPM, DiT, and TDPM—within a unified augmentation pipeline that includes quality filtering and pseudo-labeling, and benchmark them against these baselines. Using two publicly available datasets, we show that none of the generative models consistently outperforms a non-augmented, fine-tuned baseline regressor such as XGBoost or kNN, across a wide range of sparsity levels. We further show that these conclusions are robust to three potential confounders: various proportions of synthetic data, the choice of localization regressor (ruling out circularity with the pseudo-labeling model), and the dataset itself, since the findings on the first dataset replicate on a second, structurally different building. These findings suggest that reported augmentation benefits in prior work may partly reflect under-optimized baselines rather than genuine synthesis quality, and that generative augmentation should be treated as a conditional last resort rather than a universal improvement strategy.
Malikov et al. (2026) studied this question.