This study examines whether SVOD churn can be reliably predicted using only basic demographics, device usage, and subscription details when rich behavioral telemetry is unavailable. Using 5,000 anonymized Netflix-style records with 14 variables and balanced classes, we engineered features (e.g., binary churn, engagement composites) and evaluated logistic regression, decision trees, random forests, gradient boosting, and stepwise models with stratified cross-validation; although variables like device preference, region, and age were statistically significant, overall predictive power was marginal, with the best logistic model only slightly above baseline and ensembles showing similar limits and overfitting risks. The key implication is that readily available demographic-device profiles are insufficient for effective churn prediction; organizations should invest in richer behavioral signals (content interaction patterns, view depth, temporal engagement) and adopt cost-sensitive approaches that reflect asymmetric costs of false positives and false negatives, using this transparent methodology as a foundation for more advanced modeling.
Chhatrapati et al. (Wed,) studied this question.