Falls are a leading cause of injury and mortality among older adults, motivating growing interest in video-based computer vision (CV) fall detection systems. This study presents a systematic mapping of vision-based fall detection in video, synthesizing evidence from 433 primary studies published through 2025 and retrieved from five databases using explicit eligibility criteria and independent dual screening within a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-inspired workflow. We characterize the field through: (i) a structured three-level taxonomy grouping approaches into Feature Engineering, Deep Learning, and hybrid models; (ii) a quantitative analysis of commonly used algorithmic components and their reported performance across datasets and metrics; and (iii) a dedicated assessment of efficiency and deployment evidence (e.g., frames per second (FPS), latency, and hardware/platform reporting). Our findings indicate the predominance of Deep Learning pipelines-particularly convolutional neural network (CNN)-based backbones-together with a sustained prevalence of hybrid designs, while Transformers/Attention architectures show accelerated adoption in recent years. Despite frequent real-time claims, efficiency metrics and hardware specifications remain inconsistently reported, limiting reproducibility and clinical translatability. Overall, this mapping consolidates trends, benchmarks, and reporting practices, and identifies research gaps that hinder reproducibility and clinical translation.
Gomes et al. (Thu,) studied this question.