Background AI-driven healthcare monitoring systems are increasingly being developed to support elderly individuals through continuous and discreet monitoring of routine behavioral patterns. Objective This study aims to develop a multimodal framework for identifying mobility patterns by integrating inertial and vision-based sensor data to improve the accuracy and robustness of behavioral recognition. Methods Inertial data are denoised using a median filter, while video feeds are converted into grayscale frame sequences and enhanced using a Gaussian filter. Rectangular windowing is applied to inertial data, whereas U-Net is used for visual segmentation. Features are extracted using Gammatone Filter Cepstrum (GFC), Gaussian Likelihood Threshold Model (GLTM), Power Spectral Density (PSD), and temporal moments from inertial data, and Distance Transform, Gist, ORB, and SURF from visual data. The modalities are fused using shared class labels, followed by fuzzy-logic-based feature optimization and classification using a CNN-GRU hybrid model. Results The proposed multimodal framework achieves 95.20% accuracy on the Continuous Multimodal Human Action Dataset (C-MHAD) and 94.79% on the multi-modal 3D human pose estimation dataset using mmWave, RGB-D, and inertial sensors (mRI) dataset, demonstrating improved behavioral categorization compared with unimodal approaches. Conclusion The results highlight the potential of multimodal inertial and vision-based sensing for robust AI-assisted healthcare monitoring. However, evaluation is currently limited to public benchmark datasets and requires further validation in clinical settings.
No takes yet. Share an insight, caveat, or question.
Abro et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: