Abstract Objective. Wearable biosignal-based hand gesture recognition (HGR) is a key enablingtechnology for prosthetic hand control, but its reliability is often affected by limb-positionvariability and other real-world confounding factors. This study aims to systematicallycompare dual-path temporal fusion architectures for co-located surface electromyography(sEMG) and pressure-based force myography (pFMG), with emphasis on robustness,computational efficiency, and interpretability under conditions relevant to practicalprosthetic use. Approach. Three dual-path Temporal Convolutional Network (DFF-TCN)architectures were structured and investigated to integrate sEMG and pFMG usingdifferent fusion strategies: (1) a baseline concatenation-based model combiningdecision-level and feature-level fusion, (2) a decision-level cross-attention variant, and (3) afeature-level cross-attention variant. All models were evaluated under identical trainingand testing protocols using a custom dataset collected from ten participants performingnine functional hand gestures across multiple static and dynamic arm positions. Mainresults. Across all evaluated conditions, the concatenation-based DFF-TCN achievedbalanced performance, with a mean classification accuracy of 95.88%, while theattention-based variants achieved accuracies of 90.65% and 94.02%, respectively.Computational profiling showed that the concatenation-based model also achieved thelowest inference latency (1.70 ms), indicating suitability for real-time deployment.Explainable artificial intelligence analysis using Integrated Gradients revealedcomplementary contributions from sEMG (54.08%) and pFMG (45.92%), withcontribution patterns varying across gestures and subjects. Significance. The resultsdemonstrate that different fusion strategies offer distinct trade-offs between recognitionperformance, computational cost, and robustness. In particular, the concatenation-basedmodel provides a favorable balance for real-time prosthetic hand control, whileattention-based variants offer additional modeling flexibility. These findings providepractical guidance for selecting multi-modal fusion architectures in wearable HMI systemsand support the continued use of co-located sEMG-pFMG sensing in prosthetic andrehabilitation applications.
Zhang et al. (2026) studied this question.