Deep neural networks have become the dominant approach to wearable human activity recognition (HAR). Yet, their opacity creates practical barriers because clinicians cannot verify whether predictions align with physiological knowledge. Engineers lack principled guidance on which sensors to retain when cost and power constraints require system simplification. Existing explainable artificial intelligence (XAI) methods for wearable HAR typically rely on a single attribution technique, which introduces method-specific bias and reports importance at granularities such as individual channels or entire body locations that do not map cleanly to hardware design decisions. This study introduces a multi-method XAI framework that integrates counterfactual sensor group ablation, Integrated Gradients, and Shapley value sampling around a shared deep learning backbone and evaluates it using a time-distributed long short-term memory network trained on the Mobile Health dataset, which records twelve physical activities through eight sensor groups, including accelerometers, gyroscopes, magnetometers, and electrocardiogram signals at the chest, ankle, and wrist, achieving 98.2 percent accuracy and a macro-averaged F1 score of 0.98. Four coordinated experiments examine model behavior at global, sensor group, and activity-specific levels, showing that removing the ankle magnetometer reduces dynamic locomotion accuracy by 47.1 percentage points while removing the wrist accelerometer decreases confidence for static postures by more than 50 percent. Class-specific Integrated Gradients heatmaps and temporal attribution curves produce biomechanically consistent patterns such as ankle-centered signatures for gait and wrist and chest emphasis for upper body movements. Global Integrated Gradients and Shapley value rankings converge, with accelerometers accounting for 89 percent of the total attribution mass. The agreement across causal, gradient-based, and game-theoretic perspectives strengthens confidence that the identified sensor importance patterns reflect genuine model behavior, yielding sensor group-level explanations that provide actionable guidance for sensor selection, power-aware deployment, and clinically meaningful interpretation without sacrificing recognition performance. • Integrate multiple explainability methods to reveal sensor-driven health activity patterns. • Achieve highly accurate recognition for clinically relevant human activities. • Identify the most influential wearable sensors for health monitoring. • Quantify how sensor removal affects health activity recognition quality. • Provide analytics-driven guidance for power-efficient wearable health systems.
Kurniawan et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: