ABSTRACT Robust individual identity recognition serves as a crucial foundation for numerous visual applications, such as intelligent surveillance and human‐computer interaction. However, in complex real‐world scenarios, characterized by illumination variations, occlusions, and other interference factors, the performance of traditional methods relying on single biometric features degrades significantly. To address this limitation, this paper proposes a Causality‐based Algorithm for Decoupling and Fusion of Multimodal Features (CA‐DFMF). It achieves robust identity recognition through a two‐stage process. Firstly, a Causally‐driven Feature Decoupling module (CDFD) is designed. Leveraging an environment‐feature cross‐attention mechanism, this module separates confounding features induced by environmental factors from original observations, thereby extracting purified intrinsic identity features. Subsequently, a Graph Attention‐based Feature Fusion module (GAFF) is developed. This module dynamically fuses features based on the inherent correlation between facial and skeletal modalities and introduces an orthogonal decoupling loss to suppress redundant information. Our approach mitigates the impact of environmental confounders from the perspective of the data generation mechanism and enhances the complementarity of dual‐modal features. Experiments on the Market‐1501 dataset demonstrate that the proposed method improves the Rank‐1 accuracy by 7.9% and 16.1% compared with benchmark face recognition and skeleton recognition methods, respectively. On a self‐built dataset covering various occlusion and illumination challenges, our method also outperforms state‐of‐the‐art public methods, achieving superior accuracy and F1‐score by margins of 13.1% and 13.8%, respectively. The results indicate that the proposed method can effectively enhance the robustness of identity recognition in complex scenarios, providing a reliable and efficient technical solution for visual perception systems with high‐reliability requirements.
Chen et al. (2026) studied this question.