ABSTRACT Visible‐Infrared Cross‐Modal Person Re‐Identification (VI‐ReID) confronts significant challenges arising from feature misalignment across spectra and structural differences within the same identity, especially under low‐light scenarios. To address these, we propose the Frequency‐Spatial‐Channel Collaborative Enhancement Network (FSC‐Net), leveraging a spatial‐semantic collaborative mechanism and a multi‐dimensional Frequency‐Spatial‐Channel feature enhancement framework. FSC‐Net employs a Frequency Semantic Attention module for cross‐modal semantic alignment, a Gated Channel Attention module for discriminative feature enhancement, and a Spatial Transformer for improved local structural perception. Extensive experiments demonstrate the superiority of FSC‐Net, achieving state‐of‐the‐art Rank‐1/mAP accuracy of 95.07%/91.18% on RegDB, 76.90%/74.26% on SYSU‐MM01, and 56.69%/63.37% on the low‐light LLCM dataset, respectively. Especially, on the RegDB dataset, our model achieves an absolute gain of 5.2% and 8.1% in terms of Rank‐1 and mAP in infrared‐to‐visible mode.
Huang et al. (2025) studied this question.