Key points are not available for this paper at this time.
Background: Integrating human neural signals with computational vision systems offers a promising route toward more robust visual recognition, yet supporting mixed-granularity recognition, where both coarse- and fine-grained categories must be distinguished within a unified system, remains challenging due to the heterogeneous feature spaces of electroencephalography (EEG) and visual data. Methods: We propose “Align and Fuse,” a two-stage Transformer-based framework. Stage 1 constructs a shared semantic space using a hardness-aware multimodal supervised contrastive loss with Hard Negative Weighting to explicitly target confusable class pairs. Stage 2 employs a multimodal Transformer with co-attention to fuse the aligned features for classification. Results: On the 80-class EEG-ImageNet benchmark, our framework achieved 91.12% Top-1 accuracy under a temporally separated control protocol, improving over the corresponding vision-only (89.08%) and Standard Transformer (89.95%) baselines. Under the original stratified random split, it achieved 92.56% Top-1 accuracy; on the 40-class EEGCVPR dataset, accuracy reaches 95.82%. Cross-subject experiments yield 90.92% average Top-1 accuracy on four unseen subjects, and Grad-CAM analysis suggests that aligned EEG signals shift the model’s attention toward semantically relevant regions. Conclusions: Coupling hardness-aware alignment with decoupled multimodal fusion supports EEG-augmented recognition by leveraging complementary stimulus-related information under the evaluated protocols. Because EEG features are required at inference time, the framework is positioned as a human-in-the-loop EEG-augmented recognition system rather than a standalone vision model.
Zhang et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: