Abstract Missing modalities frequently occur in EEG- and eye movement-based emotion recognition due to the inherent instability of EEG signals and the complexity of multimodal data acquisition. Such modality absence leads to inefficient data utilization and substantial performance degradation, especially when missing modalities appear simultaneously across multiple datasets. This issue is further exacerbated when a substantial portion of a modality is missing within one or more datasets. Consequently, effectively exploiting the remaining multimodal information under cross-dataset and missing-modality settings remains a critical and unresolved challenge. To address these challenges, we propose a novel framework termed the Data Generation and Cross-Dataset Learning Network (DGCDLNet), which makes the first attempt to simultaneously integrate the data generation strategy and cross-dataset learning mechanism in a unified manner. DGCDLNet contains two key modules: (1) Feature Reconstruction and Fusion module, which leverages complete eye movement signals to compensate for missing EEG data and constructs discriminative multimodal features via a dual-stream attention mechanism; and (2) Cross-Dataset Learning module, which jointly learns coarse-grained representations across datasets while incorporating fine-grained features from the target-task dataset to improve classification accuracy. Extensive experiments on SEED, SEED-IV, and SEED-V demonstrate that DGCDLNet consistently outperforms state-of-the-art multimodal fusion methods and achieves satisfactory performance under various EEG missing ratios. These results indicate the potential of DGCDLNet to advance EEG-based multimodal emotion recognition beyond controlled laboratory settings toward practical real-world applications.
Li et al. (Fri,) studied this question.