In this study, a deep learning-based multimodal framework is presented for forest fire detection using RGB images, which synthetically generates night-vision-like, white-hot, and green-hot pseudo-thermal representations. The synthetic modalities are derived directly from RGB data and integrated into a hardware-independent multimodal learning pipeline to increase visual diversity without relying on additional sensing hardware. Each modality is processed using an ImageNet-pretrained convolutional backbone, and modality-specific feature vectors are combined through feature-level concatenation before classification. The proposed framework was evaluated using multiple backbone architectures, including ResNet18, EfficientNet-B0, and DenseNet121, which were assessed independently under a unified experimental protocol. Experiments were conducted on two datasets with substantially different scales and characteristics: the FLAME dataset (39,375 images, binary classification) and the FireStage dataset (791 images, three-class classification). For both datasets, stratified 80–20% training–validation splits were employed, and online stochastic data augmentation was applied exclusively to the training sets. On the FLAME dataset, the proposed framework achieved consistently high performance across different backbone and modality configurations. The best-performing models reached an accuracy of 99.66%, precision of 99.80%, recall of 99.66%, F1-score of 99.73%, and ROC AUC value of 0.9998. On the more challenging FireStage dataset, the framework demonstrated stable performance despite limited data availability, achieving an accuracy of 93.71% for RGB-only configurations and up to 93.08% for selected multimodal combinations, while macro-averaged F1-scores exceeded 0.92, and ROC AUC values reached up to 0.9919. Per-class analysis further indicates that early-stage fire (Start Fire) patterns can be discriminated, achieving ROC AUC values above 0.96, depending on the backbone and modality combination. Overall, the results suggest that synthetic-modality-based multimodal learning can provide competitive performance for both large-scale and data-limited fire detection scenarios, offering a flexible and hardware-independent alternative for forest fire monitoring applications.
Taşar et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: