This study presents an Eco-Efficient Deep Learning (EEDL) framework for cultural digitization that integrates visual data and contextual cultural metadata within a lightweight multimodal architecture. The framework is designed to balance classification performance and computational efficiency in resource-constrained cultural heritage environments. A multimodal dataset comprising 3,000 cultural heritage records was used to evaluate the proposed approach across six artifact classes. Experimental results demonstrated that the EEDL framework achieved the highest classification performance among all evaluated models, obtaining an accuracy of 0.918 and an F1-score of 0.912 while reducing parameter count, computational complexity, inference latency, and energy consumption relative to the uncompressed baseline. Comparative experiments with MobileNetV3, EfficientNet-Lite, MobileViT, TinyViT, and reproduced state-of-the-art methods confirmed the effectiveness of the proposed framework. Ablation and multimodal fusion analyses further demonstrated that contextual metadata contributes substantially to classification performance and reduces ambiguity among visually similar artifact categories. Statistical validation across repeated runs confirmed the robustness and reproducibility of the reported results. The findings suggest that combining multimodal metadata integration with eco-efficient model design can improve both predictive performance and deployment efficiency for cultural digitization applications operating under computational and energy constraints.
Jin et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: