The rapid development of digital media art (DMA) has created a demand for intelligent systems that can understand and respond to users via numerous sensory channels. However, current systems frequently rely on single-modal inputs, which limit their ability to understand user states and adjust experiences accordingly. This research seeks to create an intelligent DMA system architecture that combines multimodal perception to enhance user experience dynamically. A novel Dolphin Swarm Optimized Deep Neural Network (DSODeepNet) is proposed, where DSO dynamically tunes hyperparameters of the DeepNet to maximize multimodal emotion recognition performance, leading to improved user experience adaptation in DMA systems. The system combines physiological inputs (heart rate variability (HRV), electrodermal activity (EDA)), visual data (facial expressions), and audio cues (speech tone). A customized dataset was created from 250 individuals who interacted with digital artworks, recording synchronized biometric, visual, and audio data. Pre-processing techniques include signal denoising with band-pass filters and Wiener filtering. For feature extraction, Mel Frequency Cepstral Coefficients (MFCCs) and ResNet50 are used. Deep Canonical Correlation Analysis (DCCA) is recommended for effective fusion, alignment, and integration of information from multiple modalities. Experimental results showed that the fused multimodal model outperformed the other models in terms of emotion recognition accuracy (0.925). Emotional involvement, interpersonal satisfaction, and perceived system responsiveness all improved significantly, according to user experience evaluations. Overall, the proposed architecture effectively improves the sensitivity and adaptability of DMA systems through multimodal perception fusion, indicating a promising route for producing more immersive and individualized art experiences.
Rui Wang (Sun,) studied this question.