User evaluation study demonstrates enhanced engagement and reduced cognitive load in museum visitors, indicating the viability of real-time multimodal exhibit adaptation.
Key Points
To develop an adaptive digital museum exhibit framework that decodes visitors' real-time cognitive states and dynamically aligns complex audio-visual heritage content.
Engineered a multimodal interaction network combining visual features from Swin Transformer V2, acoustic data from Wav2Vec 2.0, and dynamic gaze/gesture gating via a trilinear attention mechanism.
Benchmarked intent recognition performance on the ICH-MModal-2024 dataset against VideoMAE, MBT, and Perceiver architectures, followed by a user experience trial with 60 participants.
The proposed network achieved 94.2% accuracy in classifying active exploration, passive browsing, and confusion, maintaining a single-frame inference latency of 24 ms.
User testing with 60 participants revealed an increase in average exhibit stay duration to 195 s alongside significant reductions in user cognitive load.