PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026PLOS Digital Health2 citationsOpen Access

AVPENet: Pain estimation from audio-visual fusion of non-speech sounds

View Full Paper
SNSami NaoualiOOOussama El Othmani

Key Points

  • The aim is to develop an objective pain assessment method for non-verbal patients using audio and visual data.
  • Multimodal deep learning framework combining audio cues and facial expressions.
  • Cross-modal attention-based fusion network integrating spectrogram audio embeddings with facial action unit features.
  • Utilized ResNet-based audio encoder and a convolutional neural network for facial expression analysis.
  • Achieved a mean absolute error of 0.89 on a 0–10 pain scale, outperforming audio-only and visual-only approaches.
  • Demonstrated robust generalization with mean absolute errors of 0.94 for neonates and 0.91 for adults.
  • Achieved 81.4% accuracy for three-class pain categorization and a Pearson correlation coefficient of 0.89.

Abstract

Pain assessment in non-verbal patients, including neonates and unconscious adults, remains a critical challenge in clinical practice. Current pain scales rely heavily on observer interpretation and may lack objectivity, introducing significant inter-rater variability. We propose a novel multimodal deep learning framework that estimates continuous pain intensity by fusing non-speech audio cues with facial expressions. Our approach addresses the critical need for objective pain assessment in vulnerable populations unable to self-report. We developed a cross-modal attention-based fusion network combining spectrogram-derived audio embeddings with facial action unit features. The model was trained and validated on 3,247 audio-visual recordings from 428 subjects, including 215 neonates and 213 adults, across three distinct pain intensity levels. We employed a ResNet-based audio encoder for mel-spectrogram processing and a facial landmark convolutional neural network for expression analysis, integrated through a transformer-based fusion module that learns complementary relationships between modalities. Our model achieved a mean absolute error of 0.89 on a 0–10 pain scale, significantly outperforming audio-only approaches (mean absolute error 1.47, 39% improvement) and visual-only baselines (mean absolute error 1.23, 28% improvement). Cross-age group validation demonstrated robust generalization with mean absolute errors of 0.94 for neonates and 0.91 for adults. The model maintained a Pearson correlation coefficient of 0.89 with ground truth annotations and achieved 81.4% accuracy for three-class pain categorization. Audio-visual fusion significantly enhances pain estimation accuracy across diverse age groups and clinical scenarios. This approach offers substantial potential for objective, automated pain monitoring in clinical settings, particularly for vulnerable populations unable to self-report pain.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Naouali et al. (2026) studied this question.

synapsesocial.com/papers/69c8c3cede0f0f753b39ed61https://doi.org/10.1371/journal.pdig.0001301
Ask AI
Helpful
Bookmark
Share
View Full Paper