PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 12, 2026Behavior Research Methods3 citationsOpen Access

Automatic pose estimation in newborn infants: Lessons from the Baby Grow study

MSMohammad Saber SotoodehOOOri OssmyGDGeorgina Donati

Key Points

  • The aim is to assess the effectiveness of pose estimation algorithms for analyzing infant movements in home settings.
  • Analyzed home-recorded videos of newborns using mobile phones at 2, 4, and 8 weeks of age.
  • Extracted and annotated 2,640 frames to create a ground truth dataset.
  • Tested various pose estimation models including MediaPipe and RTMpose under different environmental conditions.
  • Evaluated models using metrics like object keypoint similarity (OKS) and percentage of correct keypoints (PCKh).
  • RTMpose achieved the highest overall accuracy in pose estimation.
  • MediaPipe demonstrated the fastest processing speed but lower accuracy compared to others.
  • Model performance varied significantly under different environmental factors such as lighting and clothing.
  • New models outperformed legacy systems, and context impacted accuracy substantially.

Abstract

Advances in computational techniques-particularly machine learning-have expanded opportunities to analyse early infant motor repertoires, especially in naturalistic settings. The aim of this study was to evaluate the strengths, limitations, and performance of state-of-the-art pose estimation algorithms in challenging, home-based video conditions. We analysed 22 videos recorded by parents using mobile phones from eight newborns in the Baby Grow study, at 2, 4, and 8 weeks of age. The videos varied in clothing (common onesie, babygrow, vest), background (grey, black, coloured), lighting (with/without shadows), and camera angles (top, front, bottom). From these, 2,640 frames were extracted and manually annotated to serve as ground truth. We tested demo versions of MediaPipe, OpenPose, PCT, RTMpose, Sapiens, and VitPose, and evaluated performance using object keypoint similarity (OKS), percentage of correct keypoints (PCKh), speed, and accuracy. RTMpose showed the highest overall accuracy, while MediaPipe had the fastest processing speed. However, when balancing speed and accuracy at ratios of 70:30, 50:50, and 30:70, MediaPipe's speed compensated for its lower accuracy, making it a strong candidate for practical applications. Model performance varied under different environmental conditions, with RTMpose, Sapiens, and VitPose being the most robust. As infant movement research increasingly shifts to real-world environments, selecting appropriate models and ensuring video quality are essential. Our findings show that (1) new models outperform legacy tools like OpenPose, and (2) video context and model selection significantly affect pose estimation accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sotoodeh et al. (2026) studied this question.

synapsesocial.com/papers/69b2577096eeacc4fcec608bhttps://doi.org/10.3758/s13428-026-02943-z
Ask AI
Helpful
Bookmark
Share
View Full Paper