Abstract Computer vision systems (CVS) have been developed using either top-down view 3D images or side-view 2D images to predict body weight (BW) and hip height (HH). However, 3D imaging systems are often costly compared with 2D imagery setups, and side-view cameras are often difficult to deploy under commercial farm conditions due to occlusion and variation in animal distance and posture. Herein, pose estimation using top-down view 2D images offers a promising approach to automatically extract keypoint-based features that describe body biometrics and could provide information correlated with BW and HH. Additionally, the same 2D images could be used to generate depth information, such as volume and height, using monocular depth estimation (MDE). Therefore, this study aimed to (1) develop predictive models for BW and HH based on features extracted from body pose keypoints in 2D infrared images and depth images generated using MDE, and (2) compare these models with those using features extracted from depth images collected by a 3D imaging system. A total of 395 top-down view videos from 94 beef-on-dairy crossbred cattle across four experimental blocks were collected using infrared and depth sensors. BW was recorded using an electronic scale, and HH was manually measured using a measuring stick. A pose estimation model identified seven anatomical landmarks (i.e. keypoints). The same 2D infrared images were converted into 3D images using zero-shot MDE, and a pipeline extracted features including volume, area, circularity, eccentricity, as well as back heights and widths. Depth images from a 3D imaging system were processed using the same pipeline. Random Forest (RF), Partial Least Squares Regression (PLS), and Support Vector Regression (SVM) models were evaluated using a leave-one-block-out cross-validation approach. The PLS model using Euclidean distances between the keypoints as features achieved R2 values of 0.90 with a Root Mean Square Error (RMSE) of 33.1 kg. Using MDE-derived depth features, PLS achieved an R2 of 0.95 with an RMSE of 24.2 kg. For HH, PLS using keypoints achieved an R2 of 0.77 with RMSE of 3.2 cm, and models using MDE-derived depth features showed similar performance. Our findings demonstrate that biometric features extracted from top-down 2D images or MDE-derived depth features enable comparable predictive performance for BW and HH, with models using features extracted from depth images collected using a 3D imaging system.
Menezes et al. (Wed,) studied this question.