• No prior longitudinal tomato leaf-counting across growth phases under greenhouse and agrivoltaic conditions. • First tomato-focused automated leaf-counting pipeline using weekly image data from 36 crops over five weeks. • Two-stage deep learning: MobileNetV2-based CNN regression–leaf count prediction and U-Net segmentation–refinement. • Achieved accurate predictions within ± 3 leaves and R 2 between 0.7 and 0.9. • Pipeline ensures computational efficiency + visual validation, enabling practical greenhouse monitoring and edge-AI deployment. Research is limited in longitudinal leaf counting across different growth phases of tomatoes, particularly for agrivoltaics where variations in light and surrounding environmental conditions can influence plant structure and growth. Practical greenhouse monitoring requires a lightweight method for consistent, automated leaf counting that can adapt to noisy backgrounds, overlapping leaves, and variations in canopy structure. This study presents the first tomato-focused leaf counting pipeline developed and validated using weekly image data collected under greenhouse and agrivoltaic greenhouse conditions. The dataset includes 36 tomato crops over five weeks, with ground truth counts and pseudo-labeled counts where manual counting was infeasible. The pipeline is based on a two-stage deep learning approach: 1) convolutional neural network (CNN) regression model based on MobileNetV2 is trained directly to predict the leaf counts from top-view crop images and 2) U-Net segmentation model based on a subset of manually annotated images enhances predictions, especially in cases of occluded and overlapped leaves. Across the test datasets, the CNN regression model consistently produced good predictions within a ± 3 margin for actual leaf manual counts. The model’s performance was evaluated using the coefficient of determination (R2) which ranges between 0.7 and 0.9, which is comparable or exceeds the performance reported in recent leaf-counting studies. The CNN regression stage efficiently learned structural and textural features from limited data, while the U-Net refinement enhanced spatial understanding which provides segmentation-based visual validation. This combination ensured both computational efficiency and model transparency which is important for operational use in greenhouse monitoring and edge-AI deployment.
Quamer et al. (Sat,) studied this question.