Cotton fiber quality is shaped during boll development, boll opening, fluffing, and harvesting, but current assessment still relies largely on manual field inspection and postharvest laboratory testing. This limits timely harvest scheduling and plot-level quality management. To address this problem, we propose a self-supervised multimodal sensing framework for linking preharvest cotton boll status, environmental conditions, and postharvest fiber quality. First, the Cotton Boll Visual Phenotype Self-Supervised Encoding Module learns maturity-related visual representations by reconstructing masked image patches, so that boll cracking, lint exposure, and surface texture can be captured from unlabeled field images. Second, the Agricultural Sensor Temporal Masked Modeling Module reconstructs masked sensor observations to model temporal patterns in temperature, humidity, light, soil moisture, rainfall, and other environmental variables. Third, the Vision–Environment Cross-Modal Contrastive Fusion Module aligns image features with environmental features and produces a joint representation for downstream prediction. Field experiments were conducted using cotton boll images from different maturity and abnormal states, environmental sensor records, management information, and postharvest fiber quality measurements. The framework was evaluated for maturity classification, harvest-window recognition, and fiber quality prediction. The results showed that the proposed method performed consistently better than representative machine learning, single-modal deep learning, and multimodal fusion baselines, while few-shot and ablation experiments supported the value of self-supervised pretraining and multimodal fusion. These findings indicate that the proposed approach can provide useful information for preharvest cotton maturity assessment and harvest-quality management.
Luo et al. (Wed,) studied this question.