Robustness, interpretability, and generalization remain open challenges in fine-grained image classification, particularly in plant disease diagnosis where visual variability and domain shift can severely limit model reliability. This dissertation addresses these issues through two complementary directions: physics-inspired data augmentation and multimodal fusion of vision and language. Under Machine Learning Across Light, we introduce the Waveshift framework, which simulates light propagation in the Fourier domain to generate physically grounded augmentations. Waveshift 1.0 enhances subtle features via wavefront propagation, while Waveshift 2.0 incorporates aperture modulation to capture diffraction and attenuation in real imaging. An extension, Waveshift 2.5, applies region-aware augmentation using object detection to target leaves and symptomatic regions, improving robustness at both global and local scales. Under Machine Learning Across Language, we design a structured multimodal pipeline aligned with CLIP and a curated leaf-feature dictionary. By generating image-specific natural-language descriptions and converting them into semantic vectors, the system fuses semantics with visual embeddings to improve accuracy under domain shift and to link predictions to biologically meaningful traits. Together, these contributions establish a modular framework that advances generalizable and explainable AI for agriculture and other high-stakes imaging domains.
Gent Imeraj (Tue,) studied this question.