High-content assays such as the Cell Painting assay have a lot of potential to advance drug discovery projects, by using a target agnostic approach and by using a more physiologically relevant cellular model system, compared to target-based screenings. This approach captures a more holistic view of the bioactivity of treatments by generating multi-dimensional readouts of cellular responses to drug-induced perturbations. However, the primary bottleneck to leverage this high-dimensional data lies in its computational analysis, due to the complexity and vast size of the generated image data. Deep learning is a powerful computational tool to analyse image data, as it can distil the microscopy images into biologically meaningful representations that enable accurate inference of treatment bioactivity. It is important to select suitable deep learning algorithms, and developing more task-aligned algorithms plays a crucial role in creating more accurate models. Label scarcity and technical variations are among the key challenges for developing accurate and robust bioactivity prediction models, which we aimed to tackle in this thesis. To mitigate the lack of bioactivity labels for effectively training deep learning models, we developed semi-supervised contrastive learning methods to learn meaningful representations of microscopy images. This approach is inherently label-efficient, enabling the learning of task-relevant representations, without relying on large, annotated datasets. We tailored the popular supervised contrastive learning approach to be more suitable for bioactivity prediction from Cell Painting data, by developing a multi-label semi-weakly supervised contrastive learning (MuSWSupCon) algorithm. Baseline evaluation showed considerable improvements for bioactivity prediction compared to established methods from the literature. We used experimental replicates as positives, which played an important role to optimize our model towards learning representations that can mitigate technical variations. The robustness of our approach towards technical variations was demonstrated on the multi-site EU-OS bioactives dataset, in which microscopy images were acquired from laboratories across Europe. We deployed our workflow on compounds with missing annotations, to demonstrate the capacity of our method to accurately predict the bioactivity of compounds. To identify potential hits, we used ensemble virtual screening to get more reliable predictions, which is important due to the high cost of experimental validation. We achieved a hit rate of ~13% on a validation assay for the identification of HDAC2 inhibitors and ~10% for AKT1 inhibitors, in which we only screened ~30 compounds for each assay, which shows the potential efficiency improvements our method can provide for the identification of new bioactive compounds. This thesis demonstrates accurate and robust bioactivity prediction from Cell Painting image data using semi-supervised contrastive models, which can potentially be used to increase the hit rate of multiple drug discovery projects and provide a more holistic understanding of treatment bioactivity.
David Bushiri Pwesombo (Thu,) studied this question.