Integrative genomic and transcriptomic analysis of 61 patients identified significant mutations in genes such as FLNA, CST3, LGALS3, and HBA1, and an AI model predicted CVD with 95% accuracy.
Cohort (n=61)
Integrative genomic and transcriptomic analysis combined with AI/ML can identify key gene variants and predict cardiovascular disease with high accuracy.
Cardiovascular disease (CVD) is the leading cause of death in the United States and around the globe.1 Despite significant advancements in CVD diagnostics, prevention, and treatment, approximately half of the affected patients reportedly die within 5 years of receiving a diagnosis.2 The risk factors contributing to the development of CVD and response to therapy in an individual patient are highly variable. Evidence from the Framingham Heart Study suggests that CVD has a complex multifactorial aetiology including a genetic component.3 Genomics information, including high-quality sequenced DNA and RNA sequencing (RNA-seq) of transcribed genes, informs us of a CVD patient's inherent genetic makeup with the most comprehensive view of the genome.4 DNA-based gene variant detection when combined with RNA-seq-driven gene expression analysis has the potential to reveal novel and sensitive biomarkers and stratify CVD patient populations based on their disease risk.5, 6 The genetic variants predisposing to CVD span from rare and deleterious mutations that may be responsible for familial aggregation.6 Investigating differentially expressed genes (DEGs) and disease-causing variants can support finding the root cause of uncertainties in patient care.7 Understanding of the genetic basis of complex CVD can hamper genetic risk scoring which can now outperform traditional risk factors in risk prediction.7 Various genomics studies have been conducted recently to discover underlying genetic aetiology in CVD patients, especially those suspected of HF disease.7 Several CVD-related genes have been reported with significant mutations and expression differences among CVD patients.7 Application of intelligent and integrative data analysis approaches involving genomics and transcriptomics data will not only help understand the pathophysiology of CVD but also reduce heterogeneity in disease subtypes. To improve the deciphering of CVD mechanisms, here, we report our investigation of expression and variants among known genes that are responsible for the development of CVD.5, 6 Supporting this study, we developed an open-source, cross-platform, interactive, and user-friendly bioinformatics pipeline (GVViZ) for RNA-seq data pre-processing, and gene expression data analysis, annotation with relevant diseases, and heatmap visualization without requiring a strong computational background from the user.8 In addition, we have developed a new bioinformatics pipeline (JWES), for variant discovery and interpretation, and big data modelling and visualization.9 In this study, we conducted analysis of gene expression, disease-causing gene-variants, and associated phenotypes among CVD populations, with a focus on high-risk Heart Failure (HF).5, 6 We built a cohort of CVD patients, 40 male and 21 female individuals (n = 61), aged between 45 to 92, with self-described race (42 Whites, 7 Blacks or African Americans, 1 Asian, and 11 of unknown race). We collected blood samples, performed RNA-seq and gene expression analysis to generate transcriptomic profiles (Figure 1A). We performed in-depth gene expression analysis and annotation of RNA-seq data using GVViZ revealed regulation of genes known for HF (Figure 1B) and other CVDs (Figure 1C). Subsequent analyses were performed based on race and gender.5 Our analysis identified altered expression pathways of genes with gender differences in middle-aged to frail CVD patients.5 Next, we processed whole-genome sequencing (WGS) data using JWES and identified mutations among annotated genes for HF and other CVD patients (Figure 1D). We annotated these mutations to identify functional and nonfunctional mutations potentially associated with HF and other CVDs (Figure 1D). Mutation percentage in CVD genes was notably higher in HF patients compared to other CVD phenotypes (Figure 2A,B). In total, we detected 1 039 750 single nucleotide variants (SNV), insertion and deletion events. The most common mutation types in HF and CVD genes were intronic, flanking (5′ and 3′ flank) mutations.6 Mutations in these genes have been linked to aberrant expression in CVD. WGS allowed us to do in-depth analysis of CVD genes as RNA-seq cannot detect any of the variants located in noncoding DNA regions.6 We identified mutations among four genes with altered expression and having significant mutations associated with HF and other CVDs (FLNA, CST3, LGALS3, and HBA1). Notables are the missense mutations in FLNA, CT3, and LGALS3. Next, we performed splice mutation analysis (Figure 2C,D). Among most of HF and other CVD genes, we observed high frequencies in intron and 5’UTR.6 We applied the Jensen-Shannon Divergence Based Method (JS-MA) to measure the proportion distributions of genes associated with HF and other CVDs (Figure 2E,F). We observed ADRB2, NPPB, ADRB1, ADB, and NPPC genes with highest JSD scores of 0.517, 0.511, 0.498, 0.482, and 0.474 respectively, and CORIN, PLN, NR3C2, TNF, CRP, and KNG1 among with moderate genes with JSD scores of 0.438, 0.430, 0.415, 0.413, 0.406, and 0.395, respectively, for HF (Figure 2E). However, we identified HBA1 with the highest JSD score of 0.654, FADD as the second best with a JSD score of 0.473, and GMLN, FGF23, ENO2, DDX41, TAC1, CALD1, CD40LG, CD34, FLNA, SLC2A1, ZBTB805, FGF2, TEK, and GJB6 genes with moderate JSD scores for other CVDs with respective scores of 0.344, 0.333, 0.325, 0.300, 0.278, 0.277, 0.276, 0.271, 0.263, 0.256, 0.251, 0.250, 0.233, and 0.222. During our JSD analysis, we found HBA1, FADD, ADRB2, NPPB, ADRB1, ADB, and NPPC genes with the greatest variance. We found that HBA1 with high relevance and linked to multiple CVDs.6 FADD, ADRB1, ADRB2, NPPB, ADB, and NPPC are among genes that we found to have high variance to HF and other CVDs. Extending our study further, we implemented artificial intelligence (AI) and machine learning (ML) techniques to investigate genes associated with CVDs and predict disease with high accuracy.10 Visible data clusters were observed for the genes highly correlated, downregulated and with altered expression in CVD patients. During our predictive analysis, it was observed that age and gender appeared to have a high correlation in HF and other CVDs.10 Our model was able to correctly classify individuals as CVD patients and predict CVD with 95% accuracy. We observed and reported overlapping in significant results produced in gene expression, variant, phenotypic, and predictive analyses, which include genes associated with HF, and other CVDs.10
Zeeshan Ahmed (Fri,) conducted a cohort in Cardiovascular disease and Heart Failure (n=61). Genetic variants and altered gene expression was evaluated on Prediction of cardiovascular disease using AI/ML techniques based on gene expression and variant data. Integrative genomic and transcriptomic analysis of 61 patients identified significant mutations in genes such as FLNA, CST3, LGALS3, and HBA1, and an AI model predicted CVD with 95% accuracy.