Abstract The growth of RNA sequencing (RNA-seq) data accompanied by the development of novel scalable data analytic methods has revealed a deep understanding of the composition of bacterial transcriptomes. This new, first-biological-principles understanding has enabled a novel characterization of the function of the transcriptional regulatory network. Here, we present a single-strain wild-type transcriptomic knowledgebase for the model strain Escherichia coli MG1655. The associated transcriptomic compendium consists of 584 high-quality RNA-seq samples from wild-type E. coli MG1655 generated using a single protocol. These samples range over a wide condition space, including 45 carbon sources and 10 base media. Using independent component analysis, we decomposed the transcriptomic compendium to extract 115 independently modulated sets of genes (iModulons). We find that (i) iModulons explain 75% of variance in the dataset through knowledge enrichment; (ii) 67% of iModulons are associated with single/combined dominant regulators; (iii) iModulon activity profiles of samples can be utilized to elucidate patterns within the transcriptional regulatory network, such as differences in aerobicity; and (iv) the use of transcriptomic data derived from non-wild-type strains results in changes in iModulon gene membership, highlighting the malleability of the transcriptional regulatory network. Altogether, this knowledgebase serves as a resource for multi-scale knowledge mining for transcriptional regulation in E. coli MG1655.
Bajpe et al. (2026) studied this question.