Key points are not available for this paper at this time.
Long non-coding RNAs (lncRNAs), transcripts longer than 200 nucleotides with limited protein-coding potential, are key regulators of gene expression, yet their evolutionary conservation remains poorly understood due to rapid sequence divergence. We present a flexible workflow for cross-species inference of lncRNA orthology combining two synteny-based approaches with multi-species genome alignment-derived sequence conservation. The workflow relies on standardized genome annotations and one-to-one orthologous protein-coding gene relationships, retrieved here from Ensembl resources. Applied to 13 vertebrate species spanning zebrafish, birds, and mammals, chosen to capture both broad phylogenetic distances and heterogeneous genome annotation quality, and using human (18 859 lncRNAs) as reference, the approach identified on average ∼200 putative orthologs per species under stringent criteria and up to ∼5000 under relaxed criteria. At the multi-species levels, >450 human lncRNAs were conserved in at least two species under stringent conditions, and over 10 000 in at least five species under relaxed criteria. Functional downstream analyses further revealed partial conservation of expression across 17 homologous tissues between human and chicken, as well as conserved short sequence motifs detected with LncLOOM. Together, this study provides both an adaptable workflow and a multi-species atlas to investigate lncRNA conservation and prioritize candidates for functional studies.
Degalez et al. (Fri,) studied this question.