Develop DNABERT, a pre-trained bidirectional transformer framework designed to capture the syntax, semantics, and contextual representations of genomic DNA sequences.
Adapted the bidirectional encoder representations from transformers (BERT) architecture for genomic sequence processing.
Pre-trained the model on large-scale genome sequence data using nucleotide k-mer tokenization to capture bidirectional context.
Constructed a foundation model capable of learning universal representations of complex genomic sequence syntax.
Demonstrated effective feature transferability for downstream genomic prediction tasks and sequence interpretation.
Abstract
Supplementary data are available at Bioinformatics online.