DNA data storage has emerged as a promising data archive medium, which distributes data across many unordered DNA strands. Thus, retrieving large-scale data from these strands is a critical challenge, for these strands require indices and are prone to errors. Here, we propose an accompanying indexing and progressive recovery framework with specialized long composite ranging codes (LCRCs) for massive DNA data storage. Specifically, short fractions of the LCRC serve as accompanying indices for megabyte to petabyte data. Correlation with the short component codes enables rapid data recovery, while alignment to the LCRC facilitates reliable recovery under severe insertions/deletions. Simulations reveal that this scheme can be extended to the petabyte scale. Progressive error correction enables low-coverage recovery. Real-time read-by-read decoding recovered 12.87-megabyte files in ~20 min with 3.66× coverage at an error rate of ~4.9% using a nanopore sequencer. This framework provides a universal and practical strategy for DNA data storage.
Zhang et al. (Fri,) studied this question.