PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 16, 2026IEEE Transactions on Computational Biology and Bioinformatics0 citationsOpen Access

Biologically Constrained DNA Encoding with Triplet Networks for Similarity Image Retrieval

View Full Paper
TKTatsuya KoikeHAHiromitsu AwanoTSTakashi Sato

Key Points

  • The aim is to develop a DNA encoding framework that enhances image retrieval accuracy and training efficiency using deep metric learning.
  • Proposed a training framework for DNA encoding incorporating deep metric learning techniques.
  • Introduced loss functions enforcing biological constraints like homopolymer length and GC content.
  • Conducted simulations on CIFAR-10 and CIFAR-100 datasets to evaluate performance.
  • Achieved classification accuracy comparable to CNN-based methods.
  • Demonstrated a 20-fold reduction in training time compared to existing methods.
  • Maintained GC content within optimal parameters and controlled homopolymer length effectively.

Abstract

As the volume of digital data continues to grow exponentially, DNA has emerged as a promising medium for long-term data storage due to its high density and durability. For enabling data retrieval via DNA's biochemical reactions, the encoding strategy plays a critical role. This paper proposes a training framework for a DNA encoder that improves both accuracy and training efficiency in content-based image retrieval by incorporating deep metric learning. In addition, we introduce loss functions that enforce biological constraints, specifically homopolymer length and GC content, thereby improving the biochemical stability of the generated DNA sequences. To evaluate the effectiveness of the proposed method, we conduct quantitative assessments based on image classification performance. Simulations on the CIFAR- 10 and CIFAR-100 datasets demonstrate that our method achieves classification accuracy comparable to CNN-based baselines and a 20- fold speedup over the training time of the existing method. Moreover, the generated DNA sequences enable strict control of homopolymer length and maintain GC content within the optimal 40-60 improving biological feasibility compared to baseline methods. The source code is publicly available at GitHub.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Koike et al. (2026) studied this question.

synapsesocial.com/papers/69b79dce8166e15b153ab106https://doi.org/10.1109/tcbbio.2026.3673740
Ask AI
Helpful
Bookmark
Share
View Full Paper