Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Malusare, Aditya; Kothandaraman, Harish; Tamboli, Dipesh; Lanman, Nadia A; Aggarwal, Vaneet

doi:10.1093/bioadv/vbae117

Citation Details

Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Abstract SummaryThis article presents the Ensemble Nucleotide Byte-level Encoder-Decoder (ENBED) foundation model, analyzing DNA sequences at byte-level precision with an encoder–decoder Transformer architecture. ENBED uses a subquadratic implementation of attention to develop an efficient model capable of sequence-to-sequence transformations, generalizing previous genomic models with encoder-only or decoder-only architectures. We use Masked Language Modeling to pretrain the foundation model using reference genome sequences and apply it in the following downstream tasks: (i) identification of enhancers, promotors, and splice sites, (ii) recognition of sequences containing base call mismatches and insertion/deletion errors, an advantage over tokenization schemes involving multiple base pairs, which lose the ability to analyze with byte-level precision, (iii) identification of biological function annotations of genomic sequences, and (iv) generating mutations of the Influenza virus using the encoder–decoder architecture and validating them against real-world observations. In each of these tasks, we demonstrate significant improvement as compared to the existing state-of-the-art results. Availability and implementationThe source code used to develop and fine-tune the foundation model has been released on Github (https://github.itap.purdue.edu/Clan-labs/ENBED). more »

Award ID(s):: 2129097

PAR ID:: 10544572

Author(s) / Creator(s):: Malusare, Aditya; Kothandaraman, Harish; Tamboli, Dipesh; Lanman, Nadia A; Aggarwal, Vaneet

Editor(s):: Lengauer, Thomas

Publisher / Repository:: OUP

Date Published:: 2024-01-01

Journal Name:: Bioinformatics Advances

Volume:: 4

Issue:: 1

ISSN:: 2635-0041

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1093/bioadv/vbae117

More Like this