The K-mer File Format: a standardized and compact disk representation of sets of k -mers

Dufresne, Yoann (ORCID:0000000209308920); Lemane, Teo (ORCID:0000000272103178); Marijon, Pierre; Peterlongo, Pierre (ORCID:0000000307766407); Rahman, Amatur; Kokot, Marek; Medvedev, Paul (ORCID:000000033143594X); Deorowicz, Sebastian; Chikhi, Rayan; Birol, ed., Inanc

doi:10.1093/bioinformatics/btac528

Citation Details

The K-mer File Format: a standardized and compact disk representation of sets of k -mers

Abstract SummaryBioinformatics applications increasingly rely on ad hoc disk storage of k-mer sets, e.g. for de Bruijn graphs or alignment indexes. Here, we introduce the K-mer File Format as a general lossless framework for storing and manipulating k-mer sets, realizing space savings of 3–5× compared to other formats, and bringing interoperability across tools. Availability and implementationFormat specification, C++/Rust API, tools: https://github.com/Kmer-File-Format/. Supplementary informationSupplementary data are available at Bioinformatics online. more »

Award ID(s):: 1453527 1931531

PAR ID:: 10371531

Author(s) / Creator(s):: Dufresne, Yoann; Lemane, Teo; Marijon, Pierre; Peterlongo, Pierre; Rahman, Amatur; Kokot, Marek; Medvedev, Paul; Deorowicz, Sebastian; Chikhi, Rayan; Birol, ed., Inanc

Publisher / Repository:: Oxford University Press

Date Published:: 2022-07-29

Journal Name:: Bioinformatics

Volume:: 38

Issue:: 18

ISSN:: 1367-4803

Format(s):: Medium: X Size: p. 4423-4425

Size(s):: p. 4423-4425

Sponsoring Org:: National Science Foundation

Journal Article:
https://doi.org/10.1093/bioinformatics/btac528

More Like this