Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 11:00 PM ET on Thursday, August 13 until 12:00 AM ET on Friday, August 14 due to maintenance. We apologize for the inconvenience.


Title: K-mer-based Approaches to Bridging Pangenomics and Population Genetics
Abstract Many commonly studied species now have more than one chromosome-scale genome assembly, revealing a large amount of genetic diversity previously missed by approaches that map short reads to a single reference. However, many species still lack multiple reference genomes and correctly aligning references to build pangenomes can be challenging for many species, limiting our ability to study this missing genomic variation in population genetics. Here, we argue that k-mers are a very useful but underutilized tool for bridging the reference-focused paradigms of population genetics with the reference-free paradigms of pangenomics. We review current literature on the uses of k-mers for performing three core components of most population genetics analyses: identifying, measuring, and explaining patterns of genetic variation. We also demonstrate how different k-mer-based measures of genetic variation behave in population genetic simulations according to the choice of k, depth of sequencing coverage, and degree of data compression. Overall, we find that k-mer-based measures of genetic diversity scale consistently with pairwise nucleotide diversity (π) up to values of about π=0.025 (R2=0.97) for neutrally evolving populations. For populations with even more variation, using shorter k-mers will maintain the scalability up to at least π=0.1. Furthermore, in our simulated populations, k-mer dissimilarity values can be reliably approximated from counting bloom filters, highlighting a potential avenue to decreasing the memory burden of k-mer-based genomic dissimilarity analyses. For future studies, there is a great opportunity to further develop methods to identifying selected loci using k-mers.  more » « less
Award ID(s):
1934384
PAR ID:
10660851
Author(s) / Creator(s):
; ; ;
Editor(s):
Harris, Kelley
Publisher / Repository:
Oxford
Date Published:
Journal Name:
Molecular Biology and Evolution
Volume:
42
Issue:
3
ISSN:
0737-4038
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract A key prediction of neutral theory is that the level of genetic diversity in a population should scale with population size. However, as was noted by Richard Lewontin in 1974 and reaffirmed by later studies, the slope of the population size-diversity relationship in nature is much weaker than expected under neutral theory. We hypothesize that one contributor to this paradox is that current methods relying on single nucleotide polymorphisms (SNPs) called from aligning short reads to a reference genome underestimate levels of genetic diversity in many species. As a first step to testing this idea, we calculated nucleotide diversity (π) and k-mer-based metrics of genetic diversity across 112 plant species, amounting to over 205 terabases of DNA sequencing data from 27,488 individuals. After excluding 14 species with low coverage or no variant sites called, we compared how different diversity metrics correlated with proxies of population size that account for both range size and population density variation across species. We found that our population size proxies scaled anywhere from about 3 to over 20 times faster with k-mer diversity than nucleotide diversity after adjusting for evolutionary history, mating system, life cycle habit, cultivation status, and invasiveness. The relationship between k-mer diversity and population size proxies also remains significant after correcting for genome size, whereas the analogous relationship for nucleotide diversity does not. These results are consistent with the possibility that variation not captured by common SNP-based analyses explains part of Lewontin’s paradox in plants, but larger scale pangenomic studies are needed to definitively address this question. 
    more » « less
  2. Sil, Anita (Ed.)
    Aspergillus fumigatus is a deadly agent of human fungal disease where virulence heterogeneity is thought to be at least partially structured by genetic variation between strains. While population genomic analyses based on reference genome alignments offer valuable insights into how gene variants are distributed across populations, these approaches fail to capture intraspecific variation in genes absent from the reference genome. Pan-genomic analyses based on de novo assemblies offer a promising alternative to reference-based genomics with the potential to address the full genetic repertoire of a species. Here, we evaluate 260 genome sequences of A . fumigatus including 62 newly sequenced strains, using a combination of population genomics, phylogenomics, and pan-genomics. Our results offer a high-resolution assessment of population structure and recombination frequency, phylogenetically structured gene presence–absence variation, evidence for metabolic specificity, and the distribution of putative antifungal resistance genes. Although A . fumigatus disperses primarily via asexual conidia, we identified extraordinarily high levels of recombination with the lowest linkage disequilibrium decay value reported for any fungal species to date. We provide evidence for 3 primary populations of A . fumigatus , with recombination occurring only rarely between populations and often within them. These 3 populations are structured by both gene variation and distinct patterns of gene presence–absence with unique suites of accessory genes present exclusively in each clade. Accessory genes displayed functional enrichment for nitrogen and carbohydrate metabolism suggesting that populations may be stratified by environmental niche specialization. Similarly, the distribution of antifungal resistance genes and resistance alleles were often structured by phylogeny. Altogether, the pan-genome of A . fumigatus represents one of the largest fungal pan-genomes reported to date including many genes unrepresented in the Af293 reference genome. These results highlight the inadequacy of relying on a single-reference genome-based approach for evaluating intraspecific variation and the power of combined genomic approaches to elucidate population structure, genetic diversity, and putative ecological drivers of clinically relevant fungi. 
    more » « less
  3. The exceptionally large population size and cosmopolitan biogeographic distribution that distinguish many – but not all – marine zooplankton species generate similarly exceptional patterns of population genetic and genomic diversity and structure. The phylogenetic diversity of zooplankton has slowed the application of population genomic approaches, due to lack of genomic resources for closelyrelated species and diversity of genomic architecture, including highly-replicated genomes of many crustaceans. Use of numerous genomic markers, especially single nucleotide polymorphisms (SNPs), is transforming our ability to analyze population genetics and connectivity of marine zooplankton, and providing new understanding and different answers than earlier analyses, which typically used mitochondrial DNA and microsatellite markers. Population genomic approaches have confirmed that, despite high dispersal potential, many zooplankton species exhibit genetic structuring among geographic populations, especially at large ocean-basin scales, and have revealed patterns and pathways of population connectivity that do not always track ocean circulation. Genomic and transcriptomic resources are critically needed to allow further examination of micro-evolution and local adaptation, including identification of genes that show evidence of selection. These new tools will also enable further examination of the significance of small-scale genetic heterogeneity of marine zooplankton, to discriminate genetic “noise” in large and patchy populations from local adaptation to environmental conditions and change. 
    more » « less
  4. Lipka, A (Ed.)
    Abstract High-throughput sequencing-based methods for bulked segregant analysis (BSA) allow for the rapid identification of genetic markers associated with traits of interest. BSA studies have successfully identified qualitative (binary) and quantitative trait loci (QTLs) using QTL mapping. However, most require population structures that fit the models available and a reference genome. Instead, high-throughput short-read sequencing can be combined with BSA of k-mers (BSA-k-mer) to map traits that appear refractory to standard approaches. This method can be applied to any organism and is particularly useful for species with genomes diverged from the closest sequenced genome. It is also instrumental when dealing with highly heterozygous and potentially polyploid genomes without phased haplotype assemblies and for which a single haplotype can control a trait. Finally, it is flexible in terms of population structure. Here, we apply the BSA-k-mer method for the rapid identification of candidate regions related to seed spot and seed size in diploid potato. Using a mixture of F1 and F2 individuals from a cross between 2 highly heterozygous parents, candidate sequences were identified for each trait using the BSA-k-mer approach. Using parental reads, we were able to determine the parental origin of the loci. Finally, we mapped the identified k-mers to a closely related potato genome to validate the method and determine the genomic loci underlying these sequences. The location identified for the seed spot matches with previously identified loci associated with pigmentation in potato. The loci associated with seed size are novel. Both loci are relevant in future breeding toward true seeds in potato. 
    more » « less
  5. ABSTRACT Analysing large population genomic datasets requires an interdisciplinary skillset. Beyond a knowledge base in genetics and population biology, population genomic analyses involve computer science and statistics, representing a barrier for researchers without experience in those fields. PopGenHelpR seeks to lower this barrier by enabling researchers to perform population genomic analyses and generate near‐publication quality figures in a streamlined and informed fashion. PopGenHelpR allows users to estimate genetic diversity within populations as well as differentiation among populations and individuals from single nucleotide polymorphism data. PopGenHelpR includes commonly used measures such as observed heterozygosity andFST. PopGenHelpR also provides five previously unavailable measures of heterozygosity, including the proportion of heterozygous loci and homozygosity by locus. Additionally, PopGenHelpR integrates widely used visualization tools in population genomics that normally require additional software packages, such as ancestry bar charts, piechart maps, and genetic differentiation heatmaps. Moreover, PopGenHelpR is minimally dependent on other R packages, reducing its chance of being removed from public networks. The PopGenHelpR website also provides tutorials and educational resources. Altogether, PopGenHelpR provides resources for informed analyses and effective visualization, making PopGenHelpR a valuable tool for many researchers. PopGenHelpR is available on the Comprehensive R Archive Network and GitHub (https://github.com/kfarleigh/PopGenHelpR). 
    more » « less