skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Sequence‐discrete species for prokaryotes and other microbes: A historical perspective and pending issues
Abstract Whether prokaryotes, and other microorganisms, form distinct clusters that can be recognized as species remains an issue of paramount theoretical as well as practical consequence in identifying, regulating, and communicating about these organisms. In the past decade, comparisons of thousands of genomes of isolates and hundreds of metagenomes have shown that prokaryotic diversity may be predominantly organized in such sequence‐discrete clusters, albeit organisms of intermediate relatedness between the identified clusters are also frequently found. Accumulating evidence suggests, however, that the latter “intermediate” organisms show enough ecological and/or functional distinctiveness to be considered different species. Notably, the area of discontinuity between clusters often—but not always—appears to be around 85%–95% genome‐average nucleotide identity, consistently among different taxa. More recent studies have revealed remarkably similar diversity patterns for viruses and microbial eukaryotes as well. This high consistency across taxa implies a specific mechanistic process that underlies the maintenance of the clusters. The underlying mechanism may be a substantial reduction in the efficiency of homologous recombination, which mediates (successful) horizontal gene transfer, around 95% nucleotide identity. Deviations from the 95% threshold (e.g., species showing lower intraspecies diversity) may be caused by ecological differentiation that imposes barriers to otherwise frequent gene transfer. While this hypothesis that clusters are driven by ecological differentiation coupled to recombination frequency (i.e., higher recombination within vs. between groups) is appealing, the supporting evidence remains anecdotal. The data needed to rigorously test the hypothesis toward advancing the species concept are also outlined.  more » « less
Award ID(s):
2129823 1759831
PAR ID:
10526683
Author(s) / Creator(s):
Publisher / Repository:
mLife journal
Date Published:
Journal Name:
mLife
Volume:
2
Issue:
4
ISSN:
2097-1699
Page Range / eLocation ID:
341 to 349
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Cooper, Vaughn S (Ed.)
    ABSTRACT Despite the importance of intra-species variants of viruses for causing disease and/or disrupting ecosystem functioning, there is no universally applicable standard to define these. A (natural) gap in whole-genome average nucleotide identity (ANI) values around 95% is commonly used to define species, especially for bacteriophages, but whether a similar gap exists within species that can be used to define intra-species units has not been evaluated yet. Whole-genome comparisons among members of 1,016 bacteriophage (Caudoviricetes) species revealed a region of low frequency of ANI values around 99.2%–99.8%, showing threefold or fewer pairs than expected for an even distribution. This second gap is prevalent in viruses infecting various cultured or uncultured hosts from a variety of environments, although a few exceptions to this pattern were also observed (3.7% of total species) and are likely attributed to cultivation biases or other factors. Similar results were observed for a limited set of eukaryotic viruses that are adequately sampled, including SARS-CoV-2, whose ANI-based clusters matched well with the WHO-defined variants of concern, indicating that our findings from bacteriophages might be more broadly applicable and the ANI-based clusters may represent functionally and/or ecologically distinct units. These units appear to be predominantly driven by (high) ecological cohesiveness coupled to either frequent recombination for bacteriophages or selection and clonal evolution for other viruses such as SARS-CoV-2, indicating that fundamentally different underlying mechanisms could lead to similar diversity patterns. Accordingly, we propose the ANI gap approach outlined above for defining viral intra-species units, for which we propose the term genomovars. IMPORTANCEViral species are composed of an ensemble of intra-species variants whose individual dynamics may have major implications for human and animal health and/or ecosystem functioning. However, the lack of universally accepted standards to define these intra-species variants has led researchers to use different approaches for this task, creating inconsistent intra-species units across different viral families and confusion in communication. By comparing hundreds of mostly bacteriophage genomes, we show that there is a widely distributed natural gap in whole-genome average nucleotide identity values in most, but not all, of these species that can be used to define intra-species units. Therefore, these results advance the molecular toolbox for tracking viral intra-species units and should facilitate future epidemiological and environmental studies. 
    more » « less
  2. ABSTRACT Yellow monkeyflowers (Mimulus guttatuscomplex, Phrymaceae) are a powerful system for studying ecological adaptation, reproductive variation, and genome evolution. To initiate pan‐genomics in this group, we present four chromosome‐scale assemblies and annotations of accessions spanning a broad evolutionary spectrum: two from a singleM. guttatuspopulation, one from the closely related selfing speciesM. nasutus, and one from a more divergent speciesM. tilingii. All assemblies are highly complete and resolve centromeric and repetitive regions. Comparative analyses reveal such extensive structural variation in repeat‐rich, gene‐poor regions that large portions of the genome are unalignable across accessions. As a result, thisMimuluspan‐genome is primarily informative in genic regions, underscoring limitations of resequencing approaches in such polymorphic taxa. We document gene presence–absence, investigate the recombination landscape using high‐resolution linkage data, and quantify nucleotide diversity. Surprisingly, pairwise differences at fourfold synonymous sites are exceptionally high—even in regions of very low recombination—reaching ~3.2% within a singleM. guttatuspopulation, ~7% within the interfertileM. guttatusspecies complex (approximately equal to SNP divergence between great apes and Old World monkeys), and ~7.4% between that complex and the reproductively isolatedM. tilingii. Genome‐wide patterns of nucleotide variation show little evidence of linked selection, and instead suggest that the concentration of genes (and likely selected sites) in high‐recombination regions may buffer diversity loss. These assemblies, annotations, and comparative analyses provide a robust genomic foundation forMimulusresearch and offer new insights into the interplay of recombination, structural variation, and molecular evolution in highly diverse plant genomes. 
    more » « less
  3. The influence of genetic drift on population dynamics during Pleistocene glacial cycles is well understood, but the role of selection in shaping patterns of genomic variation during these events is less explored. We resequenced whole genomes to investigate how demography and natural selection interact to generate the genomic landscapes of Downy and Hairy Woodpecker, species codistributed in previously glaciated North America. First, we explored the spatial and temporal patterns of genomic diversity produced by neutral evolution. Next, we tested (i) whether levels of nucleotide diversity along the genome are correlated with intrinsic genomic properties, such as recombination rate and gene density, and (ii) whether different demographic trajectories impacted the efficacy of selection. Our results revealed cycles of bottleneck and expansion, and genetic structure associated with glacial refugia. Nucleotide diversity varied widely along the genome, but this variation was highly correlated between the species, suggesting the presence of conserved genomic features. In both taxa, nucleotide diversity was positively correlated with recombination rate and negatively correlated with gene density, suggesting that linked selection played a role in reducing diversity. Despite strong fluctuations in effective population size, the maintenance of relatively large populations during glaciations may have facilitated selection. Under these conditions, we found evidence that the individual demographic trajectory of populations modulated linked selection, with purifying selection being more efficient in removing deleterious alleles in large populations. These results highlight that while genome-wide variation reflects the expected signature of demographic change during climatic perturbations, the interaction of multiple processes produces a predictable and highly heterogeneous genomic landscape. 
    more » « less
  4. Jouline, Igor B (Ed.)
    ABSTRACT Large-scale surveys of prokaryotic communities (metagenomes), as well as isolate genomes, have revealed that their diversity is predominantly organized in sequence-discrete units that may be equated to species. Specifically, genomes of the same species commonly show genome-aggregate average nucleotide identity (ANI) >95% among themselves and ANI <90% to members of other species, while genomes showing ANI 90%–95% are comparatively rare. However, it remains unclear if such “discontinuities” or gaps in ANI values can be observed within species and thus used to advance and standardize intra-species units. By analyzing 18,123 complete isolate genomes from 330 bacterial species with at least 10 genome representatives each and available long-read metagenomes, we show that another discontinuity exists between 99.2% and 99.8% (midpoint 99.5%) ANI in most of these species. The 99.5% ANI threshold is largely consistent with how sequence types have been defined in previous epidemiological studies but provides clusters with ~20% higher accuracy in terms of evolutionary and gene-content relatedness of the grouped genomes, while strains should be consequently defined at higher ANI values (>99.99% proposed). Collectively, our results should facilitate future micro-diversity studies across clinical or environmental settings because they provide a more natural definition of intra-species units of diversity. IMPORTANCEBacterial strains and clonal complexes are two cornerstone concepts for microbiology that remain loosely defined, which confuses communication and research. Here we identify a natural gap in genome sequence comparisons among isolate genomes of all well-sequenced species that has gone unnoticed so far and could be used to more accurately and precisely define these and related concepts compared to current methods. These findings advance the molecular toolbox for accurately delineating and following the important units of diversity within prokaryotic species and thus should greatly facilitate future epidemiological and micro-diversity studies across clinical and environmental settings. 
    more » « less
  5. Most cave-obligate species (troglobionts) have small ranges due to limited dispersal ability and the isolated nature of cave habitats. The troglobiontic linyphiid spiderPhanetta subterranea(Emerton, 1875), the only member of its genus, is a notable exception to this pattern; it has been reported from more counties and caves than any other troglobiont in North America. As many troglobionts exhibit significant genetic differentiation between populations over even small geographic distances, it has been hypothesized thatPhanettamay comprise multiple, genetically distinct lineages. To test this hypothesis, we examined genetic diversity inPhanettaacross its range at the mitochondrial cytochrome c oxidase subunit I gene for 47 individuals from 40 caves, distributed across seven states and 37 counties. We found limited genetic differentiation across the species’ range with haplotypes shared by individuals collected up to 600 km apart. Intraspecific nucleotide diversity was 0.006 +/- 0.005 (mean +/- SD), and the maximum genetic p-distance observed between any two individuals was 0.022. These values are within the typical range observed for other spider species. Thus, we found no evidence of cryptic genetic diversity inPhanetta. Our observation of low genetic diversity across such a broad distribution raises the question of how these troglobiontic spiders have managed to disperse so widely. 
    more » « less