skip to main content


Title: Greengenes2 unifies microbial data in a single reference tree
Abstract

Studies using 16S rRNA and shotgun metagenomics typically yield different results, usually attributed to PCR amplification biases. We introduce Greengenes2, a reference tree that unifies genomic and 16S rRNA databases in a consistent, integrated resource. By inserting sequences into a whole-genome phylogeny, we show that 16S rRNA and shotgun metagenomic data generated from the same samples agree in principal coordinates space, taxonomy and phenotype effect size when analyzed with the same tree.

 
more » « less
Award ID(s):
1845967
NSF-PAR ID:
10435610
Author(s) / Creator(s):
; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; more » ; ; ; ; ; « less
Publisher / Repository:
Nature Publishing Group
Date Published:
Journal Name:
Nature Biotechnology
Volume:
42
Issue:
5
ISSN:
1087-0156
Format(s):
Medium: X Size: p. 715-718
Size(s):
p. 715-718
Sponsoring Org:
National Science Foundation
More Like this
  1. Kalendar, Ruslan (Ed.)

    The use of museum specimens for research in microbial evolutionary ecology remains an under-utilized investigative dimension with important potential. Despite this potential, there remain barriers in methodology and analysis to the wide-spread adoption of museum specimens for such studies. Here, we hypothesized that there would be significant differences in taxonomic prediction and related diversity among sample type (museum or fresh) and sequencing strategy (medium-depth shotgun metagenomic or 16S rRNA gene). We found dramatically higher predicted diversity from shotgun metagenomics when compared to 16S rRNA gene sequencing in museum and fresh samples, with this differential being larger in museum specimens. Broadly confirming these hypotheses, the highest diversity found in fresh samples was with shotgun sequencing using the Rep200 reference inclusive of viruses and microeukaryotes, followed by the WoL reference database. In museum-specimens, community diversity metrics also differed significantly between sequencing strategies, with the alpha-diversity ACE differential being significantly greater than the same comparisons made for fresh specimens. Beta diversity results were more variable, with significance dependent on reference databases used. Taken together, these findings demonstrate important differences in diversity results and prompt important considerations for future experiments and downstream analyses aiming to incorporate microbiome datasets from museum specimens.

     
    more » « less
  2. Abstract With advances in DNA sequencing and miniaturized molecular biology workflows, rapid and affordable sequencing of single-cell genomes has become a reality. Compared to 16S rRNA gene surveys and shotgun metagenomics, large-scale application of single-cell genomics to whole microbial communities provides an integrated snapshot of community composition and function, directly links mobile elements to their hosts, and enables analysis of population heterogeneity of the dominant community members. To that end, we sequenced nearly 500 single-cell genomes from a low diversity hot spring sediment sample from Dewar Creek, British Columbia, and compared this approach to 16S rRNA gene amplicon and shotgun metagenomics applied to the same sample. We found that the broad taxonomic profiles were similar across the three sequencing approaches, though several lineages were missing from the 16S rRNA gene amplicon dataset, likely the result of primer mismatches. At the functional level, we detected a large array of mobile genetic elements present in the single-cell genomes but absent from the corresponding same species metagenome-assembled genomes. Moreover, we performed a single-cell population genomic analysis of the three most abundant community members, revealing differences in population structure based on mutation and recombination profiles. While the average pairwise nucleotide identities were similar across the dominant species-level lineages, we observed differences in the extent of recombination between these dominant populations. Most intriguingly, the creek’s Hydrogenobacter sp . population appeared to be so recombinogenic that it more closely resembled a sexual species than a clonally evolving microbe. Together, this work demonstrates that a randomized single-cell approach can be useful for the exploration of previously uncultivated microbes from community composition to population structure. 
    more » « less
  3. We introduce Operational Genomic Unit (OGU), a metagenome analysis strategy that directly exploits sequence alignment hits to individual reference genomes as the minimum unit for assessing the diversity of microbial communities and their relevance to environmental factors. This approach is independent from taxonomic classification, granting the possibility of maximal resolution of community composition, and organizes features into an accurate hierarchy using a phylogenomic tree. The outputs are suitable for contemporary analytical protocols for community ecology, differential abundance and supervised learning while supporting phylogenetic methods, such as UniFrac and phylofactorization, that are seldomly applied to shotgun metagenomics despite being prevalent in 16S rRNA gene amplicon studies. As demonstrated in one synthetic and two real-world case studies, the OGU method produces biologically meaningful patterns from microbiome datasets. Such patterns further remain detectable at very low metagenomic sequencing depths. Compared with taxonomic unit-based analyses implemented in currently adopted metagenomics tools, and the analysis of 16S rRNA gene amplicon sequence variants, this method shows superiority in informing biologically relevant insights, including stronger correlation with body environment and host sex on the Human Microbiome Project dataset, and more accurate prediction of human age by the gut microbiomes in the Finnish population. We provide Woltka, a bioinformatics tool to implement this method, with full integration with the QIIME 2 package and the Qiita web platform, to facilitate OGU adoption in future metagenomics studies. Importance Shotgun metagenomics is a powerful, yet computationally challenging, technique compared to 16S rRNA gene amplicon sequencing for decoding the composition and structure of microbial communities. However, current analyses of metagenomic data are primarily based on taxonomic classification, which is limited in feature resolution compared to 16S rRNA amplicon sequence variant analysis. To solve these challenges, we introduce Operational Genomic Units (OGUs), which are the individual reference genomes derived from sequence alignment results, without further assigning them taxonomy. The OGU method advances current read-based metagenomics in two dimensions: (i) providing maximal resolution of community composition while (ii) permitting use of phylogeny-aware tools. Our analysis of real-world datasets shows several advantages over currently adopted metagenomic analysis methods and the finest-grained 16S rRNA analysis methods in predicting biological traits. We thus propose the adoption of OGU as standard practice in metagenomic studies. 
    more » « less
  4. ProkaryoticNostoc, one of the world's most conspicuous and widespread algal genera (similar to eukaryotic algae, plants, and animals) is known to support a microbiome that influences host ecological roles. Past taxonomic characterizations of surface microbiota (epimicrobiota) of free‐livingNostocsampled from freshwater systems employed 16S rRNA genes, typically amplicons. We compared taxa identified from 16S, 18S, 23S, and 28S rRNA gene sequences filtered from shotgun metagenomic sequence and used microscopy to illuminate epimicrobiota diversity forNostocsampled from a wetland in the northern Chilean Altiplano. Phylogenetic analysis and rRNA gene sequence abundance estimates indicated that the host was related toNostoc punctiformePCC 73102. Epimicrobiota were inferred to include 18 epicyanobacterial genera or uncultured taxa, six epieukaryotic algal genera, and 66 anoxygenic bacterial genera, all having average genomic coverage ≥90X. The epicyanobacteriaGeitlerinemia,Oscillatoria,Phormidium, and an uncultured taxon were detected only by 16S rRNA gene;GloeobacterandPseudanabaenawere detected using 16S and 23S; andPhormididesmis,Neosynechococcus,Symphothece,Aphanizomenon,Nodularia,Spirulina,Nodosilinea,Synechococcus,Cyanobium, andAnabaena(the latter corroborated by microscopy), plus two uncultured cyanobacterial taxa (JSC12, O77) were detected only by 23S rRNA gene sequences. Three chlamydomonad and two heterotrophic stramenopiles genera were inferred from 18S; the streptophyte green algaChaetosphaeridium globosumwas detected by microscopy and 28S rRNA genes, but not 18S rRNA genes. Overall, >60% of epimicrobial taxa were detected by markers other than 16S rRNA genes. Some algal taxa observed microscopically were not detected from sequence data. Results indicate that multiple taxonomic markers derived from metagenomic sequence data and microscopy increase epimicrobiota detection.

     
    more » « less
  5. Abstract

    Lake Lanier (Georgia, USA) is home to more than 11,000 microbial Operational Taxonomic Units (OTUs), many of which exhibit clear annual abundance patterns. To assess the dynamics of this microbial community, we collected time series data of 16S and 18S rRNA gene sequences, recovered from 29 planktonic shotgun metagenomic datasets. Based on these data, we constructed a dynamic mathematical model of bacterial interactions in the lake and used it to analyze changes in the abundances of OTUs. The model accounts for interactions among 14 sub-communities (SCs), which are composed of OTUs blooming at the same time of the year, and three environmental factors. It captures the seasonal variations in abundances of the SCs quite well. Simulation results suggest that changes in water temperature affect the various SCs differentially and that the timing of perturbations is critical. We compared the model results with published results from Lake Mendota (Wisconsin, USA). These comparative analyses between lakes in two very different geographical locations revealed substantially more cooperation and less competition among species in the warmer Lake Lanier than in Lake Mendota.

     
    more » « less