skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: A genome sequence for the threatened whitebark pine
Abstract Whitebark pine (WBP, Pinus albicaulis) is a white pine of subalpine regions in the Western contiguous United States and Canada. WBP has become critically threatened throughout a significant part of its natural range due to mortality from the introduced fungal pathogen white pine blister rust (WPBR, Cronartium ribicola) and additional threats from mountain pine beetle (Dendroctonus ponderosae), wildfire, and maladaptation due to changing climate. Vast acreages of WBP have suffered nearly complete mortality. Genomic technologies can contribute to a faster, more cost-effective approach to the traditional practices of identifying disease-resistant, climate-adapted seed sources for restoration. With deep-coverage Illumina short reads of haploid megagametophyte tissue and Oxford Nanopore long reads of diploid needle tissue, followed by a hybrid, multistep assembly approach, we produced a final assembly containing 27.6 Gb of sequence in 92,740 contigs (N50 537,007 bp) and 34,716 scaffolds (N50 2.0 Gb). Approximately 87.2% (24.0 Gb) of total sequence was placed on the 12 WBP chromosomes. Annotation yielded 25,362 protein-coding genes, and over 77% of the genome was characterized as repeats. WBP has demonstrated the greatest variation in resistance to WPBR among the North American white pines. Candidate genes for quantitative resistance include disease resistance genes known as nucleotide-binding leucine-rich repeat receptors (NLRs). A combination of protein domain alignments and direct genome scanning was employed to fully describe the 3 subclasses of NLRs. Our high-quality reference sequence and annotation provide a marked improvement in NLR identification compared to previous assessments that leveraged de novo-assembled transcriptomes.  more » « less
Award ID(s):
1943371
PAR ID:
10505508
Author(s) / Creator(s):
; ; ; ; ; ; ; ; ; ; ; ; ; ;
Publisher / Repository:
Oxford University Press
Date Published:
Journal Name:
G3: Genes, Genomes, Genetics
Volume:
14
Issue:
5
ISSN:
2160-1836
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. null (Ed.)
    Abstract Taro (Colocasia esculenta) is a food staple widely cultivated in the humid tropics of Asia, Africa, Pacific and the Caribbean. One of the greatest threats to taro production is Taro Leaf Blight caused by the oomycete pathogen Phytophthora colocasiae. Here we describe a de novo taro genome assembly and use it to analyze sequence data from a Taro Leaf Blight resistant mapping population. The genome was assembled from linked-read sequences (10x Genomics; ∼60x coverage) and gap-filled and scaffolded with contigs assembled from Oxford Nanopore Technology long-reads and linkage map results. The haploid assembly was 2.45 Gb total, with a maximum contig length of 38 Mb and scaffold N50 of 317,420 bp. A comparison of family-level (Araceae) genome features reveals the repeat content of taro to be 82%, >3.5x greater than in great duckweed (Spirodela polyrhiza), 23%. Both genomes recovered a similar percent of Benchmarking Universal Single-copy Orthologs, 80% and 84%, based on a 3,236 gene database for monocot plants. A greater number of nucleotide-binding leucine-rich repeat disease resistance genes were present in genomes of taro than the duckweed, ∼391 vs. ∼70 (∼182 and ∼46 complete). The mapping population data revealed 16 major linkage groups with 520 markers, and 10 quantitative trait loci (QTL) significantly associated with Taro Leaf Blight disease resistance. The genome sequence of taro enhances our understanding of resistance to TLB, and provides markers that may accelerate breeding programs. This genome project may provide a template for developing genomic resources in other understudied plant species. 
    more » « less
  2. We sequenced the genome of the North American groundhog, Marmota monax , also known as the woodchuck. Our sequencing strategy included a combination of short, high-quality Illumina reads plus long reads generated by both Pacific Biosciences and Oxford Nanopore instruments. Assembly of the combined data produced a genome of 2.74 Gbp in total length, with an N50 contig size of 1,094,236 bp. To annotate the genome, we mapped the genes from another M. monax genome and from the closely related Alpine marmot, Marmota marmota , onto our assembly, resulting in 20,559 annotated protein-coding genes and 28,135 transcripts. The genome assembly and annotation are available in GenBank under BioProject PRJNA587092 . 
    more » « less
  3. Abstract Sequencing, assembly, and annotation of the 26.5 Gbp hexaploid genome of coast redwood (Sequoia sempervirens) was completed leading toward discovery of genes related to climate adaptation and investigation of the origin of the hexaploid genome. Deep-coverage short-read Illumina sequencing data from haploid tissue from a single seed were combined with long-read Oxford Nanopore Technologies sequencing data from diploid needle tissue to create an initial assembly, which was then scaffolded using proximity ligation data to produce a highly contiguous final assembly, SESE 2.1, with a scaffold N50 size of 44.9 Mbp. The assembly included several scaffolds that span entire chromosome arms, confirmed by the presence of telomere and centromere sequences on the ends of the scaffolds. The structural annotation produced 118,906 genes with 113 containing introns that exceed 500 Kbp in length and one reaching 2 Mb. Nearly 19 Gbp of the genome represented repetitive content with the vast majority characterized as long terminal repeats, with a 2.9:1 ratio of Copia to Gypsy elements that may aid in gene expression control. Comparison of coast redwood to other conifers revealed species-specific expansions for a plethora of abiotic and biotic stress response genes, including those involved in fungal disease resistance, detoxification, and physical injury/structural remodeling and others supporting flavonoid biosynthesis. Analysis of multiple genes that exist in triplicate in coast redwood but only once in its diploid relative, giant sequoia, supports a previous hypothesis that the hexaploidy is the result of autopolyploidy rather than any hybridizations with separate but closely related conifer species. 
    more » « less
  4. The western painted turtle, Chrysemys picta bellii, has the greatest tolerance to anoxia of any tetrapod studied to date. These turtles reside in the northern United States and southern Canada, and survive months of anoxia while submerged in ice-locked ponds and bogs. Reference genomes provide an important resource for elucidating the molecular bases for such unique physiological traits. An initial reference genome for this species was published in 2013, but the assembly is highly fragmented which poses several limitations for downstream analyses and biological interpretation. In this study, we created a new and improved assembly by combining PacBio HiFi, 10x Genomics Chromium, Hi-C sequence data and BioNano optical mapping derived from a single individual to generate a new haplotype-resolved chromosome-level assembly for C. picta bellii, called SLU_Cpb5.0. The genome size of the primary assembly is 2.372 Gb with a scaffold N50 of 133.6 Mb, which is a 6.5-fold improvement over the existing assembly. Genome annotation of SLU_Cpb5.0 revealed 12,242 novel genes compared to previous assemblies. Our PacBio Iso-Seq RNA sequencing data for twelve tissues unraveled over 100,000 novel transcript isoforms and 4,325 novel genes that were not annotated by standard NCBI pipeline. We also observed distinct patterns of tissue-specific isoform expression, creating a robust foundation for future characterization of the functions of these genes. The improved genome assembly and annotation will facilitate comparative genomics studies to better understand the genetic basis of C.picta bellii's extreme physiological adaptations and other aspects of its biology. 
    more » « less
  5. Mank, Judith (Ed.)
    Abstract Urosaurus nigricaudus is a phrynosomatid lizard endemic to the Baja California Peninsula in Mexico. This work presents a chromosome-level genome assembly and annotation from a male individual. We used PacBio long reads and HiRise scaffolding to generate a high-quality genomic assembly of 1.87 Gb distributed in 327 scaffolds, with an N50 of 279 Mb and an L50 of 3. Approximately 98.4% of the genome is contained in 14 scaffolds, with 6 large scaffolds (334–127 Mb) representing macrochromosomes and 8 small scaffolds (63–22 Mb) representing microchromosomes. Using standard gene modeling and transcriptomic data, we predicted 17,902 protein-coding genes on the genome. The repeat content is characterized by a large proportion of long interspersed nuclear elements that are relatively old. Synteny analysis revealed some microchromosomes with high repeat content are more prone to rearrangements but that both macro- and microchromosomes are well conserved across reptiles. We identified scaffold 14 as the X chromosome. This microchromosome presents perfect dosage compensation where the single X of males has the same expression levels as two X chromosomes in females. Finally, we estimated the effective population size for U. nigricaudus was extremely low, which may reflect a reduction in polymorphism related to it becoming a peninsular endemic. 
    more » « less