skip to main content


Title: Improving bacterial genome assembly using a test of strand orientation
Abstract Summary The complexity of genome assembly is due in large part to the presence of repeats. In particular, large reverse-complemented repeats can lead to incorrect inversions of large segments of the genome. To detect and correct such inversions in finished bacterial genomes, we propose a statistical test based on tetranucleotide frequency (TNF), which determines whether two segments from the same genome are of the same or opposite orientation. In most cases, the test neatly partitions the genome into two segments of roughly equal length with seemingly opposite orientations. This corresponds to the segments between the DNA replication origin and terminus, which were previously known to have distinct nucleotide compositions. We show that, in several cases where this balanced partition is not observed, the test identifies a potential inverted misassembly, which is validated by the presence of a reverse-complemented repeat at the boundaries of the inversion. After inverting the sequence between the repeat, the balance of the misassembled genome is restored. Our method identifies 31 potential misassemblies in the NCBI database, several of which are further supported by a reassembly of the read data. Availability and implementation A github repository is available at https://github.com/gcgreenberg/Oriented-TNF.git. Supplementary information Supplementary data are available at Bioinformatics online.  more » « less
Award ID(s):
2046991
NSF-PAR ID:
10416295
Author(s) / Creator(s):
;
Date Published:
Journal Name:
Bioinformatics
Volume:
38
Issue:
Supplement_2
ISSN:
1367-4803
Page Range / eLocation ID:
ii34 to ii41
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Tribble, C (Ed.)
    Abstract The majority of sequenced genomes in the monocots are from species belonging to Poaceae, which include many commercially important crops. Here, we expand the number of sequenced genomes from the monocots to include the genomes of 4 related cyperids: Carex cristatella and Carex scoparia from Cyperaceae and Juncus effusus and Juncus inflexus from Juncaceae. The high-quality, chromosome-scale genome sequences from these 4 cyperids were assembled by combining whole-genome shotgun sequencing of Nanopore long reads, Illumina short reads, and Hi-C sequencing data. Some members of the Cyperaceae and Juncaceae are known to possess holocentric chromosomes. We examined the repeat landscapes in our sequenced genomes to search for potential repeats associated with centromeres. Several large satellite repeat families, comprising 3.2–9.5% of our sequenced genomes, showed dispersed distribution of large satellite repeat clusters across all Carex chromosomes, with few instances of these repeats clustering in the same chromosomal regions. In contrast, most large Juncus satellite repeats were clustered in a single location on each chromosome, with sporadic instances of large satellite repeats throughout the Juncus genomes. Recognizable transposable elements account for about 20% of each of the 4 genome assemblies, with the Carex genomes containing more DNA transposons than retrotransposons while the converse is true for the Juncus genomes. These genome sequences and annotations will facilitate better comparative analysis within monocots. 
    more » « less
  2. Mathelier, Anthony (Ed.)
    Abstract Motivation Methods to model dynamic changes in gene expression at a genome-wide level are not currently sufficient for large (temporally rich or single-cell) datasets. Variational autoencoders offer means to characterize large datasets and have been used effectively to characterize features of single-cell datasets. Here, we extend these methods for use with gene expression time series data. Results We present RVAgene: a recurrent variational autoencoder to model gene expression dynamics. RVAgene learns to accurately and efficiently reconstruct temporal gene profiles. It also learns a low dimensional representation of the data via a recurrent encoder network that can be used for biological feature discovery, and from which we can generate new gene expression data by sampling the latent space. We test RVAgene on simulated and real biological datasets, including embryonic stem cell differentiation and kidney injury response dynamics. In all cases, RVAgene accurately reconstructed complex gene expression temporal profiles. Via cross validation, we show that a low-error latent space representation can be learnt using only a fraction of the data. Through clustering and gene ontology term enrichment analysis on the latent space, we demonstrate the potential of RVAgene for unsupervised discovery. In particular, RVAgene identifies new programs of shared gene regulation of Lox family genes in response to kidney injury. Availability and implementation All datasets analyzed in this manuscript are publicly available and have been published previously. RVAgene is available in Python, at GitHub: https://github.com/maclean-lab/RVAgene; Zenodo archive: http://doi.org/10.5281/zenodo.4271097. Supplementary information Supplementary data are available at Bioinformatics online. 
    more » « less
  3. Summary

    The plastid genome (plastome), while surprisingly constant in gene order and content across most photosynthetic angiosperms, exhibits variability in several unrelated lineages. During the diversification history of the legume family Fabaceae, plastomes have undergone many rearrangements, including inversions, expansion, contraction and loss of the typical inverted repeat (IR), gene loss and repeat accumulation in both shared and independent events. While legume plastomes have been the subject of study for some time, most work has focused on agricultural species in the IR‐lacking clade (IRLC) and the plant modelMedicago truncatula. The subfamily Papilionoideae, which contains virtually all of the agricultural legume species, also comprises most of the plastome variation detected thus far in the family. In this study three non‐papilioniods were included among 34 newly sequenced legume plastomes, along with 33 publicly available sequences, to assess plastome structural evolution in the subfamily. In an effort to examine plastome variation across the subfamily, approximately 20% of the sampling represents the IRLC with the remainder selected to represent the early‐branching papilionoid clades. A number of IR‐related and repeat‐mediated changes were identified and examined in a phylogenetic context. Recombination between direct repeats associated withycf2resulted in intraindividual plastome heteroplasmy. Although loss of the IR has not been reported in legumes outside of the IRLC, one genistoid taxon was found to completely lack the typical plastome IR. The role of the IR and non‐IR repeats in the progression of plastome change is discussed.

     
    more » « less
  4. Abstract

    Although plastid genome (plastome) structure is highly conserved across most seed plants, investigations during the past two decades have revealed several disparately related lineages that experienced substantial rearrangements. Most plastomes contain a large inverted repeat and two single‐copy regions, and a few dispersed repeats; however, the plastomes of some taxa harbour long repeat sequences (>300 bp). These long repeats make it challenging to assemble complete plastomes using short‐read data, leading to misassemblies and consensus sequences with spurious rearrangements. Single‐molecule, long‐read sequencing has the potential to overcome these challenges, yet there is no consensus on the most effective method for accurately assembling plastomes using long‐read data. We generated a pipeline,plastidGenomeAssemblyUsingLong‐read data (ptGAUL), to address the problem of plastome assembly using long‐read data from Oxford Nanopore Technologies (ONT) or Pacific Biosciences platforms. We demonstrated the efficacy of the ptGAUL pipeline using 16 published long‐read data sets. We showed that ptGAUL quickly produces accurate and unbiased assemblies using only ~50× coverage of plastome data. Additionally, we deployed ptGAUL to assemble four newJuncus(Juncaceae) plastomes using ONT long reads. Our results revealed many long repeats and rearrangements inJuncusplastomes compared with basal lineages of Poales. The ptGAUL pipeline is available on GitHub:https://github.com/Bean061/ptgaul.

     
    more » « less
  5. null (Ed.)
    Finite-fault models for the 2010 M w 8.8 Maule, Chile earthquake indicate bilateral rupture with large-slip patches located north and south of the epicenter. Previous studies also show that this event features significant slip in the shallow part of the megathrust, which is revealed through correction of the forward tsunami modeling scheme used in tsunami inversions. The presence of shallow slip is consistent with the coseismic seafloor deformation measured off the Maule region adjacent to the trench and confirms that tsunami observations are particularly important for constraining far-offshore slip. Here, we benchmark the method of Optimal Time Alignment (OTA) of the tsunami waveforms in the joint inversion of tsunami (DART and tide-gauges) and geodetic (GPS, InSAR, land-leveling) observations for this event. We test the application of OTA to the tsunami Green’s functions used in a previous inversion. Through a suite of synthetic tests we show that if the bias in the forward model is comprised only of delays in the tsunami signals, the OTA can correct them precisely, independently of the sensors (DART or coastal tide-gauges) and, to the first-order, of the bathymetric model used. The same suite of experiments is repeated for the real case of the 2010 Maule earthquake where, despite the results of the synthetic tests, DARTs are shown to outperform tide-gauges. This gives an indication of the relative weights to be assigned when jointly inverting the two types of data. Moreover, we show that using OTA is preferable to subjectively correcting possible time mismatch of the tsunami waveforms. The results for the source model of the Maule earthquake show that using just the first-order modeling correction introduced by OTA confirms the bilateral rupture pattern around the epicenter, and, most importantly, shifts the inferred northern patch of slip to a shallower position consistent with the slip models obtained by applying more complex physics-based corrections to the tsunami waveforms. This is confirmed by a slip model refined by inverting geodetic and tsunami data complemented with a denser distribution of GPS data nearby the source area. The models obtained with the OTA method are finally benchmarked against the observed seafloor deformation off the Maule region. We find that all of the models using the OTA well predict this offshore coseismic deformation, thus overall, this benchmarking of the OTA method can be considered successful. 
    more » « less