skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Epidemiological associations with genomic variation in SARS-CoV-2
Abstract SARS-CoV-2 (CoV) is the etiological agent of the COVID-19 pandemic and evolves to evade both host immune systems and intervention strategies. We divided the CoV genome into 29 constituent regions and applied novel analytical approaches to identify associations between CoV genomic features and epidemiological metadata. Our results show that nonstructural protein 3 (nsp3) and Spike protein (S) have the highest variation and greatest correlation with the viral whole-genome variation. S protein variation is correlated with nsp3, nsp6, and 3′-to-5′ exonuclease variation. Country of origin and time since the start of the pandemic were the most influential metadata associated with genomic variation, while host sex and age were the least influential. We define a novel statistic—coherence—and show its utility in identifying geographic regions (populations) with unusually high (many new variants) or low (isolated) viral phylogenetic diversity. Interestingly, at both global and regional scales, we identify geographic locations with high coherence neighboring regions of low coherence; this emphasizes the utility of this metric to inform public health measures for disease spread. Our results provide a direction to prioritize genes associated with outcome predictors (e.g., health, therapeutic, and vaccine outcomes) and to improve DNA tests for predicting disease status.  more » « less
Award ID(s):
2028280 2109688
PAR ID:
10360441
Author(s) / Creator(s):
; ; ; ; ;
Publisher / Repository:
Nature Publishing Group
Date Published:
Journal Name:
Scientific Reports
Volume:
11
Issue:
1
ISSN:
2045-2322
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Cimarelli, Andrea (Ed.)
    The Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) infection causes Coronavirus Disease 2019 (COVID-19), a pandemic that seriously threatens global health. SARS-CoV-2 propagates by packaging its RNA genome into membrane enclosures in host cells. The packaging of the viral genome into the nascent virion is mediated by the nucleocapsid (N) protein, but the underlying mechanism remains unclear. Here, we show that the N protein forms biomolecular condensates with viral genomic RNA both in vitro and in mammalian cells. While the N protein forms spherical assemblies with homopolymeric RNA substrates that do not form base pairing interactions, it forms asymmetric condensates with viral RNA strands. Cross-linking mass spectrometry (CLMS) identified a region that drives interactions between N proteins in condensates, and deletion of this region disrupts phase separation. We also identified small molecules that alter the size and shape of N protein condensates and inhibit the proliferation of SARS-CoV-2 in infected cells. These results suggest that the N protein may utilize biomolecular condensation to package the SARS-CoV-2 RNA genome into a viral particle. 
    more » « less
  2. Abstract Coronavirus disease (COVID-19) is a contagious respiratory disease caused by the SARS-CoV-2 virus. The clinical phenotypes are variable, ranging from spontaneous recovery to serious illness and death. On March 2020, a global COVID-19 pandemic was declared by the World Health Organization (WHO). As of February 2023, almost 670 million cases and 6,8 million deaths have been confirmed worldwide. Coronaviruses, including SARS-CoV-2, contain a single-stranded RNA genome enclosed in a viral capsid consisting of four structural proteins: the nucleocapsid (N) protein, in the ribonucleoprotein core, the spike (S) protein, the envelope (E) protein, and the membrane (M) protein, embedded in the surface envelope. In particular, the E protein is a poorly characterized viroporin with high identity amongst all the β-coronaviruses (SARS-CoV-2, SARS-CoV, MERS-CoV, HCoV-OC43) and a low mutation rate. Here, we focused our attention on the study of SARS-CoV-2 E and M proteins, and we found a general perturbation of the host cell calcium (Ca 2+ ) homeostasis and a selective rearrangement of the interorganelle contact sites. In vitro and in vivo biochemical analyses revealed that the binding of specific nanobodies to soluble regions of SARS-CoV-2 E protein reversed the observed phenotypes, suggesting that the E protein might be an important therapeutic candidate not only for vaccine development, but also for the clinical management of COVID designing drug regimens that, so far, are very limited. 
    more » « less
  3. Through the COVID-19 pandemic, SARS-CoV-2 has gained and lost multiple mutations in novel or unexpected combinations. Predicting how complex mutations affect COVID-19 disease severity is critical in planning public health responses as the virus continues to evolve. This paper presents a novel computational framework to complement conventional lineage classification and applies it to predict the severe disease potential of viral genetic variation. The transformer-based neural network model architecture has additional layers that provide sample embeddings and sequence-wide attention for interpretation and visualization. First, training a model to predict SARS-CoV-2 taxonomy validates the architecture’s interpretability. Second, an interpretable predictive model of disease severity is trained on spike protein sequence and patient metadata from GISAID. Confounding effects of changing patient demographics, increasing vaccination rates, and improving treatment over time are addressed by including demographics and case date as independent input to the neural network model. The resulting model can be interpreted to identify potentially significant virus mutations and proves to be a robust predctive tool. Although trained on sequence data obtained entirely before the availability of empirical data for Omicron, the model can predict the Omicron’s reduced risk of severe disease, in accord with epidemiological and experimental data. 
    more » « less
  4. Abstract The ongoing COVID-19 pandemic highlights the necessity for a more fundamental understanding of the coronavirus life cycle. The causative agent of the disease, SARS-CoV-2, is being studied extensively from a structural standpoint in order to gain insight into key molecular mechanisms required for its survival. Contained within the untranslated regions of the SARS-CoV-2 genome are various conserved stem-loop elements that are believed to function in RNA replication, viral protein translation, and discontinuous transcription. While the majority of these regions are variable in sequence, a 41-nucleotide s2m element within the genome 3′ untranslated region is highly conserved among coronaviruses and three other viral families. In this study, we demonstrate that the SARS-CoV-2 s2m element dimerizes by forming an intermediate homodimeric kissing complex structure that is subsequently converted to a thermodynamically stable duplex conformation. This process is aided by the viral nucleocapsid protein, potentially indicating a role in mediating genome dimerization. Furthermore, we demonstrate that the s2m element interacts with multiple copies of host cellular microRNA (miRNA) 1307-3p. Taken together, our results highlight the potential significance of the dimer structures formed by the s2m element in key biological processes and implicate the motif as a possible therapeutic drug target for COVID-19 and other coronavirus-related diseases. 
    more » « less
  5. Prasad, Vinayaka R. (Ed.)
    ABSTRACT The ongoing coronavirus disease 2019 (COVID-19) pandemic demonstrates the threat posed by novel coronaviruses to human health. Coronaviruses share a highly conserved cell entry mechanism mediated by the spike protein, the sole product of the S gene. The structural dynamics by which the spike protein orchestrates infection illuminate how antibodies neutralize virions and how S mutations contribute to viral fitness. Here, we review the process by which spike engages its proteinaceous receptor, angiotensin converting enzyme 2 (ACE2), and how host proteases prime and subsequently enable efficient membrane fusion between virions and target cells. We highlight mutations common among severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) variants of concern and discuss implications for cell entry. Ultimately, we provide a model by which sarbecoviruses are activated for fusion competency and offer a framework for understanding the interplay between humoral immunity and the molecular evolution of the SARS-CoV-2 Spike. In particular, we emphasize the relevance of the Canyon Hypothesis (M. G. Rossmann, J Biol Chem 264:14587–14590, 1989) for understanding evolutionary trajectories of viral entry proteins during sustained intraspecies transmission of a novel viral pathogen. 
    more » « less