This content will become publicly available on October 1, 2026

Title: Predicting the trend of SARS-CoV-2 mutation frequencies using historical data
Abstract MotivationAs the SARS-CoV-2 virus rapidly evolves, predicting the trajectory of viral mutations has become a critical yet complex task. A deep understanding of future mutation patterns, in particular the mutations that will prevail in the near future, is vital in steering diagnostics, therapeutics, and vaccine strategies for disease control. ResultsIn this study, we developed a model to forecast future SARS-CoV-2 mutation surges in real-time, using historical mutation frequency data from the USA. We transformed the temporal prediction problem into a supervised learning framework using a sliding window approach. This involved breaking the time series of mutation frequencies into very short segments. Considering the time-dependent nature of the data, we focused on modeling the first-order derivative of the mutation frequency. We predicted the final derivative in each segment based on the preceding derivatives, employing various machine learning methods, including random forest, XGBoost, support vector machine, and neural network models. Empowered by the novel transformation strategy and the high capacity of machine learning models, we observed low prediction error that is confined within 0.1% and 1% when making predictions of mutation rates for the future 30 and 80 days, respectively. In addition, the method also led to a notable increase in prediction accuracy compared to traditional time-series models, as evidenced by much lower MAE (Mean Absolute Error) and MSE (Mean Squared Error) for predictions made within different time horizons. To further assess the method’s effectiveness and robustness in predicting mutation patterns for unforeseen mutations, we first designed a synthetic case where we categorized all mutations into three major patterns. The model demonstrated its robustness by accurately predicting unseen mutation patterns when training on data from two pattern categories while testing on the third pattern category, showcasing its potential in forecasting a variety of mutation trajectories. We then applied our method to prediction for a recent time frame between 1 January 2025 and 10 June 2025, for both the USA and UK, where the model training was conducted using frequency sequence data collected between 12 December 2019 and 26 January 2023 in the USA. The model demonstrated superior performance for both datasets. Availability and implementationTo enhance accessibility and utility, we built our methodology into a GitHub package (https://github.com/ZhouXY199502/SWD). Our method has the potential applicability to study other infectious diseases or forecasting tasks, thus extending its relevance beyond the current COVID pandemic.  more » « less
Award ID(s):
2528521
PAR ID:
10685758
Author(s) / Creator(s):
; ; ; ; ; ; ;
Editor(s):
Lu, Zhiyong
Publisher / Repository:
Oxford
Date Published:
Journal Name:
Bioinformatics
Volume:
41
Issue:
10
ISSN:
1367-4803
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Tremendous effort has been given to the development of diagnostic tests, preventive vaccines, and therapeutic medicines for coronavirus disease 2019 (COVID-19) caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Much of this development has been based on the reference genome collected on January 5, 2020. Based on the genotyping of 15 140 genome samples collected up to June 1, 2020, we report that SARS-CoV-2 has undergone 8309 single mutations which can be clustered into six subtypes. We introduce mutation ratio and mutation h-index to characterize the protein conservativeness and unveil that SARS-CoV-2 envelope protein, main protease, and endoribonuclease protein are relatively conservative, while SARS-CoV-2 nucleocapsid protein, spike protein, and papain-like protease are relatively nonconservative. In particular, we have identified mutations on 40% of nucleotides in the nucleocapsid gene in the population level, signaling potential impacts on the ongoing development of COVID-19 diagnosis, vaccines, and antibody and small-molecular drugs. 
    more » « less
  2. Abstract Purpose of the ReviewSARS-CoV-2 undergoes genetic mutations like many other viruses. Some mutations lead to the emergence of new Variants of Concern (VOCs), affecting transmissibility, illness severity, and the effectiveness of antiviral drugs. Continuous monitoring and research are crucial to comprehend variant behavior and develop effective response strategies, including identifying mutations that may affect current drug therapies. Recent FindingsAntiviral therapies such as Nirmatrelvir and Ensitrelvir focus on inhibiting 3CLpro, whereas Remdesivir, Favipiravir, and Molnupiravir target nsp12, thereby reducing the viral load. However, the emergence of resistant mutations in 3CLpro and nsp12 could impact the efficiency of these small molecule drug therapeutics. SummaryThis manuscript summarizes mutations in 3CLpro and nsp12, which could potentially reduce the efficacy of drugs. Additionally, it encapsulates recent advancements in small molecule antivirals targeting SARS-CoV-2 viral proteins, including their potential for developing resistance against emerging variants. 
    more » « less
  3. Liu, Shan-Lu (Ed.)
    ABSTRACT The human angiotensin-converting enzyme 2 (hACE2) is the primary receptor for the entry of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Some human alleles of ACE2 exhibit an improved affinity for the SARS-CoV-2 Spike protein. However, the impact of ACE2 polymorphisms on SARS-CoV-2 infection remains unclear. Our previous study predicted that G431 and S514 in the receptor-binding domain (RBD) of SARS-CoV-2 S1 domain are important for S protein stability, and that S protein residues G496 and F497 and ACE2 residues D355 and Y41 are critical for the RBD-ACE2 interaction. In this study, we explored the potential of hACE2-derived neutralizing peptides as a therapeutic strategy against SARS-CoV-2 and investigated how ACE2 polymorphisms affect RBD-ACE2 binding affinity. We applied computational saturation mutagenesis to systematically screen the binding affinity changes among all possible ACE2 missense mutations within the ACE2-Wuhan-S1 complex. Mutations at ACE2 residues D355 and Y41 were predicted to weaken binding affinity, whereas those at N330 and D30 enhanced it. We identified six ACE2 regions (19–49, 65–102, 320–333, 348–359, 378–395, and 552–563) to be vital for ACE2-RBD interaction. We synthesized peptides corresponding to these six regions and tested them using a pseudotyped viral particle system and dot blot assay. Three peptides were confirmed to bind with the S protein, and four exhibited inhibitory effects. We aligned ACE2-Wuhan-S1 and ACE2-Omicron-S1 complexes, conducted correlation analysis, and observed similar binding patterns, suggesting that these peptides also have the potential to neutralize Omicron strains.IMPORTANCESARS-CoV-2 continues its global spread. In this research, we identified six regions within ACE2 that are vital for interaction with the viral S receptor-binding domain and have the potential to neutralize SARS-CoV-2 infection. Among the six peptides derived from ACE2, three were confirmed to bind with the S protein of the Wuhan strain, and four exhibited inhibitory effects on the Wuhan strain SARS-CoV-2. We also found ACE2 residues D355 and Y41 as weakening affinity, and N330 and D30 as enhancing it. We also aligned this complex with the ACE2-Omicron-S1 complex, performed correlation analyses, and compared their patterns of stability changes upon mutations and obtained similar results, indicating that these peptides may also be effective against Omicron variants. These results provide insight into the role of ACE2 polymorphism in viral entry and suggest that hACE2-derived peptides may offer a promising therapeutic strategy against SARS-CoV-2, demonstrating strong consistency between our computational predictions and experimental outcomes. 
    more » « less
  4. The SARS-CoV-2 Delta variant is emerging as a globally dominant strain. Its rapid spread and high infection rate are attributed to a mutation in the spike protein of SARS-CoV-2 allowing for the virus to invade human cells much faster and with an increased efficiency. In particular, an especially dangerous mutation P681R close to the furin cleavage site has been identified as responsible for increasing the infection rate. Together with the earlier reported mutation D614G in the same domain, it offers an excellent instance to investigate the nature of mutations and how they affect the interatomic interactions in the spike protein. Here, using ultra large-scale ab initio computational modeling, we study the P681R and D614G mutations in the SD2-FP domain, including the effect of double mutation, and compare the results with the wild type. We have recently developed a method of calculating the amino-acid–amino-acid bond pairs (AABP) to quantitatively characterize the details of the interatomic interactions, enabling us to explain the nature of mutation at the atomic resolution. Our most significant finding is that the mutations reduce the AABP value, implying a reduced bonding cohesion between interacting residues and increasing the flexibility of these amino acids to cause the damage. The possibility of using this unique mutation quantifiers in a machine learning protocol could lead to the prediction of emerging mutations. 
    more » « less
  5. The detection of nucleic acids and their mutation derivatives is vital for biomedical science and applications. Although many nucleic acid biosensors have been developed, they often require pretreatment processes, such as target amplification and tagging probes to nucleic acids. Moreover, current biosensors typically cannot detect sequence-specific mutations in the targeted nucleic acids. To address the above problems, herein, we developed an electrochemical nanobiosensing system using a phenomenon comprising metal ion intercalation into the targeted mismatched double-stranded nucleic acids and a homogeneous Au nanoporous electrode array (Au NPEA) to obtain (i) sensitive detection of viral RNA without conventional tagging and amplifying processes, (ii) determination of viral mutation occurrence in a simple detection manner, and (iii) multiplexed detection of several RNA targets simultaneously. As a proof-of-concept demonstration, a SARS-CoV-2 viral RNA and its mutation derivative were used in this study. Our developed nanobiosensor exhibited highly sensitive detection of SARS-CoV-2 RNA (∼1 fM detection limit) without tagging and amplifying steps. In addition, a single point mutation of SARS-CoV-2 RNA was detected in a one-step analysis. Furthermore, multiplexed detection of several SARS-CoV-2 RNAs was successfully demonstrated using a single chip with four combinatorial NPEAs generated by a 3D printing technique. Collectively, our developed nanobiosensor provides a promising platform technology capable of detecting various nucleic acids and their mutation derivatives in highly sensitive, simple, and time-effective manners for point-of-care biosensing. 
    more » « less