Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


This content will become publicly available on March 27, 2027

Title: Fast sampling of protein conformational dynamics
Protein function often depends on dynamic transitions between conformations rather than just static structures. However, our current ability to characterize or predict such dynamics lags behind recent advances in protein structure prediction. Enhanced sampling methods can speed up molecular dynamics simulations to study protein conformational transitions but require prior knowledge of key collective motions involved. Here, we demonstrate for a series of proteins of varying complexity that the required information is encoded in anharmonic low-frequency vibrations. Using recently developed methods, we show that this information can be easily extracted from short dynamics simulations without requiring prior knowledge. Combined with enhanced sampling, we correctly predict conformational transitions in all test proteins and generate highly reproducible free energy landscapes. This allows for the rapid generation of accurate protein conformational ensembles, which is critical to unravel the complex relationship between protein sequence, structure, and dynamics.  more » « less
Award ID(s):
2154834
PAR ID:
10681023
Author(s) / Creator(s):
; ; ; ;
Publisher / Repository:
American Association for the Advancement of Science
Date Published:
Journal Name:
Science Advances
Volume:
12
Issue:
13
ISSN:
2375-2548
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Fluctuations of protein three-dimensional structures and large-scale conformational transitions are crucial for the biological function of proteins and their complexes. Experimental studies of such phenomena remain very challenging and therefore molecular modeling can be a good alternative or a valuable supporting tool for the investigation of large molecular systems and long-time events. In this minireview, we present two alternative approaches to the coarse-grained (CG) modeling of dynamic properties of protein systems. We discuss two CG representations of polypeptide chains used for Monte Carlo dynamics simulations of protein local dynamics and conformational transitions, and highly simplified structure-based elastic network models of protein flexibility. In contrast to classical all-atom molecular dynamics, the modeling strategies discussed here allow the quite accurate modeling of much larger systems and longer-time dynamic phenomena. We briefly describe the main features of these models and outline some of their applications, including modeling of near-native structure fluctuations, sampling of large regions of the protein conformational space, or possible support for the structure prediction of large proteins and their complexes. 
    more » « less
  2. Proteins perform their biological functions through motion. Although high throughput prediction of the three-dimensional static structures of proteins has proved feasible using deep-learning-based methods, predicting the conformational motions remains a challenge. Purely data-driven machine learning methods encounter difficulty for addressing such motions because available laboratory data on conformational motions are still limited. In this work, we develop a method for generating protein allosteric motions by integrating physical energy landscape information into deep-learning-based methods. We show that local energetic frustration, which represents a quantification of the local features of the energy landscape governing protein allosteric dynamics, can be utilized to empower AlphaFold2 (AF2) to predict protein conformational motions. Starting from ground state static structures, this integrative method generates alternative structures as well as pathways of protein conformational motions, using a progressive enhancement of the energetic frustration features in the input multiple sequence alignment sequences. For a model protein adenylate kinase, we show that the generated conformational motions are consistent with available experimental and molecular dynamics simulation data. Applying the method to another two proteins KaiB and ribose-binding protein, which involve large-amplitude conformational changes, can also successfully generate the alternative conformations. We also show how to extract overall features of the AF2 energy landscape topography, which has been considered by many to be black box. Incorporating physical knowledge into deep-learning-based structure prediction algorithms provides a useful strategy to address the challenges of dynamic structure prediction of allosteric proteins. 
    more » « less
  3. Abstract While significant advances have been made in predicting static protein structures, the inherent dynamics of proteins, modulated by ligands, are crucial for understanding protein function and facilitating drug discovery. Traditional docking methods, frequently used in studying protein-ligand interactions, typically treat proteins as rigid. While molecular dynamics simulations can propose appropriate protein conformations, they’re computationally demanding due to rare transitions between biologically relevant equilibrium states. In this study, we present DynamicBind, a deep learning method that employs equivariant geometric diffusion networks to construct a smooth energy landscape, promoting efficient transitions between different equilibrium states. DynamicBind accurately recovers ligand-specific conformations from unbound protein structures without the need for holo-structures or extensive sampling. Remarkably, it demonstrates state-of-the-art performance in docking and virtual screening benchmarks. Our experiments reveal that DynamicBind can accommodate a wide range of large protein conformational changes and identify cryptic pockets in unseen protein targets. As a result, DynamicBind shows potential in accelerating the development of small molecules for previously undruggable targets and expanding the horizons of computational drug discovery. 
    more » « less
  4. Abstract DNA exhibits local conformational preferences that affect its ability to adopt biologically relevant conformations, such as those required for binding proteins. Traditional methods, like Markov state models and molecular dynamics (MD) simulations, have advanced our understanding but often struggle to capture these rare conformational states due to high computational demands. Here, we introduce a novel AI framework based on dynamical graphical models (DGMs), a generative machine learning approach trained on equilibrium MD data, to predict DNA conformational transitions that are never seen in the MD ensembles. By leveraging local DNA interactions, DGMs generate a comprehensive transition matrix that captures both thermodynamic and kinetic properties of unsampled states, enabling accurate predictions of rare global conformations without the need for extensive sampling. Applying this model to the B→A transition, we demonstrate that DGMs can efficiently predict sequence-dependent A-DNA preferences, achieving results that align closely with replica exchange umbrella sampling simulations. DGMs provide new insights into DNA sequence–structure relationships, paving the way for applications in DNA sequence design and optimization. 
    more » « less
  5. Abstract Structural, regulatory and enzymatic proteins interact with DNA to maintain a healthy and functional genome. Yet, our structural understanding of how proteins interact with DNA is limited. We present MELD-DNA, a novel computational approach to predict the structures of protein–DNA complexes. The method combines molecular dynamics simulations with general knowledge or experimental information through Bayesian inference. The physical model is sensitive to sequence-dependent properties and conformational changes required for binding, while information accelerates sampling of bound conformations. MELD-DNA can: (i) sample multiple binding modes; (ii) identify the preferred binding mode from the ensembles; and (iii) provide qualitative binding preferences between DNA sequences. We first assess performance on a dataset of 15 protein–DNA complexes and compare it with state-of-the-art methodologies. Furthermore, for three selected complexes, we show sequence dependence effects of binding in MELD predictions. We expect that the results presented herein, together with the freely available software, will impact structural biology (by complementing DNA structural databases) and molecular recognition (by bringing new insights into aspects governing protein–DNA interactions). 
    more » « less