<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Sampling Conformational Ensembles of Highly Dynamic Proteins via Generative Deep Learning</title></titleStmt>
			<publicationStmt>
				<publisher>bioRxiv</publisher>
				<date>05/05/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10528862</idno>
					<idno type="doi">10.1101/2024.05.05.592587</idno>
					
					<author>Talant Ruzmetov</author><author>Ta I Hung</author><author>Saisri Padmaja Jonnalagedda</author><author>Si-han Chen</author><author>Parisa Fasihianifard</author><author>Zhefeng Guo</author><author>Bir Bhanu</author><author>Chia-en A Chang</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[<title>Abstract</title> <p>Proteins are inherently dynamic, and their conformational ensembles are functionally important in biology. Large-scale motions may govern protein structure–function relationship, and numerous transient but stable conformations of intrinsically disordered proteins (IDPs) can play a crucial role in biological function. Investigating conformational ensembles to understand regulations and disease-related aggregations of IDPs is challenging both experimentally and computationally. In this paper we first introduced an unsupervised deep learning-based model, termed Internal Coordinate Net (ICoN), which learns the physical principles of conformational changes from molecular dynamics (MD) simulation data. Second, we selected interpolating data points in the learned latent space that rapidly identify novel synthetic conformations with sophisticated and large-scale sidechains and backbone arrangements. Third, with the highly dynamic amyloid-β<sub>1-42</sub>(Aβ42) monomer, our deep learning model provided a comprehensive sampling of Aβ42’s conformational landscape. Analysis of these synthetic conformations revealed conformational clusters that can be used to rationalize experimental findings. Additionally, the method can identify novel conformations with important interactions in atomistic details that are not included in the training data. New synthetic conformations showed distinct sidechain rearrangements that are probed by our EPR and amino acid substitution studies. This approach is highly transferable and can be used for any available data for training. The work also demonstrated the ability for deep learning to utilize learned natural atomistic motions in protein conformation sampling.</p>]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Proteins are complex and have dynamic properties. Conformational ensembles of proteins are essential for performing biological processes, including enzyme activities, protein folding, protein-ligand binding, and protein aggregation <ref type="bibr">1,</ref><ref type="bibr">2</ref> . Characterization of protein conformations allows researchers to understand protein function, activity, and mechanisms. However, the tasks may be daunting and are even more challenging for investigating intrinsically disordered proteins (IDPs), for which numerous conformations can be functionally important. Some IDPs, such as amyloid-&#946;1-42 (A&#946;42), are related to a number of protein aggregation-related diseases, including  (bi, ai, &#952;i), where b, a and t are bond, angle and torsions, respectively. For example, atom 4 is presented by (b4, a4, &#952;4). The dashed light gray line shows the smooth rotation of the bond between atoms 2 and 3 (&#952;4).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ICoN model architecture, training, and validation.</head><p>Molecular representation using the vBAT coordinate. The ICoN model is a latent space-guided generative AI model trained on protein conformational ensembles to learn physical properties that govern conformational transitions in MD simulations (Figure <ref type="figure">2</ref>). Because the dihedral rotations are the major motions to determine conformations, using BAT to accurately describe the concerted motions of multiple dihedral rotations is critical. Dihedral rotations have a periodicity issue, so we used a vector called vBAT to avoid periodicity (Figure <ref type="figure">2B</ref> and details in Supplementary Material). All-atom vBAT internal coordinate representation is inherently equivariant to rotations and translations and can exactly convert vBAT vectors to Cartesian coordinates without any additional approximation. However, coordinate conversions can be time-consuming. Thus, our model implemented GPU acceleration to perform coordinate transformations; converting 5,000 frames from vBAT to Cartesian coordinates took &lt; 0.2 s for A&#946;42. Figure <ref type="figure">S2</ref> demonstrates the efficiency in coordinate transformations that successfully avoided a computation bottleneck with the use of internal coordinates. ICoN network architecture and training. The neural network is based on autoencoder architecture (Figure 2A). Classical BAT coordinates (the Z matrix) have 3N-6 DOF for presenting the internal motions of a molecule, where N is the number of atoms, and the 6 external translation and rotation DOF are eliminated. Because vBAT uses the vectors instead of one value for each bond, angle and dihedral DOF, we have 4x3x(N-3) features for each molecule. As a result, &#945;B-crystallin57-69 and A&#946;42 have 2376 and 7488 features, respectively.</p><p>The hyperparameter training in the encoder resulted in 7 layers to reduce a large dimension representation into 3D, so each molecular conformation can be presented in the 3D latent space (Figures <ref type="figure">2A</ref> and <ref type="figure">S3</ref>). The decoding process brought the representation from 3D back to the original dimensions with the same number of layers but in reverse order. Each conformation is converted from vBAT to Cartesian coordinates for further analysis. Throughout the iterative training process, in addition to learning a relationship between a vast number of features, the network also learns how to meaningfully compress the representation into the lower-dimensional 3D latent space by performing nonlinear dimensionality reduction.</p><p>Notably, although this study used 3D latent space for easy observation of conformation distribution and visualization of interpolation, the network can reduce 4x3x(N-3) features to any dimension desired by users. We used 3D to achieve efficient training and reduced memory consumption as well. It took 4 min with one NVIDIA 1080Ti GPU card to perform 15000 epochs for successful training with 10,000 frames of A&#946;42.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Validation of the models by conformation reconstruction.</head><p>A successful trained model should accurately reconstruct a protein conformation from the reduced 3D representation to the original conformation. To evaluate the accuracy, we presented conformations from our validation set using the 3D representation and then reconstructed them back to atomistic Cartesian representation and computed the root mean square deviation (RMSD) of all heavy atoms between the 2 conformations, the original and reconstructed ones. A few representative conformations are illustrated in Figure <ref type="figure">S4</ref>: &#945;B-crystallin57-69 has RMSD &lt; 0.9&#197; and A&#946;42 has RMSD &lt; 1.3&#197;. As shown in Table <ref type="table">1</ref>, with our model, both backbone and sidechains conformations can be accurately reconstructed back to the original input structure, ambient after large dimensionality reduction in the latent space. We also performed a more detailed comparisons between the original and reconstructed conformations, such as the distribution of each dihedral rotation and their correlations, to ensure that all the properties were accurately reproduced (Figures <ref type="figure">S5-S8</ref>). Notably, both &#945;B-crystallin57-69 and A&#946;42 exhibit large-scale protein motions included in our training set. The validation demonstrates the robustness of our model to learn a diverse range of conformational states with high precision, and we can describe small sidechain rotations accurately. Table <ref type="table">1</ref>. Summary of molecular dynamics (MD) simulations, numbers of MD frames used for ICoN model training and validation, and numbers of conformations generated from ICoN models in each step. The hyperparameters of A&#946;42 ICoN model were obtained from dataset MD Run1, and all Models 1 to 5 use the same hyperparameters. Energy cutoff was -200 and -400 kcal/mol for &#945;B-crystalline and A&#946;42. Two conformations are treated as repeats when the computed heavy atom RMSD is smaller than 1 &#197; in Step 3. In Step 4, the RMSD cutoff is 1 &#197; and 1.2 &#197; for &#945;Bcrystalline and A&#946;42, respectively.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Learned physics in the latent space of &#945;B-cristallin57-69</head><p>Because our ICoN model learns important dihedral rotations that determine molecular conformations, the latent space stores information that leads to smooth conformational transitions between 2 datapoints, such as various sets of concerted dihedral rotations. Moreover, we can visualize conformation transitions and clusters and can perform data analysis in the 3D</p><p>MD simulations ICoN model construction/validation Synthetic Conformations Step 1 Step 2 Step 3 Step 4 MD index force field &amp; initial structure (PDB) # of frames for training/validation Validation of reconstructed conformations, average heavy atom RMSD (backbone RMSD) Model index # of confs from initial interpolation in the latent space # of confs within energy cutoff # of distinct confs (repeat eliminated) # of novel synthetic confs (repeat with MD confs eliminated) &#945;B-crystalline (0.5 &#956;s) Seed 1 14sb (Alpha Fold2 predicted structure) 5,000/5,000 2.58 &#197; (1.84 &#197;) Model I 99990 (100%) 97594 (97.6%) 43484 (43.5%) 38425 (38.4%) Seed 2 10,000 for comparison with synthetic conformations obtained by Model I A&#946;42 (1 &#956;s for all runs) MD Run1 ff14IDPSFF (2NAO) 10,000/10,000 3.90 &#197; (3.60 &#197;) Model 1 199990 (100%) 187487 (93.7%) 150713 (75.3%) 127673 (63.8%) MD Run2 ff14IDPSFF (1Z0Q) 10,000/10,000 6.18 &#197; (5.96 &#197;) Model 2 199990 (100%) 174012 (87%) 148188 (74%) 145925 (72.9%) MD Run3 ff14SB (2NAO) 10,000/10,000 4.20 &#197; (3.91 &#197;) Model 3 199990 (100%) 189652 (94.8%) 129482 (64.7%) 90627 (45.3%) MD Run4 ff14SB (1Z0Q) 10,000/10,000 4.15 &#197; (3.82 &#197;) Model 4 199990 (100%) 194449 (94.4%) 166205 (66.2%) 148320 (48.3%) MD Run5 ESFF1 10,000/10,000 4.27 &#197; (4.08 &#197;) Model 5 199990 (100%) 166552 (83.3%) 95126 (47.6%) 87257 (43.6%) MD Run1.1 ff14IDPSFF (2NAO) 10,000 for comparison with synthetic conformations obtained by Model 1 MD Run1.2 ff14IDPSFF (2NAO) 10,000 for comparison with synthetic conformations obtained by Model 1 MD Run1.3 ff14IDPSFF (2NAO) 10,000 for comparison with synthetic conformations obtained by Model 1</p><p>latent space (Figures <ref type="figure">3</ref> and <ref type="figure">S10</ref>). We demonstrated that the DL model learned the physics governing molecular motions, and the information captured in the latent space can be effectively utilized.</p><p>Essential motions of &#945;B-cristallin57-69 revealed in the latent space. We selected 2 conformations in the latent space, Conf indexes #50 and #53, from residues were maintained in a helix secondary structure, and the middle region around Leu9</p><p>showed minor but important fluctuations to twist the center of the peptide, whereas the flexible tail (residues 7 to 13) exhibited greater dynamics. A correlated sidechain movement (i.e., sidechains of Leu9 and Met12) was observed during the transitions as well.</p><p>In the latent space, we found that the 300-ps interval MD trajectory sampled conformational jiggling during the transitions (light blue dots in Figure <ref type="figure">3D</ref>). To extract the essential motions from the conformation fluctuation, we applied PCA with BAT coordinates to analyze the major motions from the 300-ps time-period MD to quantify the motions. The first Principal Component (PC) revealed the major motions, which removed unessential fluctuations (Figure <ref type="figure">3B</ref>). Presenting these essential motions in the latent space resulted in a very smooth curve (green dots in Figure <ref type="figure">3D</ref>). Note that in the latent space, the PC motions appeared as a nonlinear curve, not a straight line.  Interpolation-generated synthetic conformations of &#945;B-cristallin57-69. Because ICoN identified DOF that determine conformational changes, we used the latent space interpolation as a conformational search engine to find more synthetic conformations. The search was achieved by interpolation between 2 dots with consecutive indexes (i.e., Conf indexes #10 and #11).</p><p>Following the essential motions, we anticipated finding both new conformations and existing ones sampled by MD. Using the same non-linear interpolation function illustrated in Figure <ref type="figure">3D</ref>, 10 points were generated for each pair, yielding 99,990 synthetic conformations from interpolating a total of 9,999 pairs of dots. We first eliminated high-energy conformations, then each new synthetic conformation was compared with its predecessors and eliminated if it was a repeat, yielding 43,484 distinct conformations (Table <ref type="table">1</ref>). We further compared these synthetic conformations with the raw MD data with 500,000 frames to identify 38,425 novel synthetic conformations.  <ref type="table">1</ref>.</p><p>We also examined the stability of the synthetic conformations quantified by their conformational energies using the molecular mechanics/generalized Born and surface areas (MM/GBSA) calculations. Conformations sampled from MD and our deep learning model have similar energy distribution (Figure <ref type="figure">4A</ref>), which validates our synthetic conformations as being thermodynamically stable. In addition, we also compared our synthetic conformations with conformations obtained by another MD run (Table <ref type="table">1</ref>,   </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Latent space interpolation-generated synthetic conformations of A&#946;42</head><p>The A&#946;42 monomer exhibits a broad spectrum of conformations, from random coil to more structured &#945;-helical and &#946;-sheet conformations. Sampling strategies that can accurately model sidechain motions are crucial because the sidechain arrangements govern the intramolecular attractions that lead to various local turns and pre-organized shapes for subsequent oligomerization. Using the same hyperparameters obtained from MD Run1, we used the same strategy for &#945;B-cristallin57-69 to interpolate 2 consecutive dots in the latent space for MD Runs 1 to 5 (Table <ref type="table">1</ref>) to obtain synthetic conformations. The 5 MD runs were initiated by a different</p><p>A&#946;42 monomer structure and/or with a different force field. To elucidate the search efficacy, we plotted the distribution with coordinates of Rg and RMSD (Figure <ref type="figure">5</ref>) for MD Run1-5 and our synthetic conformations from Model 1-5 using these MD runs for training. The search using generative AI efficiently found many new conformations without the need for lengthy MD simulations (Table <ref type="table">1</ref>).</p><p>To validate that our synthetic conformations are thermodynamically stable, we again plotted energy distribution of synthetic conformations found in Model 1 (Figure <ref type="figure">4B</ref>). We also compared the synthetic conformations to another 3 MD runs using the same initial structure and IDP force field (MD Run1.1-1.3) to examine whether our search could find those sampled by other MD simulations. Although the A&#946;42 monomer has numerous conformations, we could identify similar conformations (Figure <ref type="figure">4C</ref>). In addition to checking the total conformational energy, we compared each energy component to verify that both bonded and non-bonded terms have energy distribution similar to those modeled by MD runs (Figure <ref type="figure">S12</ref> and <ref type="figure">S13</ref>). Notably, the energy distribution from the synthetic conformations is relatively broader, thereby suggesting a greater diversity in the sampled conformations, which cover high and low energy regions. The energy comparison underscores the high quality of structural integrity of the synthetic conformations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Analysis and biological implications of novel synthetic A&#946;42 conformation ensembles</head><p>Obtaining A&#946;42 monomer conformations is a critical step for understanding the mechanisms of initial encounters and interactions between monomers, which lead to subsequent oligomerization and fibrillization. Because the oligomerization steps can be highly sensitive to different environments (e.g., different membrane, existing fibrils, or ion concentrations), substantially different monomer conformations may initiate aggregation using different mechanisms in various environments. Several experiments and modeling work also suggest that some conformations are prone to be aggregated or non-toxic <ref type="bibr">[48]</ref><ref type="bibr">[49]</ref><ref type="bibr">[50]</ref><ref type="bibr">[51]</ref><ref type="bibr">[52]</ref> . Because of various experimental results regarding the monomer structure-function relationship in A&#946;42 aggregation, we focus on structures with salt bridges and R5-A42 contacts to demonstrate the utility of the conformational search results from our generative AI model.</p><p>Of note, because our method efficiently sampled numerous low-energy monomer conformations, we used synthetic conformations found using Model 1 (Table <ref type="table">1</ref>) to demonstrate the biological importance of these conformations revealed by our ICoN model. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Major conformation turns in a global tertiary structural ensemble.</head><p>The 127,673 distinct conformations (Table <ref type="table">1</ref>) are presented in the pairwise residue contact map (Figure <ref type="figure">6A</ref>). Four local bends are marked as turn A (F4-H6), turn B (E11-H14), turn C (S26-K28), and turn D (V36-G38). We found novel synthetic conformations with all 4 turns (Figure <ref type="figure">1D</ref>). Turn C has been widely reported in the fibril structures of both A&#946;40 55 and A&#946;42 56 , and is the turn in the commonly referenced "&#946;-turn-&#946;" motif. Cryo-EM studies of brain-derived A&#946;42 fibrils <ref type="bibr">56</ref> reveal that A&#946;42 adopts an S-shaped fold with turns C and D being the two bends in the letter "S". Turns A and B are located near the N-terminal, which is highly flexible and believed to be the metal binding region under abnormal physiological conditions (residues D1 to K16) <ref type="bibr">57</ref> .</p><p>Although the conformations/turns and their population are obtained using only Model 1, they provide an overview of spatial arrangements of the A&#946;42 monomer. Notably, although A&#946;42 is intrinsically disordered, the protein sequence still can lead to many low-energy and highly populated conformations. Turn A locates in the highly flexible N-terminus, where mutation substitutions of Arg5 suppressed the aggregation of A&#946;42 49 . Our EPR measurements also showed distinct spectral features at residue 5, indicating that Arg5 may feature a distinct structure or interactions as compared with the nearby residues (Figure <ref type="figure">7</ref>). The novel synthetic conformations revealed that Arg5 can form interactions with Ala42 with or without the presence of Turn A, as illustrated in Figures <ref type="figure">6B</ref> and <ref type="figure">7C</ref>, respectively. Of note, turn D (V36-G38) is present alongside, which may help to orient and stabilize the C-terminus to form an interaction with Arg5 (red circle in Figures <ref type="figure">6B-C</ref>). Our previous work showed that substituting a nitroxide spin label compound R1 at positions G37 and G38 altered the kinetics of A&#946;42 fibril formation <ref type="bibr">58</ref> , and these synthetic conformations provide a possible monomer structure to guide future experimental design.  Existing studies showed that A&#946;42 monomer has numerous conformations, and MD simulations using initial conformation and/or force field can sample different conformations with few repeats <ref type="bibr">42,</ref><ref type="bibr">43,</ref><ref type="bibr">47,</ref><ref type="bibr">59</ref> . Using the hyperparameters obtained from MD Run 1 to train datasets with conformations different from MD Run2-5, new intermediate states important in oligomerization were identified with ICoN. For example, as illustrated in Figure <ref type="figure">6</ref> and <ref type="figure">8</ref>, novel synthetic conformations from the FF14SB force field (Model 3 in Table <ref type="table">1</ref>) reveal an intermediate state with a partial hairpin conformation which only requires minimal conformational arrangement to form and experimental determined A&#946;42 fibril from human brain 56 (Figure <ref type="figure">8A</ref>). The partial hairpin conformations pre-organize hydrophobic regions ranging from L17-F20 and I31-L34</p><p>with residues F19 or F20 rotate inward to stabilize a hydrophobic core with surrounding nonpolar residues (Figure <ref type="figure">8B-C</ref>). Interestingly, different directions of F19 and F20 could contribute to different types of brain-isolated fibril structures <ref type="bibr">56</ref> , and our ICoN model suggested the critical roles of their rearrangements. In addition, another partial hairpin conformation that has potential of forming A&#946;42 tetramer structure, which comprises a six stranded &#946;-sheet with 2 &#946;-strands in the middle and 2 &#946;-turn-&#946; on the side 60 , are also observed (Figure <ref type="figure">S14</ref>). In A&#946;42 tetramer, sidechains of F19 and F20 need to rotate outward to interact with surrounding monomers (Figure <ref type="figure">S14A</ref>). The partial hairpin conformation at Turn C leads to the formation of fibril conformations (Figure <ref type="figure">6D</ref>, <ref type="figure">8</ref>, and <ref type="figure">S14</ref>). Although the models were constructed using MD simulations for monomer, our synthetic conformations suggest a highly plausible intermediate state during fibril formation. Furthermore, novel synthetic conformations from ESFF1 force field (Model 5 in Table <ref type="table">1</ref>) show that turn C and turn D create highly packed hydrophobic regions prone to oligomerization. Specifically, we observed that V18, F20, I32, and L34 form a hydrophobic core (Figure <ref type="figure">6E</ref>). Notably, unlike hydrogen bonds which are highly specific with additional geometry restraints, hydrophobic regions usually allow fluid-like sidechain movements to retain their flexibility. This different partial hairpin conformation from Figure <ref type="figure">6D</ref> again preorganized the hydrophobic regions and preserve conformational plasticity for oligomerization.</p><p>Presence of salt bridges in toxic and less toxic monomer conformations. Salt bridges play an important role in providing intra-molecular attractions and conformational specificity that directly relate to A&#946;42 aggregation. This non-covalent interaction forms when sidechains of oppositely charged residues are close enough to each other to experience electrostatic attractions.</p><p>Therefore, protein structures must be described with precise atomistic details and sidechain motions sampled accurately. Although the D23-K28 salt bridge did not exist in the 1-&#181;s MD Run1 used in our training set, by using interpolation in the latent space, the model sampled novel synthetic conformations with a new salt bridge, D23-K28 (Figure <ref type="figure">9</ref>), which is reported to promote aggregation <ref type="bibr">48</ref> . Experiments suggested the importance of this salt bridge, but the conformation ensemble was never determined experimentally. The novel synthetic conformations revealed distinctly different rearrangements with the D23-K28 salt bridge, ranging from more extended to highly packed conformations (Figure <ref type="figure">9</ref>). The D23-K28 constraint results in Turn C, which is important for oligomerization. The synthetic conformations revealed novel structures with the salt bridge E22-K28</p><p>(Figures <ref type="figure">9D-F</ref>). This salt bridge was reported to prevent a toxic turn at positions E22 and D23, for potentially a less-toxic monomer <ref type="bibr">61</ref> . The synthetic conformation presents key features suggested by experiments, with no turn at positions 22 and 23 when the salt bridge E22-K28 exists (Figure <ref type="figure">9</ref>). Nevertheless, other publications suggested that the salt bridge E22-K28 may help form other intramolecular contacts (e.g., contacts between the C-terminus and Arg5) to stabilize a hairpin structure and promote dimer formation <ref type="bibr">49</ref> . Although which salt bridges promote or inhibit monomer aggregation is debatable, our novel synthetic conformations provide various possible sidechain rearrangements and salt-bridge conformations for further interpreting experiments in understanding aggregation mechanisms.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Comparing ICoN with Phanto-IDP</head><p>Here we compare the conformational sampling using ICoN and Phanto-IDP, a new DL model developed by the Chen's group which can efficiently sample backbone conformations of IDPs.</p><p>Using backbone only representation, Phanto-IDP utilizes a graph-based encoder and a variational layer to describe protein features. The variational layer sample neighboring area of points in latent space with mean, stand deviation, random variable temperature to yield impressive sampling results with high diversity in a recent publication <ref type="bibr">47</ref> . In this paper, the Chen group used MD Run 5 to train and generate new conformations with Phanto-IDP, where the raw data had backbone only representation. For comparison, we added sidechains to their synthetic backbone conformations and carried out minimization and subsequent steps to obtain distinct newly synthetic conformations. Using the same MD Run 5, we applied the hyperparameters obtained from our MD Run 1 to sample new synthetic conformations with ICoN (Table <ref type="table">1</ref>). Figure <ref type="figure">S15</ref> presents similar plot as </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>In this work, we demonstrated the capability of a deep learning model to directly generate new protein conformational ensembles by learning the physics of protein motions. Our approach utilized a neural network trained with protein conformations from classical MD simulations with all-atom representation in explicit solvent. We saved frames from MD trajectories without further processing, reducing protein atomic coordinates to a 3D latent space, where each conformation is a data point that can be accurately reconstructed back into a protein structure.</p><p>ICoN uses nonlinear interpolation between 2 data points in the 3D latent space to guide the search for conformational changes, identifying natural molecular motions.</p><p>The vector-based BAT (vBAT) feature overcomes problems Cartesian coordinates that fail to smoothly present dihedral rotation during conformation sampling, which is crucial in sampling large-scale motions, especially in IDPs. vBAT also effectively handles dihedral periodicity. The model is highly efficient for identifying novel synthetic A&#946;42 conformation ensembles not seen in the MD trajectory. These ensembles have different large-scale motions (i.e., new compact conformations) and detailed sidechain rearrangements such as forming a new salt bridge between D23 and K28. The synthetic A&#946;42 conformations sampled by the ICoN model provide a comprehensive view of the conformational landscape for A&#946;42 monomers. For example, the ICoN model finds salt bridges involving residues E22 or D23 (Figure 8), both of which are hotspots for familial mutations in Alzheimer's disease. The ICoN model also identifies four local bends at residue positions 4-6, 11-14, 26-28, and 36-38 (Figures <ref type="figure">1D</ref> and <ref type="figure">6</ref>). Although all 4 turns have been previously observed in isolation, it is worth noting that the experimentally determined A&#946;42 structures do not contain all these turns in the same structure. For example, turn D has been found in the brain-derived A&#946;42 fibril structures <ref type="bibr">56</ref> , but is absent in the A&#946;42 oligomers formed in the presence of detergents <ref type="bibr">60</ref> . The co-occurrence of different turns in the synthetic A&#946;42 conformations may help decode the A&#946; aggregation pathways that lead to the formation of specific oligomers and fibrils.</p><p>Training of the ICoN model requires sufficient protein conformations, so the DL can reveal natural motions of the protein to guide the search toward new low-energy structures. Since experimentally determined structures typically cover only a small proportion of the overall conformation ensembles, we rely on MD simulations. Force-field parameters used in MD may affect the performance in sampling. For example, the model trained by MD runs using a less flexible ff14SB yielded fewer novel synthetic conformations (Table <ref type="table">1</ref>). However, our training set did not need to cover the complete conformational distribution. As long as conformation fluctuation governed by key dihedral rotations can be included in the training set, the model identify these key elements for guiding the conformational search to generate novel synthetic conformations.</p><p>Conceptually, ICoN performs as a conformational search method rather than a deterministic simulation method. Therefore, the search results do not directly reflect the equilibrium distributions of protein conformations. While structured proteins have highly stable global energy minima, IDPs exhibit numerous fluctuating heterogeneous conformations with similar energy. The IDPs can be highly sensitive to their environment, making it impractical to precisely reproduce physiological conditions in experiments or simulations. Therefore, thoroughly finding IDP conformational ensembles is more practical and useful instead of aiming for the population distribution in equilibrium under specific physiological environment.</p><p>Importantly, unlike structured proteins whose native structures are functionally important, metastable conformations of IDPs can be crucial in performing the biological function. For example, experiments showed that A&#946;42 monomer aggregation has a rate-limiting nucleation-dependent polymerization process <ref type="bibr">52</ref> , which suggests that the highly populated monomer conformations might not be the ones that drive nucleation, and the monomers are not pre-organized as a readyto-aggregate conformation. Therefore, as compared with the conformational search for structure proteins, thoroughly sampling meta-stable conformations for IDPs is critical. As demonstrated in this study, the novel synthetic A&#946;42 monomer conformations provide atomistic details to support several experimental observations, which also bring insights into mechanisms of oligomerization.</p><p>Conformational changes of biomolecules follow principles of physics, which are consequences of different arrangements of atoms rotating around a set of single bonds, the internal degrees of freedom having the highest flexibility. Both PCA and NMA describe the natural motions of a molecule, which are commonly used to determine essential protein dynamics and guide conformational sampling. The two techniques are linear transforms that extract the most important elements in a data matrix, a covariance or Hessian matrix. However, real protein systems can have very complicated and higher-order correlations that have intrinsic nonlinear effects and may not be well described using standard PCA or NMA. Moreover, although PC space is commonly used for data analysis, there is no function to convert manipulated data points in PC space back to its atomistic protein structure. The activation function used in the neural network adds nonlinear effects to process nonlinear features. With use of BAT, a transition between two proteins can be achieved by interpolating two data points (conformations) in the latent space. As illustrated in Figure <ref type="figure">4</ref>, if one obtains conformations along the smooth green curve in the latent space, the conformational changes reproduce the motions following the first PC mode. As for the natural motions identified by PCA or NMA, the deep learning algorithm learns natural motions of the protein system (black dot line in Figure <ref type="figure">4</ref>), and new conformations can be found by processing the interpolations in the latent space.</p><p>Although the ICoN model is fast, taking a few minutes to train 10,000 frames of A&#946;42 and hours to perform quick energy minimization for the synthetic conformations generated from latent space interpolation using a GPU card, the training set is from a physics-based MD sampling. It typically takes days, if not weeks, to perform sufficiently long MD simulations to obtain the training set. To speed up the sampling, one can apply the ICoN model for preliminary results to select dissimilar protein conformations to seed more MD runs <ref type="bibr">19</ref> . An advantage in using BAT-based coordinates is that one can easily select a set of dihedral angles that determine the motions of interest for training, instead of using all DOF. For small proteins or peptides such as A&#946;42 and &#945;B-crystallin57-69, our classical internal BAT (Z-matrix)-based vBAT nicely presents local sidechain rearrangements and large-scale backbone motions. However, for larger (&gt;200 residues) proteins, rotations of some dihedral angles near the root atoms may result in unrealistic motions on the far end of the protein. The issue can be addressed easily by using multilayer BAT coordinates, assigning multiple fragments/chains for a protein or multi-protein complex. Also, reconnecting fragments with pseudo-DOF can avoid the accumulated dependence problem <ref type="bibr">62</ref> . Future work will implement the multilayer BAT coordinates to eliminate the protein size limitation when building the ICoN model. Various interpolation and extrapolation strategies in the latent space will be examined for generating novel synthetic conformations as well.</p></div></body>
		</text>
</TEI>
