<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Deep Learning Approach to Identify Protein’s Secondary Structure Elements</title></titleStmt>
			<publicationStmt>
				<publisher>Springer, Singapore</publisher>
				<date>07/12/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10541118</idno>
					<idno type="doi"></idno>
					
					<author>Mohammad Bataineh</author><author>Kamal Al_Nasr</author><author>Richard Mu</author><author>Mohammed Alamri</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[method for structure determination. Despite the substantial growth in deposited cryo-EM maps driven by advances in microscopy and image processing, accurately constructing models from these maps remains challenging. Extracting secondary structure information from EM maps is valuable for cryo-EM modeling. In this context, we introduce a novel deep learning secondary structure annotation framework specifically designed for intermediate-resolution cryo-EM maps, employing a three-dimensional Inception architecture. Testing it on diverse datasets, including maps with authentic intermediate resolutions, demonstrates its accuracy and robustness in identifying secondary structures in cryo-EM maps. We conducted a comparative analysis of our results against frameworks that exist in the state-of-the-art, and our framework demonstrated superior performance across nearly all secondary structure elements. We employed the F1 accuracy metric, yielding an average F1 score of 0.657 for helix, 0.712 for coil, and 0.596 for sheet predictions. Notably, certain helix and sheet predictions achieved an impressive F1 score of 0.881.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>Progress in both microscopy tools and image processing algorithms has resulted in a growing abundance of cryo-electron microscopy (cryo-EM) maps <ref type="bibr">[1]</ref><ref type="bibr">[2]</ref><ref type="bibr">[3]</ref>. Increasing the resolution of cryo-EM has opened doors to elucidating the structures of biological systems that were once considered too challenging to tackle, now achieving remarkable levels of detail <ref type="bibr">[4,</ref><ref type="bibr">5]</ref>. It is important to note that the ultimate objective of cryo-EM is not merely the acquisition of 3D maps but the precise determination of atomic structure <ref type="bibr">[6]</ref><ref type="bibr">[7]</ref><ref type="bibr">[8]</ref>.</p><p>Constructing precise structural models for cryo-EM maps poses a significant challenge <ref type="bibr">[9]</ref>. The methods typically employed, such as rigid fitting and flexible fitting, rely on pre-existing template structures for the accurate placement of atomic structures into EM maps. When template structures are lacking, the necessity for de novo modeling tools arises to construct complete atomic models within EM density maps.</p><p>Jiang et al. <ref type="bibr">[10]</ref> developed Helixhunter, software for helix identification, length, and orientation using cross-correlation search and feature extraction on density maps. They achieved over 88% helix detection accuracy on 8 &#197; resolution simulated maps, with some misclassifications and missed helices. Kong and Ma <ref type="bibr">[11]</ref> introduced Sheetminer, successfully identifying Beta-sheets in protein structures with promising results on Cryo-EM and X-ray density maps at various resolutions. Kong et al. <ref type="bibr">[12]</ref> developed two methods for Beta-sheet identification, one relying on Sheetminer's output and the other using deconvolution for improved results. Despite advancements, de novo modeling tools face accuracy limitations, leading to a gap between the quantity of cryo-EM maps and successfully reconstructed 3D structures <ref type="bibr">[13]</ref>. Machine learning models offer potential to enhance predictions based on their capabilities.</p><p>Li et al. <ref type="bibr">[14]</ref> achieved significant progress in secondary structure prediction using a CNN framework, testing it on 25 simulated 8 &#197; cryo-EM maps, achieving an average sensitivity and specificity of 71.52% and 97.86% for Alph-helix and Beta-sheet detection, surpassing SVM methods. Subramaniya et al. <ref type="bibr">[15]</ref> introduced Emap2sec, predicting secondary structure in cryo-EM maps. Their validation involved two datasets, yielding impressive results at 6 &#197; with an overall F1 score of 0.798, Alph-helix at 0.848, Beta-sheet at 0.828, and other structural elements at 0.672. At 10 &#197; resolution, results remained substantial, with Alphhelix, Beta-sheet, and other elements achieving F1 accuracy scores of 0.82, 0.75, and 0.64, respectively.</p><p>Shifting the focus to another framework, Haruspex, developed by Mostosi et al. <ref type="bibr">[16]</ref>, employed U-net architecture to predict secondary structure elements. Notably, Haruspex was primarily designed for detecting and annotating RNA/DNA and protein secondary structure elements within high-resolution cryo-EM maps. This framework showed promising results but faced challenges in 'unassigned' regions, resulting in an unbalanced classification that impacted its efficiency.</p><p>Later Wang et al. <ref type="bibr">[17]</ref> presented Emap2sec+, an updated version of Emap2sec. Where deep Residual convolutional neural network architecture was developed, ResNet. To predict the secondary structure elements, they tried to classify each voxel into one of three elements: Alph-Helix, Beta-sheet, or others. Emap2sec+ was trained and tested on the simulated and experimental datasets. In the simulated dataset, 108 non-redundant maps at 6 and 10 &#197; were used for training and testing. For the experimental dataset, 83 cryo-EM images were used for training and testing as well. In addition, Emap2sec+ outperformed Haruspex in predicting protein secondary structure in the F1 score term.</p><p>He and Huang <ref type="bibr">[18]</ref> recently introduced EMNUSS, a robust framework utilizing advanced U-net architecture (nested U-net or U-net++) with skip connectors to enhance predictive power in secondary structure element (SSE) prediction. EMNUSS showcased its capabilities across diverse datasets, spanning simulated, mid-resolution, and high-resolution maps, demonstrating significant potential in SSE prediction. It consistently outperformed other frameworks in various evaluation metrics, such as F1 scores and Q3 accuracy, making it a notable contribution to the field of SSE prediction.</p><p>Limitations, such as using improper grid intervals as seen in Haruspex and relying on small input chunks in Emap2sec and Emap2sec+, result in constraints that may not accommodate an average secondary structure element larger than those input chunks. In order to address the limitations of current methods, we've introduced an innovative deep learning framework for predicting secondary structures in authentic cryo-EM maps at intermediate resolutions. Our approach utilizes a three-dimensional (3D) Inception architecture, enabling rapid and precise prediction of protein secondary structures in cryo-EM maps of diverse dimensions. Our method has demonstrated a substantial enhancement in performance, particularly when applied to experimental maps at middle resolutions. Our approach has been tested exclusively on intermediate-resolution maps, making it specifically tailored for such data. This opens up opportunities for future research to test our approach on low-or high-resolution maps, or to develop a new model designed to handle these different resolutions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1">Dataset</head><p>In order to develop and assess the performance of our framework, we compiled a diverse and non-redundant set of intermediate-resolution electron microscopy (EM) maps for experimentation. Our initial search was conducted within the EMDataResource database, targeting EM maps with resolutions falling within the 4 to 10 Angstrom range, while also ensuring the availability of associated PDB files. However, these criteria alone did not suffice to create a dependable dataset. To prevent the inclusion of identical or highly similar EM maps, we implemented additional selection constraints. Specifically, we excluded any chains displaying the following characteristics: 1. Presence of missing residues. 2. Absence of secondary structure information. 3. A sequence identity similarity exceeding 25% with any chain already present in the dataset.</p><p>Following the application of these aforementioned selection conditions, we successfully curated a collection of 487 unique chain maps. We divided these maps into two distinct sets for training and testing purposes: 455 maps were randomly selected for the training set, while the remaining 32 maps constituted the testing set.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Network Architecture</head><p>The Inception architecture represents a significant milestone in computer vision and deep learning. In our work, we have incorporated the Inception architecture, creating our own custom variant of this 3D deep convolutional neural network (CNN) architecture. Our proposed architecture, as depicted in Fig. <ref type="figure">1</ref>, comprises several key elements. It commences with a Stem layer, responsible for preprocessing the 3D data. This is followed by a Maxpool layer, subsequently leading to the core Inception blocks layer. After that, an Up-sampling layer is applied, and the network concludes with a Final layer, which reduces the neuron count to 4 to align with the number of classes we intend to predict. A more detailed breakdown of the Inception blocks and their sublayers can be found in Fig. <ref type="figure">2</ref>. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3">Processing Training Data</head><p>To prepare our training maps for analysis, we conducted the following steps. One of our goals was to annotate the dataset for both training and testing stages effectively. Uniform Grid and Resolution: We started by implementing a standardized interval grid for all maps using trilinear interpolation. This process ensured that each interval was precisely set to 1.0 &#197;, creating consistency across the dataset. Ground Truth Assignment: Subsequently, we assigned ground truth labels to each voxel to identify secondary structure elements. For this task, we associated each voxel with the nearest backbone atom (N, C, C-alph, or O atom) within a 3.0 &#197; radius.</p><p>In cases where no backbone atoms were found within this radius, we assigned a background label to the voxel. The input size for the framework was standardized to 40 &#215; 40 &#215; 40, ensuring a consistent format for analysis. Density Value Normalization: Finally, we normalized the density values for each voxel to fall within the range of 0 to 1. This normalization process allowed us to exclude any chunks with all values equal to zero, preventing their participation in the training process. This approach contributed to the effectiveness of the training procedure. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.4">Network Training</head><p>We created three distinct Inception architectures, applying identical hyperparameters to each. The input data consisted of chunks with dimensions 40 &#215; 40 &#215; 40. We partitioned 10% of the training maps for validation. Our framework was implemented using PyTorch, with 150 epochs and a batch size of 16. We utilized the Adam optimizer and employed the cross-entropy loss function, setting the learning rate at 1e-3.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.5">Evaluation and Comparison</head><p>To assess the outcomes of our study, we employed the F1 score metric, which represents the balanced combination of precision and recall when assessing assignments. This metric was utilized to gauge our framework's performance at the voxel level. For a comprehensive evaluation against the current state of the art, we opted to compare our results with EMNUSS, a recent and widely recognized and resilient framework used for forecasting secondary structure elements in Cryo-EM maps. We computed the F1 score for both frameworks, revitalizing some of the results to provide a more insightful assessment of their performance.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Results and Discussion</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Comparison with EMNUSS</head><p>We conducted a comparative analysis of our framework and EMNUSS using a test set comprised of 32 experimental EM maps of middle-resolution, with resolutions spanning from 4.0 to 10.0 &#197;. In Fig. <ref type="figure">3</ref>, we present a visual representation of the voxel F1 score comparisons between the two methods. The figure clearly demonstrates that our framework outperformed EMNUSS significantly when applied to the middle-resolution experimental maps.</p><p>Our proposed method significantly outperforms EMNUSS across all classes in terms of voxel F1 scores, achieving an Overall F1 accuracy of 0.739 compared to EMNUSS's 0.277. For Helix prediction, the Proposed Method F1 scores 0.657 versus EMNUSS's 0.14, and for Sheet prediction, it achieves 0.596 F1 score against EMNUSS's 0.093. In Coil prediction, the Proposed Method also leads with a 0.712 F1 score compared to EMNUSS's 0.598. This demonstrates the superior reliability and effectiveness of our Proposed Method for protein structure prediction. In Fig. <ref type="figure">4</ref>, the comparison between EMD-1263H map visualizations using both the proposed method and EMNUSS reveals significant performance differences. Our framework achieves a notable overall F1 accuracy of 0.778, surpassing EMNUSS by a large margin (0.277), particularly excelling in identifying helices and strands. Additionally, our method outperforms EMNUSS in predicting secondary structure classes, including alpha helices, beta-sheets, and coils, with higher voxel F1 scores. EMNUSS struggles with accuracy across classes, mislabeling helical structures within sheet regions and misidentifying coil and sheet regions consistently. While our method exhibits slightly lower F1 scores in sheet prediction (0.59), it provides smoother and more interpretable visualizations compared to EMNUSS, with highly accurate predictions for helix and coil regions, although minor misses exist. Overall, our approach demonstrates superior performance across all structural classes.</p><p>Figure <ref type="figure">5</ref> showcases EMD 8169-C, a 6.56 &#197; map, with predictions from both our proposed method and EMNUSS. Our method notably outperforms EMNUSS, particularly excelling in accurately predicting sheet regions with a high F1 score of 0.872. Similarly, our method demonstrates strong performance in identifying coil regions, although occasional misses occur. However, helix prediction regions show lower accuracy, with some expected helical regions missed while others are incorrectly labeled as helices. In contrast, EMNUSS misclassifies segments, labeling them predominantly as coil or background and overlooking helix and sheet regions. Helix regions are entirely overlooked, with sheet regions misidentified as coil or background. The predicted sheet region by EMNUSS represents only a fraction of the actual area, located differently.</p><p>Figure <ref type="figure">6</ref> illustrates EMD-12221A, a 9.5 &#197; resolution map, lacking sheet regions in the protein chain. Our proposed framework showcases highly satisfactory results in predicting secondary structure elements, achieving an F1 score of 0.88 for helix prediction, with nearly flawless visualization despite occasional missing data. Coil prediction also demonstrates excellent visualization and an impressive F1 score. Conversely, EMNUSS performs poorly, largely failing to identify helix regions and misclassifying them as coil or background, while erroneously identifying coil regions as predominant elements. The most glaring error in EMNUSS predictions is its incorrect labeling of sheet regions despite their absence in this protein chain.    Although coil regions were relatively better predicted, many were still overlooked. Overall, our proposed method did not perform as well in predicting this protein chain compared to previous ones, exhibiting significant shortcomings. Despite EMNUSS achieving a higher overall F1 score, its visualization results in Fig. <ref type="figure">7</ref> reveal poor predictions, with most protein regions misclassified as coils and significant overlaps between different predicted classes. Specifically, EMNUSS completely missed identifying helix and sheet regions within the protein.</p><p>Figure <ref type="figure">8</ref> illustrates EMD 21156-A, portraying a complex and tangled protein structure. Our proposed method accurately predicts approximately 50% of the helices while struggling with sheet prediction, resulting in a low F1 score of 0.16, despite some correct identifications. However, it performs relatively well in predicting coil regions, with a decent F1 score of 0.6. In contrast, EMNUSS achieves higher F1 scores for both sheets and coils but produces a chaotic visualization lacking clarity, with numerous helices misidentified as coils and sheets appearing dislocated and misclassified. Despite the higher F1 scores, our proposed method's visualization offers better coherence and understandability compared to EMNUSS, indicating its superiority in presenting predictions despite similar struggles in accuracy.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Comparison with Different Architecture</head><p>To further scrutinize the effectiveness of our proposed framework, we've introduced an additional pair of Inception block architectures, we named them second and third designs. Our approach involves replacing the Inception blocks (as shown in Fig. <ref type="figure">2</ref>) within our network to explore alternative designs and understand their impact on the results. In the following sections, we'll outline these two additional Inception block designs. &#8226; Second design: The design includes four Conv3D branches also, each consisting of a sequence: a 1 &#215; 1 &#215; 1 Conv3D layer, followed by a 3 &#215; 3 &#215; 3 Conv3D layer, then a 5 &#215; 5 &#215; 5 Conv3D layer. After this, there's a pool branch with a Conv3D layer. Additionally, before each of these branches, a 1 &#215; 1 &#215; 1 kernel Conv3D sublayer this time to transform the data without big change in the kernel size. &#8226; Third design: This architecture comprises four Conv3D branches: a 1 &#215; 1 &#215; 1 Conv3D layer, followed by a 3 &#215; 3 &#215; 3 Conv3D layer, then a 5 &#215; 5 &#215; 5 Conv3D layer, and finally a pool branch with a Conv3D layer. Each branch is preceded by a 3 &#215; 3 &#215; 3 kernel Conv3D sublayer, strategically included to prepare and enhance the 3D data, ultimately leading to improved performance.</p><p>We conducted a comparative analysis of three Inception block designs, all evaluated using the same test set, consisting of 32 experimental EM maps at middle resolution. Table <ref type="table">1</ref> clearly illustrates that our primary design outperformed the second design significantly. However, the third design displayed highly competitive results, almost on par with the first design.</p><p>We attribute the excellent performance of the designs to the effective use of the 3 &#215; 3 &#215; 3 kernel Conv3D sublayer, which proved to be well-suited for Cryo-EM maps, capable of extracting meaningful information. This kernel size is particularly conducive to this type of data.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Conclusion</head><p>We created an advanced deep learning framework designed for the prediction and annotation of protein secondary structures within EM density maps. This framework underwent comprehensive testing and evaluation using a dataset of middle-resolution experimental maps. The results clearly demonstrated that our framework substantially enhanced the accuracy of secondary structure detection, surpassing existing methods. Furthermore, we introduced two additional Inception block designs to investigate their influence on the results. A promising direction for future research involves developing an improved network architecture aimed at enhancing prediction accuracy.</p></div></body>
		</text>
</TEI>
