This content will become publicly available on November 1, 2026

Title: MPBind: a multitask protein binding site predictor using protein language models and equivariant GNNs
Motivation. Proteins interact with a variety of molecules, including other proteins, DNAs, RNAs, ligands, ions, and lipids. These interactions play a crucial role in cellular communication, metabolic regulation, gene regulation, and structural integrity, making proteins fundamental to nearly all biological functions. Accurately predicting protein interaction (binding) sites is essential for understanding protein interaction and function. ResultsIn this work, we introduce MPBind, a multitask protein binding site prediction method, which integrates protein language models (PLMs) that can extract structural and functional information from sequences and equivariant graph neural networks (EGNNs) that can effectively capture geometric features of 3D protein structures. Through multitask learning, it can predict binding sites on proteins that interact with five key categories of binding partners: proteins, DNA/RNA, ligands, lipids, and ions. MPBind generalizes across the five molecular classes with state-of-the-art accuracy, achieving AUROC scores of 0.83 and 0.81 for protein–protein and protein–DNA/RNA-binding site prediction, respectively. Moreover, MPBind outperforms both general and task-specific binding site prediction methods, making it a useful, versatile tool for protein binding site prediction. Availability and implementationThe source code of MPBind is available at the GitHub repository: https://github.com/jianlin-cheng/MPBind.  more » « less
Award ID(s):
2308699 2343612
PAR ID:
10684337
Author(s) / Creator(s):
; ;
Editor(s):
Cowen, Lenore
Publisher / Repository:
Oxford University Press
Date Published:
Journal Name:
Bioinformatics
Volume:
41
Issue:
11
ISSN:
1367-4803
Subject(s) / Keyword(s):
Protein function prediction, machine learning
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Cowen, Lenore (Ed.)
    Abstract Motivationmetal-binding proteins have a central role in maintaining life processes. Nearly one-third of known protein structures contain metal ions that are used for a variety of needs, such as catalysis, DNA/RNA binding, protein structure stability, etc. Identifying metal-binding proteins is thus crucial for understanding the mechanisms of cellular activity. However, experimental annotation of protein metal-binding potential is severely lacking, while computational techniques are often imprecise and of limited applicability. Resultswe developed a novel machine learning-based method, mebipred, for identifying metal-binding proteins from sequence-derived features. This method is over 80% accurate in recognizing proteins that bind metal ion-containing ligands; the specific identity of 11 ubiquitously present metal ions can also be annotated. mebipred is reference-free, i.e. no sequence alignments are involved, and is thus faster than alignment-based methods; it is also more accurate than other sequence-based prediction methods. Additionally, mebipred can identify protein metal-binding capabilities from short sequence stretches, e.g. translated sequencing reads, and, thus, may be useful for the annotation of metal requirements of metagenomic samples. We performed an analysis of available microbiome data and found that ocean, hot spring sediments and soil microbiomes use a more diverse set of metals than human host-related ones. For human microbiomes, physiological conditions explain the observed metal preferences. Similarly, subtle changes in ocean sample ion concentration affect the abundance of relevant metal-binding proteins. These results highlight mebipred’s utility in analyzing microbiome metal requirements. Availability and implementationmebipred is available as a web server at services.bromberglab.org/mebipred and as a standalone package at https://pypi.org/project/mymetal/. Supplementary informationSupplementary data are available at Bioinformatics online. 
    more » « less
  2. Abstract Transcription factors (TFs) regulate gene expression by binding to specific DNA sites on genome, making accurate TF binding site prediction critical for understanding gene regulation and downstream phenotypes. Almost all current deep learning methods use only DNA-related information to predict TF binding sites, ignoring the fact that different TF protein sequences and structures recognize distinct DNA patterns. Not leveraging TF information not only limits prediction accuracy but also makes the methods not generalizable to predicting binding sites of new TFs that do not exist in the training data. Here, we present TransBind, a protein-aware deep learning architecture that integrates DNA sequence information with protein embeddings containing both sequence and structural information derived from a protein language model pretrained on DNA-binding proteins, to improve TF binding site prediction. Through the cross-attention, a TF embedding selectively attends to genomic regions according to its unique binding properties. Evaluated on the data of 690 ChIP-seq experiments spanning 161 TFs across 91 human cell types, TransBind achieves an AUROC of 0.9508 and AUPR of 0.3741—representing a $$\ge$$11.8% relative AUPR improvement over state-of-the-art methods including TBiNet, EPBDXDNABERT-2, DanQ, and DeepSEA. The model outperformed existing methods in $$\ge$$98% of TF–cell type combinations. It also recovered 160 known TF binding motifs in the JASPAR database, providing the biological interpretability of the model. Moreover, the approach enables label-zero-shot prediction for unseen TFs, demonstrating its potential of generalizing to new, poorly characterized TFs. The source code of TransBind is available at https://github.com/jianlin-cheng/TransBind. The version used in this work is archived at https://doi.org/10.5281/zenodo.19462292. 
    more » « less
  3. Abstract BackgroundDe novo DNA methylation by DNMT3A is a fundamental epigenetic modification for transcriptional regulation. Histone tails and regulatory proteins regulate DNMT3A, and the crosstalk between these epigenetic mechanisms ensures appropriate DNA methylation patterning. Based on findings showing thatFosecRNA inhibits DNMT3A activity in neurons, we sought to characterize the contribution of this regulatory RNA in the modulation of DNMT3A in the presence of regulatory proteins and histone tails. ResultsWe show thatFosecRNA and mRNA strongly correlate in primary cortical neurons on a single cell level and provide evidence thatFosecRNA modulation of DNMT3A at these actively transcribed sites occurs in a sequence-independent manner. Further characterization of theFosecRNA-DNMT3A interaction showed thatFos-1ecRNA binds the DNMT3A tetramer interface and clinically relevant DNMT3A substitutions that disrupt the inhibition of DNMT3A activity byFos-1ecRNA are restored by the formation of heterotetramers with DNMT3L. Lastly, using DNMT3L andFosecRNA in the presence of synthetic histone H3 tails or reconstituted polynucleosomes, we found that regulatory RNAs play dominant roles in the modulation of DNMT3A activity. ConclusionOur results are consistent with a model for RNA regulation of DNMT3A that involves localized production of short RNAs binding to a nonspecific site on the protein, rather than formation of localized RNA/DNA structures. We propose that regulatory RNAs play a dominant role in the regulation of DNMT3A catalytic activity at sites with increased production of regulatory RNAs. 
    more » « less
  4. Abstract Molecular interactions underlie nearly all biological processes, but many representation learning models either focus on single entities or are trained for a narrow set of interaction settings. Here, we introduce ATOMICA, a geometric deep learning model that learns atomic-scale representations of intermolecular interfaces across five modalities, including proteins, small molecules, metal ions, lipids, and nucleic acids. ATOMICA is trained on 2,037,972 interaction complexes to generate embeddings of interaction interfaces at the levels of atoms, chemical blocks, and molecular interfaces. The latent space is multiscale and reflects physicochemical features shared across molecular classes. On the RNAGlib 3D structure-function benchmark, ATOMICA attains the best performance across four tasks, and in protein pocket ligand classification, ATOMICA improves upon established protein pocket encoders and is comparable to protein language models. Using the shared embedding space, we embed or-thosteric PPI inhibitors and find inhibitor embeddings are more similar to interface embeddings proximal to the native binding site across protein–peptide and protein–protein complexes. We use ATOMICA to suggest putative ligands to pockets in the dark proteome, which are proteins lacking known function. In total, ligands are predicted for 2,646 dark protein pockets and heme binding is experimentally confirmed for five ATOMICA predictions. ATOMICA opens new avenues for learning representations of intermolecular interactions. 
    more » « less
  5. Abstract The binding and interaction of proteins with nucleic acids such as DNA and RNA constitutes a fundamental biochemical and biophysical process in all living organisms. Identifying and visualizing such temporal interactions in cells is key to understanding their function. To image sites of these events in cells across scales, we developed a method, named PROMPT for PROximal Molecular Probe Transfer, which is applicable to both light and correlative electron microscopy. This method relies on the transfer of a bound photosensitizer from a protein known to associate with specific nucleic acid sequence, allowing the marking of the binding site on DNA or RNA in fixed cells. The method produces a fluorescent mark at the site of their interaction, that can be made electron dense and reimaged at high resolution in the electron microscope. As proof of principle, we labeled in situ the interaction sites between the histone H2B and nuclear DNA. As an example of application for specific RNA localizations we labeled different nuclear and nucleolar fractions of the protein Fibrillarin to mark and locate where it associates with RNAs, also using electron tomography. While the current PROMPT method is designed for microscopy, with minimal variations, it can be potentially expanded to analytical techniques. 
    more » « less