Title: Transcription factor enrichment analysis (TFEA) quantifies the activity of multiple transcription factors from a single experiment
Detecting changes in the activity of a transcription factor (TF) in response to a perturbation provides insights into the underlying cellular process. Transcription Factor Enrichment Analysis (TFEA) is a robust and reliable computational method that detects positional motif enrichment associated with changes in transcription observed in response to a perturbation. TFEA detects positional motif enrichment within a list of ranked regions of interest (ROIs), typically sites of RNA polymerase initiation inferred from regulatory data such as nascent transcription. Therefore, we also introduce muMerge, a statistically principled method of generating a consensus list of ROIs from multiple replicates and conditions. TFEA is broadly applicable to data that informs on transcriptional regulation including nascent transcription (eg. PRO-Seq), CAGE, histone ChIP-Seq, and accessibility data (e.g., ATAC-Seq). TFEA not only identifies the key regulators responding to a perturbation, but also temporally unravels regulatory networks with time series data. Consequently, TFEA serves as a hypothesis-generating tool that provides an easy, rigorous, and cost-effective means to broadly assess TF activity yielding new biological insights.  more » « less
Award ID(s):
1759949
PAR ID:
10340673
Author(s) / Creator(s):
; ; ; ; ; ; ;
Date Published:
Journal Name:
Communications biology
Volume:
4
ISSN:
2399-3642
Page Range / eLocation ID:
661
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Precisecis-regulatory control of gene expression is essential for plant growth. InArabidopsis thaliana, PLANT PEPTIDE CONTAINING SULFATED TYROSINE (PSY) peptides and their receptors (PSYRs) mediate growth-stress trade-offs, yet the transcriptional regulation of these genes remains poorly understood. Here, we mapped transcription factor (TF)-promoter interactions for ninePSYand threePSYRgenes by combining high-throughput enhanced yeast one-hybrid screening with DNA Affinity Purification sequencing (DAP-seq) data, uncovering 1,207 interactions that reveal both shared and gene-specific regulatory relationships, defining the global TF-promoter interaction network of thePSY/PSYRpathway. Functional analysis of 25 TF mutants identified 12 regulators that significantly influence root growth, most acting as repressors. Of these, CYTOKININ RESPONSE FACTOR 10 (CRF10) emerged as a strong growth inhibitor. We identified a CRF10-binding motif in thePSYR3promoter using DAP-seq data and validated it using eY1H. This motif is also located in the last 3′ terminal exon ofTopoisomerase 3A(TOP3A). Guided by these insights, we used CRISPR/Cas9-mediated promoter editing to delete a small region encompassing or flanking a functional TF-binding site (TFBS). Removal of this motif, or of its surrounding region, enhanced root growth, yielding variants that retained root length comparable to thecrf10mutant. Our results suggest that the observed root growth phenotype results either from disruption of the CRF10 binding motif or from the mutation in theTOP3Aexon. 
    more » « less
  2. Gene regulatory networks (GRNs) govern gene expression and cellular identity, but accurately inferring their structure from high-dimensional single-cell RNA sequencing (scRNA-seq) data remains a major challenge. Here, we present EnsembleRegNet, a deep learning framework that infers transcription factor (TF)-target gene relationships by integrating an ensemble encoder-decoder and multilayer perceptron (MLP) architecture. EnsembleRegNet utilizes Hodges-Lehmann estimator (HLE)-based binarization, case-deletion analysis, motif enrichment using RcisTarget, and regulon activity scoring with AUCell to enhance both robustness and biological interpretability. Extensive evaluations across simulated and real scRNA-seq datasets demonstrate that EnsembleRegNet outperforms existing GRN inference methods, including SCENIC and SIGNET, in both clustering performance and regulatory accuracy. By uncovering cell-type-specific regulatory modules and enhancing interpretability, EnsembleRegNet offers a scalable and biologically grounded framework for exploring transcriptional regulation. Its demonstrated performance establishes a new benchmark for GRN inference and highlights its promise for applications in disease modeling, biomarker discovery, and cellular reprogramming. 
    more » « less
  3. Abstract Understanding how transcription factor binding site (TFBS) location influences gene regulation remains a fundamental challenge in plants. Here we integrate conserved TFBS with single-nucleus chromatin accessibility and expression datasets to resolve how TFBS location relates to regulatory function. Although TF binding patterns are conserved and frequently enriched near the TSS, positional enrichment of TFBSs poorly predicts cell type-specific gene expression. Instead, cell type-specific gene expression aligns with conserved TFBSs that are embedded in cell type-restricted chromatin, which show TF family-specific distributions across distal promoter and genic regions. In contrast, TSS-proximal TFBSs are implicated in rapid, tissue-wide transcriptional stress responses, as supported by hormone-induced gene expression analysis. Finally, distal upstream regions contain conserved TF clusters that overlap rare cell type-specific accessible chromatin and are highly enriched for genes controlling embryonic and meristem patterning, including auxin and other hormone pathway components. This positional partitioning of regulatory function indicates that TFBS location encodes distinct regulatory programs for TFs. 
    more » « less
  4. Summary In plants, the biosynthetic pathways of some specialized metabolites are partitioned into specialized or rare cell types, as exemplified by the monoterpenoid indole alkaloid (MIA) pathway ofCatharanthus roseus(Madagascar Periwinkle), the source of the anticancer compounds vinblastine and vincristine. In the leaf, theC. roseusMIA biosynthetic pathway is partitioned into three cell types with the final known steps of the pathway expressed in the rare cell type termed idioblast. How cell‐type specificity of MIA biosynthesis is achieved is poorly understood.We generated single‐cell multi‐omics data fromC. roseusleaves. Integrating gene expression and chromatin accessibility profiles across single cells, as well as transcription factor (TF)‐binding site profiles, we constructed a cell‐type‐aware gene regulatory network for MIA biosynthesis.We showcased cell‐type‐specific TFs as well as cell‐type‐specificcis‐regulatory elements. Using motif enrichment analysis, co‐expression across cell types, and functional validation approaches, we discovered a novel idioblast‐specific TF (Idioblast MYB1,CrIDM1) that activates expression of late‐stage MIA biosynthetic genes in the idioblast.These analyses not only led to the discovery of the first documented cell‐type‐specific TF that regulates the expression of two idioblast‐specific biosynthetic genes within an idioblast metabolic regulon but also provides insights into cell‐type‐specific metabolic regulation. 
    more » « less
  5. Abstract DNA–transcription factor (TF) interactions are essential for gene regulation. Fully characterizing TF recognition specificities and identifying their genomic binding targets are important to understand TF function and regulatory networks. Recently, high-throughput sequencing technology HT-SELEX (high-throughput systematic evolution of ligands by exponential enrichment) has been used to measure hundreds of TFs, providing massive datasets that comprise TF binding preferences. However, there is a need to develop comprehensive computational modeling to fully extract and characterize critical TF binding preferences and fail to distinguish genome-wide binding targets. In this study, we developed a global pairwise model called DCA-Scapes trained with experimental HT-SELEX data. Our approach uncovered high-resolution TF recognition specificity landscapes, enabled the prediction of in vivo binding sequences, and was validated with ChIP-seq (ChIP sequencing) data. In addition, the DCA-Scapes model was utilized to refine the locations of binding regions and accurately identify the binding sites within the ChIP-seq enriched peaks. Moreover, we extended our model to cover the entire human genome, uncovering potential TF target sites that exhibit tissue-specific TF recognition across various cellular environments. 
    more » « less