Transcription factors (TFs) play a central role in regulating molecular level responses of plants to external stresses such as water limiting conditions, but identification of such TFs in the genome remains a challenge. Here, we describe a network-based supervised machine learning framework that accurately predicts and ranks all TFs in the genome according to their potential association with drought tolerance. We show that top ranked regulators fall mainly into two ‘age’ groups; genes that appeared first in land plants and genes that emerged later in the Oryza clade. TFs predicted to be high in the ranking belong to specific gene families, have relatively simple intron/exon and protein structures, and functionally converge to regulate primary and secondary metabolism pathways. Repeated trials of nested cross-validation tests showed that models trained only on regulatory network patterns, inferred from large transcriptome datasets, outperform models trained on heterogenous genomic features in the prediction of known drought response regulators. A new R/Shiny based web application, called the DroughtApp, provides a primer for generation of new testable hypotheses related to regulation of drought stress response. Furthermore, to test the system we experimentally validated predictions on the functional role of the rice transcription factor OsbHLH148, using RNA sequencing of knockout mutants in response to drought stress and protein-DNA interaction assays. Our study exemplifies the integration of domain knowledge for prioritization of regulatory genes in biological pathways of well-studied agricultural traits. 
                        more » 
                        « less   
                    
                            
                            Network Biology Analyses and Dynamic Modeling of Gene Regulatory Networks under Drought Stress Reveal Major Transcriptional Regulators in Arabidopsis
                        
                    
    
            Drought is one of the most serious abiotic stressors in the environment, restricting agricultural production by reducing plant growth, development, and productivity. To investigate such a complex and multifaceted stressor and its effects on plants, a systems biology-based approach is necessitated, entailing the generation of co-expression networks, identification of high-priority transcription factors (TFs), dynamic mathematical modeling, and computational simulations. Here, we studied a high-resolution drought transcriptome of Arabidopsis. We identified distinct temporal transcriptional signatures and demonstrated the involvement of specific biological pathways. Generation of a large-scale co-expression network followed by network centrality analyses identified 117 TFs that possess critical properties of hubs, bottlenecks, and high clustering coefficient nodes. Dynamic transcriptional regulatory modeling of integrated TF targets and transcriptome datasets uncovered major transcriptional events during the course of drought stress. Mathematical transcriptional simulations allowed us to ascertain the activation status of major TFs, as well as the transcriptional intensity and amplitude of their target genes. Finally, we validated our predictions by providing experimental evidence of gene expression under drought stress for a set of four TFs and their major target genes using qRT-PCR. Taken together, we provided a systems-level perspective on the dynamic transcriptional regulation during drought stress in Arabidopsis and uncovered numerous novel TFs that could potentially be used in future genetic crop engineering programs. 
        more » 
        « less   
        
    
                            - Award ID(s):
- 2038872
- PAR ID:
- 10425001
- Date Published:
- Journal Name:
- International Journal of Molecular Sciences
- Volume:
- 24
- Issue:
- 8
- ISSN:
- 1422-0067
- Page Range / eLocation ID:
- 7349
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
- 
            
- 
            Marshall-Colon, Amy (Ed.)Abstract Gene regulatory networks (GRNs) are defined by a cascade of transcriptional events by which signals, such as hormones or environmental cues, change development. To understand these networks, it is necessary to link specific transcription factors (TFs) to the downstream gene targets whose expression they regulate. Although multiple methods provide information on the targets of a single TF, moving from groups of co-expressed genes to the TF that controls them is more difficult. To facilitate this bottom-up approach, we have developed a web application named TF DEACoN. This application uses a publicly available Arabidopsis thaliana DNA Affinity Purification (DAP-Seq) data set to search for TFs that show enriched binding to groups of co-regulated genes. We used TF DEACoN to examine groups of transcripts regulated by treatment with the ethylene precursor 1-aminocyclopropane-1-carboxylic acid (ACC), using a transcriptional data set performed with high temporal resolution. We demonstrate the utility of this application when co-regulated genes are divided by timing of response or cell-type-specific information, which provides more information on TF/target relationships than when less defined and larger groups of co-regulated genes are used. This approach predicted TFs that may participate in ethylene-modulated root development including the TF NAM (NO APICAL MERISTEM). We used a genetic approach to show that a mutation in NAM reduces the negative regulation of lateral root development by ACC. The combination of filtering and TF DEACoN used here can be applied to any group of co-regulated genes to predict GRNs that control coordinated transcriptional responses.more » « less
- 
            Abstract The regulation of gene expression is central to many biological processes. Gene regulatory networks (GRNs) link transcription factors (TFs) to their target genes and represent maps of potential transcriptional regulation. Here, we analyzed a large number of publically available maize (Zea mays) transcriptome data sets including >6000 RNA sequencing samples to generate 45 coexpression-based GRNs that represent potential regulatory relationships between TFs and other genes in different populations of samples (cross-tissue, cross-genotype, and tissue-and-genotype samples). While these networks are all enriched for biologically relevant interactions, different networks capture distinct TF-target associations and biological processes. By examining the power of our coexpression-based GRNs to accurately predict covarying TF-target relationships in natural variation data sets, we found that presence/absence changes rather than quantitative changes in TF gene expression are more likely associated with changes in target gene expression. Integrating information from our TF-target predictions and previous expression quantitative trait loci (eQTL) mapping results provided support for 68 TFs underlying 74 previously identified trans-eQTL hotspots spanning a variety of metabolic pathways. This study highlights the utility of developing multiple GRNs within a species to detect putative regulators of important plant pathways and provides potential targets for breeding or biotechnological applications.more » « less
- 
            Drought stress is a key limitation for plant growth and colonization of arid habitats. We study the evolution of gene expression response to drought stress in a wild tomato,Solanum chilense,naturally occurring in dry habitats in South America. We conduct a transcriptome analysis under standard and drought experimental conditions to identify drought‐responsive gene networks and estimate the age of the involved genes. We identify two main regulatory networks corresponding to two typical drought‐responsive strategies: cell cycle and fundamental metabolic processes. The metabolic network exhibits a more recent evolutionary origin and a more variable transcriptome response than the cell cycle network (with ancestral origin and higher conservation of the transcriptional response). We also integrate population genomics analyses to reveal positive selection signals acting at the genes of both networks, revealing that genes exhibiting selective sweeps of older age also exhibit greater connectivity in the networks. These findings suggest that adaptive changes first occur at core genes of drought response networks, driving significant network re‐wiring, which likely underpins species divergence and further spread into drier habitats. Combining transcriptomics and population genomics approaches, we decipher the timing of gene network evolution for drought stress response in arid habitats.more » « less
- 
            Nitrogen (N) and Water (W) - two resources critical for crop productivity – are becoming increasingly limited in soils globally. To address this issue, we aim to uncover the gene regulatory networks (GRNs) that regulate nitrogen use efficiency (NUE) - as a function of water availability - in Oryza sativa, a staple for 3.5 billion people. In this study, we infer and validate GRNs that correlate with rice NUE phenotypes affected by N-by-W availability in the field. We did this by exploiting RNA-seq and crop phenotype data from 19 rice varieties grown in a 2x2 N-by-W matrix in the field. First, to identify gene-to-NUE field phenotypes, we analyzed these datasets using weighted gene co-expression network analysis (WGCNA). This identified two network modules ("skyblue" & "grey60") highly correlated with NUE grain yield (NUEg). Next, we focused on 90 TFs contained in these two NUEg modules and predicted their genome-wide targets using the N-and/or-W response datasets using a random forest network inference approach (GENIE3). Next, to validate the GENIE3 TF→target gene predictions, we performed Precision/Recall Analysis (AUPR) using nine datasets for three TFs validated in planta . This analysis sets a precision threshold of 0.31, used to "prune" the GENIE3 network for high-confidence TF→target gene edges, comprising 88 TFs and 5,716 N-and/or-W response genes. Next, we ranked these 88 TFs based on their significant influence on NUEg target genes responsive to N and/or W signaling. This resulted in a list of 18 prioritized TFs that regulate 551 NUEg target genes responsive to N and/or W signals. We validated the direct regulated targets of two of these candidate NUEg TFs in a plant cell-based TF assay called TARGET, for which we also had in planta data for comparison. Gene ontology analysis revealed that 6/18 NUEg TFs - OsbZIP23 (LOC_Os02g52780), Oshox22 (LOC_Os04g45810), LOB39 (LOC_Os03g41330), Oshox13 (LOC_Os03g08960), LOC_Os11g38870, and LOC_Os06g14670 - regulate genes annotated for N and/or W signaling. Our results show that OsbZIP23 and Oshox22, known regulators of drought tolerance, also coordinate W-responses with NUEg. This validated network can aid in developing/breeding rice with improved yield on marginal, low N-input, drought-prone soils.more » « less
 An official website of the United States government
An official website of the United States government 
				
			 
					 
					
 
                                    