skip to main content


Title: Exploring plant cis‐regulatory elements at single‐cell resolution: overcoming biological and computational challenges to advance plant research
SUMMARY

Cis‐regulatory elements (CREs) are important sequences for gene expression and for plant biological processes such as development, evolution, domestication, and stress response. However, studying CREs in plant genomes has been challenging. The totipotent nature of plant cells, coupled with the inability to maintain plant cell types in culture and the inherent technical challenges posed by the cell wall has limited our understanding of how plant cell types acquire and maintain their identities and respond to the environment via CRE usage. Advances in single‐cell epigenomics have revolutionized the field of identifying cell‐type‐specific CREs. These new technologies have the potential to significantly advance our understanding of plant CRE biology, and shed light on how the regulatory genome gives rise to diverse plant phenomena. However, there are significant biological and computational challenges associated with analyzing single‐cell epigenomic datasets. In this review, we discuss the historical and foundational underpinnings of plant single‐cell research, challenges, and common pitfalls in the analysis of plant single‐cell epigenomic data, and highlight biological challenges unique to plants. Additionally, we discuss how the application of single‐cell epigenomic data in various contexts stands to transform our understanding of the importance of CREs in plant genomes.

 
more » « less
NSF-PAR ID:
10442206
Author(s) / Creator(s):
 ;  ;  ;  ;  
Publisher / Repository:
Wiley-Blackwell
Date Published:
Journal Name:
The Plant Journal
ISSN:
0960-7412
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. INTRODUCTION Genome-wide association studies (GWASs) have identified thousands of human genetic variants associated with diverse diseases and traits, and most of these variants map to noncoding loci with unknown target genes and function. Current approaches to understand which GWAS loci harbor causal variants and to map these noncoding regulators to target genes suffer from low throughput. With newer multiancestry GWASs from individuals of diverse ancestries, there is a pressing and growing need to scale experimental assays to connect GWAS variants with molecular mechanisms. Here, we combined biobank-scale GWASs, massively parallel CRISPR screens, and single-cell sequencing to discover target genes of noncoding variants for blood trait loci with systematic targeting and inhibition of noncoding GWAS loci with single-cell sequencing (STING-seq). RATIONALE Blood traits are highly polygenic, and GWASs have identified thousands of noncoding loci that map to candidate cis -regulatory elements (CREs). By combining CRE-silencing CRISPR perturbations and single-cell readouts, we targeted hundreds of GWAS loci in a single assay, revealing target genes in cis and in trans . For select CREs that regulate target genes, we performed direct variant insertion. Although silencing the CRE can identify the target gene, direct variant insertion can identify magnitude and direction of effect on gene expression for the GWAS variant. In select cases in which the target gene was a transcription factor or microRNA, we also investigated the gene-regulatory networks altered upon CRE perturbation and how these networks differ across blood cell types. RESULTS We inhibited candidate CREs from fine-mapped blood trait GWAS variants (from ~750,000 individual of diverse ancestries) in human erythroid progenitors. In total, we targeted 543 variants (254 loci) mapping to candidate CREs, generating multimodal single-cell data including transcriptome, direct CRISPR gRNA capture, and cell surface proteins. We identified target genes in cis (within 500 kb) for 134 CREs. In most cases, we found that the target gene was the closest gene and that specific enhancer-associated biochemical hallmarks (H3K27ac and accessible chromatin) are essential for CRE function. Using multiple perturbations at the same locus, we were able to distinguished between causal variants from noncausal variants in linkage disequilibrium. For a subset of validated CREs, we also inserted specific GWAS variants using base-editing STING-seq (beeSTING-seq) and quantified the effect size and direction of GWAS variants on gene expression. Given our transcriptome-wide data, we examined dosage effects in cis and trans in cases in which the cis target is a transcription factor or microRNA. We found that trans target genes are also enriched for GWAS loci, and identified gene clusters within trans gene networks with distinct biological functions and expression patterns in primary human blood cells. CONCLUSION In this work, we investigated noncoding GWAS variants at scale, identifying target genes in single cells. These methods can help to address the variant-to-function challenges that are a barrier for translation of GWAS findings (e.g., drug targets for diseases with a genetic basis) and greatly expand our ability to understand mechanisms underlying GWAS loci. Identifying causal variants and their target genes with STING-seq. Uncovering causal variants and their target genes or function are a major challenge for GWASs. STING-seq combines perturbation of noncoding loci with multimodal single-cell sequencing to profile hundreds of GWAS loci in parallel. This approach can identify target genes in cis and trans , measure dosage effects, and decipher gene-regulatory networks. 
    more » « less
  2. Abstract Epigenomics is the study of molecular signatures associated with discrete regions within genomes, many of which are important for a wide range of nuclear processes. The ability to profile the epigenomic landscape associated with genes, repetitive regions, transposons, transcription, differential expression, cis-regulatory elements, and 3D chromatin interactions has vastly improved our understanding of plant genomes. However, many epigenomic and single-cell genomic assays are challenging to perform in plants, leading to a wide range of data quality issues; thus, the data require rigorous evaluation prior to downstream analyses and interpretation. In this commentary, we provide considerations for the evaluation of plant epigenomics and single-cell genomics data quality with the aim of improving the quality and utility of studies using those data across diverse plant species. 
    more » « less
  3. Summary

    Adverse environmental conditions reduce crop productivity and often increase the load of unfolded or misfolded proteins in the endoplasmic reticulum (ER). This potentially lethal condition, known as ER stress, is buffered by the unfolded protein response (UPR), a set of signaling pathways designed to either recover ER functionality or ignite programmed cell death. Despite the biological significance of the UPR to the life of the organism, the regulatory transcriptional landscape underpinning ER stress management is largely unmapped, especially in crops. To fill this significant knowledge gap, we performed a large‐scale systems‐level analysis of the protein–DNA interaction (PDI) network in maize (Zea mays). Using 23 promoter fragments of six UPR marker genes in a high‐throughput enhanced yeast one‐hybrid assay, we identified a highly interconnected network of 262 transcription factors (TFs) associated with significant biological traits and 831 PDIs underlying the UPR. We established a temporal hierarchy of TF binding to gene promoters within the same family as well as across different families of TFs. Cistrome analysis revealed the dynamic activities of a variety ofcis‐regulatory elements (CREs) in ER stress‐responsive gene promoters. By integrating the cistrome results into a TF network analysis, we mapped a subnetwork of TFs associated with a CRE that may contribute to UPR management. Finally, we validated the role of a predicted network hub gene using the Arabidopsis system. The PDIs, TF networks, and CREs identified in our work are foundational resources for understanding transcription‐regulatory mechanisms in the stress responses and crop improvement.

     
    more » « less
  4. Summary

    Spirodela polyrhizais a fast‐growing aquatic monocot with highly reduced morphology, genome size and number of protein‐coding genes. Considering these biological features of Spirodela and its basal position in the monocot lineage, understanding its genome architecture could shed light on plant adaptation and genome evolution. Like many draft genomes, however, the 158‐Mb Spirodela genome sequence has not been resolved to chromosomes, and important genome characteristics have not been defined. Here we deployed rapid genome‐wide physical maps combined with high‐coverage short‐read sequencing to resolve the 20 chromosomes of Spirodela and to empirically delineate its genome features. Our data revealed a dramatic reduction in the number of therDNArepeat units in Spirodela to fewer than 100, which is even fewer than that reported for yeast. Consistent with its unique phylogenetic position, smallRNAsequencing revealed 29 Spirodela‐specific microRNA, with only two being shared withElaeis guineensis(oil palm) andMusa balbisiana(banana). CombiningDNAmethylation data and smallRNAsequencing enabled the accurate prediction of 20.5% long terminal repeats (LTRs) that doubled the previous estimate, and revealed a high Solo:IntactLTRratio of 8.2. Interestingly, we found that Spirodela has the lowest globalDNAmethylation levels (9%) of any plant species tested. Taken together our results reveal a genome that has undergone reduction, likely through eliminating non‐essential protein coding genes,rDNAandLTRs. In addition to delineating the genome features of this unique plant, the methodologies described and large‐scale genome resources from this work will enable future evolutionary and functional studies of this basal monocot family.

     
    more » « less
  5. Summary

    Agrobacterium tumefaciens, the causal agent of plant crown gall disease, has been widely used to genetically transform many plant species. The inter‐kingdom gene transfer capability madeAgrobacteriuman essential tool and model system to study the mechanism of exporting and integrating a segment of bacterial DNA into the plant genome. However, many biological processes such asAgrobacterium‐host recognition and interaction are still elusive. To accelerate the understanding of this important plant pathogen and further improve its capacity in plant genetic engineering, we adopted a CRISPR RNA‐guided integrase system forAgrobacteriumgenome engineering. In this work, we demonstrate thatINsertion ofTransposableElements byGuideRNA–AssistedTargEting (INTEGRATE) can efficiently generate DNA insertions to enable targeted gene knockouts. In addition, in conjunction with Cre‐loxPrecombination system, we achieved precise deletions of large DNA fragments. This work provides new genetic engineering strategies forAgrobacteriumspecies and their gene functional analyses.

     
    more » « less