skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


This content will become publicly available on August 1, 2026

Title: Uncovering methylation-dependent genetic effects on regulatory element function in diverse genomes
A major goal in evolutionary biology and biomedicine is to understand the complex interactions between genetic variants, the epigenome, and gene expression. However, the causal relationships between these factors remain poorly understood. mSTARR-seq, a methylation-sensitive massively parallel reporter assay, is capable of identifying methylation-dependent regulatory activity at many thousands of genomic regions simultaneously and allows for the testing of causal relationships between DNA methylation and gene expression on a region-by-region basis. Here, we develop a multiplexed mSTARR-seq protocol to assay naturally occurring human genetic variation from 25 individuals from 10 localities in Europe and Africa. We identify 6957 regulatory elements in either the unmethylated or methylated state, and this set was enriched for enhancer and promoter chromatin annotations, as expected. The expression of 58% of these regulatory elements is modulated by methylation, which is generally associated with decreased transcription. Within our set of regulatory elements, we use allele-specific expression analyses to identify 8020 sites with genetic effects on gene regulation; further, we find that 42.3% of these genetic effects vary in direction or magnitude between methylated and unmethylated states. Sites exhibiting methylation-dependent genetic effects are enriched for GWAS and EWAS annotations, implicating them in human disease. Compared with data sets that assay DNA from a single European ancestry individual, our multiplexed assay is able to uncover more genetic effects and methylation-dependent genetic effects, highlighting the importance of including diverse genomes in assays that aim to understand gene regulatory processes.  more » « less
Award ID(s):
2313953
PAR ID:
10649278
Author(s) / Creator(s):
; ;
Publisher / Repository:
Genome Research
Date Published:
Journal Name:
Genome Research
Volume:
35
Issue:
8
ISSN:
1088-9051
Page Range / eLocation ID:
1781 to 1793
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Gene duplication is a source of evolutionary novelty. DNA methylation may play a role in the evolution of duplicate genes (paralogs) through its association with gene expression. While this relationship has been examined to varying extents in a few individual species, the generalizability of these results at either a broad phylogenetic scale with species of differing duplication histories or across a population remains unknown. We applied a comparative epigenomic approach to 43 angiosperm species across the phylogeny and a population of 928 Arabidopsis (Arabidopsis thaliana) accessions, examining the association of DNA methylation with paralog evolution. Genic DNA methylation was differentially associated with duplication type, the age of duplication, sequence evolution, and gene expression. Whole-genome duplicates were typically enriched for CG-only gene body methylated or unmethylated genes, while single-gene duplications were typically enriched for non-CG methylated or unmethylated genes. Non-CG methylation, in particular, was a characteristic of more recent single-gene duplicates. Core angiosperm gene families were differentiated into those which preferentially retain paralogs and “duplication-resistant” families, which convergently reverted to singletons following duplication. Duplication-resistant families that still have paralogous copies were, uncharacteristically for core angiosperm genes, enriched for non-CG methylation. Non-CG methylated paralogs had higher rates of sequence evolution, higher frequency of presence–absence variation, and more limited expression. This suggests that silencing by non-CG methylation may be important to maintaining dosage following duplication and be a precursor to fractionation. Our results indicate that genic methylation marks differing evolutionary trajectories and fates between paralogous genes and have a role in maintaining dosage following duplication. 
    more » « less
  2. Bomblies, K (Ed.)
    Abstract DNA methylation in plants is depleted from cis-regulatory elements in and near genes but is present in some gene bodies, including exons. Methylation in exons solely in the CG context is called gene body methylation (gbM). Methylation in exons in both CG and non-CG contexts is called TE-like methylation (teM). Assigning functions to both forms of methylation in genes has proven to be challenging. Toward that end, we utilized recent genome assemblies, gene annotations, transcription data, and methylome data to quantify common patterns of gene methylation and their relations to gene expression in maize. We found that gbM genes exist in a continuum of CG methylation levels without a clear demarcation between unmethylated genes and gbM genes. Analysis of expression levels across diverse maize stocks and tissues revealed a weak but highly significant positive correlation between gbM and gene expression except in endosperm. gbM epialleles were associated with an approximately 3% increase in steady-state expression level relative to unmethylated epialleles. In contrast to gbM genes, which were conserved and were broadly expressed across tissues, we found that teM genes, which make up about 12% of genes, are mainly silent, are poorly conserved, and exhibit evidence of annotation errors. We used these data to flag teM genes in the 26 NAM founder genome assemblies. While some teM genes are likely functional, these data suggest that the majority are not, and their inclusion can confound the interpretation of whole-genome studies. 
    more » « less
  3. Wright, S (Ed.)
    Abstract In plants, mammals and insects, some genes are methylated in the CG dinucleotide context, a phenomenon called gene body methylation (gbM). It has been controversial whether this phenomenon has any functional role. Here, we took advantage of the availability of 876 leaf methylomes in Arabidopsis thaliana to characterize the population frequency of methylation at the gene level and to estimate the site-frequency spectrum of allelic states. Using a population genetics model specifically designed for epigenetic data, we found that genes with ancestral gbM are under significant selection to remain methylated. Conversely, ancestrally unmethylated genes were under selection to remain unmethylated. Repeating the analyses at the level of individual cytosines confirmed these results. Estimated selection coefficients were small, on the order of 4 Nes = 1.4, which is similar to the magnitude of selection acting on codon usage. We also estimated that A. thaliana is losing gbM threefold more rapidly than gaining it, which could be due to a recent reduction in the efficacy of selection after a switch to selfing. Finally, we investigated the potential function of gbM through its link with gene expression. Across genes with polymorphic methylation states, the expression of gene body methylated alleles was consistently and significantly higher than unmethylated alleles. Although it is difficult to disentangle genetic from epigenetic effects, our work suggests that gbM has a small but measurable effect on fitness, perhaps due to its association to a phenotype-like gene expression. 
    more » « less
  4. Abstract Accessible chromatin and unmethylated DNA are associated with many genes and cis-regulatory elements. Attempts to understand natural variation for accessible chromatin regions (ACRs) and unmethylated regions (UMRs) often rely upon alignments to a single reference genome. This limits the ability to assess regions that are absent in the reference genome assembly and monitor how nearby structural variants influence variation in chromatin state. In this study, de novo genome assemblies for four maize inbreds (B73, Mo17, Oh43, and W22) are utilized to assess chromatin accessibility and DNA methylation patterns in a pan-genome context. A more complete set of UMRs and ACRs can be identified when chromatin data are aligned to the matched genome rather than a single reference genome. While there are UMRs and ACRs present within genomic regions that are not shared between genotypes, these features are 6- to 12-fold enriched within regions between genomes. Characterization of UMRs present within shared genomic regions reveals that most UMRs maintain the unmethylated state in other genotypes with only ∼5% being polymorphic between genotypes. However, the majority (71%) of UMRs that are shared between genotypes only exhibit partial overlaps suggesting that the boundaries between methylated and unmethylated DNA are dynamic. This instability is not solely due to sequence variation as these partially overlapping UMRs are frequently found within genomic regions that lack sequence variation. The ability to compare chromatin properties among individuals with structural variation enables pan-epigenome analyses to study the sources of variation for accessible chromatin and unmethylated DNA. 
    more » « less
  5. Abstract BackgroundIn several eukaryotes, DNA methylation occurs within the coding regions of many genes, termed gene body methylation (GbM). Whereas the role of DNA methylation on the silencing of transposons and repetitive DNA is well understood, gene body methylation is not associated with transcriptional repression, and its biological importance remains unclear. ResultsWe report a newly discovered type of GbM in plants, which is under constitutive addition and removal by dynamic methylation modifiers in all cells, including the germline. Methylation at Dynamic GbM genes is removed by the DRDD demethylation pathway and added by an unknown source of de novo methylation, most likely the maintenance methyltransferase MET1. We show that the Dynamic GbM state is present at homologous genes across divergent lineages spanning over 100 million years, indicating evolutionary conservation. We demonstrate that Dynamic GbM is tightly associated with the presence of a promoter or regulatory chromatin state within the gene body, in contrast to other gene body methylated genes. We find Dynamic GbM is associated with enhanced gene expression plasticity across development and diverse physiological conditions, whereas stably methylated GbM genes exhibit reduced plasticity. Dynamic GbM genes exhibit reduced dynamic range indrddmutants, indicating a causal link between DNA demethylation and enhanced gene expression plasticity. ConclusionsWe propose a new model for GbM in regulating gene expression plasticity, including a novel type of GbM in which increased gene expression plasticity is associated with the activity of DNA methylation writers and erasers and the enrichment of a regulatory chromatin state. 
    more » « less