<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Deleterious Mutations Accumulate Faster in Allopolyploid Than Diploid Cotton ( &lt;i&gt;Gossypium&lt;/i&gt; ) and Unequally between Subgenomes</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>02/01/2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10383329</idno>
					<idno type="doi">10.1093/molbev/msac024</idno>
					<title level='j'>Molecular Biology and Evolution</title>
<idno>0737-4038</idno>
<biblScope unit="volume">39</biblScope>
<biblScope unit="issue">2</biblScope>					

					<author>Justin L Conover</author><author>Jonathan F Wendel</author><author>Michael Purugganan</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Abstract            Whole-genome duplication (polyploidization) is among the most dramatic mutational processes in nature, so understanding how natural selection differs in polyploids relative to diploids is an important goal. Population genetics theory predicts that recessive deleterious mutations accumulate faster in allopolyploids than diploids due to the masking effect of redundant gene copies, but this prediction is hitherto unconfirmed. Here, we use the cotton genus (Gossypium), which contains seven allopolyploids derived from a single polyploidization event 1–2 Million years ago, to investigate deleterious mutation accumulation. We use two methods of identifying deleterious mutations at the nucleotide and amino acid level, along with whole-genome resequencing of 43 individuals spanning six allopolyploid species and their two diploid progenitors, to demonstrate that deleterious mutations accumulate faster in allopolyploids than in their diploid progenitors. We find that, unlike what would be expected under models of demographic changes alone, strongly deleterious mutations show the biggest difference between ploidy levels, and this effect diminishes for moderately and mildly deleterious mutations. We further show that the proportion of nonsynonymous mutations that are deleterious differs between the two coresident subgenomes in the allopolyploids, suggesting that homoeologous masking acts unequally between subgenomes. Our results provide a genome-wide perspective on classic notions of the significance of gene duplication that likely are broadly applicable to allopolyploids, with implications for our understanding of the evolutionary fate of deleterious mutations. Finally, we note that some measures of selection (e.g., dN/dS, πN/πS) may be biased when species of different ploidy levels are compared.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Genome duplication (polyploidy) is among the most dramatic mutational processes in nature, causing myriad saltational changes at the cellular and organismal levels <ref type="bibr">(Doyle and Coate 2019;</ref><ref type="bibr">Bomblies 2020;</ref><ref type="bibr">Fernandes Gyorfy et al. 2021)</ref>, and is associated with consequential phenomena ranging from crop domestication <ref type="bibr">(Renny-Byfield and Wendel 2014;</ref><ref type="bibr">Qi et al. 2021</ref>) to cancer progression <ref type="bibr">(Matsumoto et al. 2021)</ref>. Polyploidy is especially common in the angiosperms, with all extant species having experienced at least one or more polyploidy events during their evolutionary history <ref type="bibr">(Jiao et al. 2011</ref>), and at least 30% of currently recognized species having a polyploidy event in the recent past <ref type="bibr">(One Thousand Plant Transcriptomes Initiative 2019)</ref>.</p><p>Novel evolutionary patterns created by polyploidy at the genic (e.g., neofunctionalization, subfunctionalization, and loss; <ref type="bibr">Kuzmin et al. 2021)</ref> and genomic (e.g., homoeologous recombination; Mason and Wendel 2020) levels have been well documented across taxa, including the frequent asymmetry of these responses with respect to coresident genomes in a polyploid nucleus. Nonetheless, many questions remain regarding the effects of natural selection on polyploid relative to diploid genomes <ref type="bibr">(Baduel et al. 2019;</ref><ref type="bibr">Monnahan et al. 2019)</ref> and the interplay between these novel evolutionary patterns and the long-term trajectories of genome evolution <ref type="bibr">(Qi et al. 2021)</ref> following polyploidization (e.g., biased fractionation).</p><p>One of the earliest predictions about natural selection in polyploids relative to diploids is that putatively deleterious mutations may accumulate faster due to the masking effect of completely or partially recessive deleterious mutations in duplicated genes <ref type="bibr">(Haldane 1932;</ref><ref type="bibr">Hill 1970;</ref><ref type="bibr">Bever and Felber 1992)</ref>. Only recently, however, have these predictions begun to be evaluated in young allopolyploids such as Arabidopsis kamchatica <ref type="bibr">(Paape et al. 2018)</ref> and Capsella bursa-pastoris <ref type="bibr">(Douglas et al. 2015;</ref><ref type="bibr">Kryvokhyzha, Milesi, et al. 2019;</ref><ref type="bibr">Kryvokhyzha, Salcedo, et al. 2019)</ref>, and autotetraploid Arabidopsis arenosa <ref type="bibr">(Monnahan et al. 2019)</ref>. Because the number of deleterious mutations is strongly influenced by shifts in demography and mating system <ref type="bibr">(Brandvain and Wright 2016)</ref>, which may coincide with polyploid formation <ref type="bibr">(Grant 1981;</ref><ref type="bibr">Barringer 2007)</ref>, a clear link between ploidy level and the accumulation of deleterious mutations is challenging to demonstrate in natural polyploid populations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Article</head><p>&#223; The Author(s) 2022. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (<ref type="url">https://creativecommons</ref>. org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Open Access</head><p>The cotton genus (Gossypium) represents one of the best studied allopolyploid systems <ref type="bibr">(Wendel and Grover 2015;</ref><ref type="bibr">Hu et al. 2021)</ref>. The genus includes approximately 45 currently recognized diploid species classified into eight genome groups (A-G, and K), and seven allopolyploid species resulting from a single <ref type="bibr">(Grover et al. 2012</ref>) allopolyploidization event approximately 1-2 million years ago (Mya) between members of the A and D genome groups <ref type="bibr">(Wendel 1989)</ref>. Although the most closely related extant species of these two progenitor diploids are found in southern Africa and Northern Peru, respectively, the polyploids are only found in the Americas (fig. <ref type="figure">1</ref>). Most wild populations, including those of the two independently domesticated species G. hirsutum (AD 1 ) and G. barbadense (AD 2 ), occur in small, isolated populations on islands or in coastal regions. Subsequent to their initial domestication 4,000-8,000 years ago in the Yucatan Peninsula (AD 1 ) and NW S. America (AD 2 ), respectively, the ranges of the two domesticated species rapidly expanded to encompass much of the American tropics and subtropics and then spread globally with the rise of the international cotton fiber trade <ref type="bibr">(Yuan et al. 2021)</ref>.</p><p>Here, we describe the evolutionary trajectory of deleterious mutations in two wild diploid and six wild allopolyploid cotton species (all descended from a single allopolyploidization event), with a focus on how allopolyploidization and speciation have shaped the number and genomic distribution of deleterious mutations. We use two methods to estimate the strength of selection at the amino acid and nucleotide level and show support for a nearly century-old hypothesis that polyploids accumulate mutations faster than their diploid progenitors. We demonstrate that, in agreement with this hypothesis, polyploidy has the greatest influence on strongly, rather than moderately or mildly, deleterious mutations. We also find that deleterious mutations accumulate asymmetrically between the two coresident subgenomes in the allopolyploid nucleus, indicating that these masking effects may act unequally between the subgenomes. In total, our results support theoretical predictions that allopolyploidy can lead to a faster rate of deleterious mutation accumulation through masking of recessive deleterious variants, and that the relationship of the rate of deleterious mutation accumulation between subgenomes and their progenitor diploids is complex, even when comparing identical pairs of single-copy homoeologs among lineages.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Patterns of Synonymous and Nonsynonymous Mutations</head><p>To investigate patterns of deleterious mutations, we viewed our data at three phylogenetic depths: single nucleotide mutations that occurred anywhere on the global phylogenetic tree (fig. <ref type="figure">2A-D</ref>), mutations that emerged since the divergence of each subgenome from its respective diploid progenitor  Using the curated set of 8,884 pairs of homoeologous genes, we found no evidence for differences in the rate of synonymous mutation accumulation in diploids versus polyploids at any phylogenetic depth (fig. <ref type="figure">2A</ref> and<ref type="figure">E</ref>), although differences can be found within the polyploid species (fig. <ref type="figure">2I</ref>), with G. mustelinum (AD 4 , Orange) and G. darwinii (AD 5 , Yellow) having consistently lower rates than the rest of the clade, and in both subgenomes. When viewing mutations that have accumulated since the divergence of the earliest polyploid lineage (fig. <ref type="figure">2A-I</ref>), there is an asymmetry between subgenomes in the rate of synonymous site changes, with the Dt ("t" denoting "tetraploid") subgenome containing a moderately higher number of synonymous mutations than the At subgenome for all species. This difference potentially indicates a higher mutation rate or relaxed background selection in genic regions of the Dt subgenome compared with homoeologous genic regions of the At sugenome, and is consistent with previous analysis finding that genes in the Dt subgenome are evolving faster than genes in the At subgenome in five allopolyploid cottons <ref type="bibr">(Chen et al. 2020)</ref>.  <ref type="figure">E-H</ref>) shows mutations that are variable within each subgenome and its associated progenitor diploid species. For example, in panel (F), the yellow solid circles indicate the number of derived nonsynonymous mutations in the At subgenome of AD5 (G. darwinii) that have occurred since its divergence from the A1 diploid (left half of panel), and in the Dt subgenome of AD5 (G. darwinii) that have occurred since its divergence from the D5 diploid (right half of panel). The bottom row (panels I-L) shows mutations that originated postpolyploidy and are variable within the polyploids; for example, in panel (L), the purple solid circles indicated the number of derived deleterious mutations identified by BAD_Mutations that have occurred in the At subgenome of AD3 (G. tomentosum) since its divergence from the At subgenome of AD4 (G. mustelinum; left half of panel), and in the Dt subgenome of AD3 (G. tomentosum) since its divergence from the Dt subgenome of AD4 (G. mustelinum; right half of panel). Panels (A), (B), (E), (F), (I), (J): the y axis for both synonymous and nonsynonymous represents the sum of the derived allele frequencies, interpreted as the average number of derived mutations in that category in each species. Panels (C), (G), and (K): the y axis indicates the GERP Load of each species, calculated as the sum of (derived allele frequency&#194;GERP Score) for all variant positions with GERP&gt;0. Panels (D), (H), and (L): the y axis shows the total number of deleterious mutations in each species, calculated by BAD_Mutations with Bonferroni-corrected significance (see Materials and Methods)and indicates the average number of deleterious mutations in each species at a given phylogenetic depth. For panels (E-H), comparisons between subgenomes cannot be made because the D 5 diploid is more distantly related to the D subgenome than the A 1 diploid is related to the A subgenome. Therefore, we would expect a larger number of derived mutations in D than A simply due to evolutionary history rather than to polyploidization per se.</p><p>Deleterious Mutations Accumulate Faster in Allopolyploids . doi:10.1093/molbev/msac024</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>MBE</head><p>In contrast to the relative homogeneity in rates of synonymous substitution among diploids and polyploids, rates of nonsynonymous mutation accumulation differed significantly at all phylogenetic depths. Notably, at the deepest phylogenetic depth (fig. <ref type="figure">2B</ref>), estimates for the number of derived nonsynonymous mutations in the diploids G. herbaceum (A 1 , Red) and G. raimondii (D 5 , Black) did not differ between subgenomes, indicating that any mapping biases or erroneous SNP calls in these samples were removed by our variant filtering criteria. Gossypium raimondii (D 5 ) contained more derived nonsynonymous mutations than did G. herbaceum (A 1 ), and this lineage-specific difference was reflected in the Dt and At subgenomes as well. When lineage-specific effects that arose from the long, shared ancestry between the subgenomes and their progenitor diploids were removed (fig. <ref type="figure">2F</ref>), a clear distinction between the rates of nonsynonymous mutations between diploids and their respective subgenomes in the allopolyploids becomes apparent. In all polyploids, the At subgenome contained between 25% and 58% more nonsynonymous mutations than G. herbaceum (A 1 , Red), and the Dt subgenome contained 17-36% more than G. raimondii (D 5 , Black). These results demonstrate that even though the rates of synonymous mutation accumulation did not differ significantly between the diploids and polyploids, polyploidy significantly increases the rate of nonsynonymous substitution accumulation. Finally, when restricting our attention to only those mutations that have arisen following polyploid formation (fig. <ref type="figure">2J</ref>), the lineagespecific patterns observed for nonsynonymous mutations were largely identical to the patterns of synonymous mutations. For example, G. mustelinum, (AD 4 , Orange) consistently had the lowest number of derived mutations in both subgenomes. Notably, however, the Hawaiian Island endemic G. tomentosum (AD 3 , Purple) has a higher number of derived nonsynonymous mutations than expected based on the patterns of synonymous mutations, potentially reflecting the population bottleneck associated with its origin following long-distance dispersal to the Hawaiian Islands (see Discussion). In summary, polyploidy significantly enhances the rate of nonsynonymous mutation accumulation in all Gossypium allopolyploids, and does so asymmetrically across coresident genomes.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Polyploidy Increases Rate of Deleterious Mutation Accumulation</head><p>Because the fitness effects of most nonsynonymous mutations can vary widely, from neutral to lethal, we asked if the elevated rate of nonsynonymous mutations observed in polyploid Gossypium reflects an increase in neutral or nearly neutral nonsynonymous mutations, or if instead this elevation is attributable to a greater accumulation of deleterious mutations. To address this, we used two approaches to estimate whether a mutation was deleterious: BAD_Mutations and GERP&#254;&#254;. BAD_Mutations performs a likelihood ratio test from a gene-specific multispecies alignment to determine if a mutation at a particular nonsynonymous site is potentially deleterious, whereas GERP&#254;&#254; uses a genome-wide multiple sequence alignment (i.e., agnostic to genic regions) to estimate the degree of conservation at a particular site in the genome (including noncoding and synonymous sites). Notably, because one of the hallmark long-term processes following polyploidy is pseudogenization <ref type="bibr">(Wendel 2015)</ref>, recently pseudogenized sequences may still display some degree of conservation across the multiple sequence alignment, but may not be inherently deleterious. Therefore, to avoid inflating estimates of deleterious mutations in polyploids compared with diploids, we used GERP only within the exonic regions of the 8,884 homoeologs. Additionally, although GERP can score the degree of deleteriousness of a mutation, BAD_Mutations can only classify variants into deleterious or not deleterious. Therefore, the values shown in figure <ref type="figure">2D</ref>, H, and L represent the sum of the allele frequencies of derived deleterious mutations, similar to the values for figure <ref type="figure">2A,</ref><ref type="figure">E</ref>, and I and figure <ref type="figure">2B</ref>, F, and J. For analysis with GERP, we used the GERP load, which incorporates the deleteriousness of each variant into the score, summing the frequency of each derived allele multiplied by its GERP score (see <ref type="bibr">Rodgers-Melnick et al. [2015]</ref> and <ref type="bibr">Wang et al. [2017]</ref>).</p><p>As shown in figure <ref type="figure">2</ref>, both of the foregoing analyses demonstrate that deleterious mutations accumulate in polyploids in a manner similar to nonsynonymous mutations, suggesting that the difference in nonsynonymous sites cannot be wholly attributed to putatively neutral or nearly neutral alleles. For example, there is remarkable consistency in the patterns of deleterious mutations that have accumulated since the divergence of the diploid from its respective diploid progenitor in both the count of nonsynonymous substitutions (fig. <ref type="figure">2F</ref>), the GERP load (fig. <ref type="figure">2G</ref>), and the number of deleterious mutations (fig. <ref type="figure">2H</ref>). In all three columns, the diploids show fewer accumulated alleles than the polyploids, G. tomentosum (AD 3 , Purple) shows the highest number of all the polyploids, and G. mustelinum (AD 4 , Orange) shows the fewest of all the polyploids.</p><p>An interesting pattern arises when comparing estimates of the GERP load (fig. <ref type="figure">2G</ref>) and number of deleterious mutations (fig. <ref type="figure">2H</ref>) between diploids and their closely related subgenomes: Although the total number of deleterious mutations in the At subgenomes was 52-99% higher in the polyploids than the diploids (fig. <ref type="figure">2H</ref>), the GERP load in the polyploids was only 13-42% higher (fig. <ref type="figure">2G</ref>). Similar patterns were found in the Dt subgenome, with 34-66% more deleterious mutations in the polyploids than the diploid, but only a 9-13% increase in GERP load. This discrepancy could reflect inherent differences in the types of sequences used and how deleteriousness is quantified between the two methods, suggesting that the use of multiple analytical tools for detection of genetic load may yield more nuanced insights than either method on its own (See Discussion).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Asymmetries in the Rate of Deleterious Mutation Accumulation</head><p>Although deleterious mutations are accumulating faster in polyploids relative to diploids, it is not obvious whether this increased rate is different from the increased rate of accumulation of nonsynonymous mutations. To test this, we compared, among ploidy levels, the total proportion of Conover and Wendel . doi:10.1093/molbev/msac024 MBE nonsynonymous mutations that were considered deleterious by BAD_Mutations (fig. <ref type="figure">3</ref>). For mutations that originated since the divergence of the A and D diploids (fig. <ref type="figure">3A</ref>), the proportion of nonsynonymous sites that are deleterious is roughly 2% higher in polyploids than in diploids, despite the shared evolutionary history of more than 4 My between each subgenome and their respective diploid progenitors. Notably, as similarly shown in figure <ref type="figure">2</ref>, the proportion of nonsynonymous mutations that are inferred to be deleterious in both diploids is equivalent when mapped to either subgenome, indicating that our filtering criteria did not differentially exclude deleterious or nondeleterious mutations with respect to which subgenome the diploid reads were mapped.</p><p>At shallower phylogenetic depths (fig. <ref type="figure">3B</ref>), the difference between diploids and polyploids becomes even clearer, with polyploids exhibiting 3-4% higher proportions of deleterious mutations in the Dt subgenome and 5-12% higher in the At subgenome than their respective diploid progenitors. The most unbiased and straightforward comparison of the asymmetry in strength of the masking effect of duplicate genomes between the two subgenomes of allopolyploid cottons is provided by mutations that have occurred following polyploidization (fig. <ref type="figure">3C</ref>). Here, the At subgenome of all species contain a 2-3% high proportion of deleterious mutations than the Dt subgenome, indicating that differences exist in the strength of the masking effect between the two homoeologous subgenomes that have resided in the same nucleus for over a million years. This pattern is also observed when deleterious mutations are mapped onto the phylogeny (supplementary fig. <ref type="figure">3</ref>, Supplementary Material online). Additionally, there is more variation among species in the At subgenome than in the Dt subgenome, although the patterns in this respect are not simple. The amount of subgenomic asymmetry is smallest in G. darwinii (AD 5 , Yellow) from the Galapagos Islands, and largest in the Brazilian endemic and inland species G. mustelinum (AD 4 , Orange), indicating that asymmetries between subgenomes of the same species may vary within a single clade of allopolyploids.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Disentangling Demography and Selection from Effects of Ploidy</head><p>Demography is a potential confounding factor in estimating the rate of deleterious mutation accumulation. Shifts in demography are known to complicate inferences of the strength of selection and genetic load <ref type="bibr">(Brandvain and Wright 2016)</ref>; for example, even in one of the best studied demographic shifts, the Out of Africa migration in humans, several papers <ref type="bibr">(Lohmueller et al. 2008;</ref><ref type="bibr">Gazave et al. 2013;</ref><ref type="bibr">Simons et al. 2014;</ref><ref type="bibr">Henn et al. 2016;</ref><ref type="bibr">Simons and Sella 2016)</ref> have reached seemingly contradictory conclusions on whether genetic load has increased as a result of these shifts in demography (but see <ref type="bibr">Lohmueller [2014]</ref>). The pattern of deleterious mutation accumulation has also been welldocumented in bottlenecks and population growth associated with domestication in crops such as maize <ref type="bibr">(Wang et al. 2017)</ref>, soybean and barley <ref type="bibr">(Kono et al. 2016)</ref>, sorghum <ref type="bibr">(Lozano et al. 2021)</ref>, cassava <ref type="bibr">(Ramu et al. 2017)</ref>, and rice <ref type="bibr">(Liu et al. 2017)</ref>.</p><p>Polyploidy is typically associated with a population bottleneck <ref type="bibr">(Grant 1981;</ref><ref type="bibr">Barringer 2007</ref>), but because the genetic diversity of both the diploid and polyploid species in this study is low (table 1), demographic modeling of the depth or duration of population bottlenecks and range expansion following polyploid formation is not straightforward. Generalized patterns of the effects of demography on deleterious mutations, however, can serve as a null expectation to test if our data follow the same trends observed under varying demographic scenarios, as explained in the following. Demographic shifts, including population bottlenecks and expansions, have a large influence on the accumulation of deleterious mutations. According to the nearly neutral theory <ref type="bibr">(Ohta 1992)</ref>, the fate of deleterious mutations is determined by genetic drift instead of selection when the selection coefficient (s) of deleterious mutations is less than or equal to 1/ (2 N e ), where N e is the effective population size. The reduction of N e during a population bottleneck would therefore allow weakly deleterious mutations to escape purifying selection (i.e., to behave as if they were neutral), whereas strongly deleterious mutations with a selective coefficient greater than 1/ (2N e ) would still be removed by purifying selection. On the In both demographic scenarios, we expect that mildly or moderately deleterious mutations would be most differentially affected, whereas strongly deleterious mutations would consistently be removed by purifying selection. Based on this theory, if the differences in the number of deleterious mutations we see between diploids and polyploids are due to demography, then we would expect to see most of that difference reflected in mildly, rather than strongly, deleterious mutations. In contrast, if masking of deleterious alleles in polyploids is driving a higher rate of accumulation relative to diploids, this pattern will not be observed.</p><p>To test if our data were consistent with changes in demography, we first asked if there was a correlation between the degree of deleteriousness of a mutation (as measured by GERP) and its relative increase in the polyploids compared with the diploids. To answer this question, we plotted the relative change of deleterious mutations in each subgenome relative to its most closely related diploid progenitor. We plotted this relative change for three different degrees of deleteriousness-strongly deleterious mutations (4 &lt; GERP 6), moderately deleteriousness (2 &lt; GERP 4), and mildly deleteriousness (0 &lt; GERP 2) deleterious (fig. <ref type="figure">4</ref>). We found that in both subgenomes of all six polyploids, when comparing mutations that had originated after the divergence of the diploid from its respective subgenome in the allopolyploids, strongly deleterious mutations accumulated at a faster rate relative to diploids than did moderately or mildly deleterious mutations, which is inconsistent with expectations under a demographic change model alone. We also observed this change under both an additive and recessive model of dominance (supplementary fig. <ref type="figure">5</ref>, Supplementary Material online), and also when evaluating deleteriousness using the "masked constraint" estimates generated by BAD_Mutations (supplementary fig. <ref type="figure">6</ref>, Supplementary Material online, but see <ref type="bibr">Kono et al. [2018]</ref> and <ref type="bibr">Chun and Fay [2009]</ref> for cautions on interpreting this value as deleteriousness). In total, the rate of accumulation among mutations with different inferred degrees of deleteriousness do not suggest that the patterns we see can be explained solely by demographic changes, but that the masking effect of duplicated genes may play an important role in the determining the fate of deleterious mutations in allopolyploids.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Effects of Polyploidy on Deleterious Mutation Accumulation</head><p>One of the earliest hypotheses regarding mutation accumulation in allopolyploids dates back to Haldane <ref type="bibr">(Haldane 1932)</ref> where he posits that in allopolyploids, "one gene may be altered without disadvantage, provided its functions can be performed by a gene in one of the other sets of chromosomes." Allopolyploids are therefore predicted be able to tolerate a higher mutational load than their diploid relatives, and putatively deleterious mutations may accumulate faster in polyploids than in their diploid relatives due to the masking effect of recessive or incompletely dominant deleterious alleles. Here, we demonstrate that these predictions are true in allopolyploid cottons. All polyploids in Gossypium harbor more mutations at phylogenetically conserved sites than do their closest diploid progenitors, as determined by two different methods of detecting deleterious mutations. We also find that the proportion of all nonsynonymous mutations that are inferred to be deleterious is higher in polyploids than in their diploid progenitors and that polyploidy has the greatest effect on strongly deleterious (and, inferentially, more recessive; Eyre-Walker and Keightley 2007; <ref type="bibr">Huber et al. 2018)</ref> mutations. Thus, using the power of comparative phylogenetics and genomics combined with analytical methods for detection of deleterious mutations, we For mutations that originated since the divergence of each subgenome from its diploid progenitor, we plotted the relative increase in deleterious alleles across three GERP load categories: mildly deleterious (0&lt;GERP 2; light gray), moderately deleterious (2&lt;GERP 4; gray), and strongly deleterious (4&lt;GERP 6; black). We used the diploid as the reference population, meaning that the relative increase of GERP load in the diploid is always equal to one for all categories. In both subgenomes of all polyploids, strongly deleterious mutations had the greatest relative increase compared with the diploids, followed by the moderately deleterious mutations, and finally, mildly deleterious mutations. This pattern does not fit the expected patterns under demographic models alone, where most of the changes between two populations should be seen in mildly or moderately deleterious mutations. However, under a model where recessive deleterious mutations are masked by their homoeologs, we would expect that strongly deleterious mutations would accumulate faster than moderately or mildly mutations (i.e., the pattern we see here) due to the correlation between the recessivity of a mutation (h) and its selection coefficient (s). </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Demography Alone Cannot Explain Patterns of Deleterious Mutations in Polyploids</head><p>Estimating the strength of natural selection and genetic load is notoriously challenging <ref type="bibr">(Lohmueller 2014)</ref> and is complicated by shifts in effective population size (including bottlenecks and expansions), mating systems, and effective recombination rates, among other life-history and demographic factors <ref type="bibr">(Brandvain and Wright 2016)</ref>. Here, we illuminate an additional relevant consideration, that is, wholegenome duplication. Yet many of the considerations for populations that are not in demographic equilibrium also apply to Gossypium. Diversification in the cotton tribe (Gossypieae) has been characterized by numerous long-distance dispersal events <ref type="bibr">(Grover et al. 2017</ref>), including the one from Africa to the Americas 1-2 Mya that led to the evolution of allopolyploid Gossypium. We note that in the Hawaiian Islands endemic G. tomentosum, the total number of synonymous substitutions is not significantly different from the rest of the polyploids, but the number of nonsynonymous and deleterious mutations is significantly increased, suggesting that the genetic bottleneck associated with island dispersal has elevated the number of deleterious mutations compared with the rest of the polyploids.</p><p>Although demographic changes upon polyploid formation have been shown to change the number and frequency of deleterious mutations in other systems <ref type="bibr">(Douglas et al. 2015;</ref><ref type="bibr">Paape et al. 2018;</ref><ref type="bibr">Baduel et al. 2019;</ref><ref type="bibr">Kryvokhyzha, Milesi, et al. 2019;</ref><ref type="bibr">Kryvokhyzha, Salcedo, et al. 2019)</ref>, we show here that the patterns of mutation accumulation in Gossypium cannot be explained by demography alone, and that the data are more consistent with the nearly century-old hypothesis that recessive deleterious mutations can accumulate faster in allopolyploids due to the masking effect of duplicated genes and lack of recombination between subgenomes <ref type="bibr">(Haldane 1932)</ref>. Specifically, we show that strongly (and, hence, more recessive; <ref type="bibr">Morton et al. 1956;</ref><ref type="bibr">Mukai et al. 1972;</ref><ref type="bibr">Eyre-Walker and Keightley 2007;</ref><ref type="bibr">Agrawal and Whitlock 2011;</ref><ref type="bibr">Huber et al. 2018</ref>) deleterious mutations accumulate faster in polyploids compared with diploids than moderately or mildly deleterious mutations, and that this pattern is inconsistent with demographic shifts or longterm change in population size (fig. <ref type="figure">4</ref> and supplementary fig. <ref type="figure">5</ref>, Supplementary Material online).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Asymmetry in Subgenomes in the Distribution of Deleterious Mutations</head><p>One of the elegant attributes of a clade of allopolyploid genomes derived from a single polyploidization event is that they offer a remarkable natural experiment for comparing subgenomes that have resided within the same nucleus for, in the case of Gossypium, approximately 1.5 My. Once an allopolyploid is established, each subgenome is subjected to identical external or population-level factors, including demography, mating systems, and environmental and ecological conditions, as well as internal cellular processes, including identical DNA replication and recombination machinery. These features remove many of the confounding factors that may influence the genetic load and provide a simple comparative context for revealing evolutionary forces that might differentially affect coresident genomes or homoeologs.</p><p>An unexpected finding of our analyses is the striking asymmetry in the proportion of all nonsynonymous mutations that are inferred to be deleterious between the two subgenomes of all allopolyploid species in Gossypium. We found that the At subgenome of all species contains 2-3% more nonsynonymous mutations that are inferred to be deleterious (fig. <ref type="figure">3</ref>) even when only considering mutations that have arisen following the earliest allopolyploid diversification events, and correcting for removing the biases of unequal phylogenetic distances to each subgenome's model progenitor diploid. Our work adds to a growing recognition that the two coresident subgenomes in cotton allopolyploids may be shaped asymmetrically by evolutionary processes, including interspecific introgression and selection under domestication <ref type="bibr">(Fang, Guan, et al. 2017;</ref><ref type="bibr">Fang, Wang, et al. 2017;</ref><ref type="bibr">Chen et al. 2020;</ref><ref type="bibr">Yuan et al. 2021)</ref>, and that this phenomenon also extends to other important allopolyploid crop plants, including wheat <ref type="bibr">(Pont and Salse 2017;</ref><ref type="bibr">Jiao et al. 2018)</ref> and Brassica <ref type="bibr">(Tong et al. 2020)</ref>.</p><p>Teasing apart the genesis of differential subgenomic responses to selection is rendered challenging by several factors independent of phylogeny. We note, for example, the relevant example of the recently formed allopolyploid Capsella bursa-pastoris and its diploid progenitors, where consistent asymmetries in genetic load are reported between the subgenomes <ref type="bibr">(Kryvokhyzha, Milesi, et al. 2019;</ref><ref type="bibr">Kryvokhyzha, Salcedo, et al. 2019</ref>) the differences likely reflect the dramatically different mating systems of the progenitors, in which the subgenome with the higher genetic load originated from an obligate outcrosser, C. grandiflora (N e &#188; 800,000), whereas the subgenome with the lower genetic load derives from the predominantly selfing C. orientalis (N e &#188; 5,000) <ref type="bibr">(Douglas et al. 2015)</ref>. In another recently formed (20-250 ka) allopolyploid, Arabidopsis kamchatica, no asymmetry in the distribution of fitness effects between subgenomes was found, although it was observed that each subgenome of the allopolyploid contained more neutral and fewer deleterious alleles than either of the diploid progenitors <ref type="bibr">(Paape et al. 2018)</ref>. It is unclear, however, whether this shift was due to allopolyploidy per se or if it reflects the transition from an obligate outcrossing to a mating system with some degree of inbreeding, with a concomitant purging of partially or completely recessive deleterious alleles, as shown in several other systems <ref type="bibr">(Arunkumar et al. 2015;</ref><ref type="bibr">Roessler et al. 2019)</ref>. In Gossypium, all species have similar mating systems and a canonical outcrossing floral morphology including highly exserted styles and stigmas. Population sizes often are small, however, likely leading to relatively high levels of generalized inbreeding. At present, however, no data exist that address these considerations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Polyploidy, Redundancy, and Fitness Effects</head><p>One possible interpretation of our results is that Gossypium polyploids are less fit than their closely related diploid Deleterious Mutations Accumulate Faster in Allopolyploids . doi:10.1093/molbev/msac024 MBE progenitors because they harbor more deleterious mutations in their genomes, especially mutations that have already been driven to fixation. We note, however, that the fitness effects of a mutation may change as a result of the genetic (e.g., epistasis) or environmental (e.g., local adaptation, conditional neutrality) context in which it occurs <ref type="bibr">(Huang et al. 2021)</ref>. Comparative genomics techniques used to infer deleterious mutations at phylogenetically conserved sites, as employed here, cannot identify these shifts in fitness effects <ref type="bibr">(Huber et al. 2018)</ref>, and also frequently incorrectly identify beneficial mutations as deleterious <ref type="bibr">(Chen et al. 2020)</ref>. Therefore, an additional possibility is that mutations in polyploids that occur at phylogenetically conserved sites may not actually have a deleterious effect on fitness as they do in diploids. Inferring the genetic load of a population simply by counting the number of deleterious variants also assumes that all alleles contribute independently to the total genetic load of a population. However, because of the functional overlap of duplicated genes and, in most cases, absence of recombination between homoeologous chromosomes in an allopolyploid, a recessive deleterious mutation can never be present in all four copies of a gene and thus may be invisible to selection because of the masking effect of its homoeologous partner.</p><p>An important takeaway from this study is that recessive deleterious mutations in allopolyploids, at least at some loci, may actually accumulate in a manner more similar to neutral mutations, presumably because of the lack of recombination between subgenomes and, hence, the inability of purifying selection to "see" the negative effects of these mutations. Because these recessive deleterious mutations escape the effects of purifying selection, many traditional tests for detecting positive and negative selection (e.g., dN/dS, p N /p S ) may be biased when comparing a polyploid to diploid because the polyploid would be expected to accumulate putatively deleterious sites more quickly (and maintain a higher genetic diversity at nonsynonymous sites) than their diploid relatives. This increased dN/dS value in allopolyploids compared with diploid progenitors was recently shown in five allopolyploid systems in addition to Gossypium <ref type="bibr">(Sharbrough et al. 2022)</ref>, and duplicated genes associated with an ancient polyploid event in Brassica rapa were shown to contain higher amounts of genetic variation than nonduplicated genes <ref type="bibr">(Qi et al. 2021)</ref>, indicating this phenomenon is likely not restricted to Gossypium and should be taken into consideration for other allopolyploid/diploid comparisons. Additionally, although this bias will be most notable in genome-wide comparisons, it may also be evident at the individual gene level, given that both homoeologs are still present and retain some degree of functional overlap (although the robustness of attributing this bias to a ploidy effect acting on a single gene is expected to be low).</p><p>Another important implication of this finding is that allopolyploidy (or gene duplication in general) may play an important and underrecognized role in determining how selection acts on new mutations, notwithstanding the burgeoning literature on fates of duplicate gene evolution <ref type="bibr">(Conant et al. 2014;</ref><ref type="bibr">Shi et al. 2020;</ref><ref type="bibr">Veitia and Birchler 2022)</ref>. The evolutionary trajectory of new mutations will largely be dependent on the selection coefficient (s) acting on that locus, and the dominance coefficient (h), defined as the proportion of the fitness cost that a mutation harbors when in a heterozygous state. In allopolyploids, however, the evolutionary fate of new mutations may be determined not only by allelic dominance at that locus, but also by the interaction arising from the coexistence of its homoeologous locus, a term we call "homoeologous epistatic dominance." The relationships between this homoeologous epistatic dominance, allelic dominance, and the selection coefficient are likely complicated and potentially heavily influenced by other biological considerations such as biased expression of homoeologs, sub-or neofunctionalization, and homoeologous recombination, among others. Moreover, notwithstanding these polyploidy-specific effects, even the genome-wide relationships between two of these factors, allelic dominance and the selection coefficient, have only been modeled using genomic data in the past few years <ref type="bibr">(Huber et al. 2018)</ref>.</p><p>Nonetheless, understanding how this homoeologous epistatic dominance impacts the fitness effects of new mutations is an unexplored aspect of polyploid genome evolution, and it is not yet clear whether this will equally affect advantageous and deleterious variants. How homoeologous epistatic dominance operates with respect to functional properties arising from considerations such as gene balance <ref type="bibr">(Veitia and Birchler 2022)</ref>, dosage effects <ref type="bibr">(Conant et al. 2014)</ref>, structural and functional entanglement <ref type="bibr">(Kuzmin et al. 2020</ref><ref type="bibr">(Kuzmin et al. , 2021))</ref>, and intersubgenomic cis-and trans-effects <ref type="bibr">(Bottani et al. 2018;</ref><ref type="bibr">Hu and Wendel 2019)</ref> would seem to represent important avenues for understanding how natural selection operates differently in polyploids compared with diploids. From an applied perspective, these insights could be important in agriculture, particularly because so many of our most important crop plants have a recent history that includes polyploidy (Renny-Byfield and Wendel 2014), and segregating patterns of genome fractionation have the potential to serve as targets of selection in crop improvement <ref type="bibr">(Hufford et al. 2021)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Materials and Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Plant Materials and Sequencing</head><p>We used whole-genome sequencing data from 46 individuals in Gossypium, including between two and ten individuals from each of eight species. Included in our sampling was six polyploid species originating from a single polyploidization event 1-2 Mya <ref type="bibr">(Wendel 1989;</ref><ref type="bibr">Hu et al. 2021)</ref>, two diploid species representing models of the genome donors to the allopolyploids (A and D), and three species from Australia that served as outgroups for polarizing mutations into ancestral and derived states. These sequences were previously described <ref type="bibr">(Yuan et al. 2021)</ref>, and SRA codes for all 46 resequenced individuals are listed in supplementary table 1, Supplementary Material online. For G. hirsutum, we randomly chose ten accessions that were classified in the "Wild" population from Yuan et al. <ref type="bibr">(Yuan et al. 2021)</ref>, and for the other species, we chose all accessions available that did not show evidence of being mislabeled, as determined by a PCA plot. <ref type="bibr">Conover and Wendel . doi:10</ref>.1093/molbev/msac024 MBE After the data were downloaded from NCBI, adapter sequence removal and quality score filtering of FASTQ reads were performed using Trimmomatic v0.36 <ref type="bibr">(Bolger et al. 2014)</ref> using the parameters "LEADING:28 TRAILING:28 SLIDINGWINDOW:8:28 SLIDINGWINDOW:1:10 MINLEN:65 TOPHRED33." Trimmed reads from each polyploid sample were mapped to the 26 chromosomes of the G. hirsutum reference genome <ref type="bibr">(Saski et al. 2017)</ref>, and reads from each diploid sample were mapped to each subgenome separately to avoid competitive mapping of the diploid reads against a tetraploid reference genome. Reads from the three outgroup species were separately mapped to both subgenomes to ensure that reads were not filtered out for mapping to multiple parts of the genome. All mapping was done using bwa-mem v0.7.17 <ref type="bibr">(Li and Durbin 2009)</ref> and only uniquely mapping paired reads (-F 260 flag) that were mapped in their proper orientation (-f 2 flag) were retained using Samtools v1.9 <ref type="bibr">(Li et al. 2009</ref>) before the files were sorted and converted to bam files. Using the Sentieon <ref type="bibr">(Kendig et al. 2019</ref>) SNP Calling program, gVCF files were generated, and joint genotyping was performed using the GVCFtyper algorithm (see Github repository for full scripts). Variant filtering was performed using GATK v4.0.4.0 using the filter expression "QD &lt; 2.0 k FS &gt; 60.0 k MQ &lt; 40.0 k SOR &gt; 4.0 k MQRankSum &lt; &#192;12.5 k ReadPosRankSum &lt; &#192;8.0." For each species (excluding the outgroup species, and treating G. stephensii, and G. ekmanianum as a single species), we nullified any variant call in which all individuals were heterozygous to remove any collapsed genomic region in the reference genome or paralogous regions that were not present in the reference genome. We treated G. stephensii and G. ekmanianum as a single species because we only sampled two individuals of G. stephensii, so removing any sites in which both individuals were heterozygous errantly removed real variants that were not due to paralogy mapping issues. All scripts for generating and filtering variant calls are located on our GitHub repository (<ref type="url">https://github.com/conJUSTover/ Deleterious-Mutations-Polyploidy</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Identification of Homoeologs</head><p>We used the pSONIC pipeline <ref type="bibr">(Conover et al. 2021</ref>) to identify syntenically conserved homoeologs in the G. hirsutum reference genome, and kept only homoeologous pairs that were less than 5% different in their total annotated CDS length. To remove homoeologous pairs that may have experienced homoeologous exchange events (though there is scant evidence for this in Gossypium; <ref type="bibr">Salmon et al. 2010;</ref><ref type="bibr">Flagel et al. 2012;</ref><ref type="bibr">Chen et al. 2020)</ref>, we removed any pair in which the proportion of the reads from the two progenitor diploid genomes (termed At and Dt in the allopolyploid, the "t" indicating "tetraploid") did not meet the expected 2:2 ratio, similar to previous analyses of homoeologous exchange events in other allotetraploids <ref type="bibr">(Bird et al. 2021)</ref>. Average read depth of CDS regions was determined by bedtools2 v.2.27.1 <ref type="bibr">(Quinlan and Hall 2010)</ref>. Briefly, for a single homoeologous pair, we calculated the average read depth of the At homoeolog divided by the sum of the average read depth of both homoeologs for each individual. We removed any homoeologous pair in which this fraction was less than 37.5 or greater than 62.5 in any individual. We expect any HEs that result in a 0:4 At: Dt copy number to contain 0% At reads/ total reads; HEs that result in 1:3 At: Dt copy number should have a 25% At reads/total reads; HEs that result in a 3:1 At: Dt copy number should have a 75% At reads/total reads; HEs that result in a 4:0 At: Dt copy number should have a 100% At reads/total reads; and no HE (i.e., 2:2 At: Dt copy number) would result in a 50% At reads/total reads. We used the midpoints between the "No HE" and the 1:3 and 3:1 copy numbers as cutoff points. Because we compared the read depth for homoeologous pairs within each individual, we expect there to be no difference in read depth between homoeologs, and differing read depths among individuals should not introduce any bias in our gene filtering criteria. This filtering resulted in 8,884 homoeologous pairs (17,768 genes) being analyzed further.</p><p>Nonreciprocal homoeologous exchanges (i.e., homoeologous gene conversion) could also bias the estimates of the genetic load in a way that is not related to new mutation following polyploidization or speciation. To control for positions in these non-HE homoeologs that may be influenced by gene conversion, we linked variant positions between homoeologs in the following way. We first performed pairwise alignments of the CDS sequences using MACSEv2 <ref type="bibr">(Ranwez et al. 2011</ref><ref type="bibr">(Ranwez et al. , 2018))</ref>, which aligns CDS sequences in accordance with their translated amino acid sequences, but allows for the possibility of frameshift mutations. We then used the aligned CDS sequences to identify where indels were present and found the corresponding genomic positions for every nucleotide in the alignment, inserting gaps where indels occurred. We then extracted the genomic positions for each SNP position as well as the genomic position for its aligned nucleotide. We retained only those homoeologous SNP positions in which both positions had a confidently called ancestral allele (described above) and in which the ancestral allele matched between the two homoeologs. Importantly, for homoeologs that were encoded in opposite orientations in the reference genome (i.e., one homoeolog was encoded on the forward strand of the reference genome, and the other homoeolog was encoded on the reverse complement), we ensured that the inferred ancestral states for the two SNP positions included both nucleotides of a purine/pyrimidine pair (e.g., the ancestral state for homoeologous SNP was "A," whereas the ancestral state of the other homoeologous SNP was "T"). We also removed any pair of homoeologous SNPs in which more than two alleles were present (while similarly treating homoeologous pairs encoded in opposite directions as described in the previous sentence).</p><p>In total, we only used those SNP sites that: 1) did not link to an indel in its homoeologous pair, 2) were biallelic and had consistently inferred ancestral states in the two subgenomes, 3) the derived allele was found in only one of the two subgenomes or their respective diploid progenitors, and 4) the derived allele was fixed in a diploid and segregating in its respective subgenome (or vice-versa).</p><p>Deleterious Mutations Accumulate Faster in Allopolyploids . doi:10.1093/molbev/msac024 MBE</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>Downloaded from https://academic.oup.com/mbe/article/39/2/msac024/6517786 by guest on 02 December 2022</p></note>
		</body>
		</text>
</TEI>
