<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Ecological generalism drives hyperdiversity of secondary metabolite gene clusters in xylarialean endophytes</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>02/01/2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10373352</idno>
					<idno type="doi">10.1111/nph.17873</idno>
					<title level='j'>New Phytologist</title>
<idno>0028-646X</idno>
<biblScope unit="volume">233</biblScope>
<biblScope unit="issue">3</biblScope>					

					<author>Mario E. Franco</author><author>Jennifer H. Wisecaver</author><author>A. Elizabeth Arnold</author><author>Yu‐Ming Ju</author><author>Jason C. Slot</author><author>Steven Ahrendt</author><author>Lillian P. Moore</author><author>Katharine E. Eastman</author><author>Kelsey Scott</author><author>Zachary Konkel</author><author>Stephen J. Mondo</author><author>Alan Kuo</author><author>Richard D. Hayes</author><author>Sajeet Haridas</author><author>Bill Andreopoulos</author><author>Robert Riley</author><author>Kurt LaButti</author><author>Jasmyn Pangilinan</author><author>Anna Lipzen</author><author>Mojgan Amirebrahimi</author><author>Juying Yan</author><author>Catherine Adam</author><author>Keykhosrow Keymanesh</author><author>Vivian Ng</author><author>Katherine Louie</author><author>Trent Northen</author><author>Elodie Drula</author><author>Bernard Henrissat</author><author>Huei‐Mei Hsieh</author><author>Ken Youens‐Clark</author><author>François Lutzoni</author><author>Jolanta Miadlikowska</author><author>Daniel C. Eastwood</author><author>Richard C. Hamelin</author><author>Igor V. Grigoriev</author><author>Jana M. U’Ren</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Although secondary metabolites are typically associated with competitive or pathogenic interactions, the high bioactivity of endophytic fungi in the Xylariales, coupled with their abundance and broad host ranges spanning all lineages of land plants and lichens, suggests that enhanced secondary metabolism might facilitate symbioses with phylogenetically diverse hosts.Here, we examined secondary metabolite gene clusters (SMGCs) across 96 Xylariales genomes in two clades (Xylariaceae s.l. and Hypoxylaceae), including 88 newly sequenced genomes of endophytes and closely related saprotrophs and pathogens. We paired genomic data with extensive metadata on endophyte hosts and substrates, enabling us to examine genomic factors related to the breadth of symbiotic interactions and ecological roles.All genomes contain hyperabundant SMGCs; however, Xylariaceae have increased numbers of gene duplications, horizontal gene transfers (HGTs) and SMGCs. Enhanced metabolic diversity of endophytes is associated with a greater diversity of hosts and increased capacity for lignocellulose decomposition.Our results suggest that, as host and substrate generalists, Xylariaceae endophytes experience greater selection to diversify SMGCs compared with more ecologically specialised Hypoxylaceae species. Overall, our results provide new evidence that SMGCs may facilitate symbiosis with phylogenetically diverse hosts, highlighting the importance of microbial symbioses to drive fungal metabolic diversity.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Fungal endophytes inhabit asymptomatic, living photosynthetic tissues of all major lineages of plants and lichens to form one of Earth's most prevalent groups of symbionts <ref type="bibr">(Arnold et al., 2009;</ref><ref type="bibr">Peay et al., 2016)</ref>. Known from a wide range of biomes and agroecosystems (e.g. U' <ref type="bibr">Ren et al., 2012</ref><ref type="bibr">Ren et al., , 2019))</ref>, endophytes impact plant health, productivity and evolution <ref type="bibr">(Rodriguez et al., 2009)</ref>. Although classified together due to ecologically similar patterns of colonisation, transmission and in planta biodiversity <ref type="bibr">(Rodriguez et al., 2009)</ref>, foliar fungal endophytes represent a diversity of evolutionary histories, life history strategies and functional traits <ref type="bibr">(Porras-Alfaro &amp; Bayman, 2011)</ref>. Despite the recent surge of interest in plant microbiome research <ref type="bibr">(Trivedi et al., 2020)</ref> the genomic and molecular mechanisms foliar fungal endophytes employ to establish symbiotic host associations remain largely unknown.</p><p>Global, large-scale surveys of phylogenetically diverse plant and lichen hosts have revealed that many foliar endophyte species preferentially associate with particular host species and lineages, resulting in host structured endophyte communities at local to global scales (e.g. U' <ref type="bibr">Ren et al., 2019)</ref>. However, endophytic fungi in the Xylariales (Sordariomycetes, Pezizomycotina, Ascomycota) appear unique in that they frequently have broad host ranges that span multiple lineages of land plants (e.g. angiosperms, conifers, lycophytes, ferns and mosses) as well as green algae and cyanobacteria within lichen thalli (e.g. <ref type="bibr">Arnold et al., 2009;</ref><ref type="bibr">U'Ren et al., 2016)</ref>. By contrast, described Xylariales species associate primarily with angiosperms as wood-or litter-degrading saprotrophs or woody pathogens <ref type="bibr">(Hsieh et al., 2005</ref><ref type="bibr">(Hsieh et al., , 2010))</ref>. Although the genetic factors that determine foliar endophyte host range are currently unknown, research on fungal pathogens has shown that host specificity is often determined by the presence of avirulence proteins (i.e. effectors), proteinaceous host-specific toxins and secondary metabolites (SMs) <ref type="bibr">(Li et al., 2020)</ref>. Horizontal gene transfer (HGT) of these host-determining genes frequently alters and/or expands pathogen host range <ref type="bibr">(Li et al., 2020)</ref>.</p><p>Xylariales genomes sequenced to date have revealed a rich repertoire of secondary metabolite gene clusters (SMGCs) <ref type="bibr">(Wibberg et al., 2021)</ref>, often exceeding the numbers reported for saprotrophic fungi well known for their SM production (Aspergillus, Penicillium) <ref type="bibr">(Nielsen et al., 2017;</ref><ref type="bibr">Drott et al., 2021)</ref>. Previously, it was postulated that intense competition with diverse communities of soil organisms increases selection to maintain and diversify SMGCs <ref type="bibr">(Slot, 2017)</ref>. However, the high bioactivity of Xylariales fungi (&gt; 500 SMs reported to date; <ref type="bibr">Becker &amp; Stadler, 2021)</ref>, their broad host ranges as endophytes and their ability to persist in leaf litter as saprotrophs that decompose lignocellulose (U'Ren <ref type="bibr">&amp; Arnold, 2016;</ref><ref type="bibr">U'Ren et al., 2016)</ref>, led us to hypothesise that enhanced secondary metabolism might play a role in facilitating ecological generalism in both substrate use and the phylogenetic breadth of their symbiotic associations with plants and lichens.</p><p>To test this hypothesis, we examined the genomic factors associated with endophyte host range and ecological roles (i.e. endophytic, pathogenic and saprotrophic) across 96 genomes of Xylariales, including 88 newly sequenced genomes of endophytes, saprotrophs and plant pathogens within two major clades of Xylariales [Hypoxylaceae and Xylariaceae sensu lato (s.l.)]. We paired genomic data with extensive metadata on endophyte host associations, geographic distributions and substrate usage gleaned from a collection of &gt; 6000 xylarialean endophytes isolated from phylogenetically diverse plants and lichens across North America (U' <ref type="bibr">Ren et al., 2016)</ref>, enabling us to examine for the first time the genomic factors related to the breadth of symbiotic interactions and ecological roles in this dynamic and ecologically important fungal clade.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Materials and Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Fungal strain selection</head><p>We sequenced genomes of 44 endophytic taxa (U' <ref type="bibr">Ren et al., 2012;</ref><ref type="bibr">U'Ren &amp; Arnold, 2016)</ref> and 44 named taxa of Xylariaceae s.l. and Hypoxylaceae representing c. 24 genera and 80 species, as well as two undescribed species of endophytic Xylariales (Pestalotiopsis sp. NC0098 and Xylariales sp. AK1849) included in the outgroup (Supporting Information Table <ref type="table">S1</ref>). Isolates were selected based on their phylogenetic position and ecological mode from U' <ref type="bibr">Ren et al. (2016)</ref>. Although classifying fungal ecological modes broadly as 'endophytic' or 'saprotrophic' based on the condition of the tissue from which they are cultured is often insufficient to adequately define their ecological roles, for the purposes of this study, isolates cultured from living host tissues (either plant or lichen) are referred to as endophytes even if other isolates in the same fungal operational taxonomic unit (OTU) were found in nonliving tissues as well. Isolates were defined as saprotrophs only if all isolates in the OTU were cultured from nonliving plant tissues such as senescent leaves or leaf litter (U' <ref type="bibr">Ren et al., 2016)</ref>. To minimise the effect of phylogeny when assessing the impact of ecological mode on genome evolution, we also selected 15 pairs of closely related sister taxa with contrasting ecological modes (i.e. endophyte vs nonendophyte) (U' <ref type="bibr">Ren et al., 2016)</ref>. For reference species that lacked host and substrate metadata, ecological modes were estimated based on information for that species in the literature as described in U' <ref type="bibr">Ren et al. (2016)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>DNA and RNA purification</head><p>We used two different mycelial growth and cultivation techniques to obtain DNA for either Illumina or PacBio Single-Molecule Real-Time (SMRT) sequencing. DNA isolations were performed using modified phenol/chloroform extractions (Methods S1). RNA was extracted for each isolate with the Ambion Purelink RNA Kit (Thermo Fisher Scientific, Waltham, MA, USA). DNA and RNA were quantified with a Qubit fluorometer (Invitrogen) and sample purity was assessed using NanoDrop spectrophotometer (BioNordika, Herlev, Denmark). RNA was treated with DNase (Thermo Fisher Scientific) following the manufacturer's instructions and RNA integrity was assessed on a BioAnalyser at the University of Arizona Genomics Core Facility.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Genome and transcriptome sequencing and assembly</head><p>Genomes were generated at the Department of Energy Joint Genome Institute using Illumina and PacBio technologies (Table <ref type="table">S1</ref>). For 66 isolates, Illumina standard shotgun libraries (insert sizes of 300 bp or 600 bp) were constructed and sequenced using the NovaSeq platform. Raw reads were filtered using the JGI QC pipeline. An assembly of the target genome was generated using the resulting nonorganelle reads with SPAdes <ref type="bibr">(Bankevich et al., 2012)</ref>. PacBio SMRT sequencing was performed for 22 isolates of Xylariaceae s.l. and Hypoxylaceae, as well as Xylariales spp. NC0098 and AK1849 on a PacBio SEQUEL. Library preparation was performed using either the PacBio low input 10 kb or PacBio &gt; 10 kb with AMPure bead size selection. Filtered subread data were processed with the JGI QC pipeline and de novo assembled using Falcon (SEQUEL) or Flye (SEQUEL II) systems. Stranded RNA-seq libraries were created and quantified using qPCR and transcriptome sequencing was performed on an Illumina NovaSeq S4 instrument. For both Hypoxylaceae and Xylariaceae s.l., c. 25% of genomes were sequenced with PacBio, although a higher proportion of endophyte genomes were sequenced with PacBio rather than Illumina (43% vs 28% overall). Genome completeness was assessed by benchmarking universal single-copy orthologues (BUSCO) v.2.0 using the 'eukaryota_odb9' (2016-11-02) dataset (10.1093/ bioinformatics/btv351).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Genome annotation</head><p>Gene prediction and annotation was performed using the JGI pipeline <ref type="bibr">(Grigoriev et al., 2014;</ref><ref type="bibr">Kuo et al., 2014)</ref> (Methods S1). Predicted genes were annotated using functional information from InterPro <ref type="bibr">(Mitchell et al., 2019)</ref>, PFAM <ref type="bibr">(El-Gebali et al., 2019)</ref>, gene ontology (GO) (The Gene Ontology Consortium, 2019), kyoto encyclopedia of genes and genomes (KEGG) <ref type="bibr">(Kanehisa et al., 2006)</ref>, eukaryotic orthologous groups of proteins (KOG) <ref type="bibr">(Tatusov et al., 2003)</ref>, the carbohydrate-active enzymes database (CAZy) <ref type="bibr">(Lombard et al., 2014)</ref>, MEROPS database <ref type="bibr">(Rawlings et al., 2016)</ref>, the transporter classification database (TCDB) <ref type="bibr">(Saier et al., 2016</ref><ref type="bibr">), SIGNALP v.3.0a (Nielsen, 2017)</ref> and EFFECTORP 2.0 <ref type="bibr">(Sperschneider et al., 2018)</ref>. CAZymes involved in the degradation of the plant cell wall were classified by substrate <ref type="bibr">(Kameshwar et al., 2019)</ref>. We examined repetitive elements using REPEATSCOUT <ref type="bibr">(Price et al., 2005)</ref>, which identifies novel repeats in the genomes, and REPEATMASKER (<ref type="url">http://repeatmasker.  org</ref>), which identifies known repeats based on the Repbase library <ref type="bibr">(Bao et al., 2015)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Orthogroup prediction, functional annotation and ancestral state reconstruction</head><p>For comparative analyses, data from an additional eight taxa in Xylariaceae s.l. <ref type="bibr">(Wu et al., 2017)</ref> and 23 additional genomes of Sordariomycetes were obtained from MycoCosm <ref type="bibr">(Grigoriev et al., 2014)</ref> (Table <ref type="table">S1</ref>). Orthologous gene families (i.e. orthogroups) for all 121 genomes (ingroup and outgroup) were inferred by ORTHOFINDER v.2.3.3 <ref type="bibr">(Emms &amp; Kelly, 2019)</ref>, which was executed using DIAMOND v.0.9.22 <ref type="bibr">(Buchfink et al., 2015)</ref> for the all-versus-all sequence similarity search and MAFFT v.7.427 <ref type="bibr">(Katoh &amp; Standley, 2013)</ref> for sequence alignment. Orthogroups were assigned functional annotations with KINFIN v.1.0 <ref type="bibr">(Laetsch &amp; Blaxter, 2017)</ref>, which performs a representative functional annotation of the orthogroups based on both the proportion of proteins in the group carrying a specific annotation, as well as the proportion of taxa in the cluster with such annotation. KINFIN was also used to perform network analysis of orthogroups, classify orthogroups and SMGCs into isolate-specific, clade-specific (Hypoxylaceae and Xylariaceae s.l.) and universal (i.e. orthogroups present in all taxa) categories, and to identify orthogroups that were significantly enriched or depleted in the Xylariaceae s.l. or Hypoxylaceae using the Mann-Whitney Utest. We used COUNT v.10.04 (Csur&#8364; os, 2010) with the unweighted Wagner parsimony method (gain and loss penalties both set to 1) to assess changes in the size of orthologous gene families over evolutionary time. Orthogroup annotations were also used to reconstruct the ancestral gene content for subsets of orthologous gene families corresponding to different functional categories.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Phylogenomic analysis</head><p>Protein sequences of 1526 single-copy orthogroups defined by ORTHOFINDER were aligned using MAFFT v.7.427 <ref type="bibr">(Katoh &amp; Standley, 2013)</ref>, concatenated and analysed using maximum likelihood in IQ-TREE multicore v.1.6.11 <ref type="bibr">(Nguyen et al., 2015)</ref> with the Le Gascuel (LG) substitution model. Node support was calculated with 1000 ultrafast bootstrap replicates. Additional phylogenomic analyses with different models of evolution, gene sets and outgroup taxa resulted in nearly identical topologies (Methods S1).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Metabolic gene cluster prediction</head><p>Secondary metabolite gene clusters were predicted using ANTISMASH v.5.1.0 <ref type="bibr">(Blin et al., 2019)</ref> setting the strictness to 'relaxed' and enabling 'KnownClusterBlast', 'ClusterBlast', 'SubClusterBlast', 'ActiveSiteFinder', 'Cluster Pfam analysis' and 'Pfam-based GO term annotation'. CLINKER and CLUSTERMAP.JS were used to visualise and compare SMGCs <ref type="bibr">(Gilchrist &amp; Chooi, 2021)</ref>. Sequence similarity network analysis of the SMGCs was performed using <ref type="bibr">BIG-SCAPE v.1.0.1 (Navarro-Mu&#241;oz et al., 2020)</ref>. BIG-SCAPE was executed under the hybrid mode, enabling the inclusion of singletons and the SMGCs from the MIBiG repository v.1.4 <ref type="bibr">(Medema et al., 2015)</ref>. The output from BIG-SCAPE was incorporated into KINFIN <ref type="bibr">(Laetsch &amp; Blaxter, 2017)</ref> to visualise gene content similarity as network graphs and examine SMGC distribution across clades.</p><p>We used a custom pipeline (<ref type="url">https://github.com/egluckthaler/  cluster_retrieve</ref>) to examine fungal metabolic gene clusters involved in the degradation of a broad array of plant phenylpropanoids <ref type="bibr">(Gluck-Thaler et al., 2018)</ref> (from this point forwards, catabolic gene clusters: CGCs). Cluster_retrieve searches for multiple 'cluster models' containing one of 13 anchor genes <ref type="bibr">(Gluck-Thaler et al., 2018)</ref>. Homologous genes in each locus were defined by a minimum BLASTP (v.2.2.25+) bitscore of 50, 30% amino acid identity, and target sequence alignment 50-150% of the query sequence length. Homologues of query genes were considered clustered if separated by &lt; 7 intervening genes. However, CGCs often share many gene families among classes, resulting in overlapping and adjacent clusters detected by different cluster profile searches. As the majority of CGCs have not been functionally characterised, rather than splitting loci by functional annotation alone, we empirically assessed the spatial distribution of genes in 25 contigs that contained multiple consolidated cluster predictions. Based on these results, we selected a gap size of 30 kb to define discrete clusters (i.e. clusters on the same contig were consolidated if separated by &lt; 30 kb). Homologous cluster families across genomes were inferred using a modified version of BIG-SCAPE (Navarro-Mu&#241;oz et al., 2020) (i.e. adding catabolic anchor genes to 'anchor_domains.txt' and manually tuning the 'Others' cluster type model parameters until known related clusters, such as quinate dehydrogenase clusters, merged into families). Tuning resulted in the values 0.35 for the Jaccard dissimilarity of cluster Pfams, 0.63 for Pfam sequence similarity, 0.02 adjacency index and 2.0 anchor boost.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Detection of HGT events</head><p>We used the ALIEN INDEX (AI) pipeline (<ref type="url">https://github.itap.  purdue.edu/jwisecav/phylowise</ref>) <ref type="bibr">(Wisecaver et al., 2016;</ref><ref type="bibr">Verster et al., 2019)</ref> to identify HGT candidate genes. Each predicted protein sequence was queried against a custom protein database using DIAMOND v.0.9.22.123 <ref type="bibr">(Buchfink et al., 2015)</ref>. The custom database consisted of protein sequences from NCBI RefSeq (release 98) (O' <ref type="bibr">Leary et al., 2016)</ref>, the marine microbial eukaryotic transcriptome sequencing project (MMETSP) <ref type="bibr">(Keeling et al., 2014)</ref> and the 1000 plants transcriptome sequencing project (OneKP) <ref type="bibr">(Matasci et al., 2014)</ref>. DIAMOND results were sorted based on the normalised bitscore (nbs), where nbs was calculated as the bitscore of the single best high scoring segment pair (HSP) in the hit sequence divided by the best bitscore possible for the query sequence (i.e. the bitscore of the query aligned to itself).</p><p>To identify HGT candidates, an ancestral lineage is first specified and the AI score calculated using the formula: AI = nbsO -nbsA, where nbsO is the normalised bit score of the best hit to a species outside of the ancestral lineage and nbsA is the normalised bit score of the best hit to a species within the ancestral lineage. AI scores range from &#192;1 to 1, being greater than zero if the predicted protein sequence had a better hit to species outside of the ancestral lineage and can be suggestive of either HGT or contamination <ref type="bibr">(Wisecaver et al., 2016)</ref>. To identify HGTs present in multiple species, a recipient sublineage within the larger ancestral lineage may also be specified to identify their shared HGT candidates (Fig. <ref type="figure">S1</ref>). All hits to the recipient lineage are skipped so as not to be included in the nbsA calculation. To identify candidate HGTs acquired from distant gene donors (e.g. viruses, bacteria, or plants) we performed a first AI screen using Ascomycota (NCBI: txid4890) and Xylariomycetidae (NCBI: txid 222545) as the ancestral and recipient lineages, respectively (Fig. <ref type="figure">S1</ref>). To identify candidate horizontal transfers of genes predicted by antiSMASH to be in a SMGC from more closely related donors (e.g. other filamentous fungi), we ran the AI pipeline for a second time using Xylariales (NCBI: txid 37989) as the ancestral lineage and manually curated subclades (see Table <ref type="table">S1</ref>) as recipient lineages (see Fig. <ref type="figure">S1</ref>). Genes from both the first (i.e. all genes, distant donors) and second (i.e. SMGC genes, closely related donors) were considered putative HGT candidates if they passed the following filters: (1) AI score of &gt; 0; (2) significant hits to at least 25 sequences in the custom database; and (3) at least 50% of top hits to sequences outside of the ancestral lineage.</p><p>Candidates from the first AI screen were further validated using phylogenetic analyses (described below) and designated as either high or low confidence HGT. Full-length proteins corresponding to the top &lt; 200 hits (E-value &lt; 1 9 10 -3 ) to each AI screen 1 candidate were extracted from the custom database using esl-sfetch <ref type="bibr">(Eddy, 2009)</ref>. As our initial query-based trees often lacked sufficient taxon sampling to assess HGT, we combined all orthogroup sequences with all extracted top hits to each AI candidate. Sequences were aligned using MAFFT v.7.407 using --auto <ref type="bibr">(Katoh &amp; Standley, 2013)</ref> and the number of well aligned columns was determined with TRIMAL v.1.4. rev15 using its gappyout strategy <ref type="bibr">(Capella-Guti errez et al., 2009)</ref>. Only alignments with &#8805; 50 retained columns after TRIMAL were retained for phylogenetic analysis. Phylogenetic trees were constructed with IQ-TREE v.1.6.10 <ref type="bibr">(Nguyen et al., 2015)</ref> in a single run with MODELFINDER <ref type="bibr">(Kalyaanamoorthy et al., 2017)</ref> and SH-ALRT combined with ultrafast bootstrapping analyses (1000 replicates each). Phylogenies were visualised using iTOL v.4 <ref type="bibr">(Letunic &amp; Bork, 2019)</ref>. Each phylogenetic tree was manually curated to verify HGT with either high or low confidence. High-confidence HGT events had to meet the following criteria: (1) the association between donor and recipient clades was supported by ultrafast bootstrap &#8805; 95; (2) recipient clade consisted of sequences from two or more species. If the candidate met one of the two criteria, HGT was considered lower confidence.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Statistical analyses</head><p>To assess whether genes within different functional categories are associated with endophytic ecological mode we performed phylogenetically independent contrasts (PICs) <ref type="bibr">(Felsenstein, 1985)</ref> with the function 'brunch' of the package CAPER v.1.0.1 <ref type="bibr">(Orme et al., 2012)</ref> in R v.3.6.1. All other statistical analyses were done in R v.3.6.1 or JMP v.15.1 (SAS Institute Inc., Cary, NC, USA).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results and Discussion</head><p>Genomes of 96 Xylariales taxa correspond to the previously recognised family Xylariaceae <ref type="bibr">(Ju &amp; Rogers, 1996)</ref> that was recently split into multiple families (Hypoxylaceae, Xylariaceae, Graphostromataceae, Barrmaeliaceae) <ref type="bibr">(Voglmayr et al., 2018;</ref><ref type="bibr">Wendt et al., 2018)</ref>  <ref type="figure">(Figs 1a</ref>, <ref type="figure">S2</ref>). Here, we use the term xylarialean to refer to this monophyletic clade within the Xylariales. In addition, because our analyses revealed seven undescribed endophytic isolates in five distinct clades (i.e. clades E2, E4, E5, E6 and E6; Fig. <ref type="figure">S2</ref>) nested between the Graphostromataceae and Xylariaceae sensu stricto (s.s), we refer to the sister clade to Hypoxylaceae as Xylariaceae s.l. (from this point forwards, Xylariaceae) following <ref type="bibr">Voglmayr et al. (2018)</ref> (Fig. <ref type="figure">1</ref>).</p><p>Genome sequencing yielded eukaryotic BUSCO values &#8805; 95% (Table <ref type="table">S1</ref>). Xylarialean genomes ranged in size from 33.7-60.3 Mbp (average 43.5 Mbp; Fig. <ref type="figure">S3</ref>; Table <ref type="table">S1</ref>) and contained c. 8000-15 000 predicted genes (mean 11 871; Fig. <ref type="figure">S3</ref>), congruent with average genome and proteome sizes of other Pezizomycotina New Phytologist (2022) 233: 1317-1330 <ref type="url">www.newphytologist.com</ref> &#211; 2021 The Authors New Phytologist &#211; 2021 New Phytologist Foundation. This article has been contributed to by US Government employees and their work is in the public domain in the USA. <ref type="bibr">(Shen et al., 2020)</ref>. The percentage of repetitive elements per genome ranged from &lt; 1-24% (average 1.6%; Table <ref type="table">S2</ref>), but unlike mycorrhizal fungi <ref type="bibr">(Miyauchi et al., 2020)</ref>, repeat content was not corrected with ecological mode (Fig. <ref type="figure">S3</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Xylariaceae and Hypoxylaceae genomes contain hyperdiverse metabolic gene clusters</head><p>To investigate the diversity and composition of metabolic gene clusters in xylarialean genomes, we used antiSMASH (Blin et al., 2019) to mine genomes for SMGCs, as well as a custom pipeline to examine catabolic gene clusters (CGCs) involved in fungal degradation of a broad array of plant phenylpropanoids (Gluck-Thaler et al., 2018). Across 96 xylarialean genomes we predicted a total of 6879 putative SMGCs (belonging to 3313 cluster families) and 973 putative CGCs (belonging to 190 cluster families) (Tables S3, S4). In comparison, recent large-scale analyses predicted 3399 SMGCs (in 719 cluster families) across 101 Dothideomycetes genomes (Gluck-Thaler et al., 2020) and 1110 CGCs across 341 fungal genomes (Gluck-Thaler &amp; Slot, 2018). Only 25% of predicted SMGCs (n = 1711 belonging to 816 cluster families) Eutypa lata UCREL1 Microdochium bolleyi J235TASD1 Microdochium trichocladiopsis MPI-CAGE-CH0230 Whalleya microplaca YMJ1829* Hypoxylaceae sp. FL0662B* Durotheca rogersii YMJ 92031201 Hypoxylon fuscum CBS 119018 Hypoxylon sp. FL1284 Hypoxylon sp. FL1150* Hypoxylon rubiginosum ER1909 Hypoxylon cercidicola CBS 119009* Hypoxylon rubiginosum CBS 119005 Hypoxylon fragiforme CBS 206.31 Hypoxylon crocopeplum CBS 119004 Hypoxylon sp. NC1633 Hypoxylon monticulosum FL0542 Hypoxylon submonticulosum NC0708 Daldinia eschscholtzii EC12 Daldinia eschscholtzii FL0578 Daldinia bambusicola CBS 122872 Daldinia caldariorum CBS 122874 Daldinia sp. FL1419 Daldinia vernicosa CBS 139.73 Daldinia grandis CBS 114736 Daldinia decipiens CBS 113046 Daldinia loculata AZ0526* Daldinia loculata CBS 113971* Hypoxylon trugodes CBS 135444* Hypoxylon sp. CI-4A* Hypoxylon sp. FL1857 Hypoxylon sp. FL0890 Hypoxylon sp. FL0543 Hypoxylon sp. NC0597 Hypoxylon sp. EC38 Hypoxylon sp. CO27-5 Annulohypoxylon minutellum CBS 135445 Annulohypoxylon bovei var. microspora CBS124037 Annulohypoxylon moriforme CBS 123579 Annulohypoxylon maeteangense CBS 123835 Annulohypoxylon truncatum CBS 140777 Rostrohypoxylon terebratum CBS 119137 Annulohypoxylon stygium FL0470* Annulohypoxylon nitens CBS 120705* Xylariaceae sp. FL2044 Biscogniauxia marginata CBS 124505 Biscogniauxia nummularia BnCUCC2015 Camillea tinctor CBS 203.56 Biscogniauxia sp. FL1348 Biscogniauxia mediterranea AZ0048* Biscogniauxia mediterranea CBS 129072* Xylariaceae sp. FL0641 Barrmaeliaceae sp. FL0016 Barrmaeliaceae sp. FL0804 Xylariaceae sp. FL0255 Xylariaceae sp. FL1272 Xylariaceae sp. FL1019 Xylariaceae sp. FL1651 Anthostoma avocetta NRRL 3190* Xylariaceae sp. AK1471* Xylariaceae sp. FL0594* Poronia punctata CBS 180.79* Xylaria nigripes YMJ 653 Xylaria sp. CBS 124048 Xylaria intraflava YMJ 725 Xylaria hypoxylon OSC100004 Xylaria digitata CBS 161.22 Xylaria grammica CBS 120713 Xylaria sp. FL1777 Kretzschmaria deusta IL1129* Kretzschmaria deusta CBS 826.72* Xylaria arbuscula FL1030* Xylaria bambusicola CBS 139988* Xylaria arbuscula CBS 124340 Xylaria venustula FL0490 Xylaria sp. FL1042 Xylaria sp. FL0064 Xylaria sp. FL0043 Xylaria sp. FL0933 Entoleuca mammata CFL468 Nemania sp. CBS 527.63* Nemania diffusa NC0034* Nemania abortiva FL1152 Nemania sp. FL0031 Nemania sp. FL0916 Nemania sp. NC0429 Nemania serpens AK0226 Nemania serpens AZ0576* Nemania serpens CBS 679.86* Xylaria palmicola CBS 124036 Astrocystis sublimbata CBS 130006 Xylaria telfairii CBS 121673 Xylaria cf. heliscus FL0509 Xylaria acuta CBS 122032 Xylaria longipes CBS 148.73 Xylaria cf. castorea CBS 124033 Xylaria flabelliformis CBS 116.85* Xylaria flabelliformis NC1011* Xylaria flabelliformis CBS 123580 Xylaria flabelliformis CBS 114988 Xylariaceae s.l. Hypoxylaceae No. of predicted SMGCs PKS NRPS Terpene PKS-NRP hyb rid Other PKS other RiPP 0 20 40 60 80 100 120 SMGC classification (b) Relative abundance (%) 0 20 40 60 80 100 Isolate Xylariaceae/hypoxylaceae Other Xylariaceae s.l. SMGC distribution Hypoxylaceae Specific to: (c) Salicylate hydroxylase Pterocarpan hydroxylase Naringenin 3-dioxygenase Phenol 2-monooxygenase Quinate dehydrogenase Benzoate 4-monooxygenase Epicatechin laccase Catechol dioxygenase Aromatic ring-opening dioxygenase Ferulic acid decarboxylase Ferulic acid esterase Stilbene dioxygenase Vanillyl alcohol oxidase Hybrid (families) Relative abundance (%) 0 50 100 25 75 (d) (e) Catabolic gene cluster (CGC) classification Newly sequenced genome Ecological mode Saprotrophic (dead wood, litter, fruit) Endophytic (living leaves, lichens) Saprotrophic (insect/animal dung) Pathogenic &#211; 2021 The Authors New Phytologist &#211; 2021 New Phytologist Foundation. This article has been contributed to by US Government employees and their work is in the public domain in the USA. New Phytologist (2022) 233: 1317-1330 <ref type="url">www.newphytologist.com</ref> </p><p>had BLAST hits to 168 unique MIBiG <ref type="bibr">(Medema et al., 2015)</ref> accession numbers (Table <ref type="table">S3</ref>). Total SMGCs diversity in the Xylariaceae and Hypoxylaceae is reflected in a high number of SMGCs per genome: the average number of SMGCs per genome was 71.2 (median 68), which is significantly higher than the average for other fungi in the Pezizomycotina (mean 42.8; Fig. <ref type="figure">1b</ref>). At least eight xylarialean genomes contained more than 100 predicted SMGCs, with a maximum of 119 in Anthostoma avocetta NRRL 3190 (Fig. <ref type="figure">1b</ref>; Table <ref type="table">S3</ref>). In comparison, a recent study of 24 species of Penicillium found an average of 54.9 SMGCs per genome, with a maximum number of 78 SMGCs observed in P. polonicum <ref type="bibr">(Nielsen et al., 2017)</ref>. Genomes of Xylariaceae and Hypoxylaceae contained on average 3.39 more CGCs per genome (average 10.1; Table <ref type="table">S4</ref>) compared with other genomes of Pezizomycotina (average 3.0 (Gluck-Thaler et al., 2018)).</p><p>Every xylarialean genome contained SMGCs for the production of polyketides (PK; 2871 total), nonribosomal peptides (NRP; 2482 total) and terpenes (1322 total; Fig. <ref type="figure">1b</ref>; Table <ref type="table">S3</ref>). SMGCs for ribosomally synthesised and post-translationally modified peptides (RiPPs) and hybrid NRP-PK compounds occurred less frequently (Fig. <ref type="figure">1b</ref>). The most widely distributed and abundant CGCs were pterocarpan hydroxylases (n = 93), putatively involved in isoflavonoid metabolism (Fig. <ref type="figure">1d</ref>,<ref type="figure">e</ref>; Table <ref type="table">S5</ref>). CGCs involved in the breakdown of plant salicylic acid <ref type="bibr">(Ambrose et al., 2015)</ref> (n = 251 salicylate hydroxylases) and plant flavonoids (n = 170 naringenin 3-dioxygenases) were also abundant (Fig. <ref type="figure">1d</ref>,<ref type="figure">e</ref>). CGCs classified into nine other categories (e.g. phenol 2-monooxygenase, quinate dehydrogenase) (Gluck-Thaler et al., 2018) occurred more rarely (Table <ref type="table">S4</ref>). Vanillyl alcohol oxidases, which were previously shown to be enriched in genomes of soil saprotrophs (Gluck-Thaler et al., 2018), were absent in xylarialean genomes.</p><p>Consistent with the hyperdiversity of SMGCs in the Hypoxylaceae and Xylariaceae, we observed that only c. 10% of SMGCs were shared among genomes from both Xylariaceae and Hypoxylaceae (Fig. <ref type="figure">1c</ref>), and no SMGCs were universally present in both clades (Table <ref type="table">S3</ref>). On average, 21.4% and 28.2% of SMGCs per genome were unique to either a taxon in the Hypoxylaceae or the Xylariaceae, respectively (range 0-82%; Fig. <ref type="figure">1c</ref>; Table <ref type="table">S4</ref>), but no SMGCs were universally present within either clade. For most isolates, the majority of SMGCs were unique (i.e. 'isolate specific'; Fig. <ref type="figure">1c</ref>). Isolate-specific SMGCs represented an average of 36.6% (SD AE 21.1) of the clusters per genome (range 0-85.7%; Fig. <ref type="figure">1c</ref>). Even when multiple isolates of the same species were compared (e.g. Nemania serpens clade) 30-41% of the SMGCs appeared specific to a single isolate (Fig. <ref type="figure">1b</ref>; Table <ref type="table">S3</ref>), similar to intraspecific SMGC variation in Aspergillus flavus <ref type="bibr">(Drott et al., 2021)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Impact of HGT on xylarialean genome evolution</head><p>To assess the role of HGT in shaping the genome evolution of Xylariaceae and Hypoxylaceae we performed two AI analyses <ref type="bibr">(Alexander et al., 2016;</ref><ref type="bibr">Wisecaver et al., 2016;</ref><ref type="bibr">Gonc &#223;alves et al., 2018)</ref>. The first AI screendesigned to detect candidate HGTs from more distantly related donor lineages (e.g. bacteria, plants) flagged 4262 genes representing 647 orthogroups (Table <ref type="table">S5</ref>). Using a custom phylogenetic pipeline (see the Materials and Methods section) we manually validated 168 of these genes as likely to be HGT events to Xylariaceae and Hypoxylaceae. Based on branch support and the presence of multiple xylarialean taxa in the recipient clade, we deemed 92 of these genes as highconfidence HGTs and the remaining 76 as lower confidence HGTs (Fig. <ref type="figure">2</ref>; Table <ref type="table">S5</ref>). Similar to previous studies (Marcet-Houben &amp; Gabald on, 2010; <ref type="bibr">Lawrence et al., 2011)</ref>, the majority of high-confidence HGTs were predicted to have been acquired from bacteria (n = 86) (Fig. <ref type="figure">2</ref>). Overall, 66% of genes identified as HGT from bacterial donors did not contain introns (compared with 6% of genes across 121 genomes). Other donor lineages include viruses (n = 3), Basidiomycota (n = 2) and plants (n = 1) (Fig. <ref type="figure">2</ref>; Table <ref type="table">S5</ref>). On average, xylarialean genomes had 16.2 high-confidence HGT events per genome (range: 7-30; Table <ref type="table">S5</ref>). The highest number of high-confidence HGT events per genome occurred in the genome of Xylaria flabelliformis CBS 123 580 (n = 30).</p><p>Horizontal gene transfer candidate genes were typically distributed across taxa in numerous diverse clades (n = 85 of 92 genes) rather than in monophyletic clades (Fig. <ref type="figure">2</ref>). For example, an enoyl-acyl carrier protein reductase protein (EC 1.3.1.9)a key enzyme of the type II fatty acid synthesis (FAS) system (Massengo-Tiass e &amp; Cronan, 2009)occurred in bacteria (putative donor) and four distantly related recipient taxa: Xylariales sp. PMI 506, Hypoxylon rubiginosum ER1909; H. cercidicola CBS 119 009; H. fuscum CBS 119 018 (HGT0001; Table <ref type="table">S5</ref>). Multiple evolutionary scenarios could result in patchy taxonomic distributions. For example, multiple fungi could have independently acquired the same gene from closely related bacterial donors (Marcet-Houben &amp; Gabald on, 2010). Alternatively, an initial HGT from bacteria to fungi may have been followed by fungal-fungal HGTs. In total, 38 HGT candidate genes occurred in genomes of both Sordariomycetes outgroup and Xylariales genomes, 28 were found in only Xylariales genomes, and 26 were only observed in genomes of Xylariaceae and Hypoxylaceae (Fig. <ref type="figure">2</ref>; Table <ref type="table">S5</ref>).</p><p>Functional annotation revealed that most candidate HGT genes were associated with at least one type of annotation (i.e. 95% of the highly confident and 82% of the ambiguous events; Table <ref type="table">S5</ref>). Six high-confidence HGT candidate genes were annotated as CAZymes, including three predicted plant cell walldegrading enzymes (PCWDEs) transferred from bacteria to diverse Xylariales (Fig. <ref type="figure">2</ref>). No genes predicted in CGCs were identified as candidate HGTs, consistent with convergent evolution to result in similar clustering of fungal phenolic metabolism genes (Gluck-Thaler et al., 2018). However, 43% of candidate HGT genes were predicted to be part of an SMGC (i.e. 40 of 92) (Fig. <ref type="figure">2</ref>; Tables <ref type="table">S3</ref>, <ref type="table">S5</ref>). These include 13 genes predicted to have a biosynthetic function, such as a putative FsC-acetyl coenzyme A-N 2 -transacetylase (HGT076; Table <ref type="table">S5</ref>), which is part of the siderophore biosynthetic pathway in Aspergillus implicated in fungal virulence <ref type="bibr">(Blatzer et al., 2011)</ref>.</p><p>New Phytologist (2022) 233: 1317-1330 <ref type="url">www.newphytologist.com</ref> &#211; 2021 The Authors New Phytologist &#211; 2021 New Phytologist Foundation. This article has been contributed to by US Government employees and their work is in the public domain in the USA.</p><p>Due to the high prevalence of HGT among genes predicted to be part of SMGCs, we performed a second AI screen to detect intrafungal HGT events of genes within the boundaries of SMGCs (n = 93 066 genes) (see the Materials and Methods section; Fig. <ref type="figure">S1</ref>). The second AI screen identified 1148 genes in 660 SMGCs (belonging to 594 cluster families) that were putatively transferred from other fungi to members of the Xylariales (Table <ref type="table">S5</ref>). Candidate HGT genes were primarily for polyketide HGT129</p><note type="other">HGT128 HGT005 HGT016 HGT018 HGT019 HGT027 HGT033 HGT015 HGT026 HGT035 HGT037 HGT014 HGT031 HGT030 HGT036 HGT039 HGT038 HGT040 HGT044 HGT045 HGT046 HGT047 HGT043 HGT051 HGT050 HGT055 HGT056 HGT053 HGT008 HGT057 HGT058 HGT061 HGT064 HGT013 HGT063 HGT066 HGT067 HGT071 HGT070 HGT078 HGT080 HGT003 HGT010 HGT012 HGT054 HGT069 HGT072 HGT073 HGT096 HGT009 HGT062 HGT079 HGT082 HGT089 HGT091 HGT093 HGT001 HGT094 HGT106 HGT007 HGT048 HGT049 HGT114 HGT023 HGT032 HGT074 HGT075 HGT084 HGT088 HGT095 HGT098 HGT100 HGT103 HGT104 HGT107 HGT112 HGT113 HGT110 HGT002 HGT017 HGT092 HGT011 HGT024 HGT025 HGT029 HGT041 HGT042 HGT052 HGT059 HGT081 HGT083 HGT087 HGT105 HGT108 HGT115 HGT116 HGT117 HGT121 HGT122 HGT123 HGT125</note><p>CAZymes Peptidases Transporters Signal Peptides Effectors SMGCs and nonribosomal peptide production (518 PK, 270 NRP and 180 PK-NRP hybrid clusters). In addition, &gt; 75% of hits to MIBiG contained genes identified by AI analyses as putative HGTs (see Fig. <ref type="figure">S4</ref>, bottom). SMGCs with HGT candidate genes included those with 100% similarity to MIBiG accessions from Aspergillus, Fusarium and Parastagonospora involved in mycotoxin (e.g. cyclopiazonic acid, alternariol, fusarin) and antimicrobial compound (asperlactone, koraiol) production, and clusters from Alternaria that produce host-selective toxins (e.g. ACT-Toxin II) (Tables <ref type="table">S3</ref>, <ref type="table">S5</ref>). Although the second AI analysis did not flag every gene in these clusters as potential HGTs (e.g. only four of the 19 genes in the alternariol cluster from Hypoxylon cercidicola CBS 119 009 were HGT candidates based on AI; Table <ref type="table">S5</ref>) and we were not able to further validate candidates based on the same criteria used for high-confidence HGT, the phylogenetic distribution of many of these SMGCs across Xylariales is consistent with the acquisition of SMGCs via HGT (Fig. <ref type="figure">S4</ref>).</p><p>In addition to the AI screen for HGT candidates, we identified additional putative HGTs of SMGCs to Xylariaceae and Hypoxylaceae based on their (1) high similarity to fungal MIBiG accessions from distantly related fungi; and (2) discontinuous phylogenetic distributions (Fig. <ref type="figure">S4</ref>). Putative HGT of SMGCs included xylarialean SMGCs with &gt; 70% similarity to clusters for ergoline alkaloids and their precursors (e.g. loline, ergovaline and lysergic acid production) produced by Clavicipitaceae endophytes, as well as the phytotoxin cichorine cluster from Aspergillus (Fig. <ref type="figure">S4</ref>; Table <ref type="table">S3</ref>). The griseofulvin cluster from Penicillium aethiopicum, which produces a potent antifungal compound <ref type="bibr">(Chooi et al., 2010)</ref>, also appears horizontally transferred to the clade containing X. castorea and X. flabelliformis isolates (Figs S4, S5). Although the discontinuous phylogenetic distributions of SMGCs observed here may represent unequal gene loss across taxa <ref type="bibr">(Slot, 2017;</ref><ref type="bibr">Rokas et al., 2018)</ref>, the presence of entire clusters known from Eurotiomycetes and Sordariomycetes in multiple endophytic and nonendophytic taxa provides additional support for HGTs. Overall, our first AI analysis provides the highest support for HGTs primarily from distantly related hosts such as bacteria (Fig. <ref type="figure">2</ref>) (see also <ref type="bibr">Marcet-Houben &amp; Gabald on, 2010</ref>), yet our second AI screen and comparisons of SMGCs to MIBiG within our phylogenomic framework also support fungal-fungal HGT as an important mechanism of metabolic innovation in the Xylariales, similar to pathogenic fungi <ref type="bibr">(Qiu et al., 2016)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Expansion of Xylariaceae genomes due to increased gene duplication and HGTs</head><p>Despite the close evolutionary relationship and similar ecological niches of taxa in the Xylariaceae and Hypoxylaceae, genomes of Xylariaceae were on average c. 7.2 Mbp larger than genomes of Hypoxylaceae (Fig. <ref type="figure">3a</ref>; Table <ref type="table">S6</ref>). Larger genome size was associated with higher repeat content: Xylariaceae genomes contained an average of two-fold more repetitive elements (Fig. <ref type="figure">3b</ref>; Table <ref type="table">S6</ref>) and had a higher density of repetitive elements surrounding genes (including effectors and genes identified as HGT candidates) compared with Hypoxylaceae genomes (Fig. <ref type="figure">S6</ref>).</p><p>In addition to greater repeat content, Xylariaceae genomes also contained on average 750 more protein-coding genes compared with Hypoxylaceae (P &lt; 0.0001; Table <ref type="table">S6</ref>). Ancestral state reconstructions reveal that Xylariaceae genomes have experienced significantly more gene gains (n = 472), gene duplication events (n = 136), orthogroup gains (n = 313) and orthogroup expansion events (n = 90) compared with the Hypoxylaceae clade since the radiation from their last common ancestor (Fig. <ref type="figure">3c</ref>,<ref type="figure">d</ref>), although both clades underwent similar numbers of gene losses (t 95 = 0.51, P = 0.61; Table <ref type="table">S6</ref>). Xylariaceae genomes also experienced on average c. two-fold more HGT events compared with Hypoxylaceae genomes (Fig. <ref type="figure">3e</ref>).</p><p>Horizontal gene transfer events were positively associated with increased numbers of SMGCs across both clades (Fig. <ref type="figure">3f</ref>), reflecting the fact that clustered metabolite genes in fungi are more likely to undergo HGT compared with unclustered genes <ref type="bibr">(Wisecaver et al., 2014)</ref>. Genomes of Xylariaceae contained on average c. 20 more SMGCs than Hypoxylaceae genomes (Table <ref type="table">S6</ref>) and c. two-fold greater cumulative richness of SMGCs compared with the Hypoxylaceae clade (2336 vs 1075 total; 587 vs 282 nonsingleton). Rarefaction analysis revealed that the richness of SMGCs per genome sampled also increased at a greater rate in the Xylariaceae clade (Fig. <ref type="figure">S7</ref>). Genomes of Xylariaceae also contained a greater fraction of isolate-specific SMGCs compared with Hypoxylaceae, regardless of SMGC type (Xylariaceae: 31.2 AE 16.1; Hypoxylaceae: 19.8 AE 15.3; P = 0.0007; Figs <ref type="figure">1c</ref>, <ref type="figure">S8</ref>). Yet despite the high variation of SMGCs among taxa, network analysis illustrated that the composition of SMGCs was more similar among isolates from the same clade, regardless of ecological mode (Fig. <ref type="figure">S9</ref>).</p><p>In contrast with the pattern observed for SMGCs, genomes of Hypoxylaceae contained a greater number of CGCs than Xylariaceae genomes (Xylariaceae: 9.5 AE 0.4; Hypoxylaceae: 11.0 AE 0.4; P = 0.0068; Table <ref type="table">S4</ref>) and different classes of CGCs dominated the two clades (Fig. <ref type="figure">1d</ref>,<ref type="figure">e</ref>). For example, salicylate hydroxylases were the most abundant CGCs among Hypoxylaceae, but were absent from 25% of Xylariaceae genomes (Fig. <ref type="figure">1d</ref>). Four types of CGCs were universally present across Hypoxylaceae: salicylate hydroxylase, pterocarpan hydroxylase, naringenin 3-dioxygenase, phenol 2-monooxygenase (Fig. <ref type="figure">1d</ref>). CGCs classified as pterocarpan hydroxylases were the most abundant CGC type in genomes of Xylariaceae (Fig. <ref type="figure">1d</ref>), but were not found in all Xylariaceae genomes. Only CGCs classified as naringenin 3-dioxygenases were found across all Xylariaceae genomes.</p><p>In addition to distinct metabolic gene cluster content and prevalence of HGT between clades, comparison of GO terms for shared orthogroups significantly enriched in either Xylariaceae or Hypoxylaceae (i.e. 74 and 26, respectively) revealed that Hypoxylaceae genomes had a significant increase in the number of GO terms associated with membrane transport, whereas Xylariaceae genomes had a significant increase in the number of GO terms for catalytic activities and binding (Fig. <ref type="figure">S10</ref>). Xylariaceae genomes also contained greater numbers of genes with signalling peptides, as well as genes annotated as effectors, membrane transport proteins, transcription factors, peptidases and CAZymes compared with Hypoxylaceae, even after accounting for differences in genome size (Table <ref type="table">S6</ref>). On average, genomes of Xylariaceae contained c. 50 more CAZymes than Hypoxylaceae (Xylariaceae 579.9 AE 7.7; Hypoxylaceae 529.6 AE 9.1, P &lt; 0.0001), including a significant increase in PCWDEs involved in the degradation of cellulose, hemicellulose, lignin, pectin and starch (Table <ref type="table">S6</ref>).</p><p>As genomes of fungi with saprotrophic lifestyles typically encode more CAZymes and PCWDEs compared with plant pathogens and mycorrhizal symbionts <ref type="bibr">(Knapp et al., 2018;</ref><ref type="bibr">Haridas et al., 2020;</ref><ref type="bibr">Miyauchi et al., 2020)</ref>, our genomic results are consistent with the potential for Xylariaceae fungi (including endophytes) to have greater saprotrophic abilities compared with Hypoxylaceae fungi <ref type="bibr">(Osono, 2006)</ref>. To test this prediction, we compared the abilities of 20 isolates to degrade leaves of Pinus and Quercus. Regardless of trophic mode, isolates of Xylariaceae with expanded CAZyme and PCWDE repertoires caused greater mass loss compared with taxa with fewer genes predicted to degrade lignocellulose (i.e. Hypoxylaceae and Xylariaceae from animal-dung clade; Fig. <ref type="figure">S11</ref>). In addition to increased capacity for lignocellulose degradation, Xylariaceae endophyte species associated with a greater phylogenetic diversity of plant and lichen hosts compared with species of Hypoxylaceae endophytes (t 42 = 2.25; P = 0.0294; Fig. <ref type="figure">3g</ref>). Host breadth of Xylariaceae endophytes also was positively associated with the number of total HGT events (r = 0.43, P = 0.0193), as well as the number of peptidases (r = 0.37, P = 0.0444) and NRP SMGCs (Fig. <ref type="figure">3h</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Genomic differences between endophytic and nonendophytic fungi</head><p>Both culture-based and culture-free studies of healthy photosynthetic tissues of plants and lichens demonstrate the abundance and novel diversity represented by xylarialean endophytes  <ref type="table">S6</ref>. Gene gains/losses were inferred with Wagner parsimony under a gain penalty = loss penalty = 1. (f) Relationship between the number of HGT events and secondary metabolite gene clusters (SMGCs) as a function of clade. (g) A quantile box plot showing the interquartile range and median of endophyte host breadth [measured as total number of plant families and lichen orders with which a fungal operational taxonomic unit (OTU) was cultured (U' <ref type="bibr">Ren et al., 2016)</ref>] as a function of major clade (colour). A similar pattern was observed when only the number of plant families are compared (Wilcoxon: v 2 = 4.14, P = 0.0413), but not lichen orders (Wilcoxon: v 2 = 1.77, P = 0.1834). (h) Relationship of Xylariaceae endophyte host breadth and the number of SMGCs classified as nonribosomal peptides (NRPs) per genome. For panels f and h, the shaded region indicates the 95% confidence interval of the linear fit and statistics represent Pearson's correlation coefficient (r) and P-value.</p><p>&#211; 2021 The Authors New Phytologist &#211; 2021 New Phytologist Foundation. This article has been contributed to by US Government employees and their work is in the public domain in the USA. <ref type="table">(2022) 233: 1317-1330</ref> <ref type="url">www.newphytologist.com</ref>  <ref type="bibr">(U'Ren et al., 2016)</ref>. However, some endophytes can occur in both living host tissues as well as decomposing plant materials <ref type="bibr">(Okane et al., 2008;</ref><ref type="bibr">U'Ren &amp; Arnold, 2016;</ref><ref type="bibr">U'Ren et al., 2016)</ref> and are often closely related to described species of saprotrophs and pathogens (U' <ref type="bibr">Ren et al., 2016)</ref>. This suggests that, for some species, endophytism may represent only part of a complex life cycle that blurs the lines between distinct ecological modes (U' <ref type="bibr">Ren et al., 2016;</ref><ref type="bibr">Chen et al., 2018)</ref> and few genomic signatures may be associated with the evolution of endophytism in the Xylariaceae and Hypoxylaceae.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>New Phytologist</head><p>Overall, when we analysed all ingroup genomes we observed no clear distinctions in genome size or content due to different ecological modes, even after taking phylogeny into account (Table <ref type="table">S6</ref>). One exception was the reduced genomes and CAZyme content of termite-associated Xylaria spp. (i.e. X. nigripes YMJ 653, X. sp. CBS 124 048 and X. intraflava YMJ725; Figs <ref type="figure">S3</ref>, <ref type="figure">S12</ref>) that reflects a single evolutionary transition to specialisation on termite nest substrates decomposed by a basidiomycete fungus <ref type="bibr">(Hsieh et al., 2010)</ref>. However, as evolutionary distances among taxa can impede detection of finer-scale genomic  <ref type="figure">1a</ref>). Values greater than zero indicate higher gene counts in nonendophytic taxa, whereas differences less than zero indicate higher gene counts in endophytes. Statistical differences were assessed with least squares means contrast under the null hypothesis: nonendophyte valueendophyte value = 0 (see Supporting Information Table <ref type="table">S6</ref> for summary statistics). *, P &lt; 0.05. differences due to ecological mode (e.g. <ref type="bibr">Harrington et al., 2019)</ref>, we restricted our analyses to comparisons of 15 pairs of sister taxa across both clades with contrasting ecological modes. These pairwise comparisons revealed that endophytic Hypoxylaceae genomes contained significantly fewer genes with signalling peptides, protein-coding genes, transporters, peptidases, PCWDEs (especially those involved in decomposition of cellulose and lignin), SMGCs and CGCs compared with nonendophytes (Fig. <ref type="figure">4</ref>). Yet, similar to the lack of reduced genome repertoires in some root endophytes <ref type="bibr">(Lahrmann et al., 2015;</ref><ref type="bibr">Xu et al., 2015)</ref>, no significant differences in genomic content were observed between paired endophytes and nonendophytes in the Xylariaceae clade (Fig. <ref type="figure">4</ref>; Table <ref type="table">S6</ref>). These results suggest that, compared with endophytes and saprotrophs in the Hypoxylaceae, Xylariaceae taxa have less distinct ecological modes and their increased metabolic versatility may be the result of selection maintaining diverse genes for both endophytism and saprotrophy. As saprotrophs, fungi experience strong selection to maintain highly diverse SMGCs that increase competitive abilities in diverse microbial communities <ref type="bibr">(Richards &amp; Talbot, 2013;</ref><ref type="bibr">Rokas et al., 2018;</ref><ref type="bibr">Naranjo-Ortiz &amp; Gabald on, 2020)</ref>, as well as large gene repertoires to degrade lignocellulosic compounds <ref type="bibr">(Haridas et al., 2020)</ref>. Accordingly, we observed that in genomes of nonendophytic Xylariaceae and Hypoxylaceae, SMGC abundance was positively correlated with the number of genes important for saprotrophy (e.g. CAZymes, transporters) and putative pathogenicity (e.g. signalling peptides, effectors, peptidases), even after accounting for differences among clades and genome sizes (Fig. <ref type="figure">5</ref>; Table <ref type="table">S6</ref>). By contrast, we found that endophyte SMGC abundance was decoupled from the majority of other genomic factors (Fig. <ref type="figure">5</ref>), due in part to fewer numbers of CAZymes, transporters and peptidases annotated in SMGCs (Table <ref type="table">S6</ref>). These results are consistent with different selection pressures and ecological roles of SMGCs in endophytic and nonendophytic fungi and highlight the importance of phylogenetically informed comparisons to detect genomic differences associated with endophytism, as well as the complexity of linking genotype to phenotype for complex traits, especially in dynamic genomes undergoing frequent HGT.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>SMGCs (residuals)</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Conclusions</head><p>Our analysis of 96 phylogenetically and ecologically diverse Xylariaceae and Hypoxylaceae genomes reveals that gene duplication, gene family expansion and HGT of SMGCs, effectors and peptidases from putative bacterial and fungal donors drives metabolic versatility in the Xylariaceae. Expanded metabolic diversity and secondary metabolism of Xylariaceae taxa is associated with greater ecological generalism in both substrate usage and the phylogenetic breadth of symbiotic associations compared with Hypoxylaceae taxa. Correlations between endophyte host breadth, HGT and abundance of NRPs also indicate that SMGCs may play a key role in facilitating xylarialean endophyte colonisation of diverse hosts. For example, although NRPs are known for their role as virulence factors of phytopathogenic fungi (e.g. host-selective toxins or siderophores) <ref type="bibr">(Oide &amp; Turgeon, 2020)</ref>, previous research has shown that an NRP is essential for the endophyte Neotyphodium/Epichlo&#8364; e to establish symbiosis with its host <ref type="bibr">(Johnson et al., 2007)</ref>. Overall, our results highlight the importance of plant-fungal symbioses to drive not only fungal speciation and ecological diversification <ref type="bibr">(Joy, 2013)</ref>, but vast chemical biodiversity that can be leveraged for novel pharmaceuticals and agrochemicals <ref type="bibr">(Becker &amp; Stadler, 2021;</ref><ref type="bibr">Robey et al., 2021)</ref>.     Fig. <ref type="figure">S6</ref> The density of repetitive elements surrounding genes was higher for Xylariaceae s.l. than for Hypoxylaceae genomes.      Methods S1 Additional information on strain selection, fungal growth and nucleic acid extraction, genome and transcriptome sequencing, gene prediction and genome assembly, phylogenetic analyses, comparative genomic analyses, litter decomposition assays and metabolomics.</p><p>Notes S1 List and description of appendices S1-S10 available on FigShare Repository, 10.6084/m9.figshare.c.5314025.</p><p>Table <ref type="table">S1</ref> Sequencing and assembly statistics for the 121 genomes included in this study.</p><p>Table S2 REPEATMASKER, REPEATSCOUT and RepBase Update classification of repetitive elements for 96 genomes of Xylariaceae s.l. and Hypoxylaceae. Table S3 Secondary metabolite gene clusters and cluster families for the 96 Hypoxylaceae and Xylariaceae s.l. genomes included in this study. Table S4 Catabolic gene clusters and cluster families for the 96 Hypoxylaceae and Xylariaceae s.l. genomes included in this study. Table S5 Taxonomic, phylogenetic, and functional annotation information for HGT candidate genes identified by ALIEN INDEX analyses. Table S6 Counts and statistical comparisons of genome content as a function of major clade (Xylariaceae s.l. vs Hypoxylaceae) and ecological mode (endophyte vs nonendophyte). Please note: Wiley Blackwell are not responsible for the content or functionality of any Supporting Information supplied by the authors. Any queries (other than missing material) should be directed to the New Phytologist Central Office. New Phytologist (2022) 233: 1317-1330 <ref type="url">www.newphytologist.com</ref> &#211; 2021 The Authors New Phytologist &#211; 2021 New Phytologist Foundation. This article has been contributed to by US Government employees and their work is in the public domain in the USA.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>14698137, 2022, 3, Downloaded from https://nph.onlinelibrary.wiley.com/doi/10.1111/nph.17873 by Ohio State University Ohio, Wiley Online Library on [14/10/2022]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License</p></note>
		</body>
		</text>
</TEI>
