<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Double-stranded DNA virioplankton dynamics and reproductive strategies in the oligotrophic open ocean water column</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>05/01/2020</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10250356</idno>
					<idno type="doi">10.1038/s41396-020-0604-8</idno>
					<title level='j'>The ISME Journal</title>
<idno>1751-7362</idno>
<biblScope unit="volume">14</biblScope>
<biblScope unit="issue">5</biblScope>					

					<author>Elaine Luo</author><author>John M. Eppley</author><author>Anna E. Romano</author><author>Daniel R. Mende</author><author>Edward F. DeLong</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Abstract            Microbial communities are critical to ecosystem dynamics and biogeochemical cycling in the open oceans. Viruses are essential elements of these communities, influencing the productivity, diversity, and evolution of cellular hosts. To further explore the natural history and ecology of open-ocean viruses, we surveyed the spatiotemporal dynamics of double-stranded DNA (dsDNA) viruses in both virioplankton and bacterioplankton size fractions in the North Pacific Subtropical Gyre, oneof the largest biomes on the planet. Assembly and clustering of viral genomes revealed a peak in virioplankton diversity at the base of the euphotic zone, where virus populations and host species richness both reached their maxima. Simultaneous characterization of both extracellular and intracellular viruses suggested depth-specific reproductive strategies. In particular, analyses indicated elevated lytic interactions in the mixed layer, more temporally variable temperate phage interactions at the base of the euphotic zone, and increased lysogeny in the mesopelagic ocean. Furthermore, the depth variability of auxiliary metabolic genes suggested habitat-specific strategies for viral influence on light-energy, nitrogen, and phosphorus acquisition during host infection. Most virus populations were temporally persistent over several years in this environment at the 95% nucleic acid identity level. In total, our analyses revealed variable distributional patterns and diverse reproductive and metabolic strategies of virus populations in the open-ocean water column.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Viruses represent dynamic reservoirs of unexplored genetic diversity. On average an order of magnitude more abundant than cellular organisms <ref type="bibr">[1]</ref>, viruses occur at ~10 7 /mL in the surface layer of open oceans covering ~40% of Earth <ref type="bibr">[2]</ref>. In this environment, dsDNA bacteriophages infect key microbial groups, including oxygenic photoautotrophs such as Prochlorococcus and Synechococcus (e.g., <ref type="bibr">[3,</ref><ref type="bibr">4]</ref>) and common bacterial heterotrophs such as Pelagibacter (SAR11), Puniceispirillum (SAR116), Roseobacter, and Alteromonas (e.g., <ref type="bibr">[5]</ref><ref type="bibr">[6]</ref><ref type="bibr">[7]</ref><ref type="bibr">[8]</ref>). Viruses can lyse their hosts at estimated rates of 20-40% per day <ref type="bibr">[9,</ref><ref type="bibr">10]</ref>, potentially contributing as much as 145 gigatonnes to annual global carbon flux <ref type="bibr">[11,</ref><ref type="bibr">12]</ref>. Viruses also influence the diversity and biogeochemistry of marine ecosystems by carrying auxiliary metabolic genes (AMGs) that manipulate host metabolism during infection (reviewed in <ref type="bibr">[13]</ref>).</p><p>Recent developments in high-throughput DNA sequencing have allowed for exploration of viral diversity at unprecedented scales. The continuing description of viral genetic diversity highlights the importance of further cultivation-independent in situ characterization of environmental viral populations. Metagenomic virus surveys have focused on characterizing geographic variability across surface oceans <ref type="bibr">[14]</ref><ref type="bibr">[15]</ref><ref type="bibr">[16]</ref><ref type="bibr">[17]</ref>, temporal variability at the surface ocean <ref type="bibr">[18]</ref><ref type="bibr">[19]</ref><ref type="bibr">[20]</ref><ref type="bibr">[21]</ref>, and vertical variability across depth profiles <ref type="bibr">[22]</ref><ref type="bibr">[23]</ref><ref type="bibr">[24]</ref>. Coupled studies of viral and host dynamics in both space and time at well-defined sample sites <ref type="bibr">[25]</ref><ref type="bibr">[26]</ref><ref type="bibr">[27]</ref> further have the potential to provide additional perspective on the dynamics, patterns and consequences of viral diversity.</p><p>To further explore viral diversity, environmental distributions, and dynamics in the open ocean, we characterized virus genomes from seawater in virus-enriched (0.02-0.2 &#956;m) and cell-enriched (&gt;0.2 &#956;m) size fractions over time and depth in the North Pacific Subtropical Gyre (NPSG). The approach facilitated the exploration of both extracellular virioplankton particles as well as cellassociated phages, to better characterize reproductive and metabolic strategies of viruses in the open ocean. We use the term "reproductive strategies" here to refer life history differences between strictly lytic viruses with a singular strategy of host lysis, versus temperate phages that in addition to the lytic cycle, have the potential to either integrate into their host's genome or reside as an extrachromosomal element. In this study, we analyzed samples collected over 1.5 years over 12 depths (5-500 m) to characterize dsDNA viruses in both virioplankton and bacterioplankton size fractions. This virus genome dataset (referred to here as the ALOHA 2.0 viral database) provides new perspectives on the diversity, reproductive strategies, gene content, and ecology of viruses in the open ocean.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Methods</head><p>A schematic overview of our workflow is presented in Fig. <ref type="figure">S1</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Sample collection, extraction, and sequencing</head><p>Station ALOHA (22&#176;45&#8242; N, 158&#176;W), a relatively seasonally stable environment located in the NPSG, is a wellcharacterized sampling site of the Hawaii Ocean Timeseries (HOT) program. The ALOHA 2.0 dataset contains 374 metagenomic samples collected at 12 depths <ref type="bibr">(5,</ref><ref type="bibr">24,</ref><ref type="bibr">45,</ref><ref type="bibr">75,</ref><ref type="bibr">100,</ref><ref type="bibr">125,</ref><ref type="bibr">150,</ref><ref type="bibr">175,</ref><ref type="bibr">200,</ref><ref type="bibr">225,</ref><ref type="bibr">250</ref>, and 500 m) at approximately monthly intervals of 16 timepoints spanning 1.5 years from 2014 to 2016 (Fig. <ref type="figure">S2</ref>). These collections correspond to HOT cruise numbers 267-283, for which physiochemical data are available in Table <ref type="table">S1</ref> and on Hawaii Ocean Time-Series HOT-DOGS application (<ref type="url">http://hahana.soest.hawaii.edu/hot/hot-dogs</ref>).</p><p>All samples were collected using the following procedure. 2-4 L of seawater (2 L from 5 to 175 m, 4 L from 200-500 m) were collected using CTD-attached Niskin bottles and filtered, using peristaltic pumps at a flow rate of about 6 L/h, onto a 0.2 &#956;m 25 mm Supor filter (VWR 28147-956) housed in polypropylene filter holder (Cole-Palmer EW-06623-32). These &gt;0.2 &#956;m cell-enriched samples were removed from filter holder and stored in 300 &#956;L of RNALater (Ambion AM7021, Waltham MA) at -80 &#176;C. 1-2 L (1 L from 5-175 m, 2 L from 200-500 m) of the &lt; 0.2 &#956;m filtrate were collected and filtered, using peristaltic pumps at a flow rate of about 1 L/h, onto a 0.02 &#956;m Whatman Anotop filter (VWR 28138-017, Radnor PA). These corresponding 0.2-0.02 &#956;m virus-enriched samples were stored sealed in filter housing at -80 &#176;C. DNA extraction, sequencing, and read quality-control are described in the Supplementary methods.</p><p>Quality-controlled reads, on average 9-10 million per sample (Table <ref type="table">S1</ref>), were assembled within each sample using option "-k 21,33,55,77,99,127" on metaSPAdes v3.10.1 <ref type="bibr">[28]</ref>, generating a total of 83 million contigs amongst 416 metagenomes (Table <ref type="table">S1</ref>). All sequences were used for viral reassembly and database curation. For downstream analyses, smaller metagenomes from samples with duplicate sequencing runs and samples with &lt;1 million reads were removed, resulting in a total of 374 metagenomes. Sequences were submitted to NCBI SRA under project number PRJNA352737 and assemblies can be found under BioSample SAMN12604809.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral-specific reassembly</head><p>All &gt;3 kb contigs were filtered using VIRSorter v1.03 <ref type="bibr">[29]</ref> using the virome database, under regular mode for cellenriched samples and decontamination mode for virusenriched samples. Contigs from all identified viral categories were retained (Table <ref type="table">S2</ref>). BWA-MEM v0.7.15 <ref type="bibr">(Li 2013)</ref> and msamtools <ref type="bibr">[30]</ref> was used to identify 809 million reads mapping to these putative viral contigs at &gt;95% average nucleotide identity (ANI) across &gt;45 bp (Table <ref type="table">S2</ref>). Viral reads were reassembled using metaSPAdes v3.11.1, which was chosen due to improved genome recovery and low rate of generating false apparent circularity <ref type="bibr">[31]</ref>. Reads from contigs with &gt;10 coverage, a threshold representing 99% genome recovery <ref type="bibr">[31]</ref>, were pooled across each depth for 12 reassemblies. Reads from contigs with &lt;10 coverage were pooled across all samples into a low-coverage reassembly to improve genome recovery.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>ALOHA 2.0 virus database curation</head><p>Viral contigs across the 13 reassemblies, along with viral contigs from two previous smaller datasets near Station ALOHA <ref type="bibr">[26,</ref><ref type="bibr">32]</ref>, were clustered with cd-hit-est v4.6 <ref type="bibr">[33]</ref> at &gt;95% ANI to form 1.5 million nonredundant viral populations. Populations were filtered through VIRSorter and 262 197 putative viruses from all categories were retained. Proteins were predicted using Prodigal v2.6.3 <ref type="bibr">[34]</ref> and functionally annotated using HMMer v.3.2 <ref type="bibr">[35]</ref> against the PFAM-A v30 database <ref type="bibr">[36]</ref>. Populations containing one or more known viral marker proteins were retained (bit score &gt;30 to capsid, head, neck, tail, spike, portal, terminase, clamp loader, T4 proteins, T7 proteins, Mu proteins, excisionase, phage integrase, repressor protein CI, or Cro), resulting in 56,559 high-confidence virus populations ranging from 0.5 to 366 kbp in length. To focus on full genomes or large genomic fragments, we retained only &gt;10 kbp contigs, resulting in 17,369 populations that form the ALOHA 2.0 virus database. Functional annotations were inspected to ensure that no ribosomal proteins are present, with the exception of S21 found also in a cultivated Pelagibacter phage, S33 found enriched in aquatic viruses, and L7/12 found in assembled viral contigs <ref type="bibr">[37]</ref>. One population containing L11 was retained due to the protein's proximity and interaction with L7/12 <ref type="bibr">[38]</ref>.</p><p>The high proportion of novel viral diversity in our samples precludes using reference genomes for the detection of chimeras, which are expected to occur at a frequency of ~0.5% using metaSPAdes <ref type="bibr">[31]</ref>. As a result, we inspected clusters for chimeric signature through self-alignment using LAST v756 <ref type="bibr">[39]</ref> to identify stretches of repeats at &gt;95% ANI across &gt;10 kbp. Fifteen populations displayed this signature and were noted as chimeras (Table <ref type="table">S3</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Genomic completion</head><p>We used terminal repeats (apparent circularity) to identify a nonredundant set of complete genomes <ref type="bibr">[40]</ref>, using a combination of four different methods: i. 439 were identified using Virsorter ii. 411 from check_circularity.pl <ref type="bibr">[41]</ref> iii. 790 from LAST to identify overlaps at 95% ANI across 100 bp-10 kbp within 200 bp of both ends; and iv. 1131 from NUCmer v3.1 <ref type="bibr">[42]</ref> to identify direct terminal repeats 20bp-10kbp in length within 200 bp of both ends <ref type="bibr">[43]</ref>. 1543 complete genomes were identified pooled amongst these four methods (Table <ref type="table">S3</ref>). LAST was used to assess redundant circularly permuted contigs at 95% ANI across 150 bp, yielding 961 nonredundant complete genomes (16,787 total, Fig. <ref type="figure">S1</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral and prokaryotic contribution to total DNA</head><p>Viral contribution was calculated for each sample as the proportion of reads mapping to the ALOHA 2.0 viral database (&gt;95% across &gt;45 bp). Prokaryotic contribution was calculated for each sample as the proportion of reads classified as bacterial or archaeal with Kaiju v1.6.2 <ref type="bibr">[44]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Spatiotemporal distribution and abundance</head><p>Reads from each sample were mapped using BWA-MEM to virus populations and filtered using msamtools at &gt;95% across &gt;45 bp. Anvi'o v3 <ref type="bibr">[45]</ref> was used to calculate coverage profiles for every sample, using interquartile range (IQR) coverage, which diminishes the effect of conserved or hypervariable regions in respectively over-and underestimating coverage. For analyses including all virus populations, relative nucleotides mapped (nucleotides mapped to population divided by nucleotides mapped across all populations) was used to calculate relative abundances (Tables <ref type="table">S4,</ref><ref type="table">S5</ref>). For analyses including only complete viral genomes, relative coverage (genome coverage divided by total coverage summed across all complete genomes) was used to approximate relative abundances (Tables <ref type="table">S6,</ref><ref type="table">S7</ref>).</p><p>To examine temporal persistence, reads from a previous 2010-2011 dataset from Station ALOHA <ref type="bibr">[46]</ref> were mapped to the 2014-6 ALOHA 2.0 virus database using BWA-MEM and filtered using msamtools at &gt;95% across &gt;45 bp. Populations with non-zero IQR coverage were considered present in the 2010-2011 dataset (Table <ref type="table">S3</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Characterizing cellular assemblages</head><p>COG0012, a universal single-copy marker protein, was used to generate 2568 mOTUs representing cells at the near-species level, at higher resolution than with rRNAbased OTUs <ref type="bibr">[46]</ref><ref type="bibr">[47]</ref><ref type="bibr">[48]</ref>. One Crocosphaera mOTU not previously included was curated, and relative abundances were calculated using relative coverages of reads mapping to these 2569 mOTUs (Table <ref type="table">S8</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral population identification</head><p>Taxonomy was assigned using protein LAST alignments to known phages in RefSeq84 ( <ref type="bibr">[49]</ref>; Table <ref type="table">S9</ref>), as well as five viral metagenomic databases available as of 2018: uvMED <ref type="bibr">[50]</ref>, uvDEEP <ref type="bibr">[23]</ref>, GOV <ref type="bibr">[15]</ref>, EV <ref type="bibr">[16]</ref>, and MED2017 <ref type="bibr">[24]</ref>. To avoid inflating the number of novel populations, we performed broad taxonomic assignments at &gt;60% average amino acid identity (AAI) across &gt;50% of proteins to any reference genome or contig, with the best hit assigned based on the highest AAI (Table <ref type="table">S3</ref>). A broader cut-off of &gt;50% of proteins at any AAI was used to identify phages infecting heterotrophic bacterioplankton in RefSeq84, due to lower sequence representation than picocyanophages in these databases (Table <ref type="table">S3</ref>).</p><p>Proteins were annotated using HMMsearch against PFAM (bit score &gt;30). Proteins with functional domains not found in previously reported datasets <ref type="bibr">[15,</ref><ref type="bibr">26]</ref> were considered novel viral genes (Table <ref type="table">S10</ref>).</p><p>Putative temperate phages were identified using i. functional annotations to identify phages with the genomic potential for lysogeny and ii. VIRSorter to identify integrated prophages. i. Functional annotations identified 922 populations with temperate phage markers (&gt;30 bit score to integrase, excisionase, Cro, or CI repressor). ii. VIRSorter identified prophages from original assemblies, of which 413 temperate phage populations shared significant homology (a minimum of 150 bp at 95% ANI); VIRSorter identified 73 final viral populations as prophages.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Archaeal virus identification</head><p>Putative archaeal viruses were identified in unannotated populations using archaeal protein markers and sequence similarity to Archaea or archaeal viruses in RefSeq84 using archaeal markers (PFAM bit score &gt;30). Populations carrying protein markers included one with AmoC <ref type="bibr">[51]</ref>, 16 with archaeal holiday junction resolvase, and 37 with MCM DNA helicase of archaeal origin (Table <ref type="table">S9</ref>), yielding a total of 53 archaeal virus marker populations. 632 populations were identified having &#8805;1 proteins with top hits (protein-protein LAST) to Archaea or archaeal viruses in the RefSeq84 database <ref type="bibr">[49]</ref>. We refined this set with the expectation that archaeal viruses should display high ratio of top protein hits to Archaea and archaeal viruses divided by top hits to Bacteria (AB ratio) and large proportion of proteins without RefSeq84 hits, given the scarcity of open ocean archaeal viruses in current databases. The 53 archaeal virus marker populations all displayed &gt;0.06 AB ratio and &gt;0.8 proportion of proteins without RefSeq84 hits (Fig. <ref type="figure">S3b</ref>). We used modified, conservative cut-offs (&gt;0.5 AB ratio, &gt;0.8 proportion of proteins without RefSeq84 hits) to retain 161 putative archaeal viruses (Fig. <ref type="figure">S3c</ref>). Phylum-level classifications were assigned based on number of protein top hits to Thaumarchaeota and Euryarchaeota. Populations with ties in number of top hits to these phyla were considered unclassified archaeal viruses.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Eukaryotic virus identification</head><p>Putative eukaryotic viruses were identified using the nucleocytoplasmic large DNA virus capsid marker (PFAM bit score &gt;30) and &#8805;2 proteins with top hits to eukaryotic viruses in RefSeq84 (Table <ref type="table">S9</ref>). All taxonomic identifications were cross-referenced to confirm that a population is assigned only to one group.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Crocosphaera putative phage identification</head><p>Putative Crocosphaera phages were identified through a combination of methods. First, alignments to RefSeq84 revealed one putative Crocosphaera phage (Table <ref type="table">S3</ref>) with 11 of 15 proteins hitting to Crocrosphaera, and one hit to another Cyanobacteria (Table <ref type="table">S9</ref>). For independent confirmation, this virus displayed abundance profiles expected from Crocosphaera (i.e., summer bloom in the upper ocean). 170 other potential Crocosphaera phages that displayed similar spatiotemporal distributions (Fig. <ref type="figure">S9</ref>) were also included as candidates for independent confirmation. Reads from samples capturing a Crocosphaera bloom (and presumably its phages) near Station ALOHA around the same time <ref type="bibr">[52]</ref> were mapped (&gt;95% across &gt;45 bp) to these 171 populations to confirm presence. 115 candidates that recruited reads at &gt;0 IQR coverage were retained. We then retained only candidates with higher number of top hits to non-Prochlorococcus/Synechococcus cyanobacterial proteins than top hits to Prochlorococcus or Synechococcus proteins, for a final set of one "putative" and six "potential" Crocosphaera phages (Table <ref type="table">S3</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral and prokaryotic diversity</head><p>To examine within-sample &#945;-diversity, we calculated the Shannon diversity, evenness, and richness for each sample using the vegan package in R <ref type="bibr">[53]</ref>. Respective assemblages were assessed using relative coverages of 16,787 viral populations and 2568 cellular mOTUs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>VC ratio</head><p>We define the Virus:Cell ratio (VC ratio) here as the logratio of a population's relative abundance (nucleotides mapped) in the virus-enriched versus cell-enriched size fractions. To calculate VC ratios, samtools bedcov <ref type="bibr">[54]</ref> was used to calculate nucleotides mapping to each population from all samples, normalized to the average library size of 1.3 million nucleotides per sample. Populations with zero nucleotides mapping in any sample were adjusted to one nucleotide for log-ratio calculation. Temporal variability of VC ratios was calculated for each population using the mean-normalized variance of its VC ratios within each depth.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results and discussion</head><p>A total of 374 metagenomes were generated from seawater sampled at Station ALOHA across 16 timepoints at 12 depths (5-500 m). These samples were used to prepare DNA from cell-enriched &gt;0.2 &#956;m size fractions (encompassing cellular DNA, giant viruses, active infections, prophages, and absorbed phages), and virus-enriched 0.2-0.02 &#956;m size fractions (capturing ultra-small cells, free virus particles, and ultra-small detritus). After metagenome assembly, curation, and classification (Fig. <ref type="figure">S1</ref>), the 4.2 TB of sequence data yielded 16,787 virus populations, with 8079 novel populations not represented in publicly available virus sequences. A total of 961 assembled populations were identified as complete via evidence of their terminal repeats <ref type="bibr">[55,</ref><ref type="bibr">56]</ref>. 1352 populations were identified as putative temperate phages, due to the presence of flanking cellular sequences or diagnostic temperate phage marker genes. Among novel viruses, 29 putative temperate Pelagibacter and Puniceispirillum phages were identified, along with 12 complete archaeal virus genomes, 25 genomic fragments of eukaryotic viruses, and a putative Crocosphaera phage-like element (Fig. <ref type="figure">S4,</ref><ref type="figure">S5</ref>). The curated viral populations contained 236 novel viral proteins (Table <ref type="table">S10</ref>) not previously reported from marine environments <ref type="bibr">[15,</ref><ref type="bibr">26]</ref>. These proteins included auxiliary metabolic genes in nitrogen (nitronate monooxygenase) and phosphorus cycling (transporter), phage holin, a toxin-antitoxin system (Fig. <ref type="figure">S4</ref>), and CRISPR-associated proteins.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral DNA contribution to total DNA</head><p>On average across all depths, the ALOHA 2.0 viral dataset recruited 42% of reads from the virus-enriched fraction and 8% of reads from the cell-enriched fraction (Fig. <ref type="figure">1</ref>). Assuming that viruses account for 56% of the total DNA in the &lt;0.22 &#956;m fraction <ref type="bibr">[57]</ref>, our dataset recovered on average 75% of the total viral DNA from this habitat. In the virusenriched fraction, viral contribution total DNA peaked near the base of the euphotic zone, 150-250 m. This increase in relative viral DNA contribution may reflect increased cellular turnover and viral production, increased temperate phage induction (described later), or decreased decay rates due to lower ambient UV light fluxes and temperature at these depths <ref type="bibr">[58]</ref>. In the cell-enriched fraction, relative viral contribution to total DNA peaked near the deep chlorophyll maximum (DCM) around 100 m, extending previous observations of subsurface cyanophage maxima in cellenriched fractions at Station ALOHA <ref type="bibr">[59]</ref>. Since other sources of DNA contribution are mostly cellular in origin, the increase in ratio of viral-to-bacterial DNA likely indicates a combination of increased active infections and prophages in these cell populations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Viral diversity hotspot at the base of the euphotic zone</head><p>Within-sample &#945;-diversity and richness revealed hotspots for virioplankton diversity, and confirmed it for prokaryotic diversity <ref type="bibr">[60]</ref>, at the base of the euphotic zone between 150 and 250 m (Fig. <ref type="figure">2</ref>). This peak in diversity could reflect both habitat variability and transitions in microbial metabolic diversity. For example, in this environment, photoautotrophic cyanobacteria dominate in the photic zone, while chemolithotrophic ammonium-oxidizing Thaumarchaeota and (presumptive) heterotrophic Euryarchaeota are both more abundant in deep waters. However, both cyanobacterial and archaeal groups do co-exist at transitional depths in and just below the deep chlorophyll maximum <ref type="bibr">[60]</ref>. Similarly, while cyanophages were found predominantly in shallow waters and archaeal viruses mostly below the photic zone, both cyanophages and archaeal viruses were found at transitional depths between 150 and 250 m (annotations shown on Fig. <ref type="figure">3</ref>, top bar). Previously uncharacterized viruses also increased in abundance below the upper surface waters, and represented &gt;50% of viral assemblages below 125 m (Fig. <ref type="figure">1</ref>). Elevated richness in oligotrophic picoplankton communities just below the photic zone was evident not only in Bacteria and Archaea, but also their viruses.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Depth-dependent patterns in temperate phage integration and induction</head><p>The community-wide abundances of 1352 putative temperate phages displayed depth-specific patterns (Fig. <ref type="figure">4</ref>). In cell-enriched samples, we postulate that putative temperate phages represented integrated prophages, which increased at and below the DCM, consistent with previous reports of apparent increased lysogeny in deeper waters <ref type="bibr">[26,</ref><ref type="bibr">61,</ref><ref type="bibr">62]</ref>. Depth-dependent changes in viral reproductive strategies were also evident within groups at the genus level, particularly in SAR11 phages with broad depth ranges (Fig. <ref type="figure">S7</ref>). Our results suggest that community-wide changes in reproductive strategies are driven not only by specific host groups, but also partly by environmental variability along the water column. Given that cellular host abundance generally decreases with depth (Fig. <ref type="figure">S2c</ref>), our results do not appear to support the piggyback-the-winner <ref type="bibr">[63]</ref> hypothesis in this oligotrophic pelagic habitat. It seems probable that Fig. <ref type="figure">1</ref> Depth profiles of time-averaged viral and prokaryotic contributions to total sequenced DNA in a virus-enriched and b cellenriched size fractions. Relative abundances of viral populations are colored by average amino acid identity (AAI) (&gt;60% AAI across &gt;50% genes) to known phages in RefSeq, and five other viral metagenomic datasets: uvMED <ref type="bibr">[50]</ref>, uvDEEP <ref type="bibr">[23]</ref>, Med2017 <ref type="bibr">[24]</ref>, GOV <ref type="bibr">[15]</ref>, and EV <ref type="bibr">[16]</ref>. Legend shows number of populations identified in each category. virus-host dynamics may have different trajectories in different environmental, biological, and ecological contexts, and may not be driven simply by bulk numerical host-cell and virus-particle ratios alone. In oligotrophic pelagic environments, high prokaryote cell densities in surface waters correspond to smaller average genome sizes <ref type="bibr">[46,</ref><ref type="bibr">64]</ref>. Within and below the DCM, low cell densities correspond to larger average genome sizes <ref type="bibr">[46]</ref>. Host cell genome sizes (and therefore genomic "real estate" available for prophage, genomic islands and other mobile elements), rather than cell densities, might drive viral reproductive strategies <ref type="bibr">[65]</ref>.</p><p>In virus-enriched samples, we postulate that temperate phage signal represents temperate virus particles, which peaked at 150-250 m (Fig. <ref type="figure">4</ref>). This peak may reflect increased temperate phage productivity relative to deeper waters associated with more slowly growing hosts <ref type="bibr">[66]</ref>. Temperate phage production at these intermediate depths appeared to be episodic, as indicated by temporal variability in populations' extracellular to intracellular abundance (VC) ratios (Fig. <ref type="figure">5</ref>). Low VC variability reflects phages with constant or no viral particle production. High VC variability indicates phages with more episodic virus particle production. The VC variability of temperate phages peaked at 150-250 m (Fig. <ref type="figure">5</ref>), consistent with episodic induction and production. Episodic temperate phage induction may be driven by temporal resource variability or host abundance [ <ref type="bibr">[67]</ref><ref type="bibr">[68]</ref><ref type="bibr">[69]</ref>, or increased cellular stress due to lower light for photoautotrophs at these depths (Fig. <ref type="figure">S2f</ref>).</p><p>Given the low abundance of temperate phages in the upper 75 m (Fig. <ref type="figure">4</ref>), we infer that lytic phage-host interactions were prevalent in surface waters, potentially reflecting consistently high productivity and host abundance that can favor lytic strategies <ref type="bibr">[66,</ref><ref type="bibr">[70]</ref><ref type="bibr">[71]</ref><ref type="bibr">[72]</ref>. Temperate phage induction and production appeared to peak at the base of the euphotic zone (Figs. <ref type="bibr">4,</ref><ref type="bibr">5)</ref>, potentially reflecting increased environmental variability at these depths that might favor a flexible reproductive strategy <ref type="bibr">[69,</ref><ref type="bibr">73]</ref>. Evidence for hostintegrated prophages increased in mesopelagic waters (Fig. <ref type="figure">4</ref>), possibly reflecting decreases in host productivity and abundance <ref type="bibr">[72,</ref><ref type="bibr">73]</ref>, or an increase in genome size that might better accommodate prophages <ref type="bibr">[46,</ref><ref type="bibr">65]</ref>. Using size fractionation to separate intra-and extracellular viruses, our results revealed depth-specific viral reproductive strategies from the surface to mesopelagic ocean.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>VC ratio to confirm genomic temperate phage identification</head><p>The VC ratio of each population represents its extracellular to intracellular abundance ratio. On average, temperate phages, or phages persisting intracellularly for long periods before initiating a lytic cycle, are expected to have lower VC ratios relative to strictly lytic phages. Indeed, temperate phages on average displayed lower VC ratios than that of other coexisting populations (Welsh's t test, p = 0.02). This trend provided further support for our temperate phage identifications using genomic markers. Nevertheless, intracellular temperate phage counts may reflect a variety of life history states. Intracellular temperate phage counts could represent prophage DNA still integrated in the host genome, replicating temperate phage that have have excised from the host genome, or temperate phage that entered directly into the lytic cycle immediately post-infection (in eclipse phase).</p><p>Presuming that the temperate phage signal in the virusenriched fraction represents packaged phage particles, ~98% of putative temperate phages appeared to have the ability to produce viral particles in situ (Fig. <ref type="figure">S6b</ref>, Table <ref type="table">S3</ref>). Hence, the lower VC ratios found in the temperate phages does appear to reflect their different lifestyles and reproductive strategies, relative to non-temperate phages.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Environmental distribution of viral AMGs</head><p>Viral copies of AMGs potentially encode for rate-limiting steps in host metabolism that are essential for viral propagation (reviewed in <ref type="bibr">[13]</ref>). We show here that the abundance of three key AMGs in energy and nutrient acquisition vary with depth, potentially indicating cellular host variability or energy or nutrient limitation at specific depths.</p><p>Virus-encoded copies of photosystem reaction center genes were first observed in a Synechococcus phage genome, were thought to prevent photoinhibition by supplementing declining host photosynthesis during infection <ref type="bibr">[74]</ref>, and were found to be co-expressed with high-light gene in cyanophages infecting high-light Prochlorococcus hosts <ref type="bibr">[4,</ref><ref type="bibr">75]</ref>. Photosystem genes were subsequently observed in many cyanophage genomes <ref type="bibr">[76]</ref>, yet their environmental distribution in the water column remains relatively unexplored. Here, we observed a threefold increase in abundance of phage carrying photosystem genes from the surface waters to low-light environments around the DCM (Fig. <ref type="figure">6a</ref>). Considering that cyanophages and cyanobacteria were most abundant in the surface ocean (Figs. 6a, S8a), a higher proportion of cyanophages carried photosystem genes at the DCM relative to surface waters. Our observations suggest that virus-encoded photosystem genes may be more advantageous in light-limited conditions, potentially to prevent light-energy limitation in hosts VC ratio temporal variability (10 8 ) depth (meters) other viruses Fig. <ref type="figure">5</ref> Depth profiles of temporal variabilities of VC ratios (mean &#177; SE) of 1352 inferred temperate phages (blue) and other viral populations (orange). Higher VC ratio temporal variability indicates episodic production of free viral particles, while lower variability indicates consistent production of free viral particles. Temporal variability of VC ratios is calculated for each population by pooling its VC ratios within each depth to determine the mean-normalized variance. during infection. Alternatively, longer latent periods, consistent with our observed increase in viral DNA in the cellenriched fraction at the DCM (Fig. <ref type="figure">1a</ref>), might also favor phage-encoded photosystem genes at this depth <ref type="bibr">[76]</ref>.</p><p>Viral-encoded ammonium transporter genes, which shared highest similarity to bacterial homologs (Table <ref type="table">S9</ref>), increased at and below the DCM in tandem with Thaumarchaeota abundance (Fig. <ref type="figure">6b</ref>). Ammonium-oxidizing Thaumarchaeota in the ocean have demonstrated high affinity for ammonia <ref type="bibr">[77]</ref> and therefore may compete with bacteria in ammonia-limited open-ocean environments <ref type="bibr">[78]</ref>. As a result, bacteriophage-encoded copies of ammonium transporter genes might assist in ammonium acquisition during infection, particularly at depths with higher abundances of Thaumarchaeota competitors.</p><p>PhoH genes are used during phosphorus starvation <ref type="bibr">[79]</ref>, upregulated during cyanophage infection <ref type="bibr">[80]</ref>, present in diverse groups of marine viruses <ref type="bibr">[5,</ref><ref type="bibr">7,</ref><ref type="bibr">[81]</ref><ref type="bibr">[82]</ref><ref type="bibr">[83]</ref>, and can be used as a marker gene for assessing viral assemblages <ref type="bibr">[26,</ref><ref type="bibr">84]</ref>. Viruses are richer in phosphorus relative to cells <ref type="bibr">[85]</ref>, potentially driving an enrichment of this auxiliary metabolic gene in marine viruses <ref type="bibr">[84]</ref> that inhabit relatively nutrientlimited open oceans. At Station ALOHA, up to 10% of viruses encoded copies of phoH genes, peaking at the DCM in tandem with a maximum in the total dissolved nitrogen to phosphorus ratio (Fig. <ref type="figure">6c</ref>). Our observations suggest that the phoH gene might assist in phosphorus acquisition during viral infection in relatively phosphorus-limited environments.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Spatiotemporal distributions: temporal persistence and depth distributions</head><p>Many virioplankton populations displayed temporal persistence through the 1.5-year time series (Figs. 3, S9), consistent with recent studies at Station ALOHA <ref type="bibr">[26,</ref><ref type="bibr">32]</ref>. However, some virioplankton populations occurred only at specific times of year (Fig. <ref type="figure">S10</ref>), potentially reflecting the seasonal variability of physiochemical variables and their cellular hosts. Of the 9579 viral populations that recruited reads from the ALOHA 2.0 cell-enriched samples from 2014-2016, 6959 (76%) also recruited reads from cellenriched samples from 2010-2011 (Table <ref type="table">S3</ref>). This strong temporal persistence indicates that viral genome evolution may be under stabilizing selection, constrained to very specific loci, or occur at finer resolutions than 95% ANI used here for read-mapping and defining populations. Our results show that many indigenous phage populations persist over multiannual cycles, indicative of continuous infection, replication, and lytic cycles that maintain this persistence.</p><p>Viral assemblages displayed stratified depth structure, with shifting depth ranges along the transition zone between surface and mesopelagic waters (Figs. <ref type="figure">3,</ref><ref type="figure">S9</ref>). This depth structure likely reflects biogeochemical gradients <ref type="bibr">[86]</ref> and depth partitioning of bacterial host communities <ref type="bibr">[46]</ref>. Consistent with a previous study at Station ALOHA <ref type="bibr">[26]</ref>, we found no evidence of eurybathic viruses in either the cellular or virioplankton enriched size fractions. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Depth distributions of archaeal viruses and their hosts</head><p>Several different archaeal lineages, including Euryarchaeota and Thaumarchaeota (formerly Group I marine Crenarchaeota <ref type="bibr">[87]</ref>) are commonly found along the water column at Station ALOHA <ref type="bibr">[59,</ref><ref type="bibr">88,</ref><ref type="bibr">89]</ref>. Euryarchaeota apparently function as organic matter degrading heterotrophs, while Thaumarchaeota appear to function as chemolithoautotropic ammonium oxidizers <ref type="bibr">[78,</ref><ref type="bibr">90]</ref>. We identified 161 putative archaeal virus populations, 12 of which represented complete genomes. We classified 67 of these populations as euryarchaeal and 60 as thaumarchaeal, with the latter encompassing crenarchaeal populations (Table <ref type="table">S3</ref>). In both size fractions, total archaeal virus abundance peaked at and below the DCM (Fig. <ref type="figure">S8b</ref>), with putative thaumarchaeal viruses driving most of this change. Relative abundances of archaeal viruses were consistent, though somewhat lower in magnitude, with that of archaea (Fig. <ref type="figure">S8b</ref>), suggesting that our methods of archaeal virus identification were conservative but accurate.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Putative Crocosphaera phages</head><p>The unicellular cyanobacterium Crocosphaera can fix both nitrogen and carbon, and plays an important role in fueling primary production in nutrient-poor open ocean environments <ref type="bibr">[91]</ref>. Although it is the third-most abundant cyanobacterium at Station ALOHA, after Prochlorococcus and Synechococcus <ref type="bibr">[52]</ref>, no Crocosphaera phage sequences have yet been identified in current databases. We identified the genome of a putative Crocosphaera phage or phage parasite based on multiple lines of evidence: sequence similarity to Crocosphaera in 11 of 15 proteins, similar spatiotemporal abundance profiles to that expected from Crocosphaera (summer in the upper ocean), and presence during a confirmed Crocosphaera "bloom" near Station ALOHA around the same time <ref type="bibr">[52]</ref>. This genome contains a higher GC content (41.7%) than other co-occurring viruses (37.3%), consistent with the higher GC content of Crocosphaera at ~37.4% <ref type="bibr">[92]</ref>, compared with abundant surface Prochlorococcus at ~32% <ref type="bibr">[46]</ref>. Accounting for up to 2% of viral DNA in the virus-enriched size fraction, this genome likely represents a phage, plasmid, or phageinducible chromosomal island (PICI) packaged into phage particles. PICIs are phage parasitic DNA elements that can hijack another infecting phage's packaging machinery and mobilize in infectious phage-like particles <ref type="bibr">[93,</ref><ref type="bibr">94]</ref>. The genome (11 kbp) appears complete and encodes a phage integrase (Fig. <ref type="figure">S6</ref>), characteristic of both temperate phages and PICIs. In addition to this genome, we identified six other potential Crocosphaera phages (Table <ref type="table">S3</ref>) based on similar Station ALOHA abundance profiles to Crocosphaera, their presence during a Crocosphaera "bloom" in one set of samples, and their dissimilarity from Prochlorococcus or Synechococcus gene homologs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Conclusion</head><p>Viruses impact the ecology and biogeochemistry of microbial communities across the oceans. Our study recovered a large fraction of dsDNA viruses found at Station ALOHA, revealing a hotspot for microbial diversity at the base of the euphotic zone, where the majority of viruses were distinct from those previously reported. The concurrent characterization of both intracellular and extracellular viruses provided independent support of temperate phage identification using marker genes, and revealed community-wide shifts in viral reproductive strategies. In this open-ocean environment, lytic interactions dominated the upper photic layer above the DCM, and temperate phages were most abundant in the mesopelagic ocean. Temporal variability in reproductive strategies appeared most prevalent in transitional depths at the base of the euphotic zone, marked by a peak in prophage integration and induction. Environmental distributions of viral AMGs also displayed depth-specific patterns in energy and nutrient acquisition. For example, photosystem genes were most abundant at light-limited depths around the DCM; bacteriophage-encoded ammonium transporter genes increased along with potential ammonia-oxidizing thaumarchaeal competitors; phosphorus starvation genes increased in tandem with N:P ratio. Although most viruses were temporally persistent over several years, some displayed temporal variability that appeared to reflect the seasonal distributions of specific hosts. These temporal patterns, in conjunction with other lines of evidence, led to the identification of putative Crocosphaera phages or phage parasites that have not been previously identified.</p><p>Taken together, these new data and analyses provide new insight on the spatiotemporal patterns of planktonic viral diversity, reproductive strategies and metabolic repertoires in the open ocean. Furthermore, simultaneous ennumeration of extracellular and intracellular viruses and their hosts sets the stage for delineating more specific host-virus interactions.</p><p>provided </p></div></body>
		</text>
</TEI>
