<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Yeast Surface Display of Protein Addresses Confers Robust Storage and Access of DNA-Based Data</title></titleStmt>
			<publicationStmt>
				<publisher>MDPI</publisher>
				<date>09/01/2025</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10662693</idno>
					<idno type="doi">10.3390/dna5030034</idno>
					<title level='j'>DNA</title>
<idno>2673-8856</idno>
<biblScope unit="volume">5</biblScope>
<biblScope unit="issue">3</biblScope>					

					<author>Magdelene N Lee</author><author>Gunavaran Brihadiswaran</author><author>Balaji M Rao</author><author>James M Tuck</author><author>Albert J Keung</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[<p>Background/Objectives: The potential of DNA as an information-dense storage medium has inspired a broad spectrum of creative systems. In particular, hybrid biomolecular systems that integrate new materials and chemistries with DNA could drive novel functions. In this work, we explore the potential for proteins to serve as molecular file addresses. We stored DNA-encoded data in yeast and leveraged yeast surface display to readily produce the protein addresses and make them easy to access on the cell surface. Methods: We generated yeast populations that each displayed a distinct protein on their cell surfaces. These proteins included binding partners for cognate antibodies as well as chromatin-associated proteins that bind post-translationally modified histone peptides. For each specific yeast population, we transformed a library of hundreds of DNA sequences collectively encoding a specific image file. Results: We first demonstrated that the yeast retained file-encoded DNA through multiple cell divisions without a noticeable skew in their distribution or a loss in file integrity. Second, we showed that the physical act of sorting yeast displaying a specific file address was able to recover the desired data without a loss in file fidelity. Finally, we showed that analog addresses can be achieved by using addresses that have overlapping binding specificities for target peptides. Conclusions: These results motivate further exploration into the advantages proteins may confer in molecular information storage.</p>]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>The growing interest in DNA as a potential substrate for digital data storage derives predominantly from its high theoretical information density and low resource utility. Yet, it is also intriguing to consider how DNA might confer additional advantages through its unique biophysical properties and through currently unexplored interactions with novel chemistries and other materials. As just a few examples, the physical structure of DNA itself or epigenetic modifications to the DNA could store information or confer new functionalities <ref type="bibr">[1]</ref><ref type="bibr">[2]</ref><ref type="bibr">[3]</ref><ref type="bibr">[4]</ref><ref type="bibr">[5]</ref>. Hybrid systems incorporating materials like polymers, silica particles, and soft dendritic colloids provide additional functionalities ranging from repeatable multiplexed data access, in-storage computation, and long-term storage <ref type="bibr">[6]</ref><ref type="bibr">[7]</ref><ref type="bibr">[8]</ref><ref type="bibr">[9]</ref>.</p><p>Proteins may present a powerful class of molecular materials to integrate into DNAbased information systems. They could provide diverse molecular recognition and search capabilities, evident in the natural role some proteins hold as antibodies. Their molecular recognition capabilities derive from the high degrees of three-dimensional and chemical freedom conferred by the twenty amino acid building blocks. They do face several potential limitations, including their lower stability, higher cost, and orders of magnitude lower synthesis scalability compared to DNA.</p><p>Yeast surface display presents an interesting system by which to leverage the advantages of both DNA and proteins. Yeast can be lyophilized and kept stable for up to 20 years <ref type="bibr">[10]</ref>. It can encode proteins in synthesized DNA, bypassing the need for chemical or recombinant protein synthesis and purification. It can display those proteins on its cell surface, providing a potential molecular address by which to find and manipulate specific yeast cells <ref type="bibr">[11]</ref>. Recent work also demonstrates that yeast can be transformed to hold large synthetic DNA cassettes and even whole artificial chromosomes <ref type="bibr">[12]</ref><ref type="bibr">[13]</ref><ref type="bibr">[14]</ref><ref type="bibr">[15]</ref>. In a related biomolecular system, promising work has shown file-encoded DNA stored and accessed in bacteria <ref type="bibr">[16]</ref><ref type="bibr">[17]</ref><ref type="bibr">[18]</ref><ref type="bibr">[19]</ref><ref type="bibr">[20]</ref><ref type="bibr">[21]</ref>.</p><p>Here, we probed the potential utility of yeast surface display leveraging protein addresses for DNA-based information storage. We first tracked how the distribution of yeast populations, collectively comprising data files, were maintained over multiple cell divisions. We then implemented file access via sorting based upon protein addresses. Finally, we demonstrated how proteins can confer unique types of addresses that extend beyond simple one-to-one interactions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Materials and Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">Yeast Strains</head><p>Frozen yeast strains from prior work were used for this study, and all derived from EBY100 (Table <ref type="table">S1</ref>) <ref type="bibr">[22]</ref><ref type="bibr">[23]</ref><ref type="bibr">[24]</ref>. All frozen yeast were streaked on dropout agar plates except for the parent EBY100, which was streaked on a YPD agar plate. EBY100 is BJ5465 (MATa aga1::gal1-aga1::ura3 ura3-52 trp1 leu2-delta200 his3-delta200 pep4::HIS3 prbd1.6R can1 GAL) and is an auxotroph for Leu and Trp. Single colonies were then picked and placed into a 2 mL culture of their appropriate synthetic dropout or YPD media to grow for 2 days at 30 &#8226; C, 250 rpm. From this yeast stock, a fresh SDCAA-Tryptophan (SD-trp) culture was seeded before each experiment. For an experiment, an SD-trp culture of yeast was grown for 24 h before passage into SGCAA-Tryptophan (SG-trp) induction media, where it was grown at 20 &#8226; C, 250 rpm for 16-24 h. SG-trp cultures were all induced at an OD 600nm of 1. In addition, 1 &#181;M biotin in DMSO was used in inducing the yeast strain with the FLAG tag to aid in its folding. After the file library was electroporated in the yeast strains, they were grown in SDCAA-Tryptophan-Leucine culture (SD-trp-leu) instead, with the corresponding SGCAA -Tryptophan-Leucine (SG-trp-leu) media for induction.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">Labeling Epitope Tag Yeast Before and After File Transformation with Flow Cytometry</head><p>Freshly induced yeast cultures were processed for flow cytometry by aliquoting 4 &#215; 10 5 cells into individual wells of a 96-well plate for antibody labeling. Each sample was labeled with a mixture of primary and secondary antibodies. First, 50 &#181;L of a primary antibody dilution (Chicken anti-MYC for the MYC tag and Rabbit anti-FLAG for the FLAG tag) was added to each well and allowed to incubate at room temperature, shaking at 800 rpm for 30 min. The samples were then washed of unbound primary antibody by resuspension in 200 &#181;L of 0.1% BSA 1&#215; PBS and centrifugation at 3000&#215; g for 2 min and then aspirated. The samples were then labeled with a secondary antibody corresponding to the host animal of the primary antibody (donkey anti-Chicken 647 for the MYC tag and donkey anti-rabbit 647 for the FLAG tag). The secondary antibodies were added in 50 &#181;L aliquots of 1:250 dilutions of their stock concentrations. The secondary antibodies were incubated with the samples for at least 15 min in the dark, at 4 &#8226; C. The secondary antibodies were washed in the same manner as the primary antibodies. The final sample pellets were resuspended in 200 &#181;L of 0.1% BSA 1&#215; PBS when the plate was loaded into the flow cytometer (MACSQuant &#174; VYB, Auburn, CA, USA). Flow cytometry data were analyzed with FlowJo software (FlowJo 10.10). All experiments were performed in triplicate with three biological replicates showing the same trend. In total, 50,000 cells were collected for analysis for each sample.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.">File DNA Synthesis</head><p>Each file DNA library was synthesized from a Twist DNA oligo pool containing all the oligos from the three picture files. Overhang primers were designed to include overlapping regions with the file plasmid (pCT302-SsoFe2-T2A-TOM22) to extend the strand length by 40 bp on each end (Table <ref type="table">S2</ref>). Multiple PCR reactions were set up to yield a final amount of 12 &#181;g per file library after purification. Each PCR consisted of the following reagents: 1 uL 1 &#215; 10 10 st/&#181;L template, 2 &#181;L of each 10 &#181;M primer, 5 &#181;L of the 5&#215; Q5 reaction buffer, 0.5 &#181;L of a 10 mM dNTPs mix, 0.25 &#181;L of the Q5 polymerase (New England Biolabs Cat # M0491L), and 14.25 &#181;L of nanopure water. The amplicons were run on a 1% agarose gel and extracted for gel purification. For the file plasmid, a 500 mL bacteria culture containing the pCT302-SsoFe2-T2A-TOM22 was grown overnight and spun down for plasmid extraction. The purified plasmid was then digested with restriction enzymes EagI and AvrII at 37 &#8226; C for 1 h and dephosphorylated for 10 min. The digested product was then run on a 1% agarose gel. The digested 7780 bp band was excised and extracted with a gel purification kit. The purified DNA libraries and digested vector were concentrated to a 30 &#181;L volume each and quantified using nanodrop.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.4.">Inserting Files into Selected Yeast</head><p>Each file DNA library was transformed into different yeast strains with different histone binders and epitope tags displayed. Table <ref type="table">S3</ref> notes the file DNA library for each protein address used in this manuscript. For each yeast strain, a frozen stock was streaked on an SD-trp agar plate to incubate for 2 days at 30 &#8226; C. Then, a single colony was inoculated in 5 mL of SD-trp media and grown at 30 &#8226; C to an OD 600nm of 3. This culture was passaged into 200 mL of fresh YPD media at an initial OD 600nm of 3 and grown to OD 600nm of 1.6. The culture was then spun down at 3000 rpm for 3 min. The supernatant was removed and the pellet washed twice with 100 mL of autoclaved water followed by one wash in 100 mL of autoclaved electroporation buffer (1 M sorbitol and 1 mM CaCl 2 ). The pellet was then resuspended in 40 mL of a filter-sterilized LiAC/DTT/HEPES solution (0.1 M LiAc, 50 mM HEPES, 1.5 mg/mL DTT, pH 7.5) and shaken in a 250 mL baffled flask for 30 min at 30 &#8226; C. The culture was spun down and washed with 100 mL of electroporation buffer before resuspension in 2 mL of electroporation buffer. Electroporation cuvettes (GenePulser cuvette, 0.2 cm electrode gap, Hercules, CA, USA) were chilled before 400 &#181;L aliquots of the cell mixture were added for each transformation. Two transformations were performed for each yeast strain: one with the file DNA library and one without to serve as the control. 12 &#181;g of the file DNA library and 4 &#181;g of the file plasmid vector were used for each sample, as applicable. After pipetting the sample gently, the mixture was electroporated at 2500 V, 25 F, and 200 ohms. The mixture was then transferred to a separate culture tube, and the cuvette was washed with a 1:1 sorbitol/YPD solution, with each wash transferred to the culture tube. The culture tubes were incubated at 30 &#8226; C for 1 h at 250 rpm. The cultures were then transferred to a conical tube, spun down, and washed in SD-trp-leu media. Finally, each pellet was resuspended in 5 mL of SD-trp-leu media and 10-fold serial dilutions of each transformation was plated on SD-trp-leu agar to compare the colony counts of the control and test sample of each transformation. The remaining culture containing the yeast transformations with the file inserts was suspended in 125 mL of SD-trp-leu media and grown overnight. A penicillin-strep stock solution was added to a final concentration of 1&#215; the following morning, and the cultures were used for the sorting experiments after the second passage.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.5.">Analyzing the File3 Yeast Through Multiple Cell Divisions</head><p>The transformed yeast strain containing the File3 DNA was used for the serial passaging study. In this experiment, the culture was grown for about 5 growth cycles (from OD 600nm 0.05 to OD 600nm 3) and aliquoted into 2 mL of fresh SD-trp-leu media to repeat the growth process again for 4 more times. After each growth period, 1 mL of the culture was collected for plasmid extraction. The aliquoted pellet was frozen at -80 &#8226; C so that the pellets from each passage could be processed simultaneously. The sequencing primers for File3 (Table <ref type="table">S2</ref>) were used to amplify the data payload region of the extracted plasmids using the same PCR set-up as previously described, and the amplicons were purified via gel extraction.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.6.">Yeast Sorting and File Access via Epitope Tag</head><p>After vortexing the Protein A-conjugated Dynabeads stock bottle for 30 s, a 950 uL aliquot was transferred to a separate conical tube. Each sample used 1.5 mg (50 &#181;L) Dynabeads, and the Dynabeads were prepped as a master mix. The Dynabeads were washed twice with 1 mL of a buffer containing 0.1% BSA and 0.02% Tween in PBS. This washing buffer was used throughout the experiment. After removing the supernatant by placing the tube against a magnet, the aliquot was resuspended in 1 mL of the washing buffer and divided into two portions for each antibody. The washing buffer was removed, and 50 &#181;L of each antibody stock solution (rabbit anti-FLAG or rabbit anti-MYC) was added to each aliquot and mixed with the washing buffer for a total volume of 2 mL. The tubes were then incubated with rotation for 30 min at room temperature. The antibody-bound beads were then washed twice with 1 mL of the washing buffer and resuspended in 1 mL of the washing buffer. To prep the yeast, 2 &#215; 10 7 cells of each file-transformed strain were used for each condition. The file-transformed strains were passaged and induced. Additionally, 2 &#215; 10 8 cells of EBY100 Saccharomyces cerevisiae were added for each condition. The EBY100 Saccharomyces cerevisiae was grown in YPD on the day that the file-transformed yeast strains were induced. Each sample was spun down at 4 G for 2 min. After removing the media, each cell pellet was washed twice with the washing buffer. Each cell pellet was then resuspended in 100 &#181;L of the appropriate antibody-bound bead solution and 900 &#181;L of the washing buffer. The tubes were incubated for 1 h on rotation before each mixture was washed twice with 1 mL of the washing buffer. The supernatant was then removed, and each mixture was resuspended in 2 mL of SD-trp-leu media and placed in a culture block to shake for two days at 250 rpm on 30 &#8226; C. Each mixture was then induced at an OD 600nm of 1 at 20 &#8226; C, and the magnetic bead sorting procedure was repeated upon the sorted cells for refined results. After growing the final sorted mixtures overnight, the plasmids were extracted for a multiplex amplification of the file payload regions using the Zymo Yeast Plasmid Miniprep II kit (Cat #D2004).</p><p>Specifically, primers for File2 and File3 were used with Taq polymerase to enrich the extracted DNA, and the amplicons were purified using PCR purification columns before submission to Azenta Amplicon-EZ (Table <ref type="table">S2</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.7.">Labeling the Histone Binder Yeast After File Transformation with Flow Cytometry</head><p>Freshly induced yeast cultures were processed for flow cytometry by aliquoting 4 &#215; 10 5 cells into individual wells of a 96-well plate for peptide and antibody labeling. Dilutions of a H3K9me3, H3K27me3, and H3K14ac biotinylated peptide solution were made at a final concentration of 4 uM, 4 uM, and 8 uM, respectively. A dilution of a rabbit anti-MYC solution was also prepped at a final dilution of 1:75. Equivolumes of each diluted peptide and rabbit anti-MYC solution were mixed to prep master mixes for the experiment. Then, 50 &#181;L of the corresponding peptide-myc mixture was added to each well and allowed to incubate at room temperature, shaking at 800 rpm for 30 min. The samples were then washed of unbound primary antibody by resuspension in 200 &#181;L of 0.1% BSA 1&#215; PBS, centrifugation at 3000&#215; g for 2 min, and aspiration. The samples were then stained with a mixture of streptavidin-conjugated PE staining dye and a donkey anti-rabbit 647 antibody. The staining mixture was added in 50 &#181;L aliquots of 1:250 dilutions of their stock concentrations. The staining mixture was allowed to incubate with the samples for at least 15 min in the dark, at 4 &#8226; C. The unbound staining mixture was washed in the same manner as the primary labeling step. The final sample pellets were resuspended in 200 &#181;L of 0.1% BSA 1&#215; PBS when the plate was loaded into the flow cytometer (MACSQuant &#174; VYB). Flow cytometry data were analyzed with FlowJo software. All experiments were performed in triplicate with three biological replicates showing the same trend. All fluorescent gates were created based on an unlabeled sample of the same yeast strain (with the same bound peptide when applicable), measured on the same day, as the experimental samples. In total, 50,000 cells were collected for analysis for each sample.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Data Transformed into Yeast Populations Maintain Their Fidelity Through Cell Divisions</head><p>An important first consideration for any data storage system is its fidelity over time or over repeated data access. For DNA, the highest risks to data fidelity are mutations or loss of specific sequences or strands of DNA due to chemical or physical manipulation <ref type="bibr">[25]</ref>. These risks arise in vitro when DNA strands are copied through enzymatic processes like polymerase chain reaction, and they arise in vivo when cells recombinantly assemble the file-encoded DNA and use DNA polymerase to copy their DNA during cell division. To assess these risks, we transformed yeast with the 507 DNA sequences that altogether encode for File3, an image of a muscle cell (Figure <ref type="figure">1</ref>). The encoding process is shown in Figure <ref type="figure">S1</ref>. We then cultured the yeast population for 20 cell divisions, harvesting DNA every five cell divisions. Next-generation sequencing (NGS) of the extracted DNA revealed that multiple important parameters of file fidelity were maintained. The normalized strand abundance, a metric describing the spread of the distribution of the DNA library, remained at a logtransformed median of ~0.6 copies per strand through all cell divisions with mutually similar violin plot morphologies (Figure <ref type="figure">2a</ref>, Equation (S4)). The strand distributions also remarkably matched that of the initial synthesized DNA library prior to transformation into yeast. Each distribution displayed a normal skew (skew = 0.33; 0.28; 0.37; 0.35; 0.32) and mesokurtosis (excess kurtosis = -0.14; -0.28; 0.044; -0.0; -0.18) (Figure <ref type="figure">2b</ref>).</p><p>The distribution of yeast also affects the percent of unique DNA sequences recovered. With a limited number of NGS reads, there will often be sequences that are not recovered (drop-outs), with this percentage increasing with fewer NGS reads or with greater skew in the distribution. We observed the percent of unique DNA sequences that were retained remained consistently at ~80% through all cell divisions (Figure <ref type="figure">2c</ref>). Relatedly, the manner in which the file information is encoded into DNA allows for strand dropouts through some level of redundancy and for some errors in synthesis, sequencing, or mutations through error correction <ref type="bibr">[8,</ref><ref type="bibr">26]</ref>. We also analyzed decodability and error rates to determine the efficiency and accuracy of file retrieval. To assess error rates, after filtering the sequencing reads to only include reads with length of 160 nt and normalizing the reads per condition to make sure the same number of reads were used per condition, we ran the normalized subset of sequencing reads through the FrameD sequencing analysis tool (Figure <ref type="figure">S1</ref>) <ref type="bibr">[8,</ref><ref type="bibr">26]</ref>. This tool mapped each read to its corresponding encoded strand, and the success of this mapping resulted in error rate statistics. Decodability was defined as "true" if the raw sequencing reads were successfully decoded back to the original image file and "false" if not, without prior knowledge of the file data and DNA sequences. File3 was fully decodable for all samples indicating cell divisions did not catastrophically impact strand retention or error rates (Figure <ref type="figure">2d</ref>). We also directly measured the error rate and found &lt;0.5% errors per nucleotide position in all samples through 20 cell divisions (Figure <ref type="figure">2e</ref> and Figure <ref type="figure">S2</ref>). These results indicated that yeast can stably maintain DNA-encoded image files through the transformation process, cell divisions, and DNA extraction.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Transformation of Yeast with File DNA Partially Impacts Display Effiicency but Not Labeling Specificity</head><p>We next asked how file transformation might affect the display of file address proteins. To test this, we transformed yeast that can inducibly display the MYC epitope tag with File2 and yeast that can inducibly display FLAG with File3. Each yeast population was then labeled with a MYC antibody (Figure <ref type="figure">3a</ref>) or with a FLAG antibody (Figure <ref type="figure">3b</ref>). The transformed yeast were specifically labeled by their corresponding antibody, but there was a significant decrease in epitope display efficiency upon transformation with file DNA. We attribute this decrease to two potential factors. First, there may be competition for AGA1 partners between the AGA2-epitope fusion protein and an AGA2-TOM22 protein expressed as part of the file plasmid but that was not specifically used in this work (Supplementary Materials Section S2) <ref type="bibr">[24]</ref>. Second, the retroactivity of consuming additional resources to copy, transcribe, and translate an additional plasmid could reduce the overall efficiency of surface display of proteins. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3.">Files Accessed with Specificity via Protein Addresses</head><p>After demonstrating that the transformed yeast strains preserved their address specificity, we mixed the File2 and File3 yeast populations and attempted to access each file specifically by sorting the yeast using antibodies against each displayed epitope tag. To access each file specifically, an anti-MYC or anti-FLAG antibody-bound Protein A magnetic bead was incubated with the yeast database. Bound yeast were magnetically extracted, grown for 48 h, and then magnetically sorted another time <ref type="bibr">[27]</ref>. File DNA was then extracted from the sorted yeast and submitted for NGS. The NGS data indicated that File2 and File3 were successfully enriched from the MYC-and FLAG-sorted populations, respectively (Figure <ref type="figure">4a</ref>). The strand abundances for the targeted files displayed normal skews (skew values = -0.35; -0.23) and mesokurtosis (excess kurtosis values = -0.77; -0.05) (Figure <ref type="figure">4b</ref>). Moreover, the strand diversity was retained at a high level with almost 100% of the unique sequences retained for File3 (Figure <ref type="figure">4c</ref>) and ~80% of the unique sequences retained for File2 (Figure <ref type="figure">S3</ref>). Both File2 and File3 were also decodable (Figure <ref type="figure">4d</ref>), while the unwanted file was not decodable. In addition, the sorting process did not introduce any additional errors (Figures <ref type="figure">4e</ref> and <ref type="figure">S4</ref>). </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4.">Combinatorial Peptide Binding Enables Multiplexed File Access</head><p>Data can be organized in more complex fashion than simple one-to-one addresses. For example, two emails may share a common label (e.g., inbox) but also retain their own unique labels as well (work vs. personal). Proteins may provide a platform of addresses to implement such overlapping labels. Here, we used histone binding proteins (human BRD2, UHRF1, and MPP8) as file addresses and biotinylated post-translationally modified histone peptides (H3K9me3, H3K27me3, and H3K14ac) to access files in a yeast database <ref type="bibr">[22,</ref><ref type="bibr">23]</ref>. We chose histone binders and peptides because of their known binding promiscuity. We created a database of three file-encoded yeast populations (Table <ref type="table">S1</ref>) and used each biotinylated histone-modified peptide to label yeast populations of interest. The MPP8 histone binding protein displayed on File3-encoded yeast binds to all three peptides, with highest affinity for the H3K9me3 peptide. The UHRF1 histone binding protein displayed on File2-encoded yeast binds to H3K9me3 primarily, with a weaker binding to the H3K14ac peptide. The BRD2 histone binding protein displayed on File1-encoded yeast binds to the H3K27me3 peptide slightly more than with the H3K14ac peptide. We then used streptavidin-PE to label the bound peptide-yeast complexes. The positively labeled yeast cells were analyzed using flow cytometry.</p><p>We found that the H3K9me3 peptide was able to label File2 and File3 from the database (Figure <ref type="figure">5</ref>). The H3K27me3 peptide labeled File1 and File3, and the H3K14ac peptide labeled all three files. These unique combinations therefore enable the use of one histone peptide to extract combinations of multiple files. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Discussion</head><p>In this work, we demonstrated that yeast-displayed proteins provide a physical address handle for robust storage and access of DNA-based data without requiring recombinant protein production or their sensitive chemical linkages to DNA. Here, we consider some potential limitations for future investigation and engineering.</p><p>Protein stability is important to consider. Aside from the obvious need to avoid protein degradation, the structure of the protein also needs to be maintained, as structure determines binding. There are many factors that affect protein stability, including pH, temperature, the existence of other nearby polar groups, and protein chaperones <ref type="bibr">[28]</ref>. Conditions favorable for yeast growth will likely place a limit on the types of proteins that can therefore be used with this platform (pH 7.5, 20-30 &#8226; C, etc.). However, one advantage of using inducible yeast surface display is that the proteins only need to be expressed and displayed during retrieval. During long-term storage, yeast can be maintained in a lyophilized form <ref type="bibr">[10,</ref><ref type="bibr">25]</ref>.</p><p>Storing DNA within cells has an inherent drawback compared to storage in vitro. The cell contains substantial "overhead" in the form of other cellular machinery that takes up volume. This can be partially addressed by increasing the payload each yeast cell can hold. Synthetic yeast genomes might be one avenue to increase the payload within each yeast, although the throughput of assembling such genomes would need to be improved <ref type="bibr">[13,</ref><ref type="bibr">29]</ref>. Each yeast cell could theoretically hold a maximum of 3 MB of information: the Saccharomyces cerevisiae yeast has a total genome size of ~12 megabases from 16 chromosomes and encodes for ~6000 genes, but ~5000 of them are nonessential <ref type="bibr">[30,</ref><ref type="bibr">31]</ref>. Therefore, the yeast genome could potentially be re-designed to eliminate dispensable, nonessential genes such that additional payload capacity is available to store information instead <ref type="bibr">[32]</ref>. One million yeast cells, which volumetrically fill less than a 6.75 &#215; 10 -5 cm 3 , could therefore hold 3 terabytes of information.</p><p>Finally, it is important to consider how a storage system using protein addresses could scale. One of the main potential advantages of proteins is how their diverse chemistry, from at least the twenty canonical amino acids to their three dimensional structure, can confer diverse molecular recognition. The prototypical example is how the human body can synthesize three billion different antibodies to bind diverse antigens <ref type="bibr">[33]</ref>. Additionally, there have been recent advances in accurately designing and predicting orthogonal protein-protein interactions de novo using structure guided and machine learning approaches <ref type="bibr">[34,</ref><ref type="bibr">35]</ref>. In addition, the protein address space can be further scaled exponentially by displaying combinations of multiple proteins on the yeast surface <ref type="bibr">[2]</ref>. The yeast surface display system described here can directly support future investigations into all of these approaches.</p><p>This work proposes the concept of using protein-ligand interactions to index and retrieve information from cells. The current implementation as presented here does not prove the proposed advantages of this approach, especially as it did not reach scales, densities, and speeds surpassing other DNA-based information storage systems. However, the theoretical diversity and programmable specificities of proteins and ligands motivates future exploration of this concept, in particular the practical challenges of scaling file organization and access. It is more speculative if densities and speeds could be increased, but there may be design choices that prioritize practical considerations of materials cost, economics, and storage stability and these should be investigated further.</p><p>The applications for this system could also extend past storing digital data in DNA. This system might also be used to store and report on information about each cell itself. For example, combinations of proteins could be used to tag and classify different cell types or the activities of different biological networks and mechanisms within the cells. These could facilitate autonomous organization of cells or improve the efficiency and scalability of omics methods including single-cell sequencing and the development of cell atlases <ref type="bibr">[36,</ref><ref type="bibr">37]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Supplementary Materials:</head><p>The following supporting information can be downloaded at <ref type="url">https:  //www.mdpi.com/article/10.3390/dna5030034/s1</ref>. Section S1: Materials used in this work (Table <ref type="table">S1</ref>. Main reagents and resources used in this work. Table <ref type="table">S2</ref>. Oligonucleotide and DNA sequences to assemble file-encoded plasmids and enrich file-encoded DNA regions. Table <ref type="table">S3</ref>. Breakdown of file diversity and protein addresses correlating to each file.); Section S2: File Plasmid Vector; Section S3: Data analysis; Section S4: Supplementary Figures (Figure <ref type="figure">S1</ref>. File encoding and decoding for image files. Figure <ref type="figure">S2</ref>. File-encoding yeast transformations preserve the sequence integrity without introducing nucleotide errors. Figure <ref type="figure">S3</ref>. The strand diversity of the File2-encoded DNA library is retained in the MYC-sorted yeast population. Figure <ref type="figure">S4</ref>. File-encoding yeast transformations preserve the sequence integrity without intro-ducing nucleotide errors in the header sequences.).</p></div></body>
		</text>
</TEI>
