<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>APA-Scan: Detection and Visualization of 3’-UTR APA with RNA-seq and 3’-end-seq Data</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>2020</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10181548</idno>
					<idno type="doi">10.1101/2020.02.16.951657</idno>
					<title level='j'>BioRxiv</title>
<idno></idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Naima Ahmed Fahmi</author><author>Jae-Woong Chang</author><author>Heba Nassereddeen</author><author>Khandakar Tanvir Ahmed</author><author>Deliang Fan</author><author>Jeongsik Yong</author><author>Wei Zhang</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[The eukaryotic genome is capable of producing multiple isoforms from a gene by alternative polyadenylation (APA) during pre-mRNA processing. APA in the 3’-untranslated region (3’-UTR) of mRNA produces transcripts with shorter 3’-UTR. Often, 3’-UTR serves as a binding platform for microRNAs and RNA-binding proteins, which affect the fate of the mRNA transcript. Thus, 3’-UTR APA provides a means to regulate gene expression at the post-transcriptional level and is known to promote translation. Current bioinformatics pipelines have limited capability in profiling 3’-UTR APA events due to incomplete annotations and a low-resolution analyzing power: widely available bioinformatics pipelines do not reference actionable polyadenylation (cleavage) sites but simulate 3’-UTR APA only using RNA-seq read coverage, causing false positive identifications. To overcome these limitations, we developed APA-Scan, a robust program that identifies 3’-UTR APA events and visualizes the RNA-seq short-read coverage with gene annotations. APA-Scan utilizes either predicted or experimentally validated actionable polyadenylation signals as a reference for polyadenylation sites and calculates the quantity of long and short 3’-UTR transcripts in the RNA-seq data. The performance of APA-Scan was validated by qPCR.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Implementation: APA-Scan is implemented in Python. Source code and a comprehensive user's manual are freely available at <ref type="url">https://github.com/compbiolabucf/ APA-Scan</ref> 1</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>Poly(A)-tails are added to pre-mRNA after the polyadenylation signal (PAS) during the 3'-end processing of pre-mRNA <ref type="bibr">[5]</ref>. The last exon of mRNA contains a non-coding region, 3'-untranslated region (3'-UTR), which spans from the termination codon to the polyadenylation site. 3'-UTR is a molecular scaffold for binding to microRNAs and RNA-binding proteins and functions in regulatory gene expression <ref type="bibr">[6]</ref>. In human and mouse, more than 70% of genes contain multiple PASs in their 3'-UTRs and APA using upstream PASs leads to the production of mRNA with shortened 3'-UTRs (UTR-APA) <ref type="bibr">[1]</ref>. UTR-APA is known to increase the efficiency of translation and is associated with T-cell activation, oncogene activation, and poor prognosis in many cancers <ref type="bibr">[4]</ref>.</p><p>Several bioinformatics pipelines are available for the analysis of UTR-APA using RNA-seq data <ref type="bibr">[8,</ref><ref type="bibr">7,</ref><ref type="bibr">2]</ref>. In general, all these methods measure the changes of 3'-UTR length by modeling the RNA-seq read density changes near the 3'end of mRNAs. Indeed, with the aid of these methods, RNA-seq experiments became a powerful approach to investigate UTR-APA. In many cases, however, the identified APA sites are not functionally and physiologically relevant because most pipelines do not reference actionable PASs in their UTR-APA simulation. Moreover, none of the existing pipelines can provide high-resolution read coverage plots of the APA events with an accurate annotation. We have developed APA-Scan, a bioinformatics program for the detection and visualization of genome-wide UTR-APA events. APA-Scan integrates both 3'-end-seq (an RNA-seq method with a specific enrichment of 3'-ends of mRNA) data and the location information of predicted canonical PASs with RNA-seq data to improve the quantitative definition of genome-wide UTR-APA events. APA-Scan efficiently manages large-scale alignment files and generates a comprehensive report for UTR-APA events. It is also advantageous in producing high quality plots of APA events.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Methods</head><p>APA-Scan comprises of three steps: (i) read coverage estimation; (ii) identification of polyadenylation sites and the calculation of APA; (iii) graphic illustration of UTR-APA events (Figure <ref type="figure">1</ref>). In the first step, APA-Scan takes aligned RNAseq and 3'-end-seq data in the BAM format as an input to estimate the read coverage on 3'-UTR exons and identify potential polyadenylation sites. The read coverage files are generated by SAMtools <ref type="bibr">[3]</ref>. In this step, the 3'-end-seq data is an optional input.</p><p>In the second step, all aligned reads from 3'-end-seq data are pooled together to identify peaks and the corresponding cleavage sites in 3'-UTRs, as shown in Figure <ref type="figure">1</ref>. Identified peaks in the 3'-end-seq data are considered potential cleavage and polyadenylation sites. If the 3'-end-seq data is not provided by the user, predicted PASs (AATAAA, ATTAAA) in 3'-UTRs are considered as potential cleavage sites. Next, to determine potential 3'-UTR APA events AATAAA ATTAAA</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>PAS</head><p>Step 1</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Read coverage estimation</head><p>Step 2:</p><p>Step 3: Visualization between two biological contexts (or samples), APA-Scan evaluates each experimentally proven or predicted cleavage site in the 3'-UTR of a transcript using &#967; 2 -test: it contrasts the RNA-seq short reads covering up and downstream of the candidate cleavage site between the two samples and calculates the mean coverage upstream of the site (N 1 and N 2 ) and downstream of the site (n 1 and n 2 ) as shown in at the bottom panel in Figure <ref type="figure">1</ref>, with (N 1 , n 1 ) denoting the coverage in the first sample, and (N 2 , n 2 ) denoting the coverage in the second sample. Then, the canonical 2 x 2 &#967; 2 -test is applied to report the p-value for each candidate site. All the identified events will be reported in an Excel file.</p><p>In the third step, based on the significance of 3'-UTR APA events calculated in the second step, APA-Scan can generate RNA-seq and 3'-end-seq (if provided) coverage plots with the 3'-UTR annotation for one or more user-specific events. In this step, users may specify the region of the genome locus to generate the read alignment plot. An example of this task is illustrated at the bottom panel of Figure <ref type="figure">1</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Experimental results</head><p>To validate the analysis results by APA-Scan, we conducted qPCR experiments for Srsf3 and Rpl22 transcripts from WT (wild type) and Tsc1-/-mouse embryonic fibroblasts (MEFs) based on the significant 3'-UTR APA events reported by APA-Scan. As shown in Supplementary Figure <ref type="figure">1</ref>, both Srsf3 and Rpl22 showed the increase of the short 3'-UTR transcript by APA in Tsc1-/-compared to WT MEFs, which is consistent with our observations on the RNA-seq read coverage plots. These results further confirm that APA-Scan can identify the true 3'-UTR APA events with RNA-seq and 3'-end-seq samples from two different biological contexts.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Materials and Methods</head><p>Realtime quantitative PCR (RT-qPCR) analysis and primer sequences: Total RNAs from TSC1 WT or TSC1-/-MEF cells were isolated by Trizol method according to manufacturer's protocol <ref type="url">https://assets.thermofisher. com/TFS-Assets/LSG/manuals/trizol_reagent.pdf</ref>.</p><p>Reverse transcription reaction using Oligo-d(T) priming and NxGen M-MuLV Reverse transcriptase (Lucigen) was carried out according to the manufacturer's protocol <ref type="url">https://www.lucigen.com/docs/manuals/MA115-M-MuLV. pdf</ref>. SYBR Green was used to detect and quantitate the PCR products in real-time reactions. Quantitation of the real-time PCR results was done using standard curve method for accuracy and reliability of the analysis. The primer sequences used to measure the RSI for each transcript are as follows: mRpl22 Total forward 5'-AAGTTCAC CCTGGACTGC AC-3' mRpl22 Total reverse 5'-GTGATCTT GCTCTTGCTG CG-3' The level of total, short 3'-UTR, and long 3'-UTR transcripts from Srsf3 and Rpl22 was measured by qPCR. Because it is not possible to design specific primers for the qPCR analysis of short 3'-UTR transcript, the amount of short 3'-UTR transcripts were calculated by subtracting the quantity of long 3'-UTR transcripts from total. mRPL22 Long Forward 5'-TGGGCATC TGGGCTTTTA GG-3' mRPL22 Long reverse 5'-GCTTGTTGCA GACTTGCTCA-3' mSRSF3 Total forward 5'-GCTGCCGTGTAAGAGTGGAA-3' mSRSF3 Total reverse 5'-AGGACTCCTCCTGCGGTAAT-3' mSRSF3 Long forward 5'-TGCAACAGTCTTGTGGCTTA-3' mSRSF3 Long reverse 5'-TGCAATGGCTCTTACATAGACC-3'</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Conclusion</head><p>APA-Scan offers a computational pipeline to identify transcriptome-wide 3'-UTR APA events. By integrating RNA-seq data and PAS information (experimentally verified or computationally predicted), APA-Scan can generate a comprehensive report of significant APA events and the illustration of their read coverage plots. The wet-lab approaches using qPCR experiments demonstrate that APA-Scan provides high-accuracy and quantitative profiling of 3'-UTR APA events.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>User Manual</head><p>Download APA-Scan is downloadable directly from github. Users need to have python (version 3.0 or higher) installed in their machine.  -p/-P P denotes whether the user gives the 3'-end-seq or not. If -p is initialized, the next two fields after -p will be the directories of 3' end data for two samples. If -p is not specified, APA-Scan will automatically determine APA events according to its algorithm.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Required Softwares</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>APA-Scan.py Results</head><p>APA-Scan will generate a spreadsheet in the output directory, with the following name:</p><p>&#8226; Result PAS.csv [ if the user provides the PAS data]</p><p>&#8226; Result.csv [ if only RNA-seq input is provided], which contains the potential transcript splice site for each region. APA-Scan will also generate some intermediary files in the output directory for reference purpose to the users.</p><p>The Result.csv [or Result PAS.csv] file will contain the following fields (see image below) as long as all other information necessary to compute the association among two samples.</p><p>Run Make-plots.py </p></div></body>
		</text>
</TEI>
