Title: LncRNA Subcellular Localization Across Diverse Cell Lines: An Exploration Using Deep Learning with Inexact q-mers
Background: Long non-coding Ribonucleic Acids (lncRNAs) can be localized to different cellular compartments, such as the nuclear and the cytoplasmic regions. Their biological functions are influenced by the region of the cell where they are located. Compared to the vast number of lncRNAs, only a relatively small proportion have annotations regarding their subcellular localization. It would be helpful if those few annotated lncRNAs could be leveraged to develop predictive models for localization of other lncRNAs. Methods: Conventional computational methods use q-mer profiles from lncRNA sequences and train machine learning models such as support vector machines and logistic regression with the profiles. These methods focus on the exact q-mer. Given possible sequence mutations and other uncertainties in genomic sequences and their role in biological function, a consideration of these variabilities might improve our ability to model lncRNAs and their localization. Thus, we build on inexact q-mers and use machine learning/deep learning techniques to study three specific problems in lncRNA subcellular localization, namely, prediction of lncRNA localization using inexact q-mers, the issue of whether lncRNA localization is cell-type-specific, and the notion of switching (lncRNA) genes. Results: We performed our analysis using data on lncRNA localization across 15 cell lines. Our results showed that using inexact q-mers (with q = 6) can improve the lncRNA localization prediction performance compared to using exact q-mers. Further, we showed that lncRNA localization, in general, is not cell-line-specific. We also identified a category of LncRNAs which switch cellular compartments between different cell lines (we call them switching lncRNAs). These switching lncRNAs complicate the problem of predicting lncRNA localization using machine learning models, showing that lncRNA localization is still a major challenge.  more » « less
Award ID(s):
2125872 1920920
PAR ID:
10634928
Author(s) / Creator(s):
; ; ;
Publisher / Repository:
MDPI
Date Published:
Journal Name:
Non-Coding RNA
Volume:
11
Issue:
4
ISSN:
2311-553X
Page Range / eLocation ID:
49
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Long noncoding RNAs (lncRNAs), a class of noncoding RNAs exceeding 500 nucleotides and transcribed mostly by RNA polymerase II, regulate gene transcription and chromatin organization. Acting as molecular guides, scaffolds, and structural components, lncRNAs interact with RNAs, DNA, and proteins. Repeat motifs within lncRNAs called k-mers are associated with protein interactions and chromatin complex recruitment. Understanding lncRNA mechanisms is hampered by tissue specificity, redundancy, and multifunctionality, necessitating quantitative investigation. To support the systematic investigation of repeat motifs, we optimized Golden Gate assembly with standardized 4-nucleotide linkers to assemble any four lncRNA elements into 4x k-mers without the need to design compatible overhangs. To extend the length (repeat number) of RNA constructs, a 2x BbsI Golden Gate cloning site is restored at the 3’ end of the assembled 4x k-mer, allowing the addition of more k-mers to generate an 8x k-mer, 12x k-mer, and so on. Please read the Guidelines for suggestions on how to make arrays of intermediate lengths (e.g. 1x, 2x, 3x, 5x, 6x, 7x, etc.). Beginning with k-mer identification, synthetic lncRNAs can be constructed and expressed in vitro or in cells within three weeks, offering an efficient framework to study and harness their regulatory functions. 
    more » « less
  2. Abstract The lncATLAS database quantifies the relative cytoplasmic versus nuclear abundance of long non-coding RNAs (lncRNAs) observed in 15 human cell lines. The literature describes several machine learning models trained and evaluated on these and similar datasets. These reports showed moderate performance, e.g. 72–74% accuracy, on test subsets of the data withheld from training. In all these reports, the datasets were filtered to include genes with extreme values while excluding genes with values in the middle range and the filters were applied prior to partitioning the data into training and testing subsets. Using several models and lncATLAS data, we show that this ‘middle exclusion’ protocol boosts performance metrics without boosting model performance on unfiltered test data. We show that various models achieve only about 60% accuracy when evaluated on unfiltered lncRNA data. We suggest that the problem of predicting lncRNA subcellular localization from nucleotide sequences is more challenging than currently perceived. We provide a basic model and evaluation procedure as a benchmark for future studies of this problem. 
    more » « less
  3. SEquence Evaluation throughk-mer Representation (SEEKR) is a method of sequence comparison that uses sequence substrings calledk-mers to quantify the nonlinear similarity between nucleic acid species. We describe the development of new functions within SEEKR that enable end-users to estimateP-values that ascribe statistical significance to SEEKR-derived similarities, as well as visualize different aspects ofk-mer similarity. We apply the new functions to identify chromatin-enriched lncRNAs that containXIST-like sequence features, and we demonstrate the utility of applying SEEKR on lncRNA fragments to identify potential RNA-protein interaction domains. We also highlight ways in which SEEKR can be applied to augment studies of lncRNA conservation, and we outline the best practice of visualizing RNA-seq read density to evaluate support for lncRNA annotations before their in-depth study in cell types of interest. 
    more » « less
  4. Abstract Intron retention (IR) is increasingly recognized as a feature of long noncoding RNAs (lncRNAs), yet the mechanisms that shape IR in lncRNAs and the functional consequences of this process remain largely unexplored. To investigate how IR contributes to lncRNA regulation, we performed a genome-wide screen to identify factors controlling IR in the lncRNAPURPL. This approach uncovered a prominent role for U2AF2, which promotes retention of a specific intron inPURPLthrough a weak polypyrimidine tract. IR of this intron drives nuclear enrichment ofPURPLand enhances cell proliferation, revealing biological relevance. Transcriptome-wide analyses showed that although U2AF2 broadly supports canonical splicing consistent with its well-established function in promoting splicing, it also facilitates IR within a distinct subset of RNAs, including the nuclear speckle–associated lncRNAMALAT1. Loss of U2AF2 disruptsMALAT1speckle localization and usingMALAT1knockout cells reconstituted with wild-type or intron deleted variants, we identified a single intron critical forMALAT1’s speckle localization. Deletion of this intron from endogenousMALAT1impaired speckle localization and reduced cell migration, phenocopying the loss ofMALAT1. Together, these findings reveal IR as a key regulatory mechanism governing lncRNA localization and function and uncover an unexpected role for U2AF2 in promoting IR within specific lncRNA contexts. 
    more » « less
  5. Long noncoding RNA (lncRNA) plays key roles in tumorigenesis. Misexpression of lncRNA can lead to changes in expression profiles of various target genes, which are involved in cancer initiation and progression. So, identifying key lncRNAs for a cancer would help develop the cancer therapy. Usually, to identify key lncRNAs for a cancer, expression profiles of lncRNAs for normal and cancer samples are required. But, this kind of data are not available for all cancers. In the present study, a computational framework is developed to identify cancer specific key lncRNAs using the lncRNA expression of cancer patients only. The framework consists of two state-of-the-art feature selection techniques - Recursive Feature Elimination (RFE) and Least Absolute Shrinkage and Selection Operator (LASSO); and five machine learning models - Naive Bayes, K-Nearest Neighbor, Random Forest, Support Vector Machine, and Deep Neural Network. For experiment, expression values of lncRNAs for 8 cancers - BLCA, CESC, COAD, HNSC, KIRP, LGG, LIHC, and LUAD - from TCGA are used. The combined dataset consists of 3,656 patients with expression values of 12,309 lncRNAs. Important features or key lncRNAs are identified by using feature selection algorithms RFE and LASSO. Capability of these key lncRNAs in classifying 8 different cancers is checked by the performance of five classification models. This study identified 37 key lncRNAs that can classify 8 different cancer types with an accuracy ranging from 94% to 97%. Finally, survival analysis supports that the discovered key lncRNAs are capable of differentiating between high-risk and low-risk patients. 
    more » « less