skip to main content


Search for: All records

Creators/Authors contains: "Zhou, Xiaobo"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Abstract

    Splicing factors (SFs) are the major RNA-binding proteins (RBPs) and key molecules that regulate the splicing of mRNA molecules through binding to mRNAs. The expression of splicing factors is frequently deregulated in different cancer types, causing the generation of oncogenic proteins involved in cancer hallmarks. In this study, we investigated the genes that encode RNA-binding proteins and identified potential splicing factors that contribute to the aberrant splicing applying a random forest classification model. The result suggested 56 splicing factors were related to the prognosis of 13 cancers, two SF complexes in liver hepatocellular carcinoma, and one SF complex in esophageal carcinoma. Further systematic bioinformatics studies on these cancer prognostic splicing factors and their related alternative splicing events revealed the potential regulations in a cancer-specific manner. Our analysis found high ILF2-ILF3 expression correlates with poor prognosis in LIHC through alternative splicing. These findings emphasize the importance of SFs as potential indicators for prognosis or targets for therapeutic interventions. Their roles in cancer exhibit complexity and are contingent upon the specific context in which they operate. This recognition further underscores the need for a comprehensive understanding and exploration of the role of SFs in different types of cancer, paving the way for their potential utilization in prognostic assessments and the development of targeted therapies.

     
    more » « less
  2. Abstract Objective

    The early stages of chronic disease typically progress slowly, so symptoms are usually only noticed until the disease is advanced. Slow progression and heterogeneous manifestations make it challenging to model the transition from normal to disease status. As patient conditions are only observed at discrete timestamps with varying intervals, an incomplete understanding of disease progression and heterogeneity affects clinical practice and drug development.

    Materials and Methods

    We developed the Gaussian Process for Stage Inference (GPSI) approach to uncover chronic disease progression patterns and assess the dynamic contribution of clinical features. We tested the ability of the GPSI to reliably stratify synthetic and real-world data for osteoarthritis (OA) in the Osteoarthritis Initiative (OAI), bipolar disorder (BP) in the Adolescent Brain Cognitive Development Study (ABCD), and hepatocellular carcinoma (HCC) in the UTHealth and The Cancer Genome Atlas (TCGA).

    Results

    First, GPSI identified two subgroups of OA based on image features, where these subgroups corresponded to different genotypes, indicating the bone-remodeling and overweight-related pathways. Second, GPSI differentiated BP into two distinct developmental patterns and defined the contribution of specific brain region atrophy from early to advanced disease stages, demonstrating the ability of the GPSI to identify diagnostic subgroups. Third, HCC progression patterns were well reproduced in the two independent UTHealth and TCGA datasets.

    Conclusion

    Our study demonstrated that an unsupervised approach can disentangle temporal and phenotypic heterogeneity and identify population subgroups with common patterns of disease progression. Based on the differences in these features across stages, physicians can better tailor treatment plans and medications to individual patients.

     
    more » « less
  3. Abstract

    The COVID-19 pandemic, caused by the coronavirus SARS-CoV-2, has resulted in the loss of millions of lives and severe global economic consequences. Every time SARS-CoV-2 replicates, the viruses acquire new mutations in their genomes. Mutations in SARS-CoV-2 genomes led to increased transmissibility, severe disease outcomes, evasion of the immune response, changes in clinical manifestations and reducing the efficacy of vaccines or treatments. To date, the multiple resources provide lists of detected mutations without key functional annotations. There is a lack of research examining the relationship between mutations and various factors such as disease severity, pathogenicity, patient age, patient gender, cross-species transmission, viral immune escape, immune response level, viral transmission capability, viral evolution, host adaptability, viral protein structure, viral protein function, viral protein stability and concurrent mutations. Deep understanding the relationship between mutation sites and these factors is crucial for advancing our knowledge of SARS-CoV-2 and for developing effective responses. To fill this gap, we built COV2Var, a function annotation database of SARS-CoV-2 genetic variation, available at http://biomedbdc.wchscu.cn/COV2Var/. COV2Var aims to identify common mutations in SARS-CoV-2 variants and assess their effects, providing a valuable resource for intensive functional annotations of common mutations among SARS-CoV-2 variants.

     
    more » « less
  4. Abstract

    Drug resistance poses a significant challenge in cancer treatment. Despite the initial effectiveness of therapies such as chemotherapy, targeted therapy and immunotherapy, many patients eventually develop resistance. To gain deep insights into the underlying mechanisms, single-cell profiling has been performed to interrogate drug resistance at cell level. Herein, we have built the DRMref database (https://ccsm.uth.edu/DRMref/) to provide comprehensive characterization of drug resistance using single-cell data from drug treatment settings. The current version of DRMref includes 42 single-cell datasets from 30 studies, covering 382 samples, 13 major cancer types, 26 cancer subtypes, 35 treatment regimens and 42 drugs. All datasets in DRMref are browsable and searchable, with detailed annotations provided. Meanwhile, DRMref includes analyses of cellular composition, intratumoral heterogeneity, epithelial–mesenchymal transition, cell–cell interaction and differentially expressed genes in resistant cells. Notably, DRMref investigates the drug resistance mechanisms (e.g. Aberration of Drug’s Therapeutic Target, Drug Inactivation by Structure Modification, etc.) in resistant cells. Additional enrichment analysis of hallmark/KEGG (Kyoto Encyclopedia of Genes and Genomes)/GO (Gene Ontology) pathways, as well as the identification of microRNA, motif and transcription factors involved in resistant cells, is provided in DRMref for user’s exploration. Overall, DRMref serves as a unique single-cell-based resource for studying drug resistance, drug combination therapy and discovering novel drug targets.

     
    more » « less
  5. Abstract

    StemDriver is a comprehensive knowledgebase dedicated to the functional annotation of genes participating in the determination of hematopoietic stem cell fate, available at http://biomedbdc.wchscu.cn/StemDriver/. By utilizing single-cell RNA sequencing data, StemDriver has successfully assembled a comprehensive lineage map of hematopoiesis, capturing the entire continuum from the initial formation of hematopoietic stem cells to the fully developed mature cells. Extensive exploration and characterization were conducted on gene expression features corresponding to each lineage commitment. At the current version, StemDriver integrates data from 42 studies, encompassing a diverse range of 14 tissue types spanning from the embryonic phase to adulthood. In order to ensure uniformity and reliability, all data undergo a standardized pipeline, which includes quality data pre-processing, cell type annotation, differential gene expression analysis, identification of gene categories correlated with differentiation, analysis of highly variable genes along pseudo-time, and exploration of gene expression regulatory networks. In total, StemDriver assessed the function of 23 839 genes for human samples and 29 533 genes for mouse samples. Simultaneously, StemDriver also provided users with reference datasets and models for cell annotation. We believe that StemDriver will offer valuable assistance to research focused on cellular development and hematopoiesis.

     
    more » « less
  6. Abstract

    Aging entails gradual functional decline influenced by interconnected factors. Multiple hallmarks proposed as common and conserved underlying denominators of aging on the molecular, cellular and systemic levels across multiple species. Thus, understanding the function of aging hallmarks and their relationships across species can facilitate the translation of anti-aging drug development from model organisms to humans. Here, we built AgeAnnoMO (https://relab.xidian.edu.cn/AgeAnnoMO/#/), a knowledgebase of multi-omics annotation for animal aging. AgeAnnoMO encompasses an extensive collection of 136 datasets from eight modalities, encompassing 8596 samples from 50 representative species, making it a comprehensive resource for aging and longevity research. AgeAnnoMO characterizes multiple aging regulators across species via multi-omics data, comprehensively annotating aging-related genes, proteins, metabolites, mitochondrial genes, microbiotas and age-specific TCR and BCR sequences tied to aging hallmarks for these species and tissues. AgeAnnoMO not only facilitates a deeper and more generalizable understanding of aging mechanisms, but also provides potential insights of the specificity across tissues and species in aging process, which is important to develop the effective anti-aging interventions for diverse populations. We anticipate that AgeAnnoMO will provide a valuable resource for comprehending and integrating the conserved driving hallmarks in aging biology and identifying the targetable biomarkers for aging research.

     
    more » « less
  7. Abstract

    Combination therapy is a promising strategy for confronting the complexity of cancer. However, experimental exploration of the vast space of potential drug combinations is costly and unfeasible. Therefore, computational methods for predicting drug synergy are much needed for narrowing down this space, especially when examining new cellular contexts. Here, we thus introduce CCSynergy, a flexible, context aware and integrative deep-learning framework that we have established to unleash the potential of the Chemical Checker extended drug bioactivity profiles for the purpose of drug synergy prediction. We have shown that CCSynergy enables predictions of superior accuracy, remarkable robustness and improved context generalizability as compared to the state-of-the-art methods in the field. Having established the potential of CCSynergy for generating experimentally validated predictions, we next exhaustively explored the untested drug combination space. This resulted in a compendium of potentially synergistic drug combinations on hundreds of cancer cell lines, which can guide future experimental screens.

     
    more » « less
  8. Abstract

    The coronavirus disease of 2019 pandemic has catalyzed the rapid development of mRNA vaccines, whereas, how to optimize the mRNA sequence of exogenous gene such as severe acute respiratory syndrome coronavirus 2 spike to fit human cells remains a critical challenge. A new algorithm, iDRO (integrated deep-learning-based mRNA optimization), is developed to optimize multiple components of mRNA sequences based on given amino acid sequences of target protein. Considering the biological constraints, we divided iDRO into two steps: open reading frame (ORF) optimization and 5′ untranslated region (UTR) and 3′UTR generation. In ORF optimization, BiLSTM-CRF (bidirectional long-short-term memory with conditional random field) is employed to determine the codon for each amino acid. In UTR generation, RNA-Bart (bidirectional auto-regressive transformer) is proposed to output the corresponding UTR. The results show that the optimized sequences of exogenous genes acquired the pattern of human endogenous gene sequence. In experimental validation, the mRNA sequence optimized by our method, compared with conventional method, shows higher protein expression. To the best of our knowledge, this is the first study by introducing deep-learning methods to integrated mRNA sequence optimization, and these results may contribute to the development of mRNA therapeutics.

     
    more » « less
  9. Blockchain relies on the underlying peer-to-peer (P2P) networking to broadcast and get up-to-date on the blocks and transactions. Because of the blockchain operations’ reliance on the information provided by P2P networking, it is imperative to have high P2P connectivity for the quality of the blockchain system operations and performances. High P2P networking connectivity ensures that a peer node is connected to multiple other peers providing a diverse set of observers of the current state of the blockchain and transactions. However, in a permissionless Bitcoin cryptocurrency network, using the peer identifiers – including the current approach of counting the number of distinct IP addresses and port numbers – can be ineffective in measuring the number of peer connections and estimating the networking connectivity. Such current approach is further challenged by the networking threats manipulating identities. We build a robust estimation engine for the P2P networking connectivity by sensing and processing the P2P networking traffic. We take a systematic approach to study our engine and analyze the followings: the different components of the connectivity estimation engine and how they affect the accuracy performances, the role and the effectiveness of an outlier detection to enhance the connectivity estimation, and the engine’s interplay with the Bitcoin protocol. We implement a working Bitcoin prototype connected to the Bitcoin mainnet to validate and improve our engine’s performances and evaluate the estimation accuracy and cost efficiency of our connectivity estimation engine. Our results show that our scheme effectively counters the identity-manipulations threats, achieves 96.4% estimation accuracy with a tolerance of one peer connection, and is lightweight in the overheads in the mining rate, thus making it appropriate for the miner deployment. 
    more » « less