Search for: All records

Creators/Authors contains: "Liu, Yufang"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available January 1, 2027
  2. Abstract Multimodal integration combines information from different sources or modalities to gain a more comprehensive understanding of a phenomenon. The challenges in multi-omics data analysis lie in the complexity, high dimensionality, and heterogeneity of the data, which demands sophisticated computational tools and visualization methods for proper interpretation and visualization of multi-omics data. In this paper, we propose a novel method, termed Orthogonal Multimodality Integration and Clustering (OMIC), for analyzing CITE-seq. Our approach enables researchers to integrate multiple sources of information while accounting for the dependence among them. We demonstrate the effectiveness of our approach using CITE-seq data sets for cell clustering. Our results show that our approach outperforms existing methods in terms of accuracy, computational efficiency, and interpretability. We conclude that our proposed OMIC method provides a powerful tool for multimodal data analysis that greatly improves the feasibility and reliability of integrated data. 
    more » « less
  3. Summary High‐quality genome of rosemary (Salvia rosmarinus) represents a valuable resource and tool for understanding genome evolution and environmental adaptation as well as its genetic improvement. However, the existing rosemary genome did not provide insights into the relationship between antioxidant components and environmental adaptability. In this study, by employing Nanopore sequencing and Hi‐C technologies, a total of 1.17 Gb (97.96%) genome sequences were mapped to 12 chromosomes with 46 121 protein‐coding genes and 1265 non‐coding RNA genes. Comparative genome analysis reveals that rosemary had a closely genetic relationship withSalvia splendensandSalvia miltiorrhiza, and it diverged from them approximately 33.7 million years ago (MYA), and one whole‐genome duplication occurred around 28.3 MYA in rosemary genome. Among all identified rosemary genes, 1918 gene families were expanded, 35 of which are involved in the biosynthesis of antioxidant components. These expanded gene families enhance the ability of rosemary adaptation to adverse environments. Multi‐omics (integrated transcriptome and metabolome) analysis showed the tissue‐specific distribution of antioxidant components related to environmental adaptation. During the drought, heat and salt stress treatments, 36 genes in the biosynthesis pathways of carnosic acid, rosmarinic acid and flavonoids were up‐regulated, illustrating the important role of these antioxidant components in responding to abiotic stresses by adjusting ROS homeostasis. Moreover, cooperating with the photosynthesis, substance and energy metabolism, protein and ion balance, the collaborative system maintained cell stability and improved the ability of rosemary against harsh environment. This study provides a genomic data platform for gene discovery and precision breeding in rosemary. Our results also provide new insights into the adaptive evolution of rosemary and the contribution of antioxidant components in resistance to harsh environments. 
    more » « less
  4. Abstract With the rapid advancements in large language model technology and the emergence of bioinformatics‐specific language models (BioLMs), there is a growing need for a comprehensive analysis of the current landscape, computational characteristics, and diverse applications. This survey aims to address this need by providing a thorough review of BioLMs, focusing on their evolution, classification, and distinguishing features, alongside a detailed examination of training methodologies, datasets, and evaluation frameworks. We explore the wide‐ranging applications of BioLMs in critical areas such as disease diagnosis, drug discovery, and vaccine development, highlighting their impact and transformative potential in bioinformatics. We identify key challenges and limitations inherent in BioLMs, including data privacy and security concerns, interpretability issues, biases in training data and model outputs, and domain adaptation complexities. Finally, we highlight emerging trends and future directions, offering valuable insights to guide researchers and clinicians toward advancing BioLMs for increasingly sophisticated biological and clinical applications. 
    more » « less
    Free, publicly-accessible full text available March 1, 2027