Abstract Recent advancements in spatially resolved transcriptomics (SRT) technologies have enabled the comprehensive molecular and spatial characterization of single cells, providing valuable insights into the cellular organization of tissues. SRT techniques, such as single-molecule fluorescence in situ hybridization (FISH)-based methods (e.g., seqFISH, STARmap) and next-generation sequencing (NGS)-based methods (e.g., spatial transcriptomics, 10x Visium), allow for the measurement of gene expression across large populations of cells or tissue spots. These approaches generate high-dimensional data that integrate both molecular profiles and spatial context, which is crucial for understanding tissue structure and function in areas like development, neuroscience, and cancer biology. Identifying spatially variable (SV) genes, whose expression patterns differ across spatial locations, is a key step in analyzing these complex spatial transcriptomic maps. To enhance our understanding of the spatial profiles of SV genes, we propose a Bayesian nonparametric zero-inflated Poisson (ZIP) regression model for clustering these genes. Our model explicitly accounts for zero-inflation in the data, uses non-negative matrix factorization to uncover gene expression patterns, and incorporates Moran’s I (MI) basis functions to address potential confounding. Additionally, the model infers the number of clusters directly from the data, obviating the need for pre-specifying the number of clusters. We demonstrate the utility of this approach on two SRT datasets, showing that it provides more robust and interpretable clustering of SV genes, opening new avenues for understanding complex biological processes.
more »
« less
Detecting spatially co-expressed gene clusters with functional coherence by graph-regularized convolutional neural network
Abstract Motivation Clustering spatial-resolved gene expression is an essential analysis to reveal gene activities in the underlying morphological context by their functional roles. However, conventional clustering analysis does not consider gene expression co-localizations in tissue for detecting spatial expression patterns or functional relationships among the genes for biological interpretation in the spatial context. In this article, we present a convolutional neural network (CNN) regularized by the graph of protein–protein interaction (PPI) network to cluster spatially resolved gene expression. This method improves the coherence of spatial patterns and provides biological interpretation of the gene clusters in the spatial context by exploiting the spatial localization by convolution and gene functional relationships by graph-Laplacian regularization. Results In this study, we tested clustering the spatially variable genes or all expressed genes in the transcriptome in 22 Visium spatial transcriptomics datasets of different tissue sections publicly available from 10× Genomics and spatialLIBD. The results demonstrate that the PPI-regularized CNN constantly detects gene clusters with coherent spatial patterns and significantly enriched by gene functions with the state-of-the-art performance. Additional case studies on mouse kidney tissue and human breast cancer tissue suggest that the PPI-regularized CNN also detects spatially co-expressed genes to define the corresponding morphological context in the tissue with valuable insights. Availability and implementation Source code is available at https://github.com/kuanglab/CNN-PReg. Supplementary information Supplementary data are available at Bioinformatics online.
more »
« less
- Award ID(s):
- 2042159
- PAR ID:
- 10315922
- Editor(s):
- Martelli, Pier Luigi
- Date Published:
- Journal Name:
- Bioinformatics
- Volume:
- 38
- Issue:
- 5
- ISSN:
- 1367-4803
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Kimmel, Alan (Ed.)The rapid advancement of spatially resolved transcriptomics (SRT) technology enables gene expression profiling across tissue locations while preserving spatial context. Gene co-expression analysis in SRT data provides critical insights into how genes function together within the tissue microenvironment. However, existing methods fail to effectively capture the joint influence of gene–gene interactions and spatial dependencies, limiting their biological interpretability. Here, we introduce spMOCA (SPatially informed Matrix-nOrmal model for gene Co-expression Analysis), a statistical framework for inferring gene co-expression networks while explicitly modeling spatial dependencies. By leveraging a matrix-normal model, spMOCA jointly accounts for gene–gene and spatial covariance, disentangling intrinsic co-expression relationships from spatially induced effects. Through extensive simulations, we show that spMOCA provides more accurate and unbiased estimates of gene–gene correlations than existing approaches across a range of spatial dependency levels. In applications to nine SRT datasets spanning diverse technologies, tissues, and species, spMOCA consistently identifies more experimentally validated transcription factor target genes than alternative methods. In tumors, it uncovers gene modules linked to tumorigenesis and immune pathways, revealing prognostic markers. In aging mouse brains, it captures dynamic co-expression changes associated with neurodegeneration. In cross-species analyses, it detects conserved gene modules and cell type–specific pathways in the mouse and human cortex.more » « less
-
Abstract Spatial transcripome (ST) profiling can reveal cells’ structural organizations and functional roles in tissues. However, deciphering the spatial context of gene expressions in ST data is a challenge—the high-order structure hiding in whole transcriptome space over 2D/3D spatial coordinates requires modeling and detection of interpretable high-order elements and components for further functional analysis and interpretation. This paper presents a new method GraphTucker—graph-regularized Tucker tensor decomposition for learning high-order factorization in ST data. GraphTucker is based on a nonnegative Tucker decomposition algorithm regularized by a high-order graph that captures spatial relation among spots and functional relation among genes. In the experiments on several Visium and Stereo-seq datasets, the novelty and advantage of modeling multiway multilinear relationships among the components in Tucker decomposition are demonstrated as opposed to the Canonical Polyadic Decomposition and conventional matrix factorization models by evaluation of detecting spatial components of gene modules, clustering spatial coefficients for tissue segmentation and imputing complete spatial transcriptomes. The results of visualization show strong evidence that GraphTucker detect more interpretable spatial components in the context of the spatial domains in the tissues. Availability and implementationhttps://github.com/kuanglab/GraphTucker.more » « less
-
Abstract Current clustering analysis of spatial transcriptomics data primarily relies on molecular information and fails to fully exploit the morphological features present in histology images, leading to compromised accuracy and interpretability. To overcome these limitations, we have developed a multi-stage statistical method called iIMPACT. It identifies and defines histology-based spatial domains based on AI-reconstructed histology images and spatial context of gene expression measurements, and detects domain-specific differentially expressed genes. Through multiple case studies, we demonstrate iIMPACT outperforms existing methods in accuracy and interpretability and provides insights into the cellular spatial organization and landscape of functional genes within spatial transcriptomics data.more » « less
-
Kendziorski, Christina (Ed.)Abstract MotivationThe advent of next-generation sequencing-based spatially resolved transcriptomics (SRT) techniques has reshaped genomic studies by enabling high-throughput gene expression profiling while preserving spatial and morphological context. Understanding gene functions and interactions in different spatial domains is crucial, as it can enhance our comprehension of biological mechanisms, such as cancer-immune interactions and cell differentiation in various regions. It is necessary to cluster tissue regions into distinct spatial domains and identify discriminating genes (DGs) that elucidate the clustering result, referred to as spatial domain-specific DGs. Existing methods for identifying these genes typically rely on a two-stage approach, which can lead to the phenomenon known as double-dipping. ResultsTo address the challenge, we propose a unified Bayesian latent block model that simultaneously detects a list of DGs contributing to spatial domain identification while clustering these DGs and spatial locations. The efficacy of our proposed method is validated through a series of simulation experiments, and its capability to identify DGs is demonstrated through applications to benchmark SRT datasets. Availability and implementationThe R/C++ implementation of BISON is available at https://github.com/new-zbc/BISON.more » « less
An official website of the United States government

