Title: A new Bayesian factor analysis method improves detection of genes and biological processes affected by perturbations in single-cell CRISPR screening
Abstract Clustered regularly interspaced short palindromic repeats (CRISPR) screening coupled with single-cell RNA sequencing has emerged as a powerful tool to characterize the effects of genetic perturbations on the whole transcriptome at a single-cell level. However, due to its sparsity and complex structure, analysis of single-cell CRISPR screening data is challenging. In particular, standard differential expression analysis methods are often underpowered to detect genes affected by CRISPR perturbations. We developed a statistical method for such data, called guided sparse factor analysis (GSFA). GSFA infers latent factors that represent coregulated genes or gene modules; by borrowing information from these factors, it infers the effects of genetic perturbations on individual genes. We demonstrated through extensive simulation studies that GSFA detects perturbation effects with much higher power than state-of-the-art methods. Using single-cell CRISPR data from human CD8+T cells and neural progenitor cells, we showed that GSFA identified biologically relevant gene modules and specific genes affected by CRISPR perturbations, many of which were missed by existing methods, providing new insights into the functions of genes involved in T cell activation and neurodevelopment.  more » « less
Award ID(s):
2016307
PAR ID:
10469287
Author(s) / Creator(s):
; ; ; ;
Publisher / Repository:
Springer Nature
Date Published:
Journal Name:
Nature Methods
ISSN:
1548-7091
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Immune dysfunction in cancer is enacted by multiple programs, including tumor cell-intrinsic responses to distinct immune subpopulations. A subset of these immune evasion programs can be systematically recapitulated through direct tumor-immune interactionsin vitro. Here, we present an integrated, high-throughput single-cell CRISPR screening framework focused on the protein kinome for mapping the tumor-intrinsic regulation of T cell-driven immune pressure in glioblastoma (GBM). We combine pooled CRISPR interference and activation (CRISPRi/a) with immune-matched NY-ESO-1 antigen-specific allogeneic GBM-T cell co-culture and massively multiplexed single-cell transcriptomics to systematically quantify how genetic perturbation reshapes baseline tumor state and adaptive responses across graded effector-to-target ratios. We further leverage deep generative models for analyzing pooled CRISPR screens to decipher the effects of genetic perturbations on the mechanisms of tumor resistance. This framework resolves distinct modules of immune evasion and survival, including the regulation of the antigen-presentation machinery, interferon/NF-κB signaling, oxidative stress resilience, and checkpoint/cytokine programs, while identifying perturbations that reroute the continuous tumor transcriptional trajectory induced by T cell engagement. A secondary chemical screen in patient-derived GBM cultures identified putative kinase targets of immune evasion phenotypes (e.g., EPHA2 and PDGFRA), whose inhibition leads to the blockade of evasive programs and enhances T cell-mediated GBM killing. Together, this workflow provides a scalable blueprint for comprehensive charting of the genetic control of tumor-immune interactions. 
    more » « less
  2. Abstract Single-cell CRISPR screens (perturb-seq) link genetic perturbations to phenotypic changes in individual cells. The most fundamental task in perturb-seq analysis is to test for association between a perturbation and a count outcome, such as gene expression. We conduct the first-ever comprehensive benchmarking study of association testing methods for low multiplicity-of-infection (MOI) perturb-seq data, finding that existing methods produce excess false positives. We conduct an extensive empirical investigation of the data, identifying three core analysis challenges: sparsity, confounding, and model misspecification. Finally, we develop an association testing method — SCEPTRE low-MOI — that resolves these analysis challenges and demonstrates improved calibration and power. 
    more » « less
  3. Abstract With the growing number of single-cell datasets collected under more complex experimental conditions, there is an opportunity to leverage single-cell variability to reveal deeper insights into how cells respond to perturbations. Many existing approaches rely on discretizing the data into clusters for differential gene expression (DGE), effectively ironing out any information unveiled by the single-cell variability across cell-types. In addition, DGE often assumes a statistical distribution that, if erroneous, can lead to false positive differentially expressed genes. Here, we present Cellograph: a semi-supervised framework that uses graph neural networks to quantify the effects of perturbations at single-cell granularity. Cellograph not only measures how prototypical cells are of each condition but also learns a latent space that is amenable to interpretable data visualization and clustering. The learned gene weight matrix from training reveals pertinent genes driving the differences between conditions. We demonstrate the utility of our approach on publicly-available datasets including cancer drug therapy, stem cell reprogramming, and organoid differentiation. Cellograph outperforms existing methods for quantifying the effects of experimental perturbations and offers a novel framework to analyze single-cell data using deep learning. 
    more » « less
  4. Summary CRISPR genome engineering and single-cell RNA sequencing have accelerated biological discovery. Single-cell CRISPR screens unite these two technologies, linking genetic perturbations in individual cells to changes in gene expression and illuminating regulatory networks underlying diseases. Despite their promise, single-cell CRISPR screens present considerable statistical challenges. We demonstrate through theoretical and real data analyses that a standard method for estimation and inference in single-cell CRISPR screens—“thresholded regression”—exhibits attenuation bias and a bias-variance tradeoff as a function of an intrinsic, challenging-to-select tuning parameter. To overcome these difficulties, we introduce GLM-EIV (“GLM-based errors-in-variables”), a new method for single-cell CRISPR screen analysis. GLM-EIV extends the classical errors-in-variables model to responses and noisy predictors that are exponential family-distributed and potentially impacted by the same set of confounding variables. We develop a computational infrastructure to deploy GLM-EIV across hundreds of processors on clouds (e.g. Microsoft Azure) and high-performance clusters. Leveraging this infrastructure, we apply GLM-EIV to analyze two recent, large-scale, single-cell CRISPR screen datasets, yielding several new insights. 
    more » « less
  5. Abstract Lineage tracing technology using CRISPR/Cas9 genome editing has enabled simultaneous readouts of gene expressions and lineage barcodes in single cells, which allows for inference of cell lineage and cell types at the whole organism level. While most state-of-the-art methods for lineage reconstruction utilize only the lineage barcode data, methods that incorporate gene expressions are emerging. Effectively incorporating the gene expression data requires a reasonable model of how gene expression data changes along generations of divisions. Here, we present LinRace (Lineage Reconstruction with asymmetric cell division model), which integrates lineage barcode and gene expression data using asymmetric cell division model and infers cell lineages and ancestral cell states using Neighbor-Joining and maximum-likelihood heuristics. On both simulated and real data, LinRace outputs more accurate cell division trees than existing methods. With inferred ancestral states, LinRace can also show how a progenitor cell generates a large population of cells with various functionalities. 
    more » « less