Scalable inference and identifiability of kinetic parameters for transcriptional bursting from single cell data

Gu, Junhao; Laszik, Nandor; Miles, Christopher E (ORCID:000000015494403X); Allard, Jun (ORCID:0000000227584515); Downing, Timothy L; Read, Elizabeth L (ORCID:000000029186284X)

doi:10.1093/bioinformatics/btaf581

Abstract MotivationStochastic gene expression and cell-to-cell heterogeneity have attracted increased interest in recent years, enabled by advances in single-cell measurement technologies. These studies are also increasingly complemented by quantitative biophysical modeling, often using the framework of stochastic biochemical kinetic models. However, inferring parameters for such models (i.e., the kinetic rates of biochemical reactions) remains a technical and computational challenge, particularly doing so in a manner that can leverage high-throughput single-cell sequencing data. ResultsIn this work, we develop a chemical master equation model reference library-based computational pipeline to infer kinetic parameters describing noisy mRNA distributions from single-cell RNA sequencing data, using the commonly applied stochastic telegraph model. The approach fits kinetic parameters via steady-state distributions, as measured across a population of cells in snapshot data. Our pipeline also serves as a tool for comprehensive analysis of parameter identifiability, in both a priori (studying model properties in the absence of data) and a posteriori (in the context of a particular dataset) use-cases. The pipeline can perform both of these tasks, i.e. inference and identifiability analysis, in an efficient and scalable manner, and also serves to disentangle contributions to uncertainty in inferred parameters from experimental noise versus structural properties of the model. We found that for the telegraph model, the majority of the parameter space is not practically identifiable from single-cell RNA sequencing data, and low experimental capture rates worsen the identifiability. Our methodological framework could be extended to other data types in the fitting of small biochemical network models. Availability and implementationAll code relevant to this work is available at https://github.com/Read-Lab-UCI/TelegraphLikelihoodInfer, archival DOI: https://doi.org/10.5281/zenodo.16915450.

More Like this