Search for: All records

Editors contains: "Mathelier, Anthony"

« Prev Next »

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Scalable inference and identifiability of kinetic parameters for transcriptional bursting from single cell data

https://doi.org/10.1093/bioinformatics/btaf581

Gu, Junhao; Laszik, Nandor; Miles, Christopher E; Allard, Jun; Downing, Timothy L; Read, Elizabeth L (November 2025, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract MotivationStochastic gene expression and cell-to-cell heterogeneity have attracted increased interest in recent years, enabled by advances in single-cell measurement technologies. These studies are also increasingly complemented by quantitative biophysical modeling, often using the framework of stochastic biochemical kinetic models. However, inferring parameters for such models (i.e., the kinetic rates of biochemical reactions) remains a technical and computational challenge, particularly doing so in a manner that can leverage high-throughput single-cell sequencing data. ResultsIn this work, we develop a chemical master equation model reference library-based computational pipeline to infer kinetic parameters describing noisy mRNA distributions from single-cell RNA sequencing data, using the commonly applied stochastic telegraph model. The approach fits kinetic parameters via steady-state distributions, as measured across a population of cells in snapshot data. Our pipeline also serves as a tool for comprehensive analysis of parameter identifiability, in both a priori (studying model properties in the absence of data) and a posteriori (in the context of a particular dataset) use-cases. The pipeline can perform both of these tasks, i.e. inference and identifiability analysis, in an efficient and scalable manner, and also serves to disentangle contributions to uncertainty in inferred parameters from experimental noise versus structural properties of the model. We found that for the telegraph model, the majority of the parameter space is not practically identifiable from single-cell RNA sequencing data, and low experimental capture rates worsen the identifiability. Our methodological framework could be extended to other data types in the fitting of small biochemical network models. Availability and implementationAll code relevant to this work is available at https://github.com/Read-Lab-UCI/TelegraphLikelihoodInfer, archival DOI: https://doi.org/10.5281/zenodo.16915450.
more » « less
Free, publicly-accessible full text available November 1, 2026
Nanopore decoding with speed and versatility for data storage

https://doi.org/10.1093/bioinformatics/btaf006

Volkel, Kevin D; Hook, Paul W; Keung, Albert; Timp, Winston; Tuck, James M (December 2024, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract MotivationAs nanopore technology reaches ever higher throughput and accuracy, it becomes an increasingly viable candidate for reading out DNA data storage. Nanopore sequencing offers considerable flexibility by allowing long reads, real-time signal analysis, and the ability to read both DNA and RNA. We need flexible and efficient designs that match nanopore’s capabilities, but relatively few designs have been explored and many have significant inefficiency in read density, error rate, or compute time. To address these problems, we designed a new single-read per-strand decoder that achieves low byte error rates, offers high throughput, scales to long reads, and works well for both DNA and RNA molecules. We achieve these results through a novel soft decoding algorithm that can be effectively parallelized on a GPU. Our faster decoder allows us to study a wider range of system designs. ResultsWe demonstrate our approach on HEDGES, a state-of-the-art DNA-constrained convolutional code. We implement one hard decoder that runs serially and two soft decoders that run on GPUs. Our evaluation for each decoder is applied to the same population of nanopore reads collected from a synthesized library of strands. These same strands are synthesized with a T7 promoter to enable RNA transcription and decoding. Our results show that the hard decoder has a byte error rate over 25%, while the prior state of the art soft decoder can achieve error rates of 2.25%. However, that design also suffers a low throughput of 183 s/read. Our new Alignment Matrix Trellis soft decoder improves throughput by 257× with the trade-off of a higher byte error rate of 3.52% compared to the state of the art. Furthermore, we use the faster speed of our algorithm to explore more design options. We show that read densities of 0.33 bits/base can be achieved, which is 4× larger than prior MSA-based decoders. We also compare RNA to DNA, and find that RNA has 85% as many error-free reads when compared to DNA. Availability and implementationSource code for our soft decoder and data used to generate figures is available publicly in the Github repository https://github.com/dna-storage/hedges-soft-decoder (10.5281/zenodo.11454877). All raw FAST5/FASTQ data are available at 10.5281/zenodo.11985454 and 10.5281/zenodo.12014515.
more » « less
Full Text Available
JIND: joint integration and discrimination for automated single-cell annotation

https://doi.org/10.1093/bioinformatics/btac140

Goyal, Mohit; Serrano, Guillermo; Argemi, Josepmaria; Shomorony, Ilan; Hernaez, Mikel; Ochoa, Idoia (March 2022, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Motivation An important step in the transcriptomic analysis of individual cells involves manually determining the cellular identities. To ease this labor-intensive annotation of cell-types, there has been a growing interest in automated cell annotation, which can be achieved by training classification algorithms on previously annotated datasets. Existing pipelines employ dataset integration methods to remove potential batch effects between source (annotated) and target (unannotated) datasets. However, the integration and classification steps are usually independent of each other and performed by different tools. We propose JIND (joint integration and discrimination for automated single-cell annotation), a neural-network-based framework for automated cell-type identification that performs integration in a space suitably chosen to facilitate cell classification. To account for batch effects, JIND performs a novel asymmetric alignment in which unseen cells are mapped onto the previously learned latent space, avoiding the need of retraining the classification model for new datasets. JIND also learns cell-type-specific confidence thresholds to identify cells that cannot be reliably classified. Results We show on several batched datasets that the joint approach to integration and classification of JIND outperforms in accuracy existing pipelines, and a smaller fraction of cells is rejected as unlabeled as a result of the cell-specific confidence thresholds. Moreover, we investigate cells misclassified by JIND and provide evidence suggesting that they could be due to outliers in the annotated datasets or errors in the original approach used for annotation of the target batch. Availability and implementation Implementation for JIND is available at https://github.com/mohit1997/JIND and the data underlying this article can be accessed at https://doi.org/10.5281/zenodo.6246322. Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available
VeloViz: RNA velocity-informed embeddings for visualizing cellular trajectories

https://doi.org/10.1093/bioinformatics/btab653

Atta, Lyla; Sahoo, Arpan; Fan, Jean (September 2021, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Motivation Single-cell transcriptomics profiling technologies enable genome-wide gene expression measurements in individual cells but can currently only provide a static snapshot of cellular transcriptional states. RNA velocity analysis can help infer cell state changes using such single-cell transcriptomics data. To interpret these cell state changes inferred from RNA velocity analysis as part of underlying cellular trajectories, current approaches rely on visualization with principal components, t-distributed stochastic neighbor embedding and other 2D embeddings derived from the observed single-cell transcriptional states. However, these 2D embeddings can yield different representations of the underlying cellular trajectories, hindering the interpretation of cell state changes. Results We developed VeloViz to create RNA velocity-informed 2D and 3D embeddings from single-cell transcriptomics data. Using both real and simulated data, we demonstrate that VeloViz embeddings are able to capture underlying cellular trajectories across diverse trajectory topologies, even when intermediate cell states may be missing. By considering the predicted future transcriptional states from RNA velocity analysis, VeloViz can help visualize a more reliable representation of underlying cellular trajectories. Availability and implementation Source code is available on GitHub (https://github.com/JEFworks-Lab/veloviz) and Bioconductor (https://bioconductor.org/packages/veloviz) with additional tutorials at https://JEF.works/veloviz/. Datasets used can be found on Zenodo (https://doi.org/10.5281/zenodo.4632471). Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available
RVAgene: generative modeling of gene expression time series data

https://doi.org/10.1093/bioinformatics/btab260

Mitra, Raktim; MacLean, Adam L (May 2021, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Motivation Methods to model dynamic changes in gene expression at a genome-wide level are not currently sufficient for large (temporally rich or single-cell) datasets. Variational autoencoders offer means to characterize large datasets and have been used effectively to characterize features of single-cell datasets. Here, we extend these methods for use with gene expression time series data. Results We present RVAgene: a recurrent variational autoencoder to model gene expression dynamics. RVAgene learns to accurately and efficiently reconstruct temporal gene profiles. It also learns a low dimensional representation of the data via a recurrent encoder network that can be used for biological feature discovery, and from which we can generate new gene expression data by sampling the latent space. We test RVAgene on simulated and real biological datasets, including embryonic stem cell differentiation and kidney injury response dynamics. In all cases, RVAgene accurately reconstructed complex gene expression temporal profiles. Via cross validation, we show that a low-error latent space representation can be learnt using only a fraction of the data. Through clustering and gene ontology term enrichment analysis on the latent space, we demonstrate the potential of RVAgene for unsupervised discovery. In particular, RVAgene identifies new programs of shared gene regulation of Lox family genes in response to kidney injury. Availability and implementation All datasets analyzed in this manuscript are publicly available and have been published previously. RVAgene is available in Python, at GitHub: https://github.com/maclean-lab/RVAgene; Zenodo archive: http://doi.org/10.5281/zenodo.4271097. Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available
Clustering single-cell RNA-seq data by rank constrained similarity learning

https://doi.org/10.1093/bioinformatics/btab276

Mei, Qinglin; Li, Guojun; Su, Zhengchang (May 2021, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Motivation Recent breakthroughs of single-cell RNA sequencing (scRNA-seq) technologies offer an exciting opportunity to identify heterogeneous cell types in complex tissues. However, the unavoidable biological noise and technical artifacts in scRNA-seq data as well as the high dimensionality of expression vectors make the problem highly challenging. Consequently, although numerous tools have been developed, their accuracy remains to be improved. Results Here, we introduce a novel clustering algorithm and tool RCSL (Rank Constrained Similarity Learning) to accurately identify various cell types using scRNA-seq data from a complex tissue. RCSL considers both local similarity and global similarity among the cells to discern the subtle differences among cells of the same type as well as larger differences among cells of different types. RCSL uses Spearman’s rank correlations of a cell’s expression vector with those of other cells to measure its global similarity, and adaptively learns neighbor representation of a cell as its local similarity. The overall similarity of a cell to other cells is a linear combination of its global similarity and local similarity. RCSL automatically estimates the number of cell types defined in the similarity matrix, and identifies them by constructing a block-diagonal matrix, such that its distance to the similarity matrix is minimized. Each block-diagonal submatrix is a cell cluster/type, corresponding to a connected component in the cognate similarity graph. When tested on 16 benchmark scRNA-seq datasets in which the cell types are well-annotated, RCSL substantially outperformed six state-of-the-art methods in accuracy and robustness as measured by three metrics. Availability and implementation The RCSL algorithm is implemented in R and can be freely downloaded at https://cran.r-project.org/web/packages/RCSL/index.html. Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available
Identification of cell-type-specific marker genes from co-expression patterns in tissue samples

https://doi.org/10.1093/bioinformatics/btab257

Qiu, Yixuan; Wang, Jiebiao; Lei, Jing; Roeder, Kathryn (April 2021, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Motivation Marker genes, defined as genes that are expressed primarily in a single-cell type, can be identified from the single-cell transcriptome; however, such data are not always available for the many uses of marker genes, such as deconvolution of bulk tissue. Marker genes for a cell type, however, are highly correlated in bulk data, because their expression levels depend primarily on the proportion of that cell type in the samples. Therefore, when many tissue samples are analyzed, it is possible to identify these marker genes from the correlation pattern. Results To capitalize on this pattern, we develop a new algorithm to detect marker genes by combining published information about likely marker genes with bulk transcriptome data in the form of a semi-supervised algorithm. The algorithm then exploits the correlation structure of the bulk data to refine the published marker genes by adding or removing genes from the list. Availability and implementation We implement this method as an R package markerpen, hosted on CRAN (https://CRAN.R-project.org/package=markerpen). Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available
COBRAC: a fast implementation of convex biclustering with compression

https://doi.org/10.1093/bioinformatics/btab248

Yi, Haidong; Huang, Le; Mishne, Gal; Chi, Eric C (April 2021, Bioinformatics)
Mathelier, Anthony (Ed.)
Abstract Biclustering is a generalization of clustering used to identify simultaneous grouping patterns in observations (rows) and features (columns) of a data matrix. Recently, the biclustering task has been formulated as a convex optimization problem. While this convex recasting of the problem has attractive properties, existing algorithms do not scale well. To address this problem and make convex biclustering a practical tool for analyzing larger data, we propose an implementation of fast convex biclustering called COBRAC to reduce the computing time by iteratively compressing problem size along the solution path. We apply COBRAC to several gene expression datasets to demonstrate its effectiveness and efficiency. Besides the standalone version for COBRAC, we also developed a related online web server for online calculation and visualization of the downloadable interactive results. Availability The source code and test data are available at https://github.com/haidyi/cvxbiclustr or https://zenodo.org/record/4620218. The web server is available at https://cvxbiclustr.ericchi.com. Supplementary information Supplementary data are available at Bioinformatics online.
more » « less
Full Text Available