This content will become publicly available on October 31, 2026

Title: COTree: A Statistical Framework for Deciphering Cell-Resolved Multi-Omics Trajectories
Abstract Recent advances in whole-cell modeling enable the computational tracking of the temporal evolution of thousands of molecular species across genomic, transcriptomic, proteomic, and metabolomic layers. These models provide a complementary perspective for studying cellular dynamics, offering continuous, system-wide observations that are difficult to obtain from experimental technologies, which are often destructive and yield only static measurements from limited modalities. While whole-cell models generate multi-omic simulation trajectories with high temporal resolution, analyzing and interpreting such complex data remains a major challenge that limits their potential to elucidate cellular dynamics. To address this challenge, we propose COTree, a statistical framework that learns integrated multi-omic representations and constructs a trajectory principal tree to summarize cellular progression patterns. COTree enables a broad range of downstream analyses, including cell classification, fate prediction, developmental time detection, and driver species identification, that provide new insights into how cells develop and differentiate. To demonstrate its practical utility, we apply COTree to a multi-omic trajectory dataset generated from the whole-cell model of JCVI-Syn3A, revealing cell types, characterizing long-term cellular dynamics, and identifying key driver species associated with cell death and replication.  more » « less
Award ID(s):
2243257
PAR ID:
10686250
Author(s) / Creator(s):
; ; ; ; ; ;
Publisher / Repository:
bioRxiv
Date Published:
Format(s):
Medium: X
Institution:
bioRxiv
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract In single cell biology, the complexity of tissues may hinder lineage cell mapping or tumor microenvironment decomposition, requiring digital dissociation of bulk tissues. Many deconvolution methods focus on transcriptomic assay, not easily applicable to other omics due to ambiguous cell markers and reference-to-target difference. Here, we present MOADE, a multimodal autoencoder pipeline linking multi-dimensional features to jointly predict personalized multi-omic profiles and cellular compositions, using pseudo-bulk data constructed by internal non-transcriptomic reference and external scRNA-seq data. MOADE is evaluated through rigorous simulation experiments and real multi-omic data from multiple tissue types, outperforming nine deconvolution pipelines with superior generalizability and fidelity. 
    more » « less
  2. Beiko, Robert G (Ed.)
    ABSTRACT Inflammatory bowel disease (IBD) is characterized by complex etiology and a disrupted colonic ecosystem. We provide a framework for the analysis of multi-omic data, which we apply to study the gut ecosystem in IBD. Specifically, we train and validate models using data on the metagenome, metatranscriptome, virome, and metabolome from the Human Microbiome Project 2 IBD multi-omic database, with 1,785 repeated samples from 130 individuals (103 cases and 27 controls). After splitting the participants into training and testing groups, we used mixed-effects least absolute shrinkage and selection operator regression to select features for each omic. These features, with demographic covariates, were used to generate separate single-omic prediction scores. All four single-omic scores were then combined into a final regression to assess the relative importance of the individual omics and the predictive benefits when considered together. We identified several species, pathways, and metabolites known to be associated with IBD risk, and we explored the connections between data sets. Individually, metabolomic and viromic scores were more predictive than metagenomics or metatranscriptomics, and when all four scores were combined, we predicted disease diagnosis with a Nagelkerke’sR2of 0.46 and an area under the curve of 0.80 (95% confidence interval: 0.63, 0.98). Our work supports that some single-omic models for complex traits are more predictive than others, that incorporating multiple omic data sets may improve prediction, and that each omic data type provides a combination of unique and redundant information. This modeling framework can be extended to other complex traits and multi-omic data sets. IMPORTANCEComplex traits are characterized by many biological and environmental factors, such that multi-omic data sets are well-positioned to help us understand their underlying etiologies. We applied a prediction framework across multiple omics (metagenomics, metatranscriptomics, metabolomics, and viromics) from the gut ecosystem to predict inflammatory bowel disease (IBD) diagnosis. The predicted scores from our models highlighted key features and allowed us to compare the relative utility of each omic data set in single-omic versus multi-omic models. Our results emphasized the importance of metabolomics and viromics over metagenomics and metatranscriptomics for predicting IBD status. The greater predictive capability of metabolomics and viromics is likely because these omics serve as markers of lifestyle factors such as diet. This study provides a modeling framework for multi-omic data, and our results show the utility of combining multiple omic data types to disentangle complex disease etiologies and biological signatures. 
    more » « less
  3. Abstract As microbiome research has progressed, it has become clear that most, if not all, eukaryotic organisms are hosts to microbiomes composed of prokaryotes, other eukaryotes, and viruses. Fungi have only recently been considered holobionts with their own microbiomes, as filamentous fungi have been found to harbor bacteria (including cyanobacteria), mycoviruses, other fungi, and whole algal cells within their hyphae. Constituents of this complex endohyphal microbiome have been interrogated using multi-omic approaches. However, a lack of tools, techniques, and standardization for integrative multi-omics for small-scale microbiomes (e.g., intracellular microbiomes) has limited progress towards investigating and understanding the total diversity of the endohyphal microbiome and its functional impacts on fungal hosts. Understanding microbiome impacts on fungal hosts will advance explorations of how “microbiomes within microbiomes” affect broader microbial community dynamics and ecological functions. Progress to date as well as ongoing challenges of performing integrative multi-omics on the endohyphal microbiome is discussed herein. Addressing the challenges associated with the sample extraction, sample preparation, multi-omic data generation, and multi-omic data analysis and integration will help advance current knowledge of the endohyphal microbiome and provide a road map for shrinking microbiome investigations to smaller scales. 
    more » « less
  4. BackgroundForecasting the responses of natural populations to environmental change is a key priority in the management of ecological systems. This is challenging because the dynamics of multi-species ecological communities are influenced by many factors. Populations can exhibit complex, nonlinear responses to environmental change, often over multiple temporal lags. In addition, biotic interactions, and other sources of multi-species dependence, are major contributors to patterns of population variation. Theory suggests that near-term ecological forecasts of population abundances can be improved by modelling these dependencies, but empirical support for this idea is lacking. MethodsWe test whether models that learn from multiple species, both to estimate nonlinear environmental effects and temporal interactions, improve ecological forecasts compared to simpler single species models for a semi-arid rodent community. Using dynamic generalized additive models, we analyze time series of monthly captures for nine rodent species over 25 years. ResultsModel comparisons provide strong evidence that multi-species dependencies improve both hindcast and forecast performance, as models that captured these effects gave superior predictions than models that ignored them. We show that changes in abundance for some species can have delayed, nonlinear effects on others, and that lagged, nonlinear effects of temperature and vegetation greenness are key drivers of changes in abundance for this system. ConclusionsOur findings highlight that multivariate models are useful not only to improve near-term ecological forecasts but also to ask targeted questions about ecological interactions and drivers of change. This study emphasizes the importance of jointly modelling species’ shared responses to the environment and their delayed temporal interactions when teasing apart community dynamics. 
    more » « less
  5. Abstract Recently, lineage tracing technology using CRISPR/Cas9 genome editing has enabled simultaneous readouts of gene expressions and lineage barcodes, which allows for the reconstruction of the cell division tree and makes it possible to reconstruct ancestral cell types and trace the origin of each cell type. Meanwhile, trajectory inference methods are widely used to infer cell trajectories and pseudotime in a dynamic process using gene expression data of present-day cells. Here, we present TedSim (single-cell temporal dynamics simulator), which simulates the cell division events from the root cell to present-day cells, simultaneously generating two data modalities for each single cell: the lineage barcode and gene expression data. TedSim is a framework that connects the two problems: lineage tracing and trajectory inference. Using TedSim, we conducted analysis to show that (i) TedSim generates realistic gene expression and barcode data, as well as realistic relationships between these two data modalities; (ii) trajectory inference methods can recover the underlying cell state transition mechanism with balanced cell type compositions; and (iii) integrating gene expression and barcode data can provide more insights into the temporal dynamics in cell differentiation compared to using only one type of data, but better integration methods need to be developed. 
    more » « less