NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Forseti : a mechanistic and predictive model of the splicing status of scRNA-seq reads

https://doi.org/10.1093/bioinformatics/btae207

He, Dongze; Gao, Yuan; Chan, Spencer_Skylar; Quintana-Parrilla, Natalia; Patro, Rob (June 2024, Bioinformatics)

Abstract MotivationShort-read single-cell RNA-sequencing (scRNA-seq) has been used to study cellular heterogeneity, cellular fate, and transcriptional dynamics. Modeling splicing dynamics in scRNA-seq data is challenging, with inherent difficulty in even the seemingly straightforward task of elucidating the splicing status of the molecules from which sequenced fragments are drawn. This difficulty arises, in part, from the limited read length and positional biases, which substantially reduce the specificity of the sequenced fragments. As a result, the splicing status of many reads in scRNA-seq is ambiguous because of a lack of definitive evidence. We are therefore in need of methods that can recover the splicing status of ambiguous reads which, in turn, can lead to more accuracy and confidence in downstream analyses. ResultsWe develop Forseti, a predictive model to probabilistically assign a splicing status to scRNA-seq reads. Our model has two key components. First, we train a binding affinity model to assign a probability that a given transcriptomic site is used in fragment generation. Second, we fit a robust fragment length distribution model that generalizes well across datasets deriving from different species and tissue types. Forseti combines these two trained models to predict the splicing status of the molecule of origin of reads by scoring putative fragments that associate each alignment of sequenced reads with proximate potential priming sites. Using both simulated and experimental data, we show that our model can precisely predict the splicing status of many reads and identify the true gene origin of multi-gene mapped reads. Availability and implementationForseti and the code used for producing the results are available at https://github.com/COMBINE-lab/forseti under a BSD 3-clause license.
more » « less
simpleaf : a simple, flexible, and scalable framework for single-cell data processing using alevin-fry

https://doi.org/10.1093/bioinformatics/btad614

He, Dongze; Patro, Rob; Ponty, ed., Yann (October 2023, Bioinformatics)

Abstract SummaryThe alevin-fry ecosystem provides a robust and growing suite of programs for single-cell data processing. However, as new single-cell technologies are introduced, as the community continues to adjust best practices for data processing, and as the alevin-fry ecosystem itself expands and grows, it is becoming increasingly important to manage the complexity of alevin-fry’s single-cell preprocessing workflows while retaining the performance and flexibility that make these tools enticing. We introduce simpleaf, a program that simplifies the processing of single-cell data using tools from the alevin-fry ecosystem, and adds new functionality and capabilities, while retaining the flexibility and performance of the underlying tools. Availability and implementationSimpleaf is written in Rust and released under a BSD 3-Clause license. It is freely available from its GitHub repository https://github.com/COMBINE-lab/simpleaf, and via bioconda. Documentation for simpleaf is available at https://simpleaf.readthedocs.io/en/latest/ and tutorials for simpleaf that have been developed can be accessed at https://combine-lab.github.io/alevin-fry-tutorials.
more » « less
SEESAW: detecting isoform-level allelic imbalance accounting for inferential uncertainty

https://doi.org/10.1186/s13059-023-03003-x

Wu, Euphy Y.; Singh, Noor P.; Choi, Kwangbom; Zakeri, Mohsen; Vincent, Matthew; Churchill, Gary A.; Ackert-Bicknell, Cheryl L.; Patro, Rob; Love, Michael I. (July 2023, Genome Biology)

Abstract Detecting allelic imbalance at the isoform level requires accounting for inferential uncertainty, caused by multi-mapping of RNA-seq reads. Our proposed method, SEESAW, uses Salmon and Swish to offer analysis at various levels of resolution, including gene, isoform, and aggregating isoforms to groups by transcription start site. The aggregation strategies strengthen the signal for transcripts with high uncertainty. The SEESAW suite of methods is shown to have higher power than other allelic imbalance methods when there is isoform-level allelic imbalance. We also introduce a new test for detecting imbalance that varies across a covariate, such as time.
more » « less
Fulgor: a fast and compact k-mer index for large-scale matching and color queries

https://doi.org/10.1186/s13015-024-00251-9

Fan, Jason; Khan, Jamshed; Singh, Noor_Pratap; Pibiri, Giulio_Ermanno; Patro, Rob (January 2024, Algorithms for Molecular Biology)

Abstract The problem of sequence identification or matching—determining the subset of reference sequences from a given collection that are likely to contain a short, queried nucleotide sequence—is relevant for many important tasks in Computational Biology, such as metagenomics and pangenome analysis. Due to the complex nature of such analyses and the large scale of the reference collections a resource-efficient solution to this problem is of utmost importance. This poses the threefold challenge of representing the reference collection with a data structure that is efficient to query, has light memory usage, and scales well to large collections. To solve this problem, we describe an efficientcolored de Bruijngraph index, arising as the combination of ak-mer dictionary with a compressed inverted index. The proposed index takes full advantage of the fact that unitigs in the colored compacted de Bruijn graph aremonochromatic(i.e., allk-mers in a unitig have the same set of references of origin, orcolor). Specifically, the unitigs are kept in the dictionary in color order, thereby allowing for the encoding of the map fromk-mers to their colors in as little as 1 +o(1) bits per unitig. Hence, one color per unitig is stored in the index with almost no space/time overhead. By combining this property with simple but effective compression methods for integer lists, the index achieves very small space. We implement these methods in a tool called , and conduct an extensive experimental analysis to demonstrate the improvement of our tool over previous solutions. For example, compared to —the strongest competitor in terms of index space vs. query time trade-off— requires significantly less space (up to 43% less space for a collection of 150,000Salmonella entericagenomes), is at least twice as fast for color queries, and is 2–6$$\times$$ $\times$ faster to construct.
more » « less
Fast, parallel, and cache-friendly suffix array construction

https://doi.org/10.1186/s13015-024-00263-5

Khan, Jamshed; Rubel, Tobias; Molloy, Erin; Dhulipala, Laxman; Patro, Rob (December 2024, Algorithms for Molecular Biology)

Free, publicly-accessible full text available December 1, 2025
Meta-colored Compacted de Bruijn Graphs

Pibiri, Giulio Ermanno; Fan, Jason; Patro, Rob (May 2024, Springer Nature)
Ma, Jian (Ed.)
Full Text Available
Iceberg Hashing: Optimizing Many Hash-Table Criteria at Once

https://doi.org/10.1145/3625817

Bender, Michael A.; Conway, Alex; Farach-Colton, Martín; Kuszmaul, William; Tagliavini, Guido (December 2023, Journal of the ACM)

Despite being one of the oldest data structures in computer science, hash tables continue to be the focus of a great deal of both theoretical and empirical research. A central reason for this is that many of the fundamental properties that one desires from a hash table are difficult to achieve simultaneously; thus many variants offering different trade-offs have been proposed. This article introduces Iceberg hashing, a hash table that simultaneously offers the strongest known guarantees on a large number of core properties. Iceberg hashing supports constant-time operations while improving on the state of the art for space efficiency, cache efficiency, and low failure probability. Iceberg hashing is also the first hash table to support a load factor of up to1 - o(1)while being stable, meaning that the position where an element is stored only ever changes when resizes occur. In fact, in the setting where keys are Θ (logn) bits, the space guarantees that Iceberg hashing offers, namely that it uses at most$\log \binom{|U|}{n} + O(n \log \ \text{log} n)$bits to storenitems from a universeU, matches a lower bound by Demaine et al. that applies to any stable hash table. Iceberg hashing introduces new general-purpose techniques for some of the most basic aspects of hash-table design. Notably, our indirection-free technique for dynamic resizing, which we call waterfall addressing, and our techniques for achieving stability and very-high probability guarantees, can be applied to any hash table that makes use of the front-yard/backyard paradigm for hash table design.
more » « less
Full Text Available
Fulgor: A Fast and Compact k-mer Index for Large-Scale Matching and Color Queries

https://doi.org/10.4230/LIPIcs.WABI.2023.18

Fan, Jason; Singh, Noor Pratap; Khan, Jamshed; Pibiri, Giulio Ermanno; Patro, Rob (August 2023, 23rd International Workshop on Algorithms in Bioinformatics (WABI 2023))
Belazzougui, Djamal; Ouangraoua, Aïda (Ed.)
The problem of sequence identification or matching - determining the subset of reference sequences from a given collection that are likely to contain a short, queried nucleotide sequence - is relevant for many important tasks in Computational Biology, such as metagenomics and pan-genome analysis. Due to the complex nature of such analyses and the large scale of the reference collections a resource-efficient solution to this problem is of utmost importance. This poses the threefold challenge of representing the reference collection with a data structure that is efficient to query, has light memory usage, and scales well to large collections. To solve this problem, we describe how recent advancements in associative, order-preserving, k-mer dictionaries can be combined with a compressed inverted index to implement a fast and compact colored de Bruijn graph data structure. This index takes full advantage of the fact that unitigs in the colored de Bruijn graph are monochromatic (all k-mers in a unitig have the same set of references of origin, or "color"), leveraging the order-preserving property of its dictionary. In fact, k-mers are kept in unitig order by the dictionary, thereby allowing for the encoding of the map from k-mers to their inverted lists in as little as 1+o(1) bits per unitig. Hence, one inverted list per unitig is stored in the index with almost no space/time overhead. By combining this property with simple but effective compression methods for inverted lists, the index achieves very small space. We implement these methods in a tool called Fulgor. Compared to Themisto, the prior state of the art, Fulgor indexes a heterogeneous collection of 30,691 bacterial genomes in 3.8× less space, a collection of 150,000 Salmonella enterica genomes in approximately 2× less space, is at least twice as fast for color queries, and is 2-6 × faster to construct.
more » « less
TreeTerminus —creating transcript trees using inferential replicate counts

https://doi.org/10.1016/j.isci.2023.106961

Singh, Noor Pratap; Love, Michael I.; Patro, Rob (June 2023, iScience)

Full Text Available
Spectrum Preserving Tilings Enable Sparse and Modular Reference Indexing

Jason Fan; Jamshed Khan; Giulio Ermanno Pibiri; Rob Patro (April 2023, RECOMB 2023: Research in Computational Molecular Biology)

Full Text Available

« Prev Next »

Search for: All records