Abstract MotivationPolymerase chain reaction (PCR) enables rapid, cost-effective diagnostics but requires prior identification of genomic regions that allow sensitive and specific detection of target microbial groups, herein referred to as microbial signature sequences. We introduce Seqwin, an open-source framework designed to automate microbial genome signature discovery. Tens of thousands of microbial genomes are now available for a single species, limiting the application of existing manual and automated approaches for identifying signatures. Modern approaches that are capable of leveraging all available microbial genomes will ensure sensitive and accurate DNA signature identification and enable robust pathogen detection for clinical, environmental, and public health applications. ResultsSeqwin builds weighted pan-genome minimizer graphs and uses a traversal algorithm to identify signature sequences that occur frequently in target genomes but remain rare in non-targets. Unlike earlier tools that depend on strict presence or absence of sequences, Seqwin accommodates natural sequence variation and scales to very large genome collections. When applied to genomes from C. difficile, M. tuberculosis, and S. enterica, Seqwin recovered more high-quality signatures than alternative methods with lower computational burden. Seqwin’s analysis of nearly 15,000 S. enterica genomes yielded over 200 candidate signatures in 5 minutes. Seqwin provides an open-source solution for the long-standing need for scalable microbial signature discovery and diagnostic assay design. Availability and ImplementationSeqwin is freely available for academic use (https://github.com/treangenlab/Seqwin) and can be installed via Bioconda. Benchmarking datasets, outputs, and scripts are available on Zenodohttps://doi.org/10.5281/zenodo.19176444. Contacttreangen@rice.edu,xw66@rice.edu Supplementary MaterialsProvided as separate PDF and data files.
more »
« less
iSeqSearch: incremental protein search for iBlast/iMMSeqs2/iDiamond
BackgroundThe advancement of sequencing technology has led to a rapid increase in the amount of DNA and protein sequence data; consequently, the size of genomic and proteomic databases is constantly growing. As a result, database searches need to be continually updated to account for the new data being added. However, continually re-searching the entire existing dataset wastes resources. Incremental database search can address this problem. MethodsOne recently introduced incremental search method is iBlast, which wraps the BLAST sequence search method with an algorithm to reuse previously processed data and thereby increase search efficiency. The iBlast wrapper, however, must be generalized to support better performing DNA/protein sequence search methods that have been developed, namely MMseqs2 and Diamond. To address this need, we propose iSeqsSearch, which extends iBlast by incorporating support for MMseqs2 (iMMseqs2) and Diamond (iDiamond), thereby providing a more generalized and broadly effective incremental search framework. Moreover, the previously published iBlast wrapper has to be revised to be more robust and usable by the general community. ResultsiMMseqs2 and iDiamond, which apply the incremental approach, perform nearly identical to MMseqs2 and Diamond. Notably, when comparing ranking comparison methods such as the Pearson correlation, we observe a high concordance of over 0.9, indicating similar results. Moreover, in some cases, our incremental approach, iSeqsSearch, which extends the iBlast merge function to iMMseqs2 and iDiamond, provides more hits compared to the conventional MMseqs2 and Diamond methods. ConclusionThe incremental approach using iMMseqs2 and iDiamond demonstrates efficiency in terms of reusing previously processed data while maintaining high accuracy and concordance in search results. This method can reduce resource waste in continually growing genomic and proteomic database searches. The sample codes and data are available at GitHub and Zenodo (https://github.com/EESI/Incremental-Protein-Search; DOI:10.5281/zenodo.14675319).
more »
« less
- PAR ID:
- 10586592
- Publisher / Repository:
- PeerJ
- Date Published:
- Journal Name:
- PeerJ
- Volume:
- 13
- ISSN:
- 2167-8359
- Page Range / eLocation ID:
- e19171
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Abstract BackgroundComputational cell type deconvolution enables the estimation of cell type abundance from bulk tissues and is important for understanding tissue microenviroment, especially in tumor tissues. With rapid development of deconvolution methods, many benchmarking studies have been published aiming for a comprehensive evaluation for these methods. Benchmarking studies rely on cell-type resolved single-cell RNA-seq data to create simulated pseudobulk datasets by adding individual cells-types in controlled proportions. ResultsIn our work, we show that the standard application of this approach, which uses randomly selected single cells, regardless of the intrinsic difference between them, generates synthetic bulk expression values that lack appropriate biological variance. We demonstrate why and how the current bulk simulation pipeline with random cells is unrealistic and propose a heterogeneous simulation strategy as a solution. The heterogeneously simulated bulk samples match up with the variance observed in real bulk datasets and therefore provide concrete benefits for benchmarking in several ways. We demonstrate that conceptual classes of deconvolution methods differ dramatically in their robustness to heterogeneity with reference-free methods performing particularly poorly. For regression-based methods, the heterogeneous simulation provides an explicit framework to disentangle the contributions of reference construction and regression methods to performance. Finally, we perform an extensive benchmark of diverse methods across eight different datasets and find BayesPrism and a hybrid MuSiC/CIBERSORTx approach to be the top performers. ConclusionsOur heterogeneous bulk simulation method and the entire benchmarking framework is implemented in a user friendly packagehttps://github.com/humengying0907/deconvBenchmarkingandhttps://doi.org/10.5281/zenodo.8206516, enabling further developments in deconvolution methods.more » « less
-
Abstract While space-borne optical and near-infrared facilities have succeeded in delivering a precise and spatially resolved picture of our Universe, their small survey area is known to underrepresent the true diversity of galaxy populations. Ground-based surveys have reached comparable depths but at lower spatial resolution, resulting in source confusion that hampers accurate photometry extractions. What once was limited to the infrared regime has now begun to challenge ground-based ultradeep surveys, affecting detection and photometry alike. Failing to address these challenges will mean forfeiting a representative view into the distant Universe. We introduceThe Farmer: an automated, reproducible profile-fitting photometry package that pairs a library of smooth parametric models fromThe Tractorwith a decision tree that determines the best-fit model in concert with neighboring sources. Photometry is measured by fitting the models on other bands leaving brightness free to vary. The resulting photometric measurements are naturally total, and no aperture corrections are required. Supporting diagnostics (e.g.,χ2) enable measurement validation. As fitting models is relatively time intensive,The Farmeris built with high-performance computing routines. We benchmarkThe Farmeron a set of realistic COSMOS-like images and find accurate photometry, number counts, and galaxy shapes.The Farmeris already being utilized to produce catalogs for several large-area deep extragalactic surveys where it has been shown to tackle some of the most challenging optical and near-infrared data available, with the promise of extending to other ultradeep surveys expected in the near future.The Farmeris available to download from GitHub (https://github.com/astroweaver/the_farmer) and Zenodo (https://doi.org/10.5281/zenodo.8205817).more » « less
-
dadi-cli: Automated and distributed population genetic model inference from allele frequency spectraAbstract Summarydadi is a popular software package for inferring models of demographic history and natural selection from population genomic data. But using dadi requires Python scripting and manual parallelization of optimization jobs. We developed dadi-cli to simplify dadi usage and also enable straighforward distributed computing. Availability and Implementationdadi-cli is implemented in Python and released under the Apache License 2.0. The source code is available athttps://github.com/xin-huang/dadi-cli. dadi-cli can be installed via PyPI and conda, and is also available through Cacao on Jetstream2https://cacao.jetstream-cloud.org/.more » « less
-
Abstract We review a theoretical framework for the cuprate superconductors, rooted in a fractionalized Fermi liquid (FL*) description of the intermediate-temperature pseudogap phase at low doping. The FL* theory predicted hole pockets each of fractional area at hole dopingp, in contrast to the area in spin density wave theory. Magnetotransport measurements, including observation of the Yamaji angle, show clear evidence of hole pocket quasiparticles which can tunnel coherently between square lattice layers, and are consistent with the FL* description. The FL* phase of a single-band model is described using a layer construction with a pair of ancilla qubits on each site: the Ancilla layer model (ALM). Its mean field theory yields hole pockets of area , and matches the gapped photoemission spectrum in the anti-nodal region of the Brillouin zone. Fluctuations are described by the SU(2) gauge theory of a background spin liquid with critical Dirac spinons. A Monte Carlo study of the thermal SU(2) gauge theory transforms the hole pockets into Fermi arcs in photoemission. One route to confinement of FL* upon lowering temperature yields ad-wave superconductor via a Kosterlitz–Thouless transition of vortices, with nodal Bogoliubov quasiparticles featuring anisotropic velocities and vortices surrounded by charge order halos. An alternative route yields a charge-ordered metallic state that has quantum oscillations consistent with observations. These confinement transitions are driven by the condensation of a SU(2) fundamental Higgs field, which also provides a fractionalized description of intertwined orders. Increasing doping from the FL* phase in the ALM drives a transition to a conventional FL at large doping, passing through an intermediate strange metal regime. We formulate a theory of the FL*-FL metal-metal transition without a symmetry-breaking order parameter, using a critical quantum ‘charge’ liquid of mobile electrons in the presence of disorder, developed via an extension of the Sachdev–Ye–Kitaev model to two spatial dimensions. At low temperatures, and across optimal and over doping, we address the regimes of extended non-FL behavior by Griffiths effects near quantum phase transitions in disordered metals. Partly based on lectures by S S atBoulder School 2025, Dynamics of Strongly Correlated Electrons, 14–18 July.Lecture videos.Joint ICTP-WE Heraeus School and Workshop on Advances in Quantum Matter: Pushing the Boundaries, ICTP, Trieste, 4, 6 August 2025.Lecture videos.School on Quantum Dynamics of Matter, Light and Information, ICTP, Trieste, 18, 19 August 2025.Lecture videos.Croucher Advanced Study Institute for Fractional Chern Insulators, University of Hong Kong, 4, 5 September 2025.Lecture slides.Advanced School and Conference on Quantum Matter, ICTP Trieste, 1–12 December 2025.Lecture Notes.Lecture videos.more » « less
An official website of the United States government

