NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Analysis of metagenomic data

https://doi.org/10.1038/s43586-024-00376-6

Liu, Shaopeng; Rodriguez, Judith S; Munteanu, Viorel; Ronkowski, Cynthia; Sharma, Nitesh Kumar; Alser, Mohammed; Andreace, Francesco; Blekhman, Ran; Błaszczyk, Dagmara; Chikhi, Rayan; et al (December 2025, Nature Reviews Methods Primers)

Metagenomics has revolutionized our understanding of microbial communities, offering unprecedented insights into their genetic and functional diversity across Earth’s diverse ecosystems. Beyond their roles as environmental constituents, microbiomes act as symbionts, profoundly influencing the health and function of their host organisms. Given the inherent complexity of these communities and the diverse environments where they reside, the components of a metagenomics study must be carefully tailored to yield accurate results that are representative of the populations of interest. This Primer examines the methodological advancements and current practices that have shaped the field, from initial stages of sample collection and DNA extraction to the advanced bioinformatics tools employed for data analysis, with a particular focus on the profound impact of next-generation sequencing on the scale and accuracy of metagenomics studies. We critically assess the challenges and limitations inherent in metagenomics experimentation, available technologies and computational analysis methods. Beyond technical methodologies, we explore the application of metagenomics across various domains, including human health, agriculture and environmental monitoring. Looking ahead, we advocate for the development of more robust computational frameworks and enhanced interdisciplinary collaborations. This Primer serves as a comprehensive guide for advancing the precision and applicability of metagenomic studies, positioning them to address the complexities of microbial ecology and their broader implications for human health and environmental sustainability.
more » « less
Free, publicly-accessible full text available December 1, 2026
Mixed-Up Experience Replay for Adaptive Online Condition Monitoring

https://doi.org/10.1109/TIE.2023.3260351

Russell, Matthew; Wang, Peng; Liu, Shaopeng; Jawahir, I. S. (March 2023, IEEE Transactions on Industrial Electronics)

Full Text Available
CMash: fast, multi-resolution estimation of k-mer-based Jaccard and containment indices

https://doi.org/10.1093/bioinformatics/btac237

Liu, Shaopeng; Koslicki, David (June 2022, Bioinformatics)

Abstract MotivationK-mer-based methods are used ubiquitously in the field of computational biology. However, determining the optimal value of k for a specific application often remains heuristic. Simply reconstructing a new k-mer set with another k-mer size is computationally expensive, especially in metagenomic analysis where datasets are large. Here, we introduce a hashing-based technique that leverages a kind of bottom-m sketch as well as a k-mer ternary search tree (KTST) to obtain k-mer-based similarity estimates for a range of k values. By truncating k-mers stored in a pre-built KTST with a large k=kmax value, we can simultaneously obtain k-mer-based estimates for all k values up to kmax. This truncation approach circumvents the reconstruction of new k-mer sets when changing k values, making analysis more time and space-efficient. ResultsWe derived the theoretical expression of the bias factor due to truncation. And we showed that the biases are negligible in practice: when using a KTST to estimate the containment index between a RefSeq-based microbial reference database and simulated metagenome data for 10 values of k, the running time was close to 10× faster compared to a classic MinHash approach while using less than one-fifth the space to store the data structure. Availability and implementationA python implementation of this method, CMash, is available at https://github.com/dkoslicki/CMash. The reproduction of all experiments presented herein can be accessed via https://github.com/KoslickiLab/CMASH-reproducibles. Supplementary informationSupplementary data are available at Bioinformatics online.
more » « less
Feature Space Augmentation for Long-Tailed Data

Chu, Peng; Bian, Xiao; Liu, Shaopeng; Ling, Haibin (August 2020, European Conf. on Computer Vision (ECCV))

Full Text Available
A fog computing-based framework for process monitoring and prognosis in cyber-manufacturing

https://doi.org/10.1016/j.jmsy.2017.02.011

Wu, Dazhong; Liu, Shaopeng; Zhang, Li; Terpenny, Janis; Gao, Robert X.; Kurfess, Thomas; Guzzo, Judith A. (April 2017, Journal of Manufacturing Systems)

Full Text Available

Search for: All records