Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Free, publicly-accessible full text available June 1, 2027
-
Abstract Microbiome research faces two central challenges, namely constructing reliable networks, where nodes represent microbial taxa and edges represent their associations, and identifying significant disease-associated taxa. To address the first challenge, we developed CMIMN, a novel R package that applies a Bayesian network framework based on conditional mutual information to infer microbial interaction networks. To further enhance reliability, we construct a consensus microbiome network by integrating results from CMIMN and three widely used methods, including Sparse Inverse Covariance Estimation for Ecological Association Inference (SPIEC-EASI), Semi-Parametric Rank-based correlation and partial correlation Estimation (SPRING), and Sparse Correlations for Compositional Data (SPARCC). This consensus approach, which overlays and weights edges shared across methods, reduces inconsistencies and provides a more biologically meaningful view of microbial relationships. To address the second challenge, we designed a multi-method feature selection framework that combines machine learning with network-based strategies. Our machine learning pipeline applies distinct algorithms and identifies key taxa based on their consistent importance across models. Complementing this, we employ two network-based strategies that prioritize taxa based on centrality differences between networks constructed from healthy samples and disease-affected samples, as well as a composite scoring system that ranks nodes using integrated network metrics. We applied CMIMN on soil microbiome data from potato fields affected by common scab disease. Bootstrap analysis confirmed the robustness of CMIMN, and the consensus network further improved stability and interpretability. The multi-method framework enhances confidence in identifying soil microbial taxa associated with potato disease. Notably, we identified Bacteroidota, WPS-2, and Proteobacteria at the Phylum level; Actinobacteria, AD3, Bacilli, Anaerolineae, and Ktedonobacteria at the Class level; and C0119, Defluviicoccales, Bacteroidales, and Ktedonobacterales at the Order level as key taxa associated with disease status.more » « lessFree, publicly-accessible full text available January 1, 2027
-
Abstract Microbial networks offer critical insights into community structure, ecological interactions and host–microbe dynamics. However, constructing reliable microbiome networks remains challenging due to variability among existing inference methods, limited overlap between inferred networks and the absence of a gold standard (a universally accepted reference for benchmarking) for validation.We developedCMiNet, an R package and interactive Shiny App(https://cminet.wid.wisc.edu) that enables consensus microbiome network construction by integrating up to 10 widely used inference algorithms.CMiNetsupports both correlation‐based and conditional dependence‐based methods and provides users with flexible options to construct individual or consensus networks across different approaches.CMiNetintegrates results from multiple inference methods through a voting strategy that retains edges supported by a user‐defined number of methods. To assess robustness, we complement this with a bootstrap analysis that quantifies edge stability under resampling. By jointly reporting method support and bootstrap confidence,CMiNetprovides a reproducible framework that explicitly communicates both agreement across methods and stability under perturbation.We appliedCMiNetto gut and soil microbiome datasets, constructing consensus networks that retained edges supported by multiple methods and confirmed by bootstrap reproducibility values. To identify disease‐associated taxa, we developed an integrative strategy that compared results across machine learning, differential abundance and network‐based approaches, ensuring that selected taxa were consistently recovered across methods. In the soil dataset, this analysis highlighted key taxa such asKtedonobacteria, Acidobacteriae, Vicinamibacteria, MB‐A2‐108, IgnavibacteriaandAnaerolineae, all of which were confirmed by multiple independent strategies.more » « lessFree, publicly-accessible full text available January 1, 2027
-
Birtwistle, Marc R (Ed.)High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package,HighDimMixedModels.jl.more » « less
-
ABSTRACT Freshwater lakes are dynamic ecosystems, with varying oxygen dynamics that influence microbiome structure, composition, and transcriptomic activity. In many freshwater studies, ecological function and abundance metrics are used to discover keystone species; however, it is well established that abundance does not equal activity. Despite the existence of long‐term time series spanning multiple years, no previous study has looked at how microbial community and activity (metatranscriptomics) are influenced by shifting oxygen conditions across depths at the microbial network level. In this study, we leverage metagenome‐assembled genomes and transcriptomic activity to identify keystone taxa in the ecosystem. Using theSPIEC‐EASIandCARlassomethods, we mapped key microbial associations and used permutation‐based analyses to assess the robustness of keystone identification. Our results reveal that a taxon's ecological centrality is context‐dependent and that many species identified as keystone by abundance alone do not exhibit corresponding transcriptional activity. Notably, members of Bacteroidota and other lineages emerged as keystone taxa only when both abundance and activity were considered. Our study underscores the importance of combining metagenomic and metatranscriptomic approaches for accurate identification of functionally relevant keystone species in freshwater ecosystems, providing a framework for future microbial ecology studies.more » « lessFree, publicly-accessible full text available April 1, 2027
An official website of the United States government
