Search for: All records

Award ID contains: 2144367

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. ABSTRACT Freshwater lakes are dynamic ecosystems, with varying oxygen dynamics that influence microbiome structure, composition, and transcriptomic activity. In many freshwater studies, ecological function and abundance metrics are used to discover keystone species; however, it is well established that abundance does not equal activity. Despite the existence of long‐term time series spanning multiple years, no previous study has looked at how microbial community and activity (metatranscriptomics) are influenced by shifting oxygen conditions across depths at the microbial network level. In this study, we leverage metagenome‐assembled genomes and transcriptomic activity to identify keystone taxa in the ecosystem. Using theSPIEC‐EASIandCARlassomethods, we mapped key microbial associations and used permutation‐based analyses to assess the robustness of keystone identification. Our results reveal that a taxon's ecological centrality is context‐dependent and that many species identified as keystone by abundance alone do not exhibit corresponding transcriptional activity. Notably, members of Bacteroidota and other lineages emerged as keystone taxa only when both abundance and activity were considered. Our study underscores the importance of combining metagenomic and metatranscriptomic approaches for accurate identification of functionally relevant keystone species in freshwater ecosystems, providing a framework for future microbial ecology studies. 
    more » « less
    Free, publicly-accessible full text available April 1, 2027
  2. Abstract Microbiome research faces two central challenges, namely constructing reliable networks, where nodes represent microbial taxa and edges represent their associations, and identifying significant disease-associated taxa. To address the first challenge, we developed CMIMN, a novel R package that applies a Bayesian network framework based on conditional mutual information to infer microbial interaction networks. To further enhance reliability, we construct a consensus microbiome network by integrating results from CMIMN and three widely used methods, including Sparse Inverse Covariance Estimation for Ecological Association Inference (SPIEC-EASI), Semi-Parametric Rank-based correlation and partial correlation Estimation (SPRING), and Sparse Correlations for Compositional Data (SPARCC). This consensus approach, which overlays and weights edges shared across methods, reduces inconsistencies and provides a more biologically meaningful view of microbial relationships. To address the second challenge, we designed a multi-method feature selection framework that combines machine learning with network-based strategies. Our machine learning pipeline applies distinct algorithms and identifies key taxa based on their consistent importance across models. Complementing this, we employ two network-based strategies that prioritize taxa based on centrality differences between networks constructed from healthy samples and disease-affected samples, as well as a composite scoring system that ranks nodes using integrated network metrics. We applied CMIMN on soil microbiome data from potato fields affected by common scab disease. Bootstrap analysis confirmed the robustness of CMIMN, and the consensus network further improved stability and interpretability. The multi-method framework enhances confidence in identifying soil microbial taxa associated with potato disease. Notably, we identified Bacteroidota, WPS-2, and Proteobacteria at the Phylum level; Actinobacteria, AD3, Bacilli, Anaerolineae, and Ktedonobacteria at the Class level; and C0119, Defluviicoccales, Bacteroidales, and Ktedonobacterales at the Order level as key taxa associated with disease status. 
    more » « less
    Free, publicly-accessible full text available January 1, 2027
  3. Abstract Microbial networks offer critical insights into community structure, ecological interactions and host–microbe dynamics. However, constructing reliable microbiome networks remains challenging due to variability among existing inference methods, limited overlap between inferred networks and the absence of a gold standard (a universally accepted reference for benchmarking) for validation.We developedCMiNet, an R package and interactive Shiny App(https://cminet.wid.wisc.edu) that enables consensus microbiome network construction by integrating up to 10 widely used inference algorithms.CMiNetsupports both correlation‐based and conditional dependence‐based methods and provides users with flexible options to construct individual or consensus networks across different approaches.CMiNetintegrates results from multiple inference methods through a voting strategy that retains edges supported by a user‐defined number of methods. To assess robustness, we complement this with a bootstrap analysis that quantifies edge stability under resampling. By jointly reporting method support and bootstrap confidence,CMiNetprovides a reproducible framework that explicitly communicates both agreement across methods and stability under perturbation.We appliedCMiNetto gut and soil microbiome datasets, constructing consensus networks that retained edges supported by multiple methods and confirmed by bootstrap reproducibility values. To identify disease‐associated taxa, we developed an integrative strategy that compared results across machine learning, differential abundance and network‐based approaches, ensuring that selected taxa were consistently recovered across methods. In the soil dataset, this analysis highlighted key taxa such asKtedonobacteria, Acidobacteriae, Vicinamibacteria, MB‐A2‐108, IgnavibacteriaandAnaerolineae, all of which were confirmed by multiple independent strategies. 
    more » « less
    Free, publicly-accessible full text available January 1, 2027
  4. Ouangraoua, Aida (Ed.)
    Abstract MotivationThe abundance of gene flow in the Tree of Life challenges the notion that evolution can be represented with a fully bifurcating process which cannot capture important biological realities like hybridization, introgression, or horizontal gene transfer. Coalescent-based network methods are increasingly popular, yet not scalable for big data, because they need to perform a heuristic search in the space of networks as well as numerical optimization that can be NP-hard. Here, we introduce a novel method to reconstruct phylogenetic networks based on algebraic invariants. While there is a long tradition of using algebraic invariants in phylogenetics, our work is the first to define phylogenetic invariants on concordance factors (frequencies of four-taxon splits in the input gene trees) to identify level-1 phylogenetic networks under the multispecies coalescent model. ResultsOur novel hybrid detection methodology is optimization-free as it only requires the evaluation of polynomial equations, and as such, it bypasses the traversal of network space, yielding a computational speed at least 10 times faster than the fastest-to-date network methods. We illustrate our method’s performance on simulated and real data from the genus Canis. Availability and implementationWe present an open-source publicly available Julia package PhyloDiamond.jl available at https://github.com/solislemuslab/PhyloDiamond.jl with broad applicability within the evolutionary community. 
    more » « less
  5. Abstract Phylogenetic networks encode a broader picture of evolution by the inclusion of reticulate processes such as hybridization, introgression, or horizontal gene transfer. Each hybridization event is represented by a ‘hybridization cycle’. Here, we investigate the statistical identifiability of the position of the hybrid node in a 4-node hybridization cycle in a semi-directed level-1 phylogenetic network. That is, we investigate if our model is able to detect the correct placement of the hybrid node in the hybridization cycle using quartet concordance factors as data. In the current study, we prove that the correct placement of the hybrid node in 4-node hybridization cycles, included in level-1 phylogenetic networks, is generically identifiable if the assumptions are non-restrictive such as t∈(0,∞) for all branch (or edge) lengths and γ∈(0,1) for the inheritance probability of the hybrid edges. However, simulations show that accurate detection of these cycles can be complicated by inadequate sampling, small sample size, or gene tree estimation error. We identify practical advice for evolutionary biologists on best sampling strategies to improve the detection of this type of hybridization cycle. 
    more » « less
  6. Ouangraoua, Aida (Ed.)
    Abstract Scientists world-wide are putting together massive efforts to understand how the biodiversity that we see on Earth evolved from single-cell organisms at the origin of life and this diversification process is represented through the Tree of Life. Low sampling rates and high heterogeneity in the rate of evolution across sites and lineages produce a phenomenon denoted “long branch attraction” (LBA) in which long non-sister lineages are estimated to be sisters regardless of their true evolutionary relationship. LBA has been a pervasive problem in phylogenetic inference affecting different types of methodologies from distance-based to likelihood-based. Here, we present a novel neural network model that outperforms standard phylogenetic methods and other neural network implementations under LBA settings. Furthermore, unlike existing neural network models in phylogenetics, our model naturally accounts for the tree isomorphisms via permutation invariant functions which ultimately result in lower memory and allows the seamless extension to larger trees. 
    more » « less
  7. Abstract Gene flow is increasingly recognized as an important macroevolutionary process. The many mechanisms that contribute to gene flow (e.g. introgression, hybridization, lateral gene transfer) uniquely affect the diversification of dynamics of species, making it important to be able to account for these idiosyncrasies when constructing phylogenetic models. Existing phylogenetic‐network simulators for macroevolution are limited in the ways they model gene flow.We presentSiPhyNetwork, an R package for simulating phylogenetic networks under a birth–death‐hybridization process.Our package unifies the existing birth–death‐hybridization models while also extending the toolkit for modelling gene flow. This tool can create patterns of reticulation such as hybridization, lateral gene transfer, and introgression.Specifically, we model different reticulate events by allowing events to either add, remove or keep constant the number of lineages. Additionally, we allow reticulation events to be trait dependent, creating the ability to model the expanse of isolating mechanisms that prevent gene flow. This tool makes it possible for researchers to model many of the complex biological factors associated with gene flow in a phylogenetic context. 
    more » « less
  8. Free, publicly-accessible full text available June 1, 2027
  9. Reticulate evolution has long been recognized as a key mechanism that contributes to genetic and trait diversity. With the widespread availability of genomic data, investigating historical reticulate evolution across taxa has gained significant attention, driven by the rapid development of statistical methods for detecting nontreelike patterns. Phylogenetic networks provide a biologically intuitive approach to depicting evolutionary processes such as hybrid speciation and introgressive hybridization, which result in signatures of historical gene flow. Interpreting phylogenetic networks is especially critical for groups of conservation concern that lack reference genome resources and explicit hypotheses from prior investigation, such as those based on molecular data, morphology, or species distributions. Here, we highlight recent advances in computational methods for inferring networks from genome-scale data and offer guidelines for deriving biological insights from phylogenetic networks. Particular emphasis is placed on modeling hybridization and whole-genome duplication in the context of allopolyploidization. Practical recommendations for empirical studies and the limitations of commonly used methods are discussed throughout. We anticipate that phylogenetic networks will influence conservation biology and biodiversity research, emphasizing the need for careful consideration of reticulate evolution inferred from these networks in the near future. Networks will accelerate other pressing avenues of biodiversity research, especially investigations of orphan crops and climate change resilience in natural systems. The promise of phylogenetic networks connects with broader themes in the special feature Monitoring and restoring gene flow in the increasingly fragmented ecosystems of the Anthropocene by providing an emerging probabilistic framework for inferring historical connectivity between species and populations. 
    more » « less
  10. This special collection includes topics related to the development of novel methods for reconstructing phylogenetic networks from different mathematical, statistical, and computational approaches that highlight the challenges of network reconstruction and the needs of contemporary genomic data. In addition, the collection broadcasts diverse applications of phylogenetic networks on a wide variety of organisms across the Tree of Life. 
    more » « less