Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Title: Multitask knowledge-primed neural network for predicting missing metadata and host phenotype based on human microbiome
Abstract MotivationMicrobial signatures in the human microbiome are closely associated with various human diseases, driving the development of machine learning models for microbiome-based disease prediction. Despite progress, challenges remain in enhancing prediction accuracy, generalizability, and interpretability. Confounding factors, such as host’s gender, age, and body mass index, significantly influence the human microbiome, complicating microbiome-based predictions. ResultsTo address these challenges, we developed MicroKPNN-MT, a unified model for predicting human phenotype based on microbiome data, as well as additional metadata like age and gender. This model builds on our earlier MicroKPNN framework, which incorporates prior knowledge of microbial species into neural networks to enhance prediction accuracy and interpretability. In MicroKPNN-MT, metadata, when available, serves as additional input features for prediction. Otherwise, the model predicts metadata from microbiome data using additional decoders. We applied MicroKPNN-MT to microbiome data collected in mBodyMap, covering healthy individuals and 25 different diseases, and demonstrated its potential as a predictive tool for multiple diseases, which at the same time provided predictions for the missing metadata. Our results showed that incorporating real or predicted metadata helped improve the accuracy of disease predictions, and more importantly, helped improve the generalizability of the predictive models. Availability and implementationhttps://github.com/mgtools/MicroKPNN-MT.  more » « less
Award ID(s):
2025451
PAR ID:
10647931
Author(s) / Creator(s):
; ;
Editor(s):
Lengauer, Thomas
Publisher / Repository:
Oxford Academics
Date Published:
Journal Name:
Bioinformatics Advances
Volume:
5
Issue:
1
ISSN:
2635-0041
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. The microbiota has proved to be one of the critical factors for many diseases, and researchers have been using microbiome data for disease prediction. However, models trained on one independent microbiome study may not be easily applicable to other independent studies due to the high level of variability in microbiome data. In this study, we developed a method for improving the generalizability and interpretability of machine learning models for predicting three different diseases (colorectal cancer, Crohn’s disease, and immunotherapy response) using nine independent microbiome datasets. Our method involves combining a smaller dataset with a larger dataset, and we found that using at least 25% of the target samples in the source data resulted in improved model performance. We determined random forest as our top model and employed feature selection to identify common and important taxa for disease prediction across the different studies. Our results suggest that this leveraging scheme is a promising approach for improving the accuracy and interpretability of machine learning models for predicting diseases based on microbiome data. 
    more » « less
  2. Background:Pancreatitis significantly alters the microbial composition of the oral and intestinal compartments, causing dysbiosis that may contribute to disease mechanisms and potentially serve as a basis for diagnosis or treatment. Objective:To determine whether the oral or gut microbial signature can classify chronic pancreatitis (CP). Methods:Stool samples (n=707) were collected from participants in the Prospective Evaluation of Chronic Pancreatitis for Epidemiologic and Translational Studies (PROCEED). Samples were distributed among 200 healthy (HC), 310 CP, 49 acute pancreatitis (AP), and 148 recurrent acute pancreatitis (RAP). In addition, saliva samples were collected for a subset of participants (n=156). Whole genome sequencing was performed to assess microbiome composition. Machine learning algorithms were utilized to identify a signature with microbial features predictive of CP. Results:Gut alpha diversity was significantly decreased in AP, RAP, and CP compared with HC, with CP exhibiting the lowest diversity. In contrast, oral microbial diversity showed no significant variation across groups. Beta diversity analysis revealed distinct gut microbiome compositions between HC and pancreatitis subtypes, with CP showing the most pronounced differences. Random forest models using gut microbial species demonstrated robust predictive performance for CP using a minimum of 10 species (Area under the curve—AUC: 0.834; accuracy: 0.774). Despite similarities in gut microbiome composition across pancreatitis subtypes, a unique gut microbial signature for CP was identified highlighting the microbiome’s potential in CP diagnosis. Conclusion:Our study reveals a gut microbial signature predictive of CP using machine learning models in a large US multi-institutional cohort. 
    more » « less
  3. Abstract The gut microbiome plays a fundamental role in human health and disease. Individual variations in the microbiome and the corresponding functional implications are key considerations to enhance precision health and medicine. Metaproteomics has recently revealed protein expression that might be associated with human health and disease. Existing studies focused on either human proteins or bacterial proteins that can be identified from (meta)proteomics data sets, but not both. In this study, we examined the feasibility of identifying both human and bacterial proteins that are differentially expressed between healthy and diseased individuals from metaproteomics data sets. We further evaluated different strategies of using identified peptides and proteins for building predictive models. By leveraging existing metaproteomics data sets and a tool that we have developed for metaproteomics data analysis (MetaProD), we were able to derive both human and bacterial differentially expressed proteins that could serve as potential biomarkers for all diseases we studied. We also built predictive models using identified peptides and proteins as features for prediction of human diseases. Our results showed peptide-based identifications over protein-based ones often produce the most accurate models and that feature selection can offer improvements. Prediction accuracy could be further improved, in some cases, by including bacterial identifications, but missing data in bacterial identifications remains problematic. 
    more » « less
  4. Abstract AimsElevated blood pressure (BP) is a prevalent modifiable risk factor for cardiovascular diseases and contributes to cognitive decline in late life. Despite the fact that functional changes may precede irreversible structural damage and emerge in an ongoing manner, studies have been predominantly informed by brain structure and group-level inferences. Here, we aim to delineate neurobiological correlates of BP at an individual level using machine learning and functional connectivity. Methods and resultsBased on whole-brain functional connectivity from the UK Biobank, we built a machine learning model to identify neural representations for individuals’ past (∼8.9 years before scanning, N = 35 882), current (N = 31 367), and future (∼2.4 years follow-up, N = 3 138) BP levels within a repeated cross-validation framework. We examined the impact of multiple potential covariates, as well as assessed these models’ generalizability across various contexts.The predictive models achieved significant correlations between predicted and actual systolic/diastolic BP and pulse pressure while controlling for multiple confounders. Predictions for participants not on antihypertensive medication were more accurate than for currently medicated patients. Moreover, the models demonstrated robust generalizability across contexts in terms of ethnicities, imaging centres, medication status, participant visits, gender, age, and body mass index. The identified connectivity patterns primarily involved the cerebellum, prefrontal, anterior insula, anterior cingulate cortex, supramarginal gyrus, and precuneus, which are key regions of the central autonomic network, and involved in cognition processing and susceptible to neurodegeneration in Alzheimer’s disease. Results also showed more involvement of default mode and frontoparietal networks in predicting future BP levels and in medicated participants. ConclusionThis study, based on the largest neuroimaging sample currently available and using machine learning, identifies brain signatures underlying BP, providing evidence for meaningful BP-associated neural representations in connectivity profiles. 
    more » « less
  5. Abstract BackgroundIn Alzheimer's Disease research, identifying brain regions that influence disease progression is crucial for understanding pathogenesis and developing targeted therapies. Deep learning models have advanced survival analysis in Alzheimer's Disease research, but their “black box” nature limits clinical utility. We developed Neural Additive Deep Clustering Survival Machines (NADCSM), which leverages Neural Additive Models to provide interpretable insights into brain region contributions while maintaining competitive predictive performance. This approach aims to bridge the gap between model performance and clinical utility, potentially accelerating biomarker discovery and therapeutic development. MethodThe genotyping, demographic, and imaging data used in our experiments are sourced from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. We utilized ADNI's AV45 Florbetapir PET imaging data to track MCI and early AD progression. This modality is particularly valuable for tracking the progression of mild cognitive impairment (MCI) and early Alzheimer's disease (AD). Our NADCSM framework models survival times using Weibull distributions and enhances interpretability through Neural Additive Models (NAMs)., where each input feature is processed through multiple univariate shape functions parameterized by multilayer perceptions (MLPs). Performance was evaluated using Concordance Index (C Index) for predictive ability and LogRank statistic for survival curve separation and clustering quality. ResultTable 1 compares the performance of DeepCox, DCSM, and NADCSM (ours) on the AV45 dataset using C Index and LogRank metrics. DCSM achieves the highest C Index (0.7789 ± 0.0193), followed by DeepCox (0.7781 ± 0.0306) and NADCSM (0.7772±0.0236), demonstrating strong predictive accuracy. For LogRank, DCSM leads (317.84 ± 31.89), followed by NADCSM (317.84 ± 31.89) and DeepCox (93.4171±9.7990), indicating NADCSM's competitive performance with enhanced interpretability.Figure 1 shows shape functions for the top six brain ROIs identified by NADCSM, including Fusiform (Left), Cerebelum_Crus (Right), and others. The smooth curves capture the relationships between regional amyloid burden and AD progression, highlighting key survival predictors. ConclusionOur NADCSM framework provides an interpretable risk prediction approach, uncovering significant features and explaining their effects on Alzheimer's and Related Dementias (ADRD) progression. By enhancing transparency, this study can advance precision medicine and improve understanding of ADRD‐related challenges. 
    more » « less