skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: dbCAN-PUL: a database of experimentally characterized CAZyme gene clusters and their substrates
Abstract PULs (polysaccharide utilization loci) are discrete gene clusters of CAZymes (Carbohydrate Active EnZymes) and other genes that work together to digest and utilize carbohydrate substrates. While PULs have been extensively characterized in Bacteroidetes, there exist PULs from other bacterial phyla, as well as archaea and metagenomes, that remain to be catalogued in a database for efficient retrieval. We have developed an online database dbCAN-PUL (http://bcb.unl.edu/dbCAN_PUL/) to display experimentally verified CAZyme-containing PULs from literature with pertinent metadata, sequences, and annotation. Compared to other online CAZyme and PUL resources, dbCAN-PUL has the following new features: (i) Batch download of PUL data by target substrate, species/genome, genus, or experimental characterization method; (ii) Annotation for each PUL that displays associated metadata such as substrate(s), experimental characterization method(s) and protein sequence information, (iii) Links to external annotation pages for CAZymes (CAZy), transporters (UniProt) and other genes, (iv) Display of homologous gene clusters in GenBank sequences via integrated MultiGeneBlast tool and (v) An integrated BLASTX service available for users to query their sequences against PUL proteins in dbCAN-PUL. With these features, dbCAN-PUL will be an important repository for CAZyme and PUL research, complementing our other web servers and databases (dbCAN2, dbCAN-seq).  more » « less
Award ID(s):
1933521
PAR ID:
10234854
Author(s) / Creator(s):
; ; ; ; ; ; ;
Date Published:
Journal Name:
Nucleic Acids Research
Volume:
49
Issue:
D1
ISSN:
0305-1048
Page Range / eLocation ID:
D523 to D528
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Carbohydrate active enzymes (CAZymes) are made by various organisms for complex carbohydrate metabolism. Genome mining of CAZymes has become a routine data analysis in (meta-)genome projects, owing to the importance of CAZymes in bioenergy, microbiome, nutrition, agriculture, and global carbon recycling. In 2012, dbCAN was provided as an online web server for automated CAZyme annotation. dbCAN2 (https://bcb.unl.edu/dbCAN2) was further developed in 2018 as a meta server to combine multiple tools for improved CAZyme annotation. dbCAN2 also included CGC-Finder, a tool for identifying CAZyme gene clusters (CGCs) in (meta-)genomes. We have updated the meta server to dbCAN3 with the following new functions and components: (i) dbCAN-sub as a profile Hidden Markov Model database (HMMdb) for substrate prediction at the CAZyme subfamily level; (ii) searching against experimentally characterized polysaccharide utilization loci (PULs) with known glycan substates of the dbCAN-PUL database for substrate prediction at the CGC level; (iii) a majority voting method to consider all CAZymes with substrate predicted from dbCAN-sub for substrate prediction at the CGC level; (iv) improved data browsing and visualization of substrate prediction results on the website. In summary, dbCAN3 not only inherits all the functions of dbCAN2, but also integrates three new methods for glycan substrate prediction. 
    more » « less
  2. Abstract Carbohydrate Active EnZymes (CAZymes) are significantly important for microbial communities to thrive in carbohydrate rich environments such as animal guts, agricultural soils, forest floors, and ocean sediments. Since 2017, microbiome sequencing and assembly have produced numerous metagenome assembled genomes (MAGs). We have updated our dbCAN-seq database (https://bcb.unl.edu/dbCAN_seq) to include the following new data and features: (i) ∼498 000 CAZymes and ∼169 000 CAZyme gene clusters (CGCs) from 9421 MAGs of four ecological (human gut, human oral, cow rumen, and marine) environments; (ii) Glycan substrates for 41 447 (24.54%) CGCs inferred by two novel approaches (dbCAN-PUL homology search and eCAMI subfamily majority voting) (the two approaches agreed on 4183 CGCs for substrate assignments); (iii) A redesigned CGC page to include the graphical display of CGC gene compositions, the alignment of query CGC and subject PUL (polysaccharide utilization loci) of dbCAN-PUL, and the eCAMI subfamily table to support the predicted substrates; (iv) A statistics page to organize all the data for easy CGC access according to substrates and taxonomic phyla; and (v) A batch download page. In summary, this updated dbCAN-seq database highlights glycan substrates predicted for CGCs from microbiomes. Future work will implement the substrate prediction function in our dbCAN2 web server. 
    more » « less
  3. Abstract MotivationCarbohydrate-active enzymes (CAZymes) are extremely important to bioenergy, human gut microbiome, and plant pathogen researches and industries. Here we developed a new amino acid k-mer-based CAZyme classification, motif identification and genome annotation tool using a bipartite network algorithm. Using this tool, we classified 390 CAZyme families into thousands of subfamilies each with distinguishing k-mer peptides. These k-mers represented the characteristic motifs (in the form of a collection of conserved short peptides) of each subfamily, and thus were further used to annotate new genomes for CAZymes. This idea was also generalized to extract characteristic k-mer peptides for all the Swiss-Prot enzymes classified by the EC (enzyme commission) numbers and applied to enzyme EC prediction. ResultsThis new tool was implemented as a Python package named eCAMI. Benchmark analysis of eCAMI against the state-of-the-art tools on CAZyme and enzyme EC datasets found that: (i) eCAMI has the best performance in terms of accuracy and memory use for CAZyme and enzyme EC classification and annotation; (ii) the k-mer-based tools (including PPR-Hotpep, CUPP and eCAMI) perform better than homology-based tools and deep-learning tools in enzyme EC prediction. Lastly, we confirmed that the k-mer-based tools have the unique ability to identify the characteristic k-mer peptides in the predicted enzymes. Availability and implementationhttps://github.com/yinlabniu/eCAMI and https://github.com/zhanglabNKU/eCAMI. Supplementary informationSupplementary data are available at Bioinformatics online. 
    more » « less
  4. Abstract Chemoenzymatic approaches using carbohydrate‐active enzymes (CAZymes) offer a promising avenue for the synthesis of glycans like oligosaccharides. Here, we report a novel chemoenzymatic route for cellodextrins synthesis employed by chimeric CAZymes, akin to native glycosyltransferases, involving the unprecedented participation of a “non‐catalytic” lectin‐like domain or carbohydrate‐binding modules (CBMs) in the catalytic step for glycosidic bond synthesis using β‐cellobiosyl donor sugars as activated substrates. CBMs are often thought to play a passive substrate targeting role in enzymatic glycosylation reactions mostly via overcoming substrate diffusion limitations for tethered catalytic domains (CDs) but are not known to participate directly in any nucleophilic substitution mechanisms that impact the actual glycosyl transfer step. This study provides evidence for the direct participation of CBMs in the catalytic reaction step for β‐glucan glycosidic bonds synthesis enhancing activity for CBM‐based CAZyme chimeras by >140‐fold over CDs alone. Dynamic intradomain interactions that facilitate this poorly understood reaction mechanism were further revealed by small‐angle X‐ray scattering structural analysis along with detailed mutagenesis studies to shed light on our current limited understanding of similar transglycosylation‐type reaction mechanisms. In summary, our study provides a novel strategy for engineering similar CBM‐based CAZyme chimeras for the synthesis of bespoke oligosaccharides using simple activated sugar monomers. 
    more » « less
  5. na (Ed.)
    Long Terminal Repeat (LTR) retrotransposons are a class of repetitive elements that are widespread in the genomes of plants and many fungi. LTR retrotransposons have been associated with rapidly evolving gene clusters in plants and virulence factor transfer in fungal-plant parasite-host interactions. We report here the abundance and transcriptional activity of LTR retrotransposons across several species of the early-branching Neocallimastigomycota, otherwise known as the anaerobic gut fungi (AGF). The ubiquity of LTR retrotransposons in these genomes suggests key evolutionary roles in these rumen-dwelling biomass degraders, whose genomes also contain many enzymes that are horizontally transferred from other rumen-dwelling prokaryotes. Up to 10% of anaerobic fungal genomes consist of LTR retrotransposons, and the mapping of sequences from LTR retrotransposons to transcriptomes shows that the majority of clusters are transcribed, with some exhibiting expression greater than 104 reads per kilobase million mapped reads (rpkm). Many LTR retrotransposons are strongly differentially expressed upon heat stress during fungal cultivation, with several exhibiting a nearly three-log10 fold increase in expression, whereas growth substrate variation modulated transcription to a lesser extent. We show that some LTR retrotransposons contain carbohydrate-active enzymes (CAZymes), and the expansion of CAZymes within genomes and among anaerobic fungal species may be linked to retrotransposon activity. We further discuss how these widespread sequences may be a source of promoters and other parts towards the bioengineering of anaerobic fungi. 
    more » « less