skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: An Evolving View of Phylogenetic Support
Abstract If all nucleotide sites evolved at the same rate within molecules and throughout the history of lineages, if all nucleotides were in equal proportion, if any nucleotide or amino acid evolved to any other with equal probability, if all taxa could be sampled, if diversification happened at well-spaced intervals, and if all gene segments had the same history, then tree building would be easy. But of course, none of those conditions are true. Hence, the need for evaluating the information content and accuracy of phylogenetic trees. The symposium for which this historical essay and presentation were developed focused on the importance of phylogenetic support, specifically branch support for individual clades. Here, I present a timeline and review significant events in the history of systematics that set the stage for the development of the sophisticated measures of branch support and examinations of the information content of data highlighted in this symposium. [Bayes factors; bootstrap; branch support; concordance factors; internode certainty; posterior probabilities; spectral analysis; transfer bootstrap expectation.]  more » « less
Award ID(s):
1655891
PAR ID:
10208107
Author(s) / Creator(s):
Editor(s):
Lanfear, Robert
Date Published:
Journal Name:
Systematic Biology
ISSN:
1063-5157
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. null (Ed.)
    Genome-scale data have greatly facilitated the resolution of recalcitrant nodes that Sanger-based datasets have been unable to resolve. However, phylogenomic studies continue to use traditional methods such as bootstrapping to estimate branch support; and high bootstrap values are still interpreted as providing strong support for the correct topology. Furthermore, relatively little attention has been given to assessing discordances between gene and species trees, and the underlying processes that produce phylogenetic conflict. We generated novel genomic datasets to characterize and determine the causes of discordance in Old World treefrogs (Family: Rhacophoridae)—a group that is fraught with conflicting and poorly supported topologies among major clades. Additionally, a suite of data filtering strategies and analytical methods were applied to assess their impact on phylogenetic inference. We showed that incomplete lineage sorting was detected at all nodes that exhibited high levels of discordance. Those nodes were also associated with extremely short internal branches. We also clearly demonstrate that bootstrap values do not reflect uncertainty or confidence for the correct topology and, hence, should not be used as a measure of branch support in phylogenomic datasets. Overall, we showed that phylogenetic discordances in Old World treefrogs resulted from incomplete lineage sorting and that species tree inference can be improved using a multi-faceted, total-evidence approach, which uses the most amount of data and considers results from different analytical methods and datasets. 
    more » « less
  2. Abstract The hemipteran suborder Auchenorrhyncha is a highly diverse, ecologically and agriculturally important group of primarily phytophagous insects which has been a source of phylogenetic contention for many years. Here, we have used transcriptome sequencing to assemble 2139 orthologues from 84 auchenorrhynchan species representing 27 families; this is the largest and most taxonomically comprehensive phylogenetic dataset for this group to date. We used both maximum likelihood and multispecies coalescent analyses to reconstruct the evolutionary history in this group using amino acid, nucleotide, and degeneracy‐coded nucleotide orthologue data. Although many relationships at the superfamily level were consistent between analyses, several differing, highly supported topologies were recovered using different datasets and reconstruction methods, most notably the differential placement of Cercopoidea as sister to either Cicadoidea or Membracoidea. To further interrogate the recovered topologies, we explored the contribution of genes as partitioned by third‐codon‐position guanine‐cytosine (GC) content and heterogeneity. We found consistent support for several relationships, including Cercopoidea + Cicadoidea, most often in genes that would be expected to be enriched for the true species tree if recombination‐based dynamics in GC content have contributed to the observed GC heterogeneity. Our results provide a generally well‐supported framework for future studies of auchenorrhynchan phylogeny and suggest that transcriptome sequencing is likely to be a fruitful source of phylogenetic data for resolving its clades. However, we caution that future work should account for the potential effects of GC content heterogeneity on relationships recovered in this group. 
    more » « less
  3. Abstract Background The 16S mitochondrial rRNA gene is the most widely sequenced molecular marker in amphibian systematic studies, making it comparable to the universal CO1 barcode that is more commonly used in other animal groups. However, studies employ different primer combinations that target different lengths/regions of the 16S gene ranging from complete gene sequences (~ 1500 bp) to short fragments (~ 500 bp), the latter of which is the most ubiquitously used. Sequences of different lengths are often concatenated, compared, and/or jointly analyzed to infer phylogenetic relationships, estimate genetic divergence ( p -distances), and justify the recognition of new species (species delimitation), making the 16S gene region, by far, the most influential molecular marker in amphibian systematics. Despite their ubiquitous and multifarious use, no studies have ever been conducted to evaluate the congruence and performance among the different fragment lengths. Results Using empirical data derived from both Sanger-based and genomic approaches, we show that full-length 16S sequences recover the most accurate phylogenetic relationships, highest branch support, lowest variation in genetic distances (pairwise p -distances), and best-scoring species delimitation partitions. In contrast, widely used short fragments produce inaccurate phylogenetic reconstructions, lower and more variable branch support, erratic genetic distances, and low-scoring species delimitation partitions, the numbers of which are vastly overestimated. The relatively poor performance of short 16S fragments is likely due to insufficient phylogenetic information content. Conclusions Taken together, our results demonstrate that short 16S fragments are unable to match the efficacy achieved by full-length sequences in terms of topological accuracy, heuristic branch support, genetic divergences, and species delimitation partitions, and thus, phylogenetic and taxonomic inferences that are predicated on short 16S fragments should be interpreted with caution. However, short 16S fragments can still be useful for species identification, rapid assessments, or definitively coupling complex life stages in natural history studies and faunal inventories. While the full 16S sequence performs best, it requires the use of several primer pairs that increases cost, time, and effort. As a compromise, our results demonstrate that practitioners should utilize medium-length primers in favor of the short-fragment primers because they have the potential to markedly improve phylogenetic inference and species delimitation without additional cost. 
    more » « less
  4. Billions of years ago, the Earth’s atmosphere had very little oxygen. It was only after some bacteria and early plants evolved to harness energy from sunlight that oxygen began to fill the Earth’s environment. Oxygen is highly reactive and can interfere with enzymes and other molecules that are essential to life. Organisms living at this point in history therefore had to adapt to survive in this new oxygen-rich world. An ancient family of enzymes known as ribonucleotide reductases are used by all free-living organisms and many viruses to repair and replicate their DNA. Because of their essential role in managing DNA, these enzymes have been around on Earth for billions of years. Understanding how they evolved could therefore shed light on how nature adapted to increasing oxygen levels and other environmental changes at the molecular level. One approach to study how proteins evolved is to use computational analysis to construct a phylogenetic tree. This reveals how existing members of a family are related to one another based on the chain of molecules (known as amino acids) that make up each protein. Despite having similar structures and all having the same function, ribonucleotide reductases have remarkably diverse sequences of amino acids. This makes it computationally very demanding to build a phylogenetic tree. To overcome this, Burnim, Spence, Xu et al. created a phylogenetic tree using structural information from a part of the enzyme that is relatively similar in many modern-day ribonucleotide reductases. The final result took seven continuous months on a supercomputer to generate, and includes over 6,000 members of the enzyme family. The phylogenetic tree revealed a new distinct group of ribonucleotide reductases that may explain how one adaptation to increasing levels of oxygen emerged in some family members, while another adaptation emerged in others. The approach used in this work also opens up a new way to study how other highly diverse enzymes and other protein families evolved, potentially revealing new insights about our planet’s past. 
    more » « less
  5. Abstract Aholein a graph is an induced cycle of length at least four, and a ‐multiholein is the union of pairwise disjoint and nonneighbouring holes. It is well known that if does not contain any holes then its chromatic number is equal to its clique number. In this paper we show that, for any integer , if does not contain a ‐multihole, then its chromatic number is at most a polynomial function of its clique number. We show that the same result holds if we ask for all the holes to be odd or of length four; and if we ask for the holes to be longer than any fixed constant or of length four. This is part of a broader study of graph classes that are polynomially ‐bounded. 
    more » « less