Search for: All records

Creators/Authors contains: "Hahn, Matthew W"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Abstract Quantitative traits provide insights into how phenotypes evolve across species. However, standard comparative methods often assume a single species tree and overlook the discordant gene tree histories that may underlie complex traits. Here, we develop a model and software (Spinney) that explicitly incorporate gene tree heterogeneity into rate estimation. Spinney finds the optimal rate of evolution and ancestral states by jointly maximizing the likelihoods across a set of gene trees. Using simulated data, we compare rate estimates from Spinney with those using the species tree alone, test Spinney’s ability to distinguish true rate variation from spurious signals caused by gene tree discordance, and evaluate ancestral state reconstruction. Spinney consistently produced more accurate rate estimates and reduced incorrect inferences of rate variation. This method therefore provides a flexible framework to integrate gene tree heterogeneity into comparative methods and to produce reliable inferences of quantitative trait evolution, regardless of the source of discordance. 
    more » « less
    Free, publicly-accessible full text available January 1, 2027
  2. Knowles, L Lacey; Kirkpatrick, Mark (Ed.)
    Abstract Historically, phylogenetic datasets had relatively few loci but were over-represented for cytoplasmic sequences (mitochondria and chloroplast) because of their ease of amplification and large numbers of informative sites. Under those circumstances, it made sense to contrast individual gene tree topologies obtained from cytoplasmic loci and nuclear loci, with the goal of detecting differences between them—so-called cytonuclear discordance. In the current age of phylogenomics and ubiquitous gene tree discordance among thousands of loci, it is important to distinguish between simply observing discordance between cytoplasmic trees and a species tree inferred from many nuclear loci and identifying the cause of discordance. Here, we examine what inferences one can make from trees representing different genomic compartments. While topological discordance can be caused by multiple factors, the end goal of many studies is to determine whether the compartments have different evolutionary histories: what we refer to as “cytonuclear dissonance.” Answering this question is more complex than simply asking whether there is discordance, requiring additional analyses to determine whether genetic exchange has affected only (or mostly) one compartment. Furthermore, even when these histories differ, expectations about why they differ are not always clear. We conclude by pointing to current research and future opportunities that may help to shed light on topological variation across the multiple genomes contained within a single eukaryotic cell. 
    more » « less
    Free, publicly-accessible full text available October 7, 2026
  3. Abstract Standard methods for estimating the population recombination parameter, ρ, are dependent on sampling individual genotypes and calculating various types of disequilibria. However, recent machine learning (ML) approaches to estimating recombination have used pooled sequencing data, which does not sample individual genotypes and cannot be used to calculate disequilibria beyond the length of a single sequence read. Motivated by these results, this study examines the “black box” of such ML methods to understand what signals are being used to infer recombination rates. We find that it is indeed possible to estimate recombination solely using the allele frequency spectrum, and we provide a genealogical interpretation of these results. We further show that even a simplified representation of the allele frequency spectrum can be used to estimate recombination. We demonstrate the accuracy of such inferences using both simulations and data from humans. These results offer a new way to understand the effects of recombination on patterns of sequence data, as well as providing an example of how the internal workings of ML methods can give insight into biological processes. 
    more » « less
  4. Ralph, P (Ed.)
    Abstract Detecting introgression between closely related populations or species is a fundamental objective in evolutionary biology. Existing methods for detecting migration and inferring migration rates from population genetic data often assume a neutral model of evolution. Growing evidence of the pervasive impact of selection on large portions of the genome across diverse taxa suggests that this assumption is unrealistic in most empirical systems. Further, ignoring selection has previously been shown to negatively impact demographic inferences (e.g. of population size histories). However, the impacts of biologically realistic selection on inferences of migration remain poorly explored. Here, we simulate data under models of background selection, selective sweeps, balancing selection, and adaptive introgression. We show that ignoring selection sometimes leads to false inferences of migration in popularly used methods that rely on the site frequency spectrum. Specifically, balancing selection and some models of background selection result in the rejection of isolation-only models in favor of isolation-with-migration models and lead to elevated estimates of migration rates. BPP, a method that analyzes sequence data directly, showed false positives for all conditions at recent divergence times, but balancing selection also led to false positives at medium-divergence times. Our results suggest that such methods may be unreliable in some empirical systems, such that new methods that are robust to selection need to be developed. 
    more » « less
  5. Machine learning has increasingly been applied to a wide range of questions in phylogenetic inference. Supervised machine learning approaches that rely on simulated training data have been used to infer tree topologies and branch lengths, to select substitution models, and to perform downstream inferences of introgression and diversification. Here, we review how researchers have used several promising machine learning approaches to make phylogenetic inferences. Despite the promise of these methods, several barriers prevent supervised machine learning from reaching its full potential in phylogenetics. We discuss these barriers and potential paths forward. In the future, we expect that the application of careful network designs and data encodings will allow supervised machine learning to accommodate the complex processes that continue to confound traditional phylogenetic methods. 
    more » « less
  6. Hurst, Laurence D (Ed.)
    Every mammal studied to date has been found to have a male mutation bias: male parents transmit more de novo mutations to offspring than female parents, contributing increasingly more mutations with age. Although male-biased mutation has been studied for more than 75 years, its causes are still debated. One obstacle to understanding this pattern is its near universality—without variation in mutation bias, it is difficult to find an underlying cause. Here, we present new data on multiple pedigrees from two primate species: aye-ayes (Daubentonia madagascariensis), a member of the strepsirrhine primates, and olive baboons (Papio anubis). In stark contrast to the pattern found across mammals, we find a much larger effect of maternal age than paternal age on mutation rates in the aye-aye. In addition, older aye-aye mothers transmit substantially more mutations than older fathers. We carry out both computational and experimental validation of our results, contrasting them with results from baboons and other primates using the same methodologies. Further, we analyze a set of DNA repair and replication genes to identify candidate mutations that may be responsible for the change in mutation bias observed in aye-ayes. Our results demonstrate that mutation bias is not an immutable trait, but rather one that can evolve between closely related species. Further work on aye-ayes (and possibly other lemuriform primates) should help to explain the molecular basis for sex-biased mutation. 
    more » « less
  7. Schwartz, Russell (Ed.)
    The application of machine learning approaches in phylogenetics has been impeded by the vast model space associated with inference. Supervised machine learning approaches require data from across this space to train models. Because of this, previous approaches have typically been limited to inferring relationships among unrooted quartets of taxa, where there are only three possible topologies. Here, we explore the potential of generative adversarial networks (GANs) to address this limitation. GANs consist of a generator and a discriminator: at each step, the generator aims to create data that is similar to real data, while the discriminator attempts to distinguish generated and real data. By using an evolutionary model as the generator, we use GANs to make evolutionary inferences. Since a new model can be considered at each iteration, heuristic searches of complex model spaces are possible. Thus, GANs offer a potential solution to the challenges of applying machine learning in phylogenetics. ResultsWe developed phyloGAN, a GAN that infers phylogenetic relationships among species. phyloGAN takes as input a concatenated alignment, or a set of gene alignments, and infers a phylogenetic tree either considering or ignoring gene tree heterogeneity. We explored the performance of phyloGAN for up to 15 taxa in the concatenation case and 6 taxa when considering gene tree heterogeneity. Error rates are relatively low in these simple cases. However, run times are slow and performance metrics suggest issues during training. Future work should explore novel architectures that may result in more stable and efficient GANs for phylogenetics. 
    more » « less
  8. Summary White oak (Quercus alba) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important.We report the first haplotype‐resolved chromosome‐scale genome assembly ofQ. albaand conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data.Despite strongly conserved chromosome synteny and genome size acrossQuercus, certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found thatQ. albahas high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus.The white oak genome represents a major new resource for studying genome diversity and evolution inQuercus. Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution inQuercus. 
    more » « less
  9. Matschiner, Michael (Ed.)
    Hundreds or thousands of loci are now routinely used in modern phylogenomic studies. Concatenation approaches to tree inference assume that there is a single topology for the entire dataset, but different loci may have different evolutionary histories due to incomplete lineage sorting (ILS), introgression, and/or horizontal gene transfer; even single loci may not be treelike due to recombination. To overcome this shortcoming, we introduce an implementation of a multi-tree mixture model that we call mixtures across sites and trees (MAST). This model extends a prior implementation by Boussau et al. (2009) by allowing users to estimate the weight of each of a set of pre-specified bifurcating trees in a single alignment. The MAST model allows each tree to have its own weight, topology, branch lengths, substitution model, nucleotide or amino acid frequencies, and model of rate heterogeneity across sites. We implemented the MAST model in a maximum-likelihood framework in the popular phylogenetic software, IQ-TREE. Simulations show that we can accurately recover the true model parameters, including branch lengths and tree weights for a given set of tree topologies, under a wide range of biologically realistic scenarios. We also show that we can use standard statistical inference approaches to reject a single-tree model when data are simulated under multiple trees (and vice versa). We applied the MAST model to multiple primate datasets and found that it can recover the signal of ILS in the Great Apes, as well as the asymmetry in minor trees caused by introgression among several macaque species. When applied to a dataset of 4 Platyrrhine species for which standard concatenated maximum likelihood (ML) and gene tree approaches disagree, we observe that MAST gives the highest weight (i.e., the largest proportion of sites) to the tree also supported by gene tree approaches. These results suggest that the MAST model is able to analyze a concatenated alignment using ML while avoiding some of the biases that come with assuming there is only a single tree. We discuss how the MAST model can be extended in the future. 
    more » « less