Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Search for: All records

Award ID contains: 2410565

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Abstract Secondary contact between previously allopatric lineages offers a test of reproductive isolating mechanisms that may have accrued in isolation. Such instances of contact can produce stable hybrid zones—where reproductive isolation can further develop via reinforcement or phenotypic displacement—or result in the lineages merging. Ongoing secondary contact is most visible in continental systems, where steady input from parental taxa can occur readily. In oceanic island systems, however, secondary contact between closely related species of birds is relatively rare. When observed on sufficiently small islands, relative to population size, secondary contact likely represents a recent phenomenon. Here, we examine the dynamics of a group of birds whose apparent widespread hybridization influenced Ernst Mayr’s foundational work on allopatric speciation: the whistlers of Fiji (Aves: Pachycephala). We demonstrate two clear instances of secondary contact within the Fijian archipelago, one resulting in a hybrid zone on a larger island, and the other resulting in a wholly admixed population on a smaller island. We leveraged low genome-wide divergence in the hybrid zone to pinpoint a single genomic region associated with observed phenotypic differences. We use genomic data to present a new hypothesis that emphasizes rapid plumage evolution and post-divergence gene flow. 
    more » « less
  2. Free, publicly-accessible full text available September 1, 2027
  3. Abstract: Pacific island taxa have long informed our understanding of speciation and biogeography. However, until genomic data became commonplace, many explosive radiations have remained unresolved. One underappreciated radiation is of the genus Aplonis, a clade of starlings distributed across the Pacific from continental Sundaland to the far-flung Cook Islands. Of note are the high levels of secondary sympatry in Aplonis relative to other geographic radiations in the region. Here, we attempt to resolve relationships in this group by sampling all extant and three extinct taxa. We sequenced ultraconserved elements from 141 ingroup samples (approximately 2 per subspecies), of which 104 were derived from historic toepad samples. We were able to infer a strongly supported phylogeny, which we used to infer the biogeographic history of the genus. Aplonis, which largely followed two one-way colonization routes via a stepping stone pattern from west to east. However, there was evidence for three independent, long-distance colonizations of Micronesia, and two cases of back-colonization from more distant to more nearby islands. Furthermore, based on patterns of sympatry and rules of monophyly, we found evidence for three species-level splits, including two in the widespread and polyphyletic Asian Glossy Starling (Aplonis panayensis). Finally, the subspecies-level sampling we performed allowed us to better understand the diversification dynamics in this group, with a stable accumulation of lineages over time when considering species or secondary contact as our operational units. TechnicalInfo: # Data and code from: Phylogeny of Pacific starlings (genus *Aplonis*) reveals cryptic diversity and diverse biogeographic patterns Repository for scripts and input files used in a subspecies-level phylogeny of the genus Aplonis. This README describes the scripts and data files used for the aforementioned project. Each section corresponds to a zipped directory in the Dryad repository and a folder on GitHub. *** denotes files in the Dryad only (mostly large alignment files or similar). #### Read processing (01_read_processing.zip) Raw reads are housed at NCBI's Sequencing Read Archive (PRJNA1336315). aplonis-illumiprocessor.conf - Configuration file for illumiprocessor, made following standard Phyluce specifications. illumiprocessor.slurm - Slurm script for running illumiprocessor to get cleaned reads. Originally run as part of a longer script that was not used for final analyses, so subset to only the relevant command. read_count.slurm - Script for counting the number of cleaned reads in each sample. sample_list - Simple text list of all samples. #### Tissue pipeline to generate reference (02_tissue_pipeline.zip) Note that this was run for all tissue samples, but only one was used. Toepad samples were too fragmented for timely assemblies, so this naturally led to mostly tissue samples being considered. Only one (A_panayensis_panayensis_KU121758) was used as a reference. velvet_parallel.slurm - Slurm script for running Phyluce's wrapper around velvet per sample in parallel (default runs many at a time, this allowed more flexibility). Note that this was re-run at least once, although better samples finished faster. sample_list_tissue - Simple text list of all tissue samples. A_panayensis_panayensis_KU121758_ref.fasta - Fasta of the best assembly, used for pseudoreference downstream. #### Reference-based pipeline (03_reference_pipeline.zip) This is the stage that actually made input for all analyses, on the basis of the high-quality reference in the de novo assembly step. This section is a bit complicated, so it has a sub-README. In short, reference_pipeline.slurm is run remotely from the main working directory, then 2_trim_ambig.py is run from a directory with consensus sequences (script from [https://datadryad.org/dataset/doi:10.5061/dryad.n5tb2rbsp](https://datadryad.org/dataset/doi:10.5061/dryad.n5tb2rbsp), trims consensus sequence based on ambiguous bases), then further phyluce scripts are run from the main directory. README_toepad - Supplemental README describing how it is run. reference_pipeline.slurm - Pipeline for making consensus sequence, note prior note about trimming with the python script from Smith et al. 2020. Pipeline itself is heavily inspired by Smith et al. 2020, and compensates for the difficulties default phyluce has with toepads, as de novo assemblies are a lot harder. Note that the threshold used for the trimming script is 40%, so it was run like "python 2_trim_ambig.py 0.4". calc_contigs_stats.sh - Script for calculating contig stats, run on the consensus sequence, use the amples list in step 1. phyluce_thresh40_trim.slurm - Script for running phyluce, side-stepping some of it for trimming with trimal. phyluce_aplonis76.slurm - The same as above, but subset for select analyses using one per subspecies-level unit. \*** aplonis_*phylip - Phylip files output by phyluce for scripts noted above. astral.slurm - Slurm script for running ASTRAL. #### Mitogenome pipeline (04_mitogenome.zip) run_mitofinder.slurm - Script used for tissue and toepad mitogenomes. Originally run separately to ease computing. Note that these use a singularity image of mitofinder. Didn't work out how to use Singularity parallelized across nodes. I'm sure there's a way to do it, but this was fast enough, and I had time to spare. Both use NCBI OK542103.1 (Acridotheres tristis) as a reference. sample_list* - Sample lists used for dating subsets or full (goodmito) analyses. gene_list.txt and gene_list_focal.txt - Simple text files of gene names used for dating (_focal) and phylogenetic analyses. split_genes.sh - Shell script used for splitting out gene-specific fastas with all species in them. A "monolith" of all of these is then spit out for Phyluce to use. phyluce_mito*.slurm - Scripts for hijacking phyluce for making concatenated mitochondrial datasets, either for dating (time*) or phylogenetic analyses. aplonis_mito_time*_nexus.nexus - Two nexus files with cytB and ND2 for use in data analyses. aplonis_mito136_concat_75per.phylip - Phylip used for tree building (not divergence times). #### Phylogenetics (05_phylogenies.zip) sed_crcm.sh - Script for fixing some character issues with the CRCM samples. iqtree.slurm - Script for "basic" iqtree runs, all the same command for different alignments. make_genetrees.slurm - Script for making genetrees for Astral/iqtree. iqtree_concordance.slurm - Script for iqtree concordance factor runs, takes in any phylogeny, sequence alignment, and optionally gene tree sets (latter doesn't work well with UCE target capture onto toepads). *.stat - Concordance factor stat output, used for examining in detail manually or with evaluate_concordance_factors.R. *tree, *treefile, and *tre (other than aplonis149_thresh40_trim_genetrees_75.treefile)- Assorted tree file outputs, as described in their name, with cf denoting concordance factors attached to the nodes. \*** aplonis149_thresh40_trim_genetrees_75.treefile - A file of gene trees for the focal dataset, used for concordance factors and astral. evaluate_concordance_factors.R - R script for analyzing concordance factors, inspired by or fully drawing from this post by Rob Lanfear: [https://www.robertlanfear.com/blog/files/concordance_factors.html](https://www.robertlanfear.com/blog/files/concordance_factors.html). \*** aplonis76_concat_75per.nexus - Input nexus file used for SVDQuartets, using the charsets file output by Phyluce to generate partitions. [Analysis credit to Michael Andersen] aplonis76_concat_75per.charsets - Charsets file of subspecies-level dataset. svdq_standardBoot.log - Log file for SVDQuartets run. [Analysis credit to Michael Andersen] #### Divergence dating (06_divergence_dating.zip) Divergence dating using a constraint of the IQtree topology for given datasets and mtDNA for rates. Note that only 75 of 76 subspecies made it to the subspecies-level dataset, but all 34 species-level tips made it. aplonis_mito_time*.xml - XML files for running BEAST, including the mtDNA alignments and constraint tree. [Analysis credit to Michael Andersen] aplonis_mito_time75-tree_logcombined.tre - Time calibrated tree from BEAST used for LTT plots. ltt_plot.R - R script for plotting LTT data across taxonomic scales. #### Biogeography (07_biogeography.zip) BioGeoBear analysis was run for two levels, subspecies (75 tips) and species (34 tips). All analysis credit to Michael Andersen for this section. aplonis_mito_time*-tree.logcombined.newick - Phylogenies used for running BioGeoBEARS, time calibrated based on IQTree topology and mtDNA rates. aplonis_*_BioGeoBear.R - R scripts for running species and subspecies level BioGeoBEARS analyses. time*_geography.txt - Geography files for running BioGeoBEARS. Codings are in order as follows: Continental (C), Wallacean (W), Sahul (S), Near Melanesia (N), Far Melanesia (F), Polynesia (P), and Micronesia (M). 
    more » « less
  4. Biogeography relies on estimates of phylogeny and gene flow to test hypotheses. Established theory posits: (i) gene flow varies non-randomly with respect to geography and (ii) gene flow can mislead phylogenetic inference. These findings, often viewed in isolation, together suggest that spatial impacts on gene flow could bias biogeographic inference. For example, whenever a ‘peripheral’ population is isolated and inferred as sister to a ‘core’ clade of adjacent populations that exchange migrants, it might reflect elevated gene flow of ‘core’ populations rather than macroevolutionary processes (e.g. direction of colonization). Given this ubiquitous but often unacknowledged biogeographic setting, it remains unknown how spatial influences on gene flow impact metrics used to assess biogeographic hypotheses (e.g. branch length or monophyly). We simulate gene flow that varies non-randomly regarding geography and found that relatively low levels of gene flow lead to monophyly of adjacent populations, with longer branches for peripheral populations, regardless of true divergence history. We highlight an empirical example wherein antbirds (Thamnophilidae) are young in Amazonia, but older at Amazonia’s periphery. Although it is unclear if the simulated bias applies to antbirds, our results suggest a distinguishability problem for any nodes stemming from radiations in geographically heterogeneous environments. 
    more » « less
    Free, publicly-accessible full text available April 22, 2027
  5. Abstract: Biogeography relies on estimates of phylogeny and gene flow to test hypotheses. Established theory posits (1) gene flow varies non-randomly with respect to geography and (2) gene flow can mislead phylogenetic inference. These findings, often viewed in isolation, together suggest that spatial impacts on gene flow could bias biogeographic inference. For example, whenever a “peripheral” population is isolated and inferred as sister to a “core” clade of adjacent populations that exchange migrants, it might reflect elevated gene flow of “core” populations rather than macroevolutionary processes (e.g., direction of colonization). Given this ubiquitous but often unacknowledged biogeographic setting, it remains unknown how spatial influences on gene flow impact metrics used to assess biogeographic hypotheses (e.g., branch length or monophyly). We simulate gene flow that varies nonrandomly regarding geography and found that relatively low levels of gene flow lead to monophyly of adjacent populations, with longer branches for peripheral populations, regardless of true divergence history. We highlight an empirical example wherein antbirds (Thamnophilidae) are young in Amazonia, but older at Amazonia’s periphery. Although it is unclear if the simulated bias applies to antbirds, our results suggest a distinguishability problem for any nodes stemming from radiations in geographically heterogeneous environments. TechnicalInfo: # Data and code from: Spatially heterogeneous gene flow may hinder linking phylogeographic data to macroevolutionary patterns Last updated 4 March 2026 Repository for scripts and input files used in a study that investigates the potential role of spatially heterogeneous gene flow in shaping phylogenetic inference, with an eye towards a biogeographic pattern (the core-periphery effect). Sections 1 through 3 are written by Ethan Gyllenhaal; Section 4 is written by Lukas Musher. This README describes the scripts and data files used for the aforementioned project. Each section corresponds to a zipped directory in the Dryad repository. ### Coalescent Simulations (01_simulations.zip) core_phylo_noanc.py - Python script used to run msprime simulations, taking in a set of parameters defined in the batch script and outputting fasta format files for phylogenetic inference. Takes in parameters for output path, replicate number, number of migrants per generation (generally <1), and per population effective pop. size, time to run after colonization completes, interval of colonization events, and a string split order. Useage like python $dir/core_trial_noanc.py -r [int replicate] -m [float Nm] -n [int/float effective pop size] -e [int/float post-split evolution time] -i [int/float split interval in gens] -s [string split interval: CAB, ACB, or ABC] -o [output stem, ideally includes all variables used]. variant_fasta.sh - Custom script for reducing output fasta files to variant sites only, for space efficiency. Usage: sh variant_fasta.sh in.fasta out.fasta. full_coresim_run.slurm - SLURM bash script used for running replicates. It uses GNU parallel to iterate over replicates. Note that this did not finish in the 48-hour walltime, and had to be resumed several times. Uploaded script reflects the original run. sim_tree_output.zip - Zipped directory of all output treefiles from simulations. Used as input for the next section. ### Simulation processing (02_processing.zip) summarize_branch.py - Python script for summarizing output phylogenies run in a batch script with output from the msprime Python script (both in section 1). This uses ETE3 to process trees and output topologies and branch lengths. Takes a big list of tree files as input. Usage example: python summarize_branch.py -i 'sim_tree_output/*treefile' -o coresim_output_branch.csv coresim_output_branch.csv - Output from summarize_branch.py. Columns are as follows, (NOTE, after the first use of Anest, Bnest, ABnest, Cmono, Other, Non-monophyletic, and Poor-resolution, the full set are refered to with a *): * First column is the stem name for the parameter combination * Anest - Number of topologies with population A but not B nested in the core (C) populations * Bnest - Number of topologies with population B but not A nested in the core (C) populations * ABnest - Number of topologies with population A and B nested in the core (C) populations * Cmono - Number of topologies with a monophyletic core group (all 4 populations) * Other - number of topologies not matching the prior set, never found here, not sure it can happen but a catch for if it does * Non-monophyletic - Number of topologies where A or B were non-monophyletic (none in this set) * Poor-resolution - Number of poorly resolved topologies without a well-supported monophyletic core * *total - The SUMMED tree heights for each category, notably not a mean, needs to be divided by the count per category * *_core - The summed mean branch lengths of core populations per category * *_a - The summed branch lengths of population a per category * *_b - The summed branch lengths of population b per category * total - The summed tree heights across categories * core - The summed mean core branch lengths across categories * a - The summed population a branch lengths across categories * b - The summed population b branch lengths across categories * name - Same as the first column, stem name of files * disp - Simulated core gene flow level in number of migrants WITHOUT THE DECIMAL (IN THIS CASE DIVIDED BY 10) for easier parsing * ne - Simulated per population effective population size * time - Simulated post-divergence time in generations * int - Simulated split interval * history - Simulated split history, named based on what orders divergences occurred in (A = population a, B = population b, C = core radiation) plot_topo_branch_heatmap.R - R script for plotting summarized phylogenetic information, primarily heat maps. Also includes code for making input used in 04_divergence_dates (Fig 3A) ### Richness Map (03_mapping.zip) plot_antbird_richness.R - R script that uses BirdLife data to plot species richness map. Note that you'll have to get your own BirdLife data: [https://datazone.birdlife.org/contact-us/request-our-data](https://datazone.birdlife.org/contact-us/request-our-data). Use the species list (below) to subset the massive input dataframe. species.txt - List of antbird species used here, based on BirdLife taxonomy. ### Divergence date estimation (04_divergence_dates.zip) area.thamno.txt - Output from R code categorizing biogeographic scores into categories for plot in Figure 3B. violin_full.csv - Input used for violin plots, generated by plot_topo_branch_heatmap.R (step 2). Most values from coresim_output_branch.csv, with some additional math from R. Columns are below: * Combination - Arbitrary number for each combination of parameters for each population's branch length * disp - Simulated core gene flow level in number of migrants WITHOUT THE DECIMAL (IN THIS CASE DIVIDED BY 10) for easier parsing * ne - Simulated per population effective population size * time - Simulated post-divergence time in generations * int - Simulated split interval * a - The summed population a branch lengths across categories * b - The summed population b branch lengths across categories * total - The summed tree heights across categories * core - The summed mean core branch lengths across categories * core_per - The summed mean core branch length per parameter combination divided by the smaller of summed population a or b branch lengths (if doing per category, would have to divide by the category count first, not relevant here) * branch.length - The mean branch length for the given population (core, A, or B) * name - Stem name of the file the data was calculated from * history - Simulated split history, named based on what orders divergences occurred in (A = population a, B = population b, C = core radiation) * count - Sum of replicates per parameter combo (20 for all here) * pop - The population these statistics were computed for, either all core populations combined or one individual peripheral population. * migration - Binary variable describing if any gene flow was included (does not separate by migration rates) T400F_Thamnophilidae.tre - Phylogeny used for divergence dates and plotting. Thamnophil_biogeographic_scoring.csv - CSV used as input for biogeographic assignment code. Columns are below: * Taxon_name_code - Name of taxon in phylogeny * Family - Family of taxon * genus - Genus of taxon * species - Species of taxon * subspecies_if_mult_sampled - Subspecies of taxon if multiple were sampled * specimen locality - Locality of specimen (only for one) * real_elevational_range - Elevational range of taxon if known (not used here) * n_america to islands - Biogeographic range of taxon (reduced to summary in area.thamno.txt). All of these variables are scored based on LJM's assessment of ranges based on available literature. In order, these refer to: North America, West Mexico, East Mexico, Central America, Northwest South America, Northern Andes, Southern Andes, Western Amazonia, Northeast Amazonia, Southeast Amazonia, Atlantic Forest of South America, Cerrado savanna ecoregion, Chaco region of South America, Caatinga biome of Brazil, Beni savanna ecoregion, Llanos grasslands, Tepui highlands, Parana river basin, Amazonian white-sand forest habitat, Temperate regions of South America, Atacama desert, Caribbean region, general lowland habitats under 250m elevation, general upland habitat from 250 to 1500m, general montane habitat from 1500m to 3500m, high montant habitat over 3500m, Amazonian flooded forest, Amazonian terra firme forest, general arid habitats, general humid habitats, and general oceanic islands. * Notes - additional notes if relevant thamnophilidae_ages.R - R code for generating the phylogeny and box plots in Figure 3 (A and B). Uses the biogeographic scoring CSV and tree, outputs area.thamno.txt 
    more » « less
  6. Abstract: Allopatric divergence is a fundamental component of most traditional models of biogeography and community assembly. Gene flow between allopatric populations should be influenced by the nature of geographic barriers and can have a profound impact on adaptation, the speciation process, and phylogenetic inference. Superspecies—monophyletic groups of taxa with species-level differences in phenotype or genotype that are found exclusively in allopatry or parapatry—present an opportunity to characterize the effects of gene flow on the divergence process. Here we investigate patterns of gene flow, population structure, and inferred phylogenetic relationships for members of an avian superspecies, the Solomons Monarchs (Aves: Symposiachrus barbatus complex) occupying the Solomon Islands. We found that gene flow among allopatric species matches predictions based on geography, but phylogenetic relationships were not concordant with the most likely colonization history based on a stepping-stone colonization model. Notably, the most isolated island, Makira, has a species that was inferred to be sister to the taxa on all other islands in concatenated phylogenetic analyses, despite Makira being farthest from the presumed original source of immigrants. We use population genetic simulations to demonstrate that such a result could be driven by bias resulting from low levels of gene flow, reflecting a challenge in phylogeographic inference that results when one population is differentially isolated. These simulated findings demonstrate a distinguishability issue in phylogeographic inference, where gene flow and colonization history can be difficult to disentangle. Methods: The dataset is mostly a RAD-seq dataset, generating 1000s of loci across the genome, which were then sequenced on an Illumina sequencer. These were then processed with Stacks, with input generated for downstream analyses using several different scripts outlined in the document. Additionally, we ran two sets of simulations to demonstrate the theorhetical basis of two of our major claims (colonization and the impact of geographically mediated gene flow on phylogenetic inference). Other: There is a README.txt file included which describes the script and datasets included. Datasets are divided by mode of inference, with scripts associated with given input files (e.g., vcf, phylip, or nexus files) in the same .zip files. I did my best to rename input files in a meaningful way, and change my scripts accordingly, but most scripts aren't designed to run everything with the same directory structure/naming format as the zipped directory (i.e., some alteration needed). Scripts are also uploaded individually to Zenodo. Scripts, accessory files, and trees are also located on my GitHub: https://github.com/ethangyllenhaal/SolomonsSymposRad. Most figures were edited in Adobe Illustrator to change, e.g., text and line format, but no data were fundamentally altered. TechnicalInfo: README by: Ethan Gyllenhaal Last updated: 10 Sept 2025 Repository for scripts and input files used in a study investigating gene flow and its impacts on phylogeography in a radiation of Symposiachrus monarch flycatchers in the Solomon islands. Everything is in subfolders within a zipped file. If you use something here, please cite us, or contact EFG on who is best to cite (and remind him if he forgets to update the citation): This README describes the scripts and data files used for the afforementioned project. The data were collected using Illumina sequencing for RAD-seq libraries, largely following the Stacks reference-based pipeline. Scripts and input data file are in the same folder for a given analysis. ## Useage The majority of the data included here is output from a Stacks bioinformatics pipeline, which was run as a Torque script on a high performancce computing center. Individual VCFs were then processed with command line tools (e.g., vcftools) or R scripts for preparing figures, tables, etc for the paper. Data file formats (e.g. vcf, phylip) are standard except when noted. Additionally, there are two sets of simulation scripts and output (SLiM and msprime). Access reccomendations per data type are below. All data types except .xlsx and compressed files can be viewed in any plain text editor, but certain files are large enough that may not be advisable depending on the editor. ### Supplemental Tables out of appendix **TableS1_Sampling_table_SolSym.xlsx:** Table of samples used for RAD-seq dataset. From left to right, columns are species name, subspecies (when relevant), island, country, voucher number, tissue number, sample name used in scripts, sample name used for SRA accession, number of reads, which datasets the sample was included in, and Arctos database links. The numbers in the dataset column correspond to: 1) full dataset, 2) all Solomons samples, 3) all New Georgia samples, 4) all Bukida samples, 5) samples used in DSuite tests for gene flow, and 6) samples used in SNAPP. **TableS2_UCE_sampling.xlsx:** Table of samples used for UCE dataset. From left to right, columns are species name, subspecies (when relevant), island, country, voucher number, tissue number (when relevant), sample name used in scripts, sample name used for SRA accession, and number of recovered loci. ### Data and other files * .xlsx - Excel files, in our case manipulated with microsoft excel, but LibreOffice Calc and similar open source software can be used as well (and formulas transfers for at least LibreOffice). * .vcf - Variant Call Format files of SNP data. These contain a standard set of information including position on the reference, genotype info, and depth. Full specifications can be found here: [https://samtools.github.io/hts-specs/VCFv4.2.pdf](https://samtools.github.io/hts-specs/VCFv4.2.pdf). They can be viewed in any text editor, and I tend to use VCFtools to manipulate/filter them. * .tre and .treefile - Newick format phylogeny. Best viewed with FigTree: [https://tree.bio.ed.ac.uk/software/figtree/](https://tree.bio.ed.ac.uk/software/figtree/) * .stat - File output by IQTree detailing concordance factor statistics, viewed in text editors. * .nex - Nexus format phylogeny. Best viewed with FigTree: [https://tree.bio.ed.ac.uk/software/figtree/](https://tree.bio.ed.ac.uk/software/figtree/). Note that this is technically the same format as .nexus file, but they contain different fields. * .phylip and .nexus - Two allignment file formats used for IQ-TREE and SVDQUarters. These both contain concatenated sequence information, and although they are just text files, they are best used as input for other programs. * .fa - Allignment file format used only for storing per-locus gene trees used for locus-level IQ-TREE analyses. Can be viewed in any text editor, note that only first allele of each set was used. * .nwk - Newick format phylogeny generated in a text editor. Can be viewed in FigTree, but note that it isn't an empirical phylogeny per say (i.e., just made based off of topology from empirical trees). * .XML - Markdown file generated by BEAUti ([https://beast.community/beauti](https://beast.community/beauti)) for use with BEAST ([https://beast.community/index.html](https://beast.community/index.html)). Contains UCE allignments and details for BEAST to run analyses on the file. Can be viewed in a text editor raw or in BEAUti for user-friendly viewing/editing. * .ped and .map - Two plink format files for genotypes and site information, respectively. Formats are outlined here: [https://www.cog-genomics.org/plink/1.9/formats](https://www.cog-genomics.org/plink/1.9/formats). In short, .ped includes population ID, individual ID, four unused columns (value of 0 or -9), and two columns of SNP information per genotype. The other, .map, includes columns for chromosome code, variant ID, an unused column, and coordinator in basepairs. Used as input for admixture. * .trees - Nexus file of trees from the postior of SNAPP runs, used as input for DensiTree and TreeAnnotator. * .conf - File used for configuration of the PHYLUCE pipeline for UCE data. * .tar.gz - Compressed tar archive, can be decompressed with "tar -zxvf". * .gz - Directory compressed with gzip, can be decompressed with "gzip". * .jpeg or .png - Image files, used for photos of specimens, viewable in any image viewing software. * unspecified (no extension) - These are plain text files, visible with text editor of viewer of choice. ### Scripts * .pbs and .slurm - Batch scripts submitted from command line on UNM's cluster. * .R - R script designed to be run interactively in RStudio. * .sh - Shell scripts used for various tasks. Run on command line locally or on the cluster. * .py - Python scripts, run either interactively or as part of a pipeline. * .slim - Eidos scripts for running SLiM simulations, run on the cluster, but with slight modifications can be viewed and run in SLiM's GUI as well. ## Supplemental data These are excel files that act as supplemental tables for the paper either explicitly or to make explicit how I went about correcting tests of gene flow per-category. ## Zipped directories The numbered directories each contain files used for the bioinformatics pipeline and other analyses ### Stacks pipeline (01_stacks.zip) run_stacks_pipeline.pbs - Torque submission script used to run the pipeline. It was very quick, using only a single 8-core node for <10 hours. Made to run on UNM's Center for Advanced Research Computing's Wheeler cluster. popmaps directory - Directory of population maps used to make subset input files. Individual files are below, with the first column denoting sample names and second population assignment: *popmap_NG - New Georgia Group samples *popmap_bukida - Bukida Group samples *popmap_full - All samples *popmap_full_indiv - All samples, but with a dummy population used to output per-individual phylip files *popmap_solo - All Solomon Island samples *snapp_even - SNAPP input with three samples per island group and two samples from an outgroup *snapp_five_island - SNAPP input with up to five islands per island group (two per island if group have more than one) and two samples from an outgroup *snapp_two_island - SNAPP input with up to two islands per island group (two per island if group have more than one) and two samples from an outgroup ### Adegenet (02_adegenet.zip) SoloSympos_PCA.R - R script for the 3 PCAs in the paper (resulting in 4 figures), with the first one walking through the process in detail. The full and Bukida PCAs are in Figure 2. The New Georgia PCAs are in Figure S1. solsym_*_75.vcf - 75% complete VCF files for subsets of the data (all Solomons birds, all Bukida birds, and all New Georgia birds, respectively). ### Admixture (03_admixture.zip) solsym_admixture_plot.R - R script for plotting Admixture results. Output used in Figure 2. run_admixture_NG.sh - Shell script for running Admixture. Assumes a conda environment with Plink and Admixture is loaded (or similar ways of making "plink" and "admixture" call their respective programs). solsym_NG_75.* - Plink input files from Stacks, with 75% completeness for New Georgia Samples. solsym_NG_75_12.* - Plink files from prior line converted to 12 format for use in Admixture. ### Confirming intergrade (04_confirm_intergrade.zip) keep_interspecific - Sample list of individuals to keep for fixed SNP analysis (only one is used, rest are there to confirm we do indeed have fixed SNPs) kolo_etc_samples - Sample list of Kolombangara and connected islands. rano_vella_samples - Sample list of Vella Lavella and Ranongga. solsym_solomons_75.vcf - 75% complete VCF file for Solomons subset of data. run_interspecific_het.sh - Main script, first uses VCFtools and unix commands to find fixed SNPs (between the combined Kolombangara etc and Vella/Ranongga populations), then limits input VCF to those SNPs, then does some unix math so calculate proportion of heterozygotes for the focal individual at those sites. Comparable to what the program introgression calculates. ### Pairwise Fst (05_pairwise_Fst.zip) pairwise_Fst_vcftools.sh - Shell script used for calculating pairwise Fst between all Solomons population using VCFtools. Output used in Table 1. pops directory - Files used to specify populations. solsym_solomons_75.vcf - 75% complete VCF file for Solomons subset of data. ### IQTree and SVDQuartet (06_iqtree_svdq.zip) NOTE: SVDQuartets analyses were run with the PAUP GUI. Output in Figure 3 and Figure S2. iqtree.slurm - SLURM batch script for running IQTree for inference on RAD and UCE data, as well as site concordance factor calculation for RAD data. full_rad_75.phylip - Input for IQTree, generated with Stacks with slight manual modifications. full_rad_75.phylip.treefile - Output from IQtree, without concordance factors. solsym_cf_full_75.cf.tree - Phylogeny for RAD dataset with concordance factors on nodes. solsym_cf_full_75.cf.stat - Full concordance factor data for RAD dataset. solomons_triv_SVDQ_75.tre - SVDQuartets phylogeny for ingroup and S. trvirgatus. Tips are grouped together. solomons_triv_SVDQ_75.nex - SVDQuartets input modified from IQTree phylip without a script. Removed all outgroups other than trvirgatus. ### SNAPP (07_SNAPP.zip) NOTE: Some steps were run with the Beauti GUI. Output in Figure 4. snapp_input_maker.sh - Driver script for converting a VCF to diploid SNAPP input (i.e. Nexus input for processing in Beauti). Converts from VCF to haploid SNAPP using [https://github.com/BEAST2-Dev/SNAPP/tree/master/script](https://github.com/BEAST2-Dev/SNAPP/tree/master/script), then to diploid SNAPP using a custom python script (SNAPP_haploid_to_diploid.py), and finally adding lines for a new nexus file using sed commands. Now takes command line arguments (relative to original version). generate_snapp_input.sh - Simple script with commands for each SNAPP input file. SNAPP_haploid_to_diploid.py - Script for converting a tab-delimited set of haploid nexus SNP data to diploid data. Takes and outputs "middle" lines, header and footer done by "snapp_input_maker.sh". Note that it's "dumb" and needs the input to be named "snapp_dip.txt" and outputs "snapp_out.txt". snapp100_*.vcf - 100% complete VCF for SNAPP samples of a given set: even = 3 samples from 1 island per group; group = up to 4 samples per island group, with 2 islands (2 samples each) for Bukida and New Georgia groups; island = up to 10 samples per island group, with 5 islands (2 samples each) for Bukida and New Georgia groups. snapp100_*.nex - Nexus file used to make Beast XML for SNAPP. snapp100_*.xml - Beast XML input file. Simply run with the line: beast -threads 32 solsym_limit100.xml. (adding --resume if needed) snapp*.trees - The first N trees from a given SNAPP run (resuming continues past limit, to if limit was 10 million we pruned extra trees past that). treeanno_*.tre - TreeAnnotator summary trees. ### UCEs Phylogenies (08_UCEs.zip) NOTE: SVDQuartets analyses were run with the PAUP GUI. Output in Figure 3. IQTREE script is in section 06. run_phyluce_solo_triv.sh - Shell script for running PHYLUCE pipeline. Starts with cleaned reads and outputs Nexus and Phyllip files. solsym_uce.conf - Taxon set .conf file. Used in several phyluce calls. solsym_spades.conf - Spades .conf file, note paths are changed to generic ones up until clean reads folder. Used in phyluce_assembly_assemblo_spades. solsym_uce_contigs.tar.gz - All of the contigs.fa outputs from spades, tarballed and gzipped. sympos_uce_nexus_90.nexus - 90% complete nexus file used for SVDQuartets analysis. sympos_uce_SVDQ_90.tre - Tree produced by SVDQ, combining samples by population. sympos_uce_raxml_90.phylip - 90% complete phylip file used for RAxML analysis. sympos_uce_raxml.tre - Tree produced by RAxML, combining samples by population. ### DSuite (09_DSuite.zip) NOTE: The topologies and popmaps have 6 consistent name corresponding to subfigures in Figure S3. B = main, C = noMal, D = conservative, E = conservative_noMal, F = malbar G = malbar_conservative. run_D.sh - Shell script used for running Dsuite. The first one is used for the main ABBA/BABA results (Figure 5). The first and the rest are used for fbranch (Figure S3). fbranch_topologies/* - Topologies used for fbranch, outlined below: *topo_main.nwk - The main topology. *topo_noMal.nwk - Malaita removed. *topo_conservative.nwk - Certain populations removed to avoid uncertainty within island groups. *topo_conservative_noMal.nwk - As above, but also with Malaita removed. *topo_malbar.nwk - New Georgia samples removed. *topo_malbar_conservative.nwk - New Georgia samples and one uncertain Bukida sample removed. pop_maps/* - Population maps used for the runs. All but two downsampled ones match the above topologies. These downsampled maps match the main topology: *map_downsample3 - Population map with a maximum of 3 samples per population. *map_downsample2 - Population map with a maximum of 2 samples per population. ### Genetrees (10_genetrees.zip) all_genetrees.slurm - Slurm batch script for making locus-specific fastas and running gene trees with given numbers of parsimony informative sites. check_monophyly_all.py - Python script for outputting proportion and branch length stats for gene trees. locus_fastas.fa - Fasta of locus-specific fastas. popmap_genetrees - Population map for gene tree samples (one per island group). *.zip - Zipped directories of locus fastas and IQ-TREE output for fastas with given number of parsimony informative sites, as listed below: *all.zip - At least one PIS *all_PIS2.zip - At least two PIS *all_PIS3.zip - At least three PIS ### Photos of intergrade (11_photos.zip) All of these are photos of the intergrade (center), Ranongga samples (left two), and Kolombangara samples (right two). They show the view over the underparts (Belly_View_Full.jpeg), sides (Side_View_Full.jpeg), and the main one used in Figure 2 (Main_hybrid_comp.png). ### Colonization simulations (12_colonization_sims.zip) solsym_nonWF_carc_loui_v2.slim - Slim script used to run colonization simulations. run_loui_colonize.slurm- Slurm script used to run simulations in parallel. Requires a conda environment with SLiM. drive_summarize.sh - Script used to modify output then run python script. summarize_loui.py - Script used to summarize the output from slim. summarized_*.txt - Text files output by summarize script. Each row is for a given dispersal value. Three files correspond to the proportion of simulations where a given island group met specific conditions (columns 2, 3, 4, and 5 correspond to New Georgia, Malaita, Makira, and Bukida respectively): "first" means which was colonized first, "last" means which was colonized last, and "success" means which islands were successfully colonized **of the simulations where the Solomons were colonzied**. The path file is colonization paths, with columns corresponding to NLKB, NLBK, NKLB, NKBL, NBLK, NBKL, LNKB, LNBK, LKNB, LKBN, LBNK, LBKN, KNLB, KNBL, KLNB, KLBN, KBNL, KBLN, BNLK, BNKL, BLNK, BLKN, BKNL, then BKLN (B=Bukida, N=New Georgia, L=Malaita, K=Makira). slim_sankey.R - R script for plotting colonization paths in a Sankey diagram. Note that paths are converted from proportions to numbers and summed across all simulations, then assigned manually in R (couldn't figure out a smart way to do it). ### Phylogenetic simulations (13_phylo_sims.zip) new_solo_phylo_msp.py - Script for running msprime simulations, made to be run in parallel. run_solomons_msp.slurm - Slurm batch script for running replicates of msprime followed by IQTREE. Requires a conda environment with msprime and IQTREE (plus other python packages). variant_fasta.sh - Script for removing invariant sites to keep alignment files smaller for storage purposes. combined_sumstats.sh - Simple script for combining summary stat files in parallel. summarize_*.txt - Tables summarizing topologies from simulations that started in Bukida (_buk) or Makira (_mak). Columns are the row name (stem name), 15 possible topologies (values are counts), count of non-monophyletic trees, count of poorly resolved trees, the stem name of the simulation set, dispersal distance, effective population size, post-split time, and split interval. summarize_trees_bs.py - Script for parsing and classifying topologies from set of trees. summarize_trees.slurm - Slurm script for running summarize_trees_bs.py. plot_msp_heatmap.R - Script for plotting heatmap of phylogenetic simulation results. output_trees.zip - Zipped directory of trees output by IQTREE. The "d" corresponds to dispersal distance (0.0001 acts as 0), n to total population time, t for post-split time, i for split interval, and r is the replicate number. output_sim.zip - Zipped directory of fastas and summary stats output by msprime. The "d" corresponds to dispersal distance (0.0001 acts as 0), n to total population time, t for post-split time, i for split interval, and r is the replicate number. ### Additional summary simulation files (14_bonus_sim_stats.zip) plot_sumstats.R - Script for plotting a lot of data exploration plots not used in the paper. source_d* - Source population data from SLiM simulations for a range of dispersal values. From left to right, columns are the source population values for the New Georgia Group, Malaita, Makira, and the Bukida Group. sumstat_d* - Summary stat (pi and private alleles) data for SLiM simulations for a range of dispersal values. From left to right, the first four columns are for pi and second four for private allele count, each in the order of: New Georgia Group, Malaita, Makira, Bukida Group. sumstat_combined - Msprime output for all simulations, first with parameters then with population-specific summary stats. The first 6 columns are: dispersal distance, total Ne, post-split divergence time, split interval, simulation start point, and replicate number. Next four are pi (for Bukida, New Georgia, Malaita, Makira) and then number of private alleles (in the same order). 
    more » « less
  7. Carstens, Bryan (Ed.)
    Allopatric divergence is a fundamental component of most traditional models of biogeography and community assembly. Gene flow between allopatric populations should be influenced by the nature of geographic barriers and can have a profound impact on adaptation, the speciation process, and phylogenetic inference. Superspecies—monophyletic groups of taxa with species-level differences in phenotype or genotype that are found exclusively in allopatry or parapatry—present an opportunity to characterize the effects of gene flow on the divergence process. Here, we investigate patterns of gene flow, population structure, and inferred phylogenetic relationships for members of an avian superspecies, the Solomons Monarchs (Aves: Symposiachrus barbatus complex) occupying the Solomon Islands. We found that gene flow among allopatric species matches predictions based on geography, but phylogenetic relationships were not concordant with the most likely colonization history based on a stepping-stone colonization model. Notably, the most isolated island, Makira, has a species that was inferred to be sister to the taxa on all other islands in concatenated phylogenetic analyses, despite Makira being farthest from the presumed original source of immigrants. We use population genetic simulations to demonstrate that such a result could be driven by bias resulting from low levels of gene flow, reflecting a challenge in phylogeographic inference that results when one population is differentially isolated. These simulated findings demonstrate a distinguishability issue in phylogeographic inference, where gene flow and colonization history can be difficult to disentangle. 
    more » « less
    Free, publicly-accessible full text available November 12, 2026
  8. Islands have long represented natural laboratories for studying many aspects of ecology and evolutionary biology, from speciation to community assembly. One aspect that has been well documented is the correlation between island size and taxonomic diversity, likely due to decreased complexity and population size on small islands. This same logic can apply to genetic diversity, which should predictably decrease with effective population size. The island size–diversity correlation has received support over the years but often focuses on single metrics of genetic diversity. Here, we useZosteropswhite-eyes in the Solomon Islands to study the correlation between island size and various metrics related to genetic diversity, including runs of homozygosity and fixation of transposable elements. We find that almost all these metrics strongly correlate with island size, and in turn with each other. We infer that island size is independently correlated with these different variables, demonstrating that population size impacts genomic metrics of diversity in a variety of ways across temporal and hierarchical scales. 
    more » « less