Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 11:00 PM ET on Thursday, August 13 until 12:00 AM ET on Friday, August 14 due to maintenance. We apologize for the inconvenience.


Title: OceanGreensFunctionMethods.jl: Julia software package with pedagogical box model
{"Abstract":["Released to complement a revised version of Haine et al. 2024.\n\nJulia package to complement "A Review of Green's Function Methods in Ocean Circulation Models," by Haine et al. (2024). See https://essopenarchive.org/users/528978/articles/1215807-a-review-of-green-s-function-methods-for-tracer-timescales-and-pathways-in-ocean-models. This package has two goals.\n\n1. One of the stated goals of the manuscript is to make Green's Function methods accessible for learning purposes. A computational notebook is provided with this package and detailed below.\n\n2. Here, we also aim to make a Julia package that is useful and computationally efficient for research purposes. The codes are contained in the source directory (`src`) and can be used and imported by other Julia projects. \n\n "]}  more » « less
Award ID(s):
2103049
PAR ID:
10677567
Author(s) / Creator(s):
; ; ; ;
Publisher / Repository:
Zenodo
Date Published:
Edition / Version:
v1.0
Format(s):
Medium: X
Right(s):
MIT License
Sponsoring Org:
National Science Foundation
More Like this
  1. {"Abstract":["This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package."]} 
    more » « less
  2. {"Abstract":["This release archives the Julia code and curated data products supporting the paper "Universality of stochastic control of quantum chaos with measurement and feedback".\n\nControlCAT provides reproducible workflows for the controlled quantum Arnold cat-map calculations in the paper, including exact quantum trajectories, truncated-Wigner approximation simulations, quantum/TWA comparison plots, and the (p, theta) TWA phase diagram. It also includes inverted harmonic oscillator simulation and steady-state analysis scripts used for the paper's Fig. 2 and Fig. 3 results.\n\nHighlights\n\n\n\nImplements the quantum Arnold cat map and probabilistic measurement-feedback control simulations.\n\nProvides CTWA dynamics for large effective system sizes and phase-diagram sweeps.\n\nIncludes production and smoke-test commands for regenerating paper-scale outputs.\n\nMaps scripts and archived data products to the paper's main figures.\n\nUses Julia 1.12 and the companion ControlMap.jl package.\n\nCode is MIT licensed; data, documentation, figures, and research artifacts are CC BY 4.0 licensed.\n\n\nGenerated .jld2, .csv, .png, and .pdf artifacts are intentionally not tracked in the main source tree, so regeneration is part of the archive workflow."]} 
    more » « less
  3. Bhattacharya, Sayan; Nanongkai, Danupon; Benedikt, Michael; Puppis, Gabriele (Ed.)
    Abstract: The relative-error property testing model was introduced in [Chen et al., 2024] to facilitate the study of property testing for "sparse" Boolean-valued functions, i.e. ones for which only a small fraction of all input assignments satisfy the function. In this framework, the distance from the unknown target function f that is being tested to a function g is defined as Vol(f△g)/Vol(f), where the numerator is the fraction of inputs on which f and g disagree and the denominator is the fraction of inputs that satisfy f. Recent work [Chen et al., 2026] has shown that over the Boolean domain {0,1}ⁿ, any relative-error testing algorithm for the fundamental class of {halfspaces} (i.e. linear threshold functions) must make Ω(log n) oracle calls. In this paper we complement the [Chen et al., 2026] lower bound by showing that halfspaces can be relative-error tested over ℝⁿ under the standard N(0,I_n) Gaussian distribution using a sublinear number of oracle calls - in particular, substantially fewer than would be required for learning. Our results use a wide range of tools including Hermite analysis, Gaussian isoperimetric inequalities, and geometric results on noise sensitivity and surface area. 
    more » « less
  4. Abstract: Structural variants (SV) are major drivers of evolutionary processes such as adaptation and speciation, yet their complexity and dynamics in wild populations remain largely unexplored. Avian diversity is highest in the Neotropics, primarily due to the suboscine passerine radiation; however, despite this diversity, genomic resources and studies of SVs in suboscines are scarce compared to their sister clade, the oscine passerines (“songbirds”). Here, we used long-read and chromatin conformation capture sequencing to assemble a high-quality scaffolded reference genome and construct a population-scale pangenome from 5 individuals of the Pearly-vented Tody-Tyrant (Hemitriccus margaritaceiventer), a suboscine bird with plumage variation across its distribution in South American dry forests. Our pangenome graph reveals extensive structural variation, with the chromosomal distribution of SVs strongly predicted by simple and low-complexity repeats – highlighting how specific repeat architecture may influence genome evolution. We discovered intraspecific copy number variation in multigene families, with the most complex instance including beta-keratin genes. Lastly, weidentified a 306 kb inversion spanning several melanin pigmentation-associated genes (e.g. MREG, MLPH, RAB17), making it a potential candidate SV for known intraspecific plumage variation. Our study establishes a population-scale pangenome resource for a suboscine bird, enabling characterization of the genome-wide abundance, diversity, and distribution of SVs within this species. TechnicalInfo: # Pangenomes reveal extensive structural variation in a suboscine passerine bird, the Pearly-vented Tody-Tyrant (Hemitriccus margaritaceiventer) ## Overview This repository contains data files related to the pangenome analysis of Hemitriccus margaritaceiventer genomes. The data includes: * **PGGB Pangenome Graph Files (`pggb_graphs.tar.gz`)**: The pangenome graphs are stored in [GFA v1](https://gfa-spec.github.io/GFA-spec/) (Graphical Fragment Assembly) format and compressed with gzip (`*.gfa.gz`). Each GFA file encodes a bidirected sequence graph for a single scaffold (e.g. `scaffold_20.pan.fa.gz.gfaffix.unchop.Ygs.view.gfa.gz`). - **`S`** **lines (segments)** define graph nodes and their DNA sequence. * Example:\ `S 19 TTTGCCCCCAGTCTGACTCCCAGTTTGCCCCTCAGTTTGGCCCCAGTTGCCCTCCCAGTTTGCCCCCAATTTCACTCCCCATTT`\ Here, `19` is the segment (node) identifier and the third field is the nucleotide sequence assigned to that node. - **`L`** **lines (links)** define edges between segments, including orientation. * Example:\ `L 19 + 21 + 0M`\ links the end of segment `19` in the forward orientation (`+`) to the start of segment `21` in the forward orientation (`+`) with an overlap descriptor `0M` (no shared bases; end‑to‑end adjacency as produced by PGGB/`gfaffix`). - Segment identifiers (e.g. `17`, `18`, `19`, …) are local to each scaffold graph and correspond to chopped pieces of the underlying multiple sequence alignment; they can be traversed along graph paths to reconstruct haplotypes and reference‑like sequences. There is one graph per scaffold (31 total: `scaffold_1`–`scaffold_34`). We excluded the scaffolds corresponding to sex chromosomes, as well as scaffold 30, because its assignment to a specific community (groups of contigs of all the pseudo-haplotypes which best correspond to the H. margaritaceiventer reference scaffold) was inconsistent across samples; rather than forming a distinct community, it was variably grouped with different autosomal communities, unlike all other scaffolds, which were reproducibly assigned. Therefore, we excluded this scaffold due to these community placement ambiguities * **Final PGGB Variation File (`pggb_variation_overlaps_final.tab.gz`)**: A gzipped Variant Call Format (VCF)-like tab file containing variant information decomposed from the pangenome graphs (see Supplementary Methods in the paper for variant identification and classification details). ## Usage notes and recommended software All file types in this repository can be opened and processed using free and open-source software. * **GFA graph files (`*.gfa.gz`)** * Can be visualized and/or processed with: * **vg** (Graph Genome Toolkit): [https://github.com/vgteam/vg](https://github.com/vgteam/vg) * **odgi**: [https://github.com/pangenome/odgi](https://github.com/pangenome/odgi) * **Bandage / BandageNG** for graphical exploration of assembly graphs: [https://rrwick.github.io/Bandage/](https://rrwick.github.io/Bandage/) * Example to inspect graph statistics with `odgi`: ```bash zcat scaffold_20.pan.fa.gz.gfaffix.unchop.Ygs.view.gfa.gz \ | odgi build -g - -o scaffold_20.og odgi stats -i scaffold_20.og ``` ### Processed Variant Information (`pggb_variation_overlaps_final.tab.gz`) #### Column Descriptions: * **chrom**: Chromosome identifier for the variant position (with PanSN-spec naming wih "HemMar#1#" prefix before scaffold name; see Methods). * **bedStart**: Start position of the variant on the chromosome (0-based). - **bedEnd**: End position of the variant on the chromosome (1-based). - **type**: Bcftools type of structural variant, one of SNP, MNP, INDEL, or OTHER. - **overlap**: Genomic region(s) overlapped by variant. Either intergenic, cds, intron, or none. The overlap is based on coordinates from the annotated H. margaritaceiventer reference genome. - **repeat**: Comma-separated list of repeatmasker annotations overlapped by variant (otherwise none); with specific repeat name. Overlap is based on coordinates from the annotated the H. margaritaceiventer reference genome. - **repeat_family**: Comma-separated list of the repeat family of the repeat overlapped by variant (DNA, Simple_repeat, LTR, Low_complexity, LINE, SINE, Unknown, Satellite, Other, or none). - **subtype**: Specific subtype of variant. Either SNP, SV, SVINS, SVDEL, INDEL, DEL, INS, and/or including Complex types of each variant. - **ref**: Reference allele nucleotide(s) at the variant position; based on the H. margaritaceiventer reference genome. - **alt**: Alternative allele nucleotide(s) at the variant position; based on the H. margaritaceiventer reference genome. - **aa**: Ancestral allele nucleotide(s) at the variant position; based on the two pseudo-haplotype genomes of the outgroup Pyrocephalus rubinus. - **inv**: Inversion present? Identified by PGGB; not the other inversion detection programs (Methods) - **polarized**: Is the variant polarized? False if aa == ".", True otherwise contains allele nucleotide(s) of the outgroup ancestral allele (aa). - **base_allele_len**: length of reference allele if aa == ".", otherwise length of aa allele. - **alt_len_max**: Longest non-base allele length. - **alt_len_min**: Shortest non-base allele length. - **allele_count**: Total Allele Count out of 10 total pseudo-haplotype assemblies. Missing data reduces this count. - **HMRG_DAC**: Derived Allele Count out of 10 total pseudo-haplotype assemblies. Missing data reduces this count. - **HMRG_AN**: Number of alleles out of 10 total pseudo-haplotype samples. Missing data reduces this count. - **HMRG_MISS**: Missing genotype counts. Count of missing ('.') genotype alleles out of 10 total pseudo-haplotype assemblies. - **HMRG_6371**: Genotype for HMRG_6371. Either 1|1, 1|0, 0|1, 0|0, and/or including missing data as indicated by '.'. - **HMRG_6386**: Genotype for HMRG_6386. Either 1|1, 1|0, 0|1, 0|0, and/or including missing data as indicated by '.'. - **HMRG_6388**: Genotype for HMRG_6388. Either 1|1, 1|0, 0|1, 0|0, and/or including missing data as indicated by '.'. - **HMRG_6431**: Genotype for HMRG_6431. Either 1|1, 1|0, 0|1, 0|0, and/or including missing data as indicated by '.'. - **HMRG_6433**: Genotype for HMRG_6433. Either 1|1, 1|0, 0|1, 0|0, and/or including missing data as indicated by '.'. ## Methods: Genomes for the Pearly-vented Tody-Tyrant (Hemitriccus margaritaceiventer) were assembled using Hifiasm ([https://github.com/chhylp123/hifiasm](https://github.com/chhylp123/hifiasm)). For pangenome graph construction, we used the PanGenome Graph Builder (PGGB, [https://github.com/pangenome/pggb](https://github.com/pangenome/pggb); Garrison et al. 2024) via the nf-core/nextflow pipeline ([https://nf-co.re/pangenome/1.0.0/](https://nf-co.re/pangenome/1.0.0/)). Contigs were first assigned to chromosomes using wfmash mapping and split-mapping for unmapped contigs. PGGB graph induction and normalization used wfmash, seqwish, smoothxg, and gfaffix. Variant decomposition from the resulting graphs was performed using vg deconstruct (v1.40.0, [https://github.com/vgteam/vg](https://github.com/vgteam/vg)), with filtering for large variants with vcfbub ([https://github.com/pangenome/vcfbub](https://github.com/pangenome/vcfbub)) and conversion to primitive alleles using vcfwave (Garrison et al., 2022). Variant files were merged and processed with bcftools, and additional annotation and summary tables of variants were generated using BEDtools and custom Python scripts. Full details of all software, versions, and parameter settings are described in the supplementary methods and GitHub repository ([https://github.com/kelsiealopez/Hemitriccus-Pangenome](https://github.com/kelsiealopez/Hemitriccus-Pangenome)). Garrison, E., Guarracino, A., Heumos, S., Villani, F., Bao, Z., Tattini, L., Hagmann, J., Vorbrugg, S., Marco-Sola, S., Kubica, C., Ashbrook, D. G., Thorell, K., Rusholme-Pilcher, R. L., Liti, G., Rudbeck, E., Golicz, A. A., Nahnsen, S., Yang, Z., Mwaniki, M. N., … Prins, P. (2024). Building pangenome graphs. Nature Methods, 21(11), 2008–2012. [https://doi.org/10.1038/s41592-024-02430-3](https://doi.org/10.1038/s41592-024-02430-3) Garrison, E., Kronenberg, Z. N., Dawson, E. T., Pedersen, B. S., & Prins, P. (2022). A spectrum of free software tools for processing the VCF variant call format: Vcflib, bio-vcf, cyvcf2, hts-nim and slivar. PLOS Computational Biology, 18(5), e1009123. [https://doi.org/10.1371/journal.pcbi.1009123](https://doi.org/10.1371/journal.pcbi.1009123) 
    more » « less
  5. {"Abstract":["This dataset contains machine learning and volunteer classifications from the Gravity Spy project. It includes glitches from observing runs O1, O2, O3a and O3b that received at least one classification from a registered volunteer in the project. It also indicates glitches that are nominally retired from the project using our default set of retirement parameters, which are described below. See more details in the Gravity Spy Methods paper. <\/p>\n\nWhen a particular subject in a citizen science project (in this case, glitches from the LIGO datastream) is deemed to be classified sufficiently it is "retired" from the project. For the Gravity Spy project, retirement depends on a combination of both volunteer and machine learning classifications, and a number of parameterizations affect how quickly glitches get retired. For this dataset, we use a default set of retirement parameters, the most important of which are: <\/p>\n\nA glitches must be classified by at least 2 registered volunteers<\/li>Based on both the initial machine learning classification and volunteer classifications, the glitch has more than a 90% probability of residing in a particular class<\/li>Each volunteer classification (weighted by that volunteer's confusion matrix) contains a weight equal to the initial machine learning score when determining the final probability<\/li><\/ol>\n\nThe choice of these and other parameterization will affect the accuracy of the retired dataset as well as the number of glitches that are retired, and will be explored in detail in an upcoming publication (Zevin et al. in prep). <\/p>\n\nThe dataset can be read in using e.g. Pandas: \n```\nimport pandas as pd\ndataset = pd.read_hdf('retired_fulldata_min2_max50_ret0p9.hdf5', key='image_db')\n```\nEach row in the dataframe contains information about a particular glitch in the Gravity Spy dataset. <\/p>\n\nDescription of series in dataframe<\/strong><\/p>\n\n['1080Lines', '1400Ripples', 'Air_Compressor', 'Blip', 'Chirp', 'Extremely_Loud', 'Helix', 'Koi_Fish', 'Light_Modulation', 'Low_Frequency_Burst', 'Low_Frequency_Lines', 'No_Glitch', 'None_of_the_Above', 'Paired_Doves', 'Power_Line', 'Repeating_Blips', 'Scattered_Light', 'Scratchy', 'Tomte', 'Violin_Mode', 'Wandering_Line', 'Whistle']\n\tMachine learning scores for each glitch class in the trained model, which for a particular glitch will sum to unity<\/li><\/ul>\n\t<\/li>['ml_confidence', 'ml_label']\n\tHighest machine learning confidence score across all classes for a particular glitch, and the class associated with this score<\/li><\/ul>\n\t<\/li>['gravityspy_id', 'id']\n\tUnique identified for each glitch on the Zooniverse platform ('gravityspy_id') and in the Gravity Spy project ('id'), which can be used to link a particular glitch to the full Gravity Spy dataset (which contains GPS times among many other descriptors)<\/li><\/ul>\n\t<\/li>['retired']\n\tMarks whether the glitch is retired using our default set of retirement parameters (1=retired, 0=not retired)<\/li><\/ul>\n\t<\/li>['Nclassifications']\n\tThe total number of classifications performed by registered volunteers on this glitch<\/li><\/ul>\n\t<\/li>['final_score', 'final_label']\n\tThe final score (weighted combination of machine learning and volunteer classifications) and the most probable type of glitch<\/li><\/ul>\n\t<\/li>['tracks']\n\tArray of classification weights that were added to each glitch category due to each volunteer's classification<\/li><\/ul>\n\t<\/li><\/ul>\n\n <\/p>\n\n```\nFor machine learning classifications on all glitches in O1, O2, O3a, and O3b, please see Gravity Spy Machine Learning Classifications on Zenodo<\/p>\n\nFor the most recently uploaded training set used in Gravity Spy machine learning algorithms, please see Gravity Spy Training Set on Zenodo.<\/p>\n\nFor detailed information on the training set used for the original Gravity Spy machine learning paper, please see Machine learning for Gravity Spy: Glitch classification and dataset on Zenodo. <\/p>"]} 
    more » « less