Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 10:00 PM ET on Thursday, July 16 until 12:00 AM ET on Friday 17 due to maintenance. We apologize for the inconvenience.


Search for: All records

Award ID contains: 1835860

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Abstract We propose and unify classes of different models for information propagation over graphs. In a first class, propagation is modelled as a wave, which emanates from a set ofknownnodes at an initial time, to all otherunknownnodes at later times with an ordering determined by the arrival time of the information wave front. A second class of models is based on the notion of a travel time along paths between nodes. The time of information propagation from an initialknownset of nodes to a node is defined as the minimum of a generalised travel time over subsets of all admissible paths. A final class is given by imposing a local equation of an eikonal form at eachunknownnode, with boundary conditions at theknownnodes. The solution value of the local equation at a node is coupled to those of neighbouring nodes with lower values. We provide precise formulations of the model classes and prove equivalences between them. Finally, we apply the front propagation models on graphs to semi-supervised learning via label propagation and information propagation on trust networks. 
    more » « less
    Free, publicly-accessible full text available October 1, 2026
  2. Abstract This paper presents the dynamical core of the Climate Modeling Alliance (CliMA) atmosphere model, designed for efficient simulation of a wide range of atmospheric flows across scales. The core uses the nonhydrostatic equations of motion for a deep atmosphere, discretized with a hybrid approach that combines a spectral element method (SEM) in the horizontal and a staggered finite‐difference method in a height‐based, terrain‐following coordinate in the vertical. This approach leverages the high‐order accuracy and scalability of the SEM, while maintaining the computational efficiency and stability of finite differences on a staggered grid. The model's coordinate‐independent equation set allows for simulations in a variety of geometries and planetary configurations. The use of the specific total energy of moist air as a prognostic variable, along with a consistent thermodynamic formulation, ensures the conservation of energy, air mass, and water mass, even in moist atmospheres and in the presence of subgrid‐scale parameterizations, without ad hoc fixers. A horizontally explicit, vertically implicit (HEVI) timestepping strategy treats fast vertical processes implicitly and further enhances computational efficiency by allowing larger timesteps. The model demonstrates excellent strong and weak scaling on CPUs and GPUs, making it well‐suited for high‐resolution simulations on modern supercomputing architectures, including those on the cloud, which widens access to climate models. 
    more » « less
    Free, publicly-accessible full text available March 1, 2027
  3. Abstract Cloud microphysics is a critical aspect of the Earth's climate system, which involves processes at the nano‐ and micrometer scales of droplets and ice particles. In climate modeling, cloud microphysics is commonly represented by bulk models, which contain simplified process rates that require calibration. This study presents a framework for calibrating warm‐rain bulk schemes using high‐fidelity super‐droplet simulations that provide a more accurate and physically based representation of cloud and precipitation processes. The calibration framework employs ensemble Kalman methods including Ensemble Kalman Inversion and Unscented Kalman Inversion to calibrate bulk microphysics schemes with probabilistic super‐droplet simulations. We demonstrate the framework's effectiveness by calibrating a single‐moment bulk scheme, resulting in a reduction of data‐model mismatch by more than 75% compared to the model with initial parameters. Thus, this study demonstrates a powerful tool for enhancing the accuracy of bulk microphysics schemes in atmospheric models and improving climate modeling. 
    more » « less
  4. Abstract We present a method to downscale idealized geophysical fluid simulations using generative models based on diffusion maps. By analyzing the Fourier spectra of fields drawn from different data distributions, we show how a diffusion bridge can be used as a transformation between a low resolution and a high resolution dataset, allowing for new sample generation of high-resolution fields given specific low resolution features. The ability to generate new samples allows for the computation of any statistic of interest, without any additional calibration or training. Our unsupervised setup is also designed to downscale fields without access to paired training data; this flexibility allows for the combination of multiple source and target domains without additional training. We demonstrate that the method enhances resolution and corrects context-dependent biases in geophysical fluid simulations, including in extreme events. We anticipate that the same method can be used to downscale the output of climate simulations, including temperature and precipitation fields, without needing to train a new model for each application and providing a significant computational cost savings. 
    more » « less
  5. Physical Review Different approaches to using data-driven methods for subgrid-scale closure modeling of geophysical turbulence have emerged recently. Most of these approaches are data hungry and lack interpretability and out-of-distribution generalizability. Here, we use a hybrid approach that combines turbulence theory, physics-based modeling, and data-driven methods to overcome these challenges. Specifically, we address the parametric uncertainty of well-known physics-based large-eddy simulation (LES) closures: the Smagorinsky (Smag) and Leith eddy-viscosity models (one free parameter) and the Jansen-Held (JH) backscattering model (two free parameters). For various cases of two-dimensional turbulence, optimal parameters are first learned online from data via ensemble Kalman inversion (EKI), such that for each case, the LES energy spectrum matches that of direct numerical simulation (DNS). We quantify the uncertainties on these parameters using a modern machine-learning-accelerated Bayesian workflow, “, , ” Only a small training dataset is needed (to calculate the DNS spectra); i.e., the approach is data-efficient. We find the optimized parameter(s) and their associated uncertainty for each closure to be constant across broad flow regimes that differ in dominant length scales, eddy/jet structures, and dynamics, suggesting that these closures are generalizable. Next, we show that the online learned constants agree with the predictions of a recent semianalytical derivation, providing further interpretability. In both and tests that include examining the extreme events, LES with optimized closures, especially with JH, outperforms the baselines (LES with standard Smag, dynamic Smag, or Leith). This work shows the promise of combining advances in theory, physics-based modeling (e.g., JH), and data-driven modeling (e.g., online learning with EKI) to develop data-efficient frameworks for accurate, interpretable, and generalizable closures for geophysical turbulence, with ultimate applications in weather and climate prediction. 
    more » « less
    Free, publicly-accessible full text available May 1, 2027
  6. The filtering distribution in hidden Markov models evolves according to the law of a mean-field model in state–observation space. The ensemble Kalman filter (EnKF) approximates this mean-field model with an ensemble of interacting particles, employing a Gaussian ansatz for the joint distribution of the state and observation at each observation time. These methods are robust, but the Gaussian ansatz limits accuracy. Here this shortcoming is addressed by using machine learning to map the joint predicted state and observation to the updated state estimate. The derivation of methods from a mean field formulation of the true filtering distribution suggests a single parametrization of the algorithm that can be deployed at different ensemble sizes. And we use a mean field formulation of the ensemble Kalman filter as an inductive bias for our architecture. To develop this perspective, in which the mean-field limit of the algorithm and finite interacting ensemble particle approximations share a common set of parameters, a novel form of neural operator is introduced, taking probability distributions as input: a measure neural mapping (MNM). A MNM is used to design a novel approach to filtering, the MNM-enhanced ensemble filter (MNMEF), which is defined in both the mean-field limit and for interacting ensemble particle approximations. The ensemble approach uses empirical measures as input to the MNM and is implemented using the set transformer, which is invariant to ensemble permutation and allows for different ensemble sizes. In practice fine-tuning of a small number of parameters, for specific ensemble sizes, further enhances the accuracy of the scheme. The promise of the approach is demonstrated by its superior root-mean-square-error performance relative to leading methods in filtering the Lorenz ‘96 and Kuramoto-Sivashinsky models. 
    more » « less
    Free, publicly-accessible full text available February 1, 2027
  7. Ensemble Kalman methods, introduced in 1994 in the context of ocean state estimation, are now widely used for state estimation and parameter estimation (inverse problems) in many arenae. Their success stems from the fact that they take an underlying computational model as a black box to provide a systematic, derivative-free methodology for incorporating observations; furthermore the ensemble approach allows for sensitivities and uncertainties to be calculated. Analysis of the accuracy of ensemble Kalman methods, especially in terms of uncertainty quantification, is lagging behind empirical success; this paper provides a unifying mean-field-based framework for their analysis. Both state estimation and parameter estimation problems are considered, and formulations in both discrete and continuous time are employed. For state estimation problems, both the control and filtering approaches are considered; analogously for parameter estimation problems, the optimization and Bayesian perspectives are both studied. As well as providing an elegant framework, the mean-field perspective also allows for the derivation of a variety of methods used in practice. In addition it unifies a wide-ranging literature in the field and suggests open problems. 
    more » « less
    Free, publicly-accessible full text available July 1, 2026
  8. Randomized algorithms exploit stochasticity to reduce computational complexity. One important example is random feature regression (RFR) that accelerates Gaussian process regression (GPR). RFR approximates an unknown function with a random neural network whose hidden weights and biases are sampled from a probability distribution. Only the final output layer is fit to data. In randomized algorithms like RFR, the hyperparameters that characterize the sampling distribution greatly impact performance, yet are not directly accessible from samples. This makes optimization of hyperparameters via standard (gradient-based) optimization tools inapplicable. Inspired by Bayesian ideas from GPR, this paper introduces a random objective function that is tailored for hyperparameter tuning of vector-valued random features. The objective is minimized with ensemble Kalman inversion (EKI). EKI is a gradient-free particle-based optimizer that is scalable to high-dimensions and robust to randomness in objective functions. A numerical study showcases the new black-box methodology to learn hyperparameter distributions in several problems that are sensitive to the hyperparameter selection: two global sensitivity analyses, integrating a chaotic dynamical system, and solving a Bayesian inverse problem from atmospheric dynamics. The success of the proposed EKI-based algorithm for RFR suggests its potential for automated optimization of hyperparameters arising in other randomized algorithms. 
    more » « less
  9. We propose a sampling method based on an ensemble approximation of second order Langevin dynamics. The log target density is appended with a quadratic term in an auxiliary momentum variable and damped-driven Hamiltonian dynamics introduced; the resulting stochastic differential equation is invariant to the Gibbs measure, with marginal on the position coordinates given by the target. A preconditioner based on covariance under the law of position coordinates under the dynamics does not change this invariance property, and is introduced to accelerate convergence to the Gibbs measure. The resulting mean-field dynamics may be approximated by an ensemble method; this results in a gradient-free and affine-invariant stochastic dynamical system with desirable provably uniform convergence properties across the class of all Gaussian targets. Numerical results demonstrate the potential of the method as the basis for a numerical sampler in Bayesian inverse problems, beyond the Gaussian setting. 
    more » « less
  10. This work integrates machine learning into an atmospheric parameterization to target uncertain mixing processes while maintaining interpretable, predictive, and well‐established physical equations. We adopt an eddy‐diffusivity mass‐flux (EDMF) parameterization for the unified modeling of various convective and turbulent regimes. To avoid drift and instability that plague offline‐trained machine learning parameterizations that are subsequently coupled with climate models, we frame learning as an inverse problem: Data‐driven models are embedded within the EDMF parameterization and trained online in a one‐dimensional vertical global climate model (GCM) column. Training is performed against output from large‐eddy simulations (LES) forced with GCM‐simulated large‐scale conditions in the Pacific. Rather than optimizing subgrid‐scale tendencies, our framework directly targets climate variables of interest, such as the vertical profiles of entropy and liquid water path. Specifically, we use ensemble Kalman inversion to simultaneously calibrate both the EDMF parameters and the parameters governing data‐driven lateral mixing rates. The calibrated parameterization outperforms existing EDMF schemes, particularly in tropical and subtropical locations of the present climate, and maintains high fidelity in simulating shallow cumulus and stratocumulus regimes under increased sea surface temperatures from AMIP4K experiments. The results showcase the advantage of physically constraining data‐driven models and directly targeting relevant variables through online learning to build robust and stable machine learning parameterizations. 
    more » « less