Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Abstract Machine learning (ML) algorithms have emerged in many meteorological applications. However, these algorithms struggle to extrapolate beyond the data they were trained on, i.e., they may adopt faulty strategies that lead to catastrophic failures. These failures are difficult to predict due to the opaque nature of ML algorithms. In high-stakes applications, such as severe weather forecasting, it is crucial to avoid such failures. One approach to address this issue is to develop more interpretable ML algorithms. The primary goal of this work is to illustrate the use of a specific interpretable ML algorithm that has not yet found much use in meteorology, explainable boosting machines (EBMs). We demonstrate that EBMs are particularly suitable to implement human-guided strategies in an ML algorithm. As a guiding example, we show how to develop an EBM to detect overshooting tops (OTs) in satellite imagery. EBMs require input features to be scalar. We use techniques from knowledge-guided machine learning to first extract scalar features from meteorological imagery. For the application of identifying OTs, this includes extracting cloud texture from satellite imagery using gray-level co-occurrence matrices. Once trained, the EBM was examined and minimally altered to more closely match strategies used by domain scientists to identify OTs. The result of our efforts is a fully interpretable ML algorithm developed in a human–machine collaboration that uses human-guided strategies. While the final model does not reach the accuracy of more complex approaches, it performs reasonably well, and we hope it paves the way for building more interpretable ML algorithms for this and other meteorological applications. Significance StatementThe purpose of this work is to introduce the interpretable machine learning method of explainable boosting machines (EBMs) to an atmospheric science audience by closely examining how they can be used to detect the location of overshooting cloud tops, which have been associated with the occurrence of severe weather, in satellite imagery. Interpretable machine learning methods are important in high-risk situations such as overshooting top detection as they allow forecasters to have a better understanding of how exactly an identification is made. We walk through how to build, interpret, and modify the machine learning algorithm for use on this task and discuss other applications that could benefit from this approach.more » « lessFree, publicly-accessible full text available July 1, 2027
-
Abstract Improving the skill of medium-range (3–8 day) severe weather prediction is crucial for mitigating societal impacts. This study introduces a novel approach leveraging decoder-only transformer networks to post-process AI-based weather forecasts, specifically from the Pangu-Weather model, for improved severe weather guidance. Unlike traditional post-processing methods that use a dense neural network to predict the probability of severe weather using discrete forecast samples, our method treats forecast lead times as sequential “tokens”, enabling the transformer to learn complex temporal relationships within the evolving atmospheric state. We compare this approach against post-processing of the Global Forecast System (GFS) using both a traditional dense neural network and our transformer, as well as configurations that exclude convective parameters to fairly evaluate the impact of using the Pangu-Weather AI model. Results demonstrate that the transformer-based post-processing significantly enhances forecast skill compared to dense neural networks. Furthermore, AI-driven forecasts, particularly Pangu-Weather initialized from high resolution analysis, exhibit superior performance to GFS in the medium-range, even without explicit convective parameters. Our approach offers improved accuracy, and reliability, which also provides interpretability through feature attribution analysis, advancing medium-range severe weather prediction capabilities.more » « lessFree, publicly-accessible full text available November 25, 2026
-
Abstract An ensemble postprocessing method is developed to improve the probabilistic forecasts of extreme precipitation events across the conterminous United States (CONUS). The method combines a 3D vision transformer (ViT) for bias correction with a latent diffusion model (LDM), a generative artificial intelligence (AI) method, to postprocess 6-hourly precipitation ensemble forecasts and produce an enlarged generative ensemble that contains spatiotemporally consistent precipitation trajectories. These trajectories are expected to improve the characterization of extreme precipitation events and offer skillful multiday accumulated and 6-hourly precipitation guidance. The method is tested using the Global Ensemble Forecast System (GEFS) precipitation forecasts out to day 6 and is verified against the Climatology-Calibrated Precipitation Analysis (CCPA) data. Verification results indicate that the method generated skillful ensemble members with improved continuous ranked probabilistic skill scores (CRPSSs) and Brier skill scores (BSSs) over the raw operational GEFS and a multivariate statistical postprocessing baseline. It showed skillful and reliable probabilities for events at extreme precipitation thresholds. Explainability studies were further conducted, which revealed the decision-making process of the method and confirmed its effectiveness on ensemble member generation. This work introduces a novel, generative AI–based approach to address the limitation of small numerical ensembles and the need for larger ensembles to identify extreme precipitation events. Significance StatementWe use a new artificial intelligence (AI) technique to improve extreme precipitation forecasts from a numerical weather prediction ensemble, generating more scenarios that better characterize extreme precipitation events. This AI-generated ensemble improved the accuracy of precipitation forecasts and probabilistic warnings for extreme precipitation events. The study explores AI methods to generate precipitation forecasts and explains the decision-making mechanisms of such AI techniques to prove their effectiveness.more » « less
-
Abstract As an increasing number of machine learning (ML) products enter the research-to-operations (R2O) pipeline, researchers have anecdotally noted a perceived hesitancy by operational forecasters to adopt this relatively new technology. One explanation often cited in the literature is that this perceived hesitancy derives from the complex and opaque nature of ML methods. Because modern ML models are trained to solve tasks by optimizing a potentially complex combination of mathematical weights, thresholds, and nonlinear cost functions, it can be difficult to determine how these models reach a solution from their given input. However, it remains unclear to what degree a model’s transparency may influence a forecaster’s decision to use that model or if that impact differs between ML and more traditional (i.e., non-ML) methods. To address this question, a survey was offered to forecaster and researcher participants attending the 2021 NOAA Hazardous Weather Testbed (HWT) Spring Forecasting Experiment (SFE) with questions about how participants subjectively perceive and compare machine learning products to more traditionally derived products. Results from this study revealed few differences in how participants evaluated machine learning products compared to other types of guidance. However, comparing the responses between operational forecasters, researchers, and academics exposed notable differences in what factors the three groups considered to be most important for determining the operational success of a new forecast product. These results support the need for increased collaboration between the operational and research communities. Significance StatementParticipants of the 2021 Hazardous Weather Testbed Spring Forecasting Experiment were surveyed to assess how machine learning products are perceived and evaluated in operational settings. The results revealed little difference in how machine learning products are evaluated compared to more traditional methods but emphasized the need for explainable product behavior and comprehensive end-user training.more » « less
-
Abstract Studies suggest a strong link between low‐frequency sea level variability in the South Atlantic Bight (SAB) and open ocean dynamics. However, the mechanisms driving this connection remain unclear. By analyzing a high‐resolution, three‐dimensional baroclinic ocean reanalysis, we identify a pathway that links open ocean dynamics to SAB coastal sea level variability through the shelf edge near Cape Hatteras. Gulf Stream meanders in this region induce sea level fluctuations that propagate along the entire SAB shelf. Using an idealized barotropic model, we further demonstrate that topographic waves mediate the propagation of the Gulf Stream signal onto the shelf. Moreover, the Gulf Stream variability is driven by zonal wind stress in the Northwest Atlantic, which is likely modulated by the North Atlantic Oscillation. These findings offer new insights into regional sea level prediction and contribute to broader climate research efforts.more » « less
-
Abstract Building upon recent advancements in AI‐driven atmospheric emulation, we present a novel framework for AI‐based ocean emulation, downscaling, and bias correction, with a specific focus on high‐resolution modeling of the regional ocean in the Gulf of Mexico. Emulating regional ocean dynamics poses distinct challenges due to intricate bathymetry, complex lateral boundary conditions, and inherent limitations of deep learning models, including instability and the potential for hallucinations. In this study, we introduce a deep learning framework that autoregressively integrates ocean surface variables at 8 km spatial resolution over the Gulf of Mexico, maintaining physical consistency over decadal time scales. Simultaneously, the framework downscales and bias‐corrects the outputs to 4 km resolution using a physics‐informed generative model. Our approach demonstrates short‐term predictive skill comparable to high‐resolution physics‐based simulations, while also accurately capturing long‐term statistical properties, including temporal mean and variability.more » « less
-
Abstract The benefits of collaboration between the research and operational communities during the research-to-operations (R2O) process have long been documented in the scientific literature. Operational forecasters have a practiced, expert insight into weather analysis and forecasting but typically lack the time and resources for formal research and development. Conversely, many researchers have the resources, theoretical knowledge, and formal experience to solve complex meteorological challenges but lack an understanding of operation procedures, needs, requirements, and authority necessary to effectively bridge the R2O gap. Collaboration then serves as the most viable strategy to further a better understanding and improved prediction of atmospheric processes via ongoing multi-disciplinary knowledge transfer between the research and operational communities. However, existing R2O processes leave room for improvement when it comes to collaboration throughout a new product’s development cycle. This study assesses the subjective importance of collaboration at various stages of product development via a survey presented to participants of the 2021 Hazardous Weather Testbed Spring Forecasting Experiment. This feedback is then applied to create a proposed new R2O workflow that combines components from existing R2O procedures and modern co-production philosophies.more » « less
-
Abstract Pure artificial intelligence (AI)-based weather prediction (AIWP) models have made waves within the scientific community and the media, claiming superior performance to numerical weather prediction (NWP) models. However, these models often lack impactful output variables such as precipitation. One exception is Google DeepMind’s GraphCast model, which became the first mainstream AIWP model to predict precipitation, but performed only limited verification. We present an analysis of the ECMWF’s Integrated Forecasting System (IFS)-initialized (GRAPIFS) and the NCEP’s Global Forecast System (GFS)-initialized (GRAPGFS) GraphCast precipitation forecasts over the contiguous United States and compare to results from the GFS and IFS models using 1) grid-based, 2) neighborhood, and 3) object-oriented metrics verified against the fifth major global reanalysis produced by ECMWF (ERA5) and the NCEP/Environmental Modeling Center (EMC) stage IV precipitation analysis datasets. We affirmed that GRAPGFSand GRAPIFSperform better than the GFS and IFS in terms of root-mean-square error and stable equitable errors in probability space, but the GFS and IFS precipitation distributions more closely align with the ERA5 and stage IV distributions. Equitable threat score also generally favored GraphCast, particularly for lower accumulation thresholds. Fractions skill score for increasing neighborhood sizes shows greater gains for the GFS and IFS than GraphCast, suggesting the NWP models may have a better handle on intensity but struggle with the location. Object-based verification for GraphCast found positive area biases at low accumulation thresholds and large negative biases at high accumulation thresholds. GRAPGFSsaw similar performance gains to GRAPIFSwhen compared to their NWP counterparts, but initializing with the less familiar GFS conditions appeared to lead to an increase in light precipitation. Significance StatementPure artificial intelligence (AI)-based weather prediction (AIWP) has exploded in popularity with promises of better performance and faster run times than numerical weather prediction (NWP) models. However, less attention has been paid to their capability to predict impactful, sensible weather like precipitation, precipitation type, or specific meteorological features. We seek to address this gap by comparing precipitation forecast performance by an AI model called GraphCast to the Global Forecast System (GFS) and the Integrated Forecasting System (IFS) NWP models. While GraphCast does perform better on many verification metrics, it has some limitations for intense precipitation forecasts. In particular, it less frequently predicts intense precipitation events than the GFS or IFS. Overall, this article emphasizes the promise of AIWP while at the same time stresses the need for robust verification by domain experts.more » « less
-
Abstract Subseasonal‐to‐decadal atmospheric prediction skill attained from initial conditions is typically limited by the chaotic nature of the atmosphere. However, for some atmospheric phenomena, prediction skill on subseasonal‐to‐decadal timescales is increased when the initial conditions are in a particular state. In this study, we employ machine learning to identify sea surface temperature (SST) regimes that enhance prediction skill of North Atlantic atmospheric circulation. An ensemble of artificial neural networks is trained to predict anomalous, low‐pass filtered 500 mb height at 7–8 weeks lead using SST. We then use self‐organizing maps (SOMs) constructed from 9 regions within the SST domain to detect state‐dependent prediction skill. SOMs are built using the entire SST time series, and we assess which SOM units feature confident neural network predictions. Four regimes are identified that provide skillful seasonal predictions of 500 mb height. Our findings demonstrate the importance of extratropical decadal SST variability in modulating downstream ENSO teleconnections to the North Atlantic. The methodology presented could aid future forecasting on subseasonal‐to‐decadal timescales.more » « less
-
Abstract AI-based algorithms are emerging in many meteorological applications that produce imagery as output, including for global weather forecasting models. However, the imagery produced by AI algorithms, especially by convolutional neural networks (CNNs), is often described as too blurry to look realistic, partly because CNNs tend to represent uncertainty as blurriness. This blurriness can be undesirable since it might obscure important meteorological features. More complex AI models, such as Generative AI models, produce images that appear to be sharper. However, improved sharpness may come at the expense of a decline in other performance criteria, such as standard forecast verification metrics. To navigate any trade-off between sharpness and other performance metrics it is important to quantitatively assess those other metrics along with sharpness. While there is a rich set of forecast verification metrics available for meteorological images, none of them focus on sharpness. This paper seeks to fill this gap by 1) exploring a variety of sharpness metrics from other fields, 2) evaluating properties of these metrics, 3) proposing the new concept of Gaussian Blur Equivalence as a tool for their uniform interpretation, and 4) demonstrating their use for sample meteorological applications, including a CNN that emulates radar imagery from satellite imagery (GREMLIN) and an AI-based global weather forecasting model (GraphCast).more » « less
An official website of the United States government
