Abstract AI models have emerged as potential complements to physics‐based models, but their skill in capturing observed regional trends with important societal impacts remains unexplored. Here, we benchmark satellite‐era regional thermodynamic trends, including extremes, in an AI emulator (ACE2) and a hybrid model (NeuralGCM), against physics‐based models and ERA5. Both AI models capture regional temperature trends such as satellite‐era Arctic warming. ACE2 outperforms other models in capturing midlatitude vertical temperature trends. However, the AI models do not capture trends in heat extremes over the US Southwest. Furthermore, they do not capture drying trends in arid regions, but generally outperform physics‐based models. Our results show that a data‐driven AI emulator can perform comparably to, or better than, hybrid and physics‐based models in capturing regional thermodynamic trends. We also find that ACE2 learns much of the signal from .
more »
« less
This content will become publicly available on February 28, 2027
Benchmarking Atmospheric Circulation Variability in an AI Emulator, ACE2, and a Hybrid Model, NeuralGCM
Abstract Physics‐based atmosphere‐land models with prescribed sea surface temperature have notable successes but also biases in their ability to represent atmospheric variability compared to observations. Recently, AI emulators and hybrid models have emerged with the potential to overcome these biases, but still require systematic evaluation against metrics grounded in fundamental atmospheric dynamics. We evaluate the representation of four atmospheric variability benchmarking metrics in a fully data‐driven AI emulator (ACE2‐ERA5) and hybrid model (NeuralGCM). The hybrid model and emulator can capture the spectra of large‐scale tropical waves and extratropical eddy‐mean flow interactions, including critical levels. However, both struggle to capture the timescales associated with quasi‐biennial oscillation (QBO, ∼28 months) and Southern annular mode propagation (∼150 days). These dynamical metrics serve as an initial benchmarking tool to inform AI model development and understand their limitations, which may be essential for out‐of‐distribution applications (e.g., extrapolating to unseen climates).
more »
« less
- Award ID(s):
- 2544065
- PAR ID:
- 10671943
- Publisher / Repository:
- Geophysical Research Letters
- Date Published:
- Journal Name:
- Geophysical Research Letters
- Volume:
- 53
- Issue:
- 4
- ISSN:
- 0094-8276
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Abstract The current generation of earth system models (ESMs) exhibit substantial biases and inter-model spread in representing ocean deoxygenation and the expansion of tropical oxygen minimum zones (OMZs). Improving dissolved oxygen simulation is challenging as it depends on uncertain physical and biogeochemical processes and their interactions. A machine learning-based emulator, O2EMU, aims to reduce model biases and inter-model spread by replacing ESMs’ biogeochemical parameterizations with the learned relationships between dissolved oxygen and physical variables. The emulator first learns from the historical shipboard and autonomous observations of dissolved oxygen and temperature and salinity from ocean reanalyses. The learned relationships are then applied to the temperature and salinity outputs from an ensemble of ESM historical simulations. The results show significantly improved climatological OMZ structure and long-term oxygen trends including the shoaling and expansion of the tropical Atlantic OMZs. Application of O2EMU to ESM simulations can reduce the inter-model spread by 70%–75% and the root mean square error by 50%–60%, offering a computationally efficient and scalable alternative to standard biogeochemical models, and providing a stepping stone towardshybridbiogeochemical projections that blend mechanistic ESM dynamics with observationally constrained tracer distributions.more » « less
-
Abstract This study presents a comprehensive climatological benchmarking of tropical cyclones (TCs) generated by AI‐based global weather prediction models. Using all TC events from the North Atlantic and Western Pacific basins between 2020 and 2025, we assess the ability of two AI models (Pangu‐Weather and Aurora) to reproduce observed TC track density, climatology of storm characteristics, and physical consistency with TC theory. By comparing AI‐simulated TCs with ERA5 reanalysis, we benchmark the distributions of intensity, size, forward speed, and evaluate the model's ability to credibly simulate extratropical transition. Results show that both Pangu and Aurora perform well in reproducing storm track density, forward speed distribution, and outer size distribution. Aurora shows an improved performance in simulating storm intensity compared to Pangu, with less bias in the distribution of minimum central pressure and maximum wind speed. However, both models overestimate the distribution of storm inner size (radius of maximum winds), especially for extreme events. AI models capture the relative frequency and temporal evolution of extratropical transition patterns with reasonable accuracy. The AI‐simulated TCs are also less likely to conform to gradient wind balance compared to ERA5, indicating that the AI TCs may not be physically realistic in many cases. This benchmarking identifies systematic biases that can guide future corrections and support extended applications of AI models for TC hazard and risk assessment. Our work establishes a foundation for future studies using AI weather models in the context of TC climatological and hazard research.more » « less
-
Abstract Stellar spectra emulators often rely on large grids and tend to reach a plateau in emulation accuracy, leading to significant systematic errors when inferring stellar properties. Our study explores the use of Transformer models to capture long-range information in spectra, comparing their performance to the Payne emulator (a fully connected multilayer perceptron), an expanded version of The Payne, and a convolutional-based emulator. We tested these models on synthetic spectral grids, evaluating their performance by analyzing emulation residuals and assessing the quality of spectral parameter inference. The newly introduced TransformerPayne emulator outperformed all other tested models, achieving a mean absolute error (MAE) of approximately 0.15% when trained on the full grid. The most significant improvements were observed in grids containing between 1000 and 10,000 spectra, with TransformerPayne showing 2–5 times better performance than the scaled-up version of The Payne. Additionally, TransformerPayne demonstrated superior fine-tuning capabilities, allowing for pretraining on one spectral model grid before transferring to another. This fine-tuning approach enabled up to a 10-fold reduction in training grid size compared to models trained from scratch. Analysis of TransformerPayne's attention maps revealed that they encode interpretable features common across many spectral lines of chosen elements. While scaling up The Payne to a larger network reduced its MAE from 1.2% to 0.3% when trained on the full data set, TransformerPayne consistently achieved the lowest MAE across all tests. The inductive biases of the TransformerPayne emulator enhance accuracy, data efficiency, and interpretability for spectral emulation compared to existing methods.more » « less
-
Abstract Pure artificial intelligence (AI)-based weather prediction (AIWP) models have made waves within the scientific community and the media, claiming superior performance to numerical weather prediction (NWP) models. However, these models often lack impactful output variables such as precipitation. One exception is Google DeepMind’s GraphCast model, which became the first mainstream AIWP model to predict precipitation, but performed only limited verification. We present an analysis of the ECMWF’s Integrated Forecasting System (IFS)-initialized (GRAPIFS) and the NCEP’s Global Forecast System (GFS)-initialized (GRAPGFS) GraphCast precipitation forecasts over the contiguous United States and compare to results from the GFS and IFS models using 1) grid-based, 2) neighborhood, and 3) object-oriented metrics verified against the fifth major global reanalysis produced by ECMWF (ERA5) and the NCEP/Environmental Modeling Center (EMC) stage IV precipitation analysis datasets. We affirmed that GRAPGFSand GRAPIFSperform better than the GFS and IFS in terms of root-mean-square error and stable equitable errors in probability space, but the GFS and IFS precipitation distributions more closely align with the ERA5 and stage IV distributions. Equitable threat score also generally favored GraphCast, particularly for lower accumulation thresholds. Fractions skill score for increasing neighborhood sizes shows greater gains for the GFS and IFS than GraphCast, suggesting the NWP models may have a better handle on intensity but struggle with the location. Object-based verification for GraphCast found positive area biases at low accumulation thresholds and large negative biases at high accumulation thresholds. GRAPGFSsaw similar performance gains to GRAPIFSwhen compared to their NWP counterparts, but initializing with the less familiar GFS conditions appeared to lead to an increase in light precipitation. Significance StatementPure artificial intelligence (AI)-based weather prediction (AIWP) has exploded in popularity with promises of better performance and faster run times than numerical weather prediction (NWP) models. However, less attention has been paid to their capability to predict impactful, sensible weather like precipitation, precipitation type, or specific meteorological features. We seek to address this gap by comparing precipitation forecast performance by an AI model called GraphCast to the Global Forecast System (GFS) and the Integrated Forecasting System (IFS) NWP models. While GraphCast does perform better on many verification metrics, it has some limitations for intense precipitation forecasts. In particular, it less frequently predicts intense precipitation events than the GFS or IFS. Overall, this article emphasizes the promise of AIWP while at the same time stresses the need for robust verification by domain experts.more » « less
An official website of the United States government
