Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Free, publicly-accessible full text available March 1, 2027
-
Computational models play an increasingly vital role in scientific research by enabling the numerical simulation of complex processes. Such models are also fundamental in geosciences. For instance, they offer critical insights into the impacts of global change on the Earth system today and in the future. Beyond their value as research tools, models are also software products and should therefore adhere to certain established software engineering standards. However, scientists are rarely trained as software developers, which can lead to potential deficiencies in software quality like unreadable, inefficient, or erroneous code. The complexity of models, coupled with their integration into broader workflows, also often makes it challenging to reproduce results, evaluate processes, and build upon them. In this paper, we review the state and current practices of the development processes of the state-of-the-art land surface models used by the Global Carbon Budget. We combine the experience of modelers from the respective research groups with the expertise of software engineers from tech companies to outline key principles and tools for improving software quality in research. We explore four main areas: (1) model testing and validation, (2) scientific, technical, and user documentation, (3) version control, continuous integration, and code review, and (4) the portability and reproducibility of workflows. Our review reveals that while modeling communities are incorporating many best practices, significant room for improvement remains in areas such as automated testing, automated documentation, and reproducibility. Therefore, we here identify and promote essential software engineering practices, including numerous examples of practices from within the community that can serve as guidelines for other models and could help streamline processes across the entire community. We conclude with an open-source example implementation of these principles, demonstrating portable and reproducible data flows, a continuous integration setup, and web-based visualizations. This example may serve as a practical resource for model developers, users, and all scientists engaged in scientific programming.more » « lessFree, publicly-accessible full text available January 1, 2027
-
Abstract This paper summarizes the open community conventions developed by the Ecological Forecasting Initiative (EFI) for the common formatting and archiving of ecological forecasts and the metadata associated with these forecasts. Such open standards are intended to promote interoperability and facilitate forecast communication, distribution, validation, and synthesis. For output files, we first describe the convention conceptually in terms of global attributes, forecast dimensions, forecasted variables, and ancillary indicator variables. We then illustrate the application of this convention to the two file formats that are currently preferred by the EFI, netCDF (network common data form), and comma‐separated values (CSV), but note that the convention is extensible to future formats. For metadata, EFI's convention identifies a subset of conventional metadata variables that are required (e.g., temporal resolution and output variables) but focuses on developing a framework for storing information about forecast uncertainty propagation, data assimilation, and model complexity, which aims to facilitate cross‐forecast synthesis. The initial application of this convention expands upon the Ecological Metadata Language (EML), a commonly used metadata standard in ecology. To facilitate community adoption, we also provide a Github repository containing a metadata validator tool and several vignettes in R and Python on how to both write and read in the EFI standard. Lastly, we provide guidance on forecast archiving, making an important distinction between short‐term dissemination and long‐term forecast archiving, while also touching on the archiving of code and workflows. Overall, the EFI convention is a living document that can continue to evolve over time through an open community process.more » « less
-
Abstract Navigating uncertainty is a critical challenge in all fields of science, especially when translating knowledge into real-world policies or management decisions. However, the wide variance in concepts and definitions of uncertainty across scientific fields hinders effective communication. As a microcosm of diverse fields within Earth Science, NASA’s Carbon Monitoring System (CMS) provides a useful crucible in which to identify cross-cutting concepts of uncertainty. The CMS convened the Uncertainty Working Group (UWG), a group of specialists across disciplines, to evaluate and synthesize efforts to characterize uncertainty in CMS projects. This paper represents efforts by the UWG to build a heuristic framework designed to evaluate data products and communicate uncertainty to both scientific and non-scientific end users. We consider four pillars of uncertainty: origins, severity, stochasticity versus incomplete knowledge, and spatial and temporal autocorrelation. Using a common vocabulary and a generalized workflow, the framework introduces a graphical heuristic accompanied by a narrative, exemplified through contrasting case studies. Envisioned as a versatile tool, this framework provides clarity in reporting uncertainty, guiding users and tempering expectations. Beyond CMS, it stands as a simple yet powerful means to communicate uncertainty across diverse scientific communities.more » « less
-
Abstract. Monitoring leaf phenology tracks the progression ofclimate change and seasonal variations in a variety of organismal andecosystem processes. Networks of finite-scale remote sensing, such as thePhenoCam network, provide valuable information on phenological state at hightemporal resolution, but they have limited coverage. Satellite-based data withlower temporal resolution have primarily been used to more broadly measurephenology (e.g., 16 d MODIS normalizeddifference vegetation index (NDVI) product). Recent versions of the GeostationaryOperational Environmental Satellites (GOES-16 and GOES-17) can monitor NDVI attemporal scales comparable to that of PhenoCam throughout most of thewestern hemisphere. Here we begin to examine the current capacity of thesenew data to measure the phenology of deciduous broadleaf forests for thefirst 2 full calendar years of data (2018 and 2019) by fittingdouble-logistic Bayesian models and comparing the transition dates of the start, middle, and end of theseason to those obtained from PhenoCam and MODIS 16 dNDVI and enhanced vegetation index (EVI) products. Compared to these MODIS products, GOES was morecorrelated with PhenoCam at the start and middle of spring but had a largerbias (3.35 ± 0.03 d later than PhenoCam) at the end of spring.Satellite-based autumn transition dates were mostly uncorrelated with thoseof PhenoCam. PhenoCam data produced significantly more certain (allp values ≤0.013) estimates of all transition dates than any of thesatellite sources did. GOES transition date uncertainties were significantlysmaller than those of MODIS EVI for all transition dates (all p values ≤0.026), but they were only smaller (based on p value <0.05) than thosefrom MODIS NDVI for the estimates of the beginning and middle of spring. GOES willimprove the monitoring of phenology at large spatial coverages and providesreal-time indicators of phenological change even when the entire springtransition period occurs within the 16 d resolution of these MODISproducts.more » « less
An official website of the United States government
