Abstract We propose the multiple changepoint isolation (MCI) method for detecting multiple changes in the mean and covariance of a functional process. We first introduce a pair of projections to represent the variability “between” and “within” the functional observations. We then present an augmented fused lasso procedure to split the projections into multiple regions robustly. These regions act to isolate each changepoint away from the others so that the powerful univariate CUSUM statistic can be applied region‐wise to identify the changepoints. Simulations show that our method accurately detects the number and locations of changepoints under many different scenarios. These include light and heavy tailed data, data with symmetric and skewed distributions, sparsely and densely sampled changepoints, and mean and covariance changes. We show that our method outperforms a recent multiple functional changepoint detector and several univariate changepoint detectors applied to our proposed projections. We also show that MCI is more robust than existing approaches and scales linearly with sample size. Finally, we demonstrate our method on a large time series of water vapor mixing ratio profiles from atmospheric emitted radiance interferometer measurements.
more »
« less
Good Practices and Common Pitfalls in Climate Time Series Changepoint Techniques: A Review
Abstract Climate changepoint (homogenization) methods abound today, with a myriad of techniques existing in both the climate and statistics literature. Unfortunately, the appropriate changepoint technique to use remains unclear to many. Further complicating issues, changepoint conclusions are not robust to perturbations in assumptions; for example, allowing for a trend or correlation in the series can drastically change changepoint conclusions. This paper is a review of the topic, with an emphasis on illuminating the models and techniques that allow the scientist to make reliable conclusions. Pitfalls to avoid are demonstrated via actual applications. The discourse begins by narrating the salient statistical features of most climate time series. Thereafter, single- and multiple-changepoint problems are considered. Several pitfalls are discussed en route and good practices are recommended. While most of our applications involve temperatures, a sea ice series is also considered. Significance StatementThis paper reviews the methods used to identify and analyze the changepoints in climate data, with a focus on helping scientists make reliable conclusions. The paper discusses common mistakes and pitfalls to avoid in changepoint analysis and provides recommendations for best practices. The paper also provides examples of how these methods have been applied to temperature and sea ice data. The main goal of the paper is to provide guidance on how to effectively identify the changepoints in climate time series and homogenize the series.
more »
« less
- Award ID(s):
- 2143550
- PAR ID:
- 10473093
- Publisher / Repository:
- American Meteorological Society
- Date Published:
- Journal Name:
- Journal of Climate
- Volume:
- 36
- Issue:
- 23
- ISSN:
- 0894-8755
- Format(s):
- Medium: X Size: p. 8041-8057
- Size(s):
- p. 8041-8057
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Abstract BackgroundClimate change is warming the Arctic faster than the rest of the planet. Shifts in whale migration timing have been linked to climate change in temperate and sub-Arctic regions, and evidence suggests Bering–Chukchi–Beaufort (BCB) bowhead whales (Balaena mysticetus) might be overwintering in the Canadian Beaufort Sea. MethodsWe used an 11-year timeseries (spanning 2009–2021) of BCB bowhead whale presence in the southern Chukchi Sea (inferred from passive acoustic monitoring) to explore relationships between migration timing and sea ice in the Chukchi and Bering Seas. ResultsFall southward migration into the Bering Strait was delayed in years with less mean October Chukchi Sea ice area and earlier in years with greater sea ice area (p = 0.04, r2 = 0.40). Greater mean October–December Bering Sea ice area resulted in longer absences between whales migrating south in the fall and north in the spring (p < 0.01, r2 = 0.85). A stepwise shift after 2012–2013 shows some whales are remaining in southern Chukchi Sea rather than moving through the Bering Strait and into the northwestern Bering Sea for the winter. Spring northward migration into the southern Chukchi Sea was earlier in years with less mean January–March Chukchi Sea ice area and delayed in years with greater sea ice area (p < 0.01, r2 = 0.82). ConclusionsAs sea ice continues to decline, northward spring-time migration could shift earlier or more bowhead whales may overwinter at summer feeding grounds. Changes to bowhead whale migration could increase the overlap with ships and impact Indigenous communities that rely on bowhead whales for nutritional and cultural subsistence.more » « less
-
null (Ed.)Online algorithms for detecting changepoints, or abrupt shifts in the behavior of a time series, are often deployed with limited resources, e.g., to edge computing settings such as mobile phones or industrial sensors. In these scenarios it may be beneficial to trade the cost of collecting an environmental measurement against the quality or "fidelity" of this measurement and how the measurement affects changepoint estimation. For instance, one might decide between inertial measurements or GPS to determine changepoints for motion. A Bayesian approach to changepoint detection is particularly appealing because we can represent our posterior uncertainty about changepoints and make active, cost-sensitive decisions about data fidelity to reduce this posterior uncertainty. Moreover, the total cost could be dramatically lowered through active fidelity switching, while remaining robust to changes in data distribution. We propose a multi-fidelity approach that makes cost-sensitive decisions about which data fidelity to collect based on maximizing information gain with respect to changepoints. We evaluate this framework on synthetic, video, and audio data and show that this information-based approach results in accurate predictions while reducing total cost.more » « less
-
Summary This paper deals with the detection and identification of changepoints among covariances of high-dimensional longitudinal data, where the number of features is greater than both the sample size and the number of repeated measurements. The proposed methods are applicable under general temporal-spatial dependence. A new test statistic is introduced for changepoint detection, and its asymptotic distribution is established. If a changepoint is detected, an estimate of the location is provided. The rate of convergence of the estimator is shown to depend on the data dimension, sample size, and signal-to-noise ratio. Binary segmentation is used to estimate the locations of possibly multiple changepoints, and the corresponding estimator is shown to be consistent under mild conditions. Simulation studies provide the empirical size and power of the proposed test and the accuracy of the changepoint estimator. An application to a time-course microarray dataset identifies gene sets with significant gene interaction changes over time.more » « less
-
Summary Many non‐homogeneous Poisson process software reliability growth models (SRGM) are characterized by a single continuous curve. However, failures are driven by factors such as the testing strategy and environment, integration testing and resource allocation, which can introduce one or more changepoint into the fault detection process. Some researchers have proposed non‐homogeneous Poisson process SRGM, but only consider a common failure distribution before and after changepoints. This paper proposes a heterogeneous single changepoint framework for SRGM, which can exhibit different failure distributions before and after the changepoint. Combinations of two simple and distinct curves including an exponential and S‐shaped curve are employed to illustrate the concept. Ten data sets are used to compare these heterogeneous models against their homogeneous counterparts. Experimental results indicate that heterogeneous changepoint models achieve better goodness‐of‐fit measures on 60% and 80% of the data sets with respect to the Akaike information criterion and predictive sum of squares measures.more » « less