This content will become publicly available on March 9, 2027

Title: Deep Spatiotemporal Point Processes: Advances and New Directions
Spatiotemporal point processes model discrete events distributed in space and time, with applications in criminology, seismology, epidemiology, and social networks. Classical models rely on parametric kernels, limiting their ability to capture heterogeneous, nonstationary dynamics. Recent advances integrate deep neural architectures, either by modeling the conditional intensity directly or by learning flexible, data-driven influence kernels. This article reviews the deep influence kernel approach, which balances statistical interpretability by retaining explicit kernels to capture event propagation, with expressive power from neural architectures. We outline key components, including functional basis decomposition, graph neural networks for encoding spatial or network structures, and both likelihood-based and likelihood-free estimation methods, while addressing scalability for large data. We also highlight theoretical results on kernel identifiability. Applications in crime analysis, earthquake aftershock prediction, and sepsis modeling demonstrate the framework's effectiveness. We conclude with promising directions for developing explainable and scalable deep kernel point processes.  more » « less
Award ID(s):
2220387 2237842
PAR ID:
10690317
Author(s) / Creator(s):
 ;  ;  
Publisher / Repository:
ANNUAL REVIEW OF STATISTICS AND ITS APPLICATION
Date Published:
Journal Name:
Annual Review of Statistics and Its Application
Volume:
13
Issue:
1
ISSN:
2326-8298
Page Range / eLocation ID:
201 to 224
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. null (Ed.)
    Gaussian processes offer an attractive framework for predictive modeling from longitudinal data, i.e., irregularly sampled, sparse observations from a set of individuals over time. However, such methods have two key shortcomings: (i) They rely on ad hoc heuristics or expensive trial and error to choose the effective kernels, and (ii) They fail to handle multilevel correlation structure in the data. We introduce Longitudinal deep kernel Gaussian process regression (L-DKGPR) to overcome these limitations by fully automating the discovery of complex multilevel correlation structure from longitudinal data. Specifically, L-DKGPR eliminates the need for ad hoc heuristics or trial and error using a novel adaptation of deep kernel learning that combines the expressive power of deep neural networks with the flexibility of non-parametric kernel methods. L-DKGPR effectively learns the multilevel correlation with a novel additive kernel that simultaneously accommodates both time-varying and the time-invariant effects. We derive an efficient algorithm to train L-DKGPR using latent space inducing points and variational inference. Results of extensive experiments on several benchmark data sets demonstrate that L-DKGPR significantly outperforms the state-of-the-art longitudinal data analysis (LDA) methods. 
    more » « less
  2. Deep neural networks have become essential for numerous applications due to their strong empirical performance such as vision, RL, and classification. Unfortunately, these networks are quite difficult to interpret, and this limits their applicability in settings where interpretability is important for safety, such as medical imaging. One type of deep neural network is neural tangent kernel that is similar to a kernel machine that provides some aspect of interpretability. To further contribute interpretability with respect to classification and the layers, we develop a new network as a combination of multiple neural tangent kernels, one to model each layer of the deep neural network individually as opposed to past work which attempts to represent the entire network via a single neural tangent kernel. We demonstrate the interpretability of this model on two datasets, showing that the multiple kernels model elucidates the interplay between the layers and predictions. 
    more » « less
  3. Deep neural networks have been increasingly used in real-world applications, making it critical to ensure their ability to adapt to new, unseen data. In this paper, we study the generalization capability of neural networks trained with (stochastic) gradient flow. We establish a new connection between the loss dynamics of gradient flow and general kernel machines by proposing a new kernel, called loss path kernel. This kernel measures the similarity between two data points by evaluating the agreement between loss gradients along the path determined by the gradient flow. Based on this connection, we derive a new generalization upper bound that applies to general neural network architectures. This new bound is tight and strongly correlated with the true generalization error. We apply our results to guide the design of neural architecture search (NAS) and demonstrate favorable performance compared with state-of-the-art NAS algorithms through numerical experiments. 
    more » « less
  4. null (Ed.)
    The standard approach to fitting an autoregressive spike train model is to maximize the likelihood for one-step prediction. This maximum likelihood estimation (MLE) often leads to models that perform poorly when generating samples recursively for more than one time step. Moreover, the generated spike trains can fail to capture important features of the data and even show diverging firing rates. To alleviate this, we propose to directly minimize the divergence between neural recorded and model generated spike trains using spike train kernels. We develop a method that stochastically optimizes the maximum mean discrepancy induced by the kernel. Experiments performed on both real and synthetic neural data validate the proposed approach, showing that it leads to well-behaving models. Using different combinations of spike train kernels, we show that we can control the trade-off between different features which is critical for dealing with model-mismatch. 
    more » « less
  5. Abstract Machine learning methods have recently begun to be used for fitting and comparing cognitive models, yet they have mainly focused on methods for dealing with models that lack tractable likelihoods. Evaluating how these approaches compare to traditional likelihood-based methods is critical to understanding the utility of machine learning for modeling and determining what role it might play in the development of new models and theories. In this paper, we systematically benchmark neural network approaches against likelihood-based approaches to model fitting and comparison, focusing on intertemporal choice modeling as an illustrative application. By applying each approach to intertemporal choice data from participants with substance use problems, we show that there is convergence between neural network and Bayesian methods when it comes to making inferences about latent processes and related substance use outcomes. For model comparison, however, classification networks significantly outperformed likelihood-based metrics. Next, we explored two extensions of this approach, using recurrent layers to allow them to fit data with variable stimuli and numbers of trials, and using dropout layers to allow for posterior sampling. We ultimately suggest that neural networks are better suited to fast parameter estimation and posterior sampling, applications to large data sets, and model comparison, while Bayesian MCMC methods should be preferred for flexible applications to smaller data sets featuring many conditions or experimental designs. 
    more » « less