Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Title: Realized regression with asynchronous and noisy high frequency and high dimensional data
We develop regression for high frequency data. This regression is novel in that it can be for both fixed and increasing dimension. Also, the data may have microstructure noise, and observations (trades, or quotes) can be asynchronous, (i.e., the observations do not need to be synchronized across dimensions). As is customary for high-frequency inference methods, we refer to our method as “realized” regression. In our methodology, spot beta becomes a key quantity in the nonparametric framework of high frequency econometrics. The central contribution of this paper is a feasible estimator of spot beta, which is robust to noise and asynchronicity. With the help of the spot-version of the Smoothed TSRV estimator, spot beta can be consistently estimated. There are two direct applications of the spot beta estimates in the current paper. In the first application, the integrated beta can be consistently estimated by aggregating the spot beta estimates. After a bias-correction procedure, a fixed dimension central limit theorem is established for the bias-corrected estimator, with convergence rate which may be arbitrarily close to Op(n^{1/4}). In the second application we assume time-varying factor structure and conditional sparsity. The spot beta matrix estimator enables the estimation of high dimensional spot covariance and precision matrices. The latter is obtained by thresholding the spot residual covariance estimates, and convergence rates derived. As an empirical application, this paper explores the hourly change in beta around earnings announcements of the S&P 100 constituents.  more » « less
Award ID(s):
2015530 2015544 2413953 2413952
PAR ID:
10553584
Author(s) / Creator(s):
; ;
Publisher / Repository:
Elsevier
Date Published:
Journal Name:
Journal of Econometrics
Volume:
239
Issue:
2
ISSN:
0304-4076
Page Range / eLocation ID:
105446
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract This paper considers binary classification of high-dimensional features under a postulated model with a low-dimensional latent Gaussian mixture structure and nonvanishing noise. A generalized least-squares estimator is used to estimate the direction of the optimal separating hyperplane. The estimated hyperplane is shown to interpolate on the training data. While the direction vector can be consistently estimated, as could be expected from recent results in linear regression, a naive plug-in estimate fails to consistently estimate the intercept. A simple correction, which requires an independent hold-out sample, renders the procedure minimax optimal in many scenarios. The interpolation property of the latter procedure can be retained, but surprisingly depends on the way the labels are encoded. 
    more » « less
  2. For a multidimensional Itô semimartingale, we consider the problem of estimating integrated volatility functionals. Jacod and Rosenbaum (2013,The Annals of Statistics41(3), 1462–1484) studied a plug-in type of estimator based on a Riemann sum approximation of the integrated functional and a spot volatility estimator with a forward uniform kernel. Motivated by recent results that show that spot volatility estimators with general two-sided kernels of unbounded support are more accurate, in this article, an estimator using a general kernel spot volatility estimator as the plug-in is considered. A biased central limit theorem for estimating the integrated functional is established with an optimal convergence rate. Central limit theorems for properly de-biased estimators are also obtained both at the optimal convergence regime for the bandwidth and when applying undersmoothing. Our results show that one can significantly reduce the estimator’s bias by adopting a general kernel instead of the standard uniform kernel. Our proposed bias-corrected estimators are found to maintain remarkable robustness against bandwidth selection in a variety of sampling frequencies and functions. 
    more » « less
  3. Abstract We propose a sparse deep ReLU network (SDRN) estimator of the regression function obtained from regularized empirical risk minimization with a Lipschitz loss function. Our framework can be applied to a variety of regression and classification problems. We establish novel nonasymptotic excess risk bounds for our SDRN estimator when the regression function belongs to a Sobolev space with mixed derivatives. We obtain a new, nearly optimal, risk rate in the sense that the SDRN estimator can achieve nearly the same optimal minimax convergence rate as one-dimensional nonparametric regression with the dimension involved in a logarithm term only when the feature dimension is fixed. The estimator has a slightly slower rate when the dimension grows with the sample size. We show that the depth of the SDRN estimator grows with the sample size in logarithmic order, and the total number of nodes and weights grows in polynomial order of the sample size to have the nearly optimal risk rate. The proposed SDRN can go deeper with fewer parameters to well estimate the regression and overcome the overfitting problem encountered by conventional feedforward neural networks. 
    more » « less
  4. We study the construction of a confidence interval (CI) for a simulation output performance measure that accounts for input uncertainty when the input models are estimated from finite data. In particular, we focus on performance measures that can be expressed as a ratio of two dependent simulation outputs’ means. We adopt the parametric bootstrap method to mimic input data sampling and construct the percentile bootstrap CI after estimating the ratio at each bootstrap sample. The standard estimator, which takes the ratio of two sample means, tends to exhibit large finite-sample bias and variance, leading to overcoverage of the percentile bootstrap CI. To address this, we propose two new ratio estimators that replace the sample means with pooled mean estimators via the k-nearest neighbor (kNN) regression: the kNN estimator and the kLR estimator. The kNN estimator performs well in low dimensions, but its estimation error converges more slowly as the dimension increases. The kLR estimator combines the likelihood ratio (LR) method with the kNN regression, leveraging the strengths of both while mitigating their weaknesses; the LR method removes dependence of the error convergence rate on the dimension, whereas the kNN method controls the variance of the kLR estimator to be asymptotically bounded. From the asymptotic analyses and finite-sample heuristics, we propose an experiment design for the ratio estimators and demonstrate their superior empirical performances over the standard ratio estimator using three examples, including one in the enterprise risk management application. History: Accepted by Bruno Tuffin, Area Editor for Simulation. Funding: This work was supported by the National Science Foundation [Grants CAREER CMMI-2246281 and CMMI-2417616] and the Natural Sciences and Engineering Research Council of Canada [Grant RGPIN-2018-03755]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0914 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2024.0914 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . 
    more » « less
  5. In this article, we exploit the spiked covariance structure of the clutter plus noise covariance matrix for radar signal processing. Using state-of-the-art techniques high dimensional statistics, we propose a nonlinear shrinkage-based rotation invariant spiked covariance ma- trix estimator. We state the convergence of the estimated spiked eigen- values. We use a dataset generated from the high-fidelity, site-specific physics-based radar simulation software RFView to compare the proposed algorithm against the existing rank constrained maximum likelihood (RCML)-expected likelihood (EL) covariance estimation 
    more » « less