Traditional methods for handling incomplete data, including Multiple Imputation and Maximum Likelihood, require that the data be Missing At Random (MAR). In most cases, however, missingness in a variable depends on the underlying value of that variable. In this work, we devise model-based methods to consistently estimate mean, variance and covariance given data that are Missing Not At Random (MNAR). While previous work on MNAR data require variables to be discrete, we extend the analysis to continuous variables drawn from Gaussian distributions. We demonstrate the merits of our techniques by comparing it empirically to state of the art software packages.
more » « less- Award ID(s):
- 1704932
- NSF-PAR ID:
- 10098282
- Date Published:
- Journal Name:
- Proceedings of the International Joint Conferences on Artificial Intelligence Organization
- Page Range / eLocation ID:
- 5082 to 5088
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
Missing data are ubiquitous in many domain such as healthcare. When these data entries are not missing completely at random, the (conditional) independence relations in the observed data may be different from those in the complete data generated by the underlying causal process.Consequently, simply applying existing causal discovery methods to the observed data may lead to wrong conclusions. In this paper, we aim at developing a causal discovery method to recover the underlying causal structure from observed data that are missing under different mechanisms, including missing completely at random (MCAR),missing at random (MAR), and missing not at random (MNAR). With missingness mechanisms represented by missingness graphs (m-graphs),we analyze conditions under which additional correction is needed to derive conditional independence/dependence relations in the complete data. Based on our analysis, we propose Miss-ing Value PC (MVPC), which extends the PC algorithm to incorporate additional corrections.Our proposed MVPC is shown in theory to give asymptotically correct results even on data that are MAR or MNAR. Experimental results on both synthetic data and real healthcare applications illustrate that the proposed algorithm is able to find correct causal relations even in the general case of MNAR.more » « less
-
Implicit feedback is widely used in collaborative filtering methods for recommendation. It is well known that implicit feedback contains a large number of values that are missing not at random (MNAR); and the missing data is a mixture of negative and unknown feedback, making it difficult to learn users’ negative preferences. Recent studies modeled exposure, a latent missingness variable which indicates whether an item is exposed to a user, to give each missing entry a confidence of being negative feedback. However, these studies use static models and ignore the information in temporal dependencies among items, which seems to be an essential underlying factor to subsequent missingness. To model and exploit the dynamics of missingness, we propose a latent variable named “user intent” to govern the temporal changes of item missingness, and a hidden Markov model to represent such a process. The resulting framework captures the dynamic item missingness and incorporate it into matrix factorization (MF) for recommendation. We also explore two types of constraints to achieve a more compact and interpretable representation of user intents. Experiments on real-world datasets demonstrate the superiority of our method against state-of-the-art recommender systems.more » « less
-
Abstract Missing data is a prevalent problem in bioarchaeological research and imputation could provide a promising solution. This work simulated missingness on a control dataset (481 samples × 41 variables) in order to explore imputation methods for mixed data (qualitative and quantitative data). The tested methods included Random Forest (RF), PCA/MCA, factorial analysis for mixed data (FAMD), hotdeck, predictive mean matching (PMM), random samples from observed values (RSOV), and a multi-method (MM) approach for the three missingness mechanisms (MCAR, MAR, and MNAR) at levels of 5%, 10%, 20%, 30%, and 40% missingness. This study also compared single imputation with an adapted multiple imputation method derived from the R package “mice”. The results showed that the adapted multiple imputation technique always outperformed single imputation for the same method. The best performing methods were most often RF and MM, and other commonly successful methods were PCA/MCA and PMM multiple imputation. Across all criteria, the amount of missingness was the most important parameter for imputation accuracy. While this study found that some imputation methods performed better than others for the control dataset, each imputation method has advantages and disadvantages. Imputation remains a promising solution for datasets containing missingness; however when making a decision it is essential to consider dataset structure and research goals.
-
The aim of the article is twofold: (i) present a pivotal setting where using an extra experiment for restoring information lost due to missing not at random (MNAR) is practically feasible; (ii) attract attention to a wide spectrum of new research topics created by the proposed methodology of exploring the missing mechanism. It is well known that if the likelihood of missing an observation depends on its value, then the missing is MNAR, no consistent estimation is possible, and the only way to recover destroyed information is to study the likelihood of missing via an extra experiment. One of the main practical issues with an extra‐sample approach is as follows. Let
n andm be the numbers of observations in a MNAR time series and in an extra sample exploring the likelihood of missing respectively. An oracle, that knows the likelihood of missing, can estimate the spectral density of an ARMA‐type spectral density with the MISE proportional to , while a differentiable likelihood may be estimated only with the MISE proportional tom −2/3. On first glance, these familiar facts yield that the proposed approach is impractical becausem must be in order larger thann to match the oracle. Surprisingly, the article presents the theory and a numerical study indicating thatm may be in order smaller thann and still the statistician can match performance of the oracle. The proposed methodology is used for the analysis of MNAR time series of systolic blood pressure of a person with immunoglobulin D multiple myeloma. A number of possible extensions and future research topics are outlined. -
Abstract Many surveys collect information on discrete characteristics and continuous variables, that is, mixed-type variables. Small-area statistics of interest include means or proportions of the response variables as well as their domain means, which are the mean values at each level of a different categorical variable. However, item nonresponse in survey data increases the complexity of small-area estimation. To address this issue, we propose a multivariate mixed-effects model for mixed-type response variables subject to item nonresponse. We apply this method to two data structures where the data are missing completely at random by design. We use empirical data from two separate studies: a survey of pet owners and a dataset from the National Resources Inventory. In these applications, our proposed method leads to improvements relative to a direct estimator and a predictor based on a univariate model.