Statistical Analysis with Linked Data

Han, Ying  (ORCID:0000000300825654); Lahiri, Partha

doi:10.1111/insr.12295

Citation Details

Statistical Analysis with Linked Data

Summary Computerised Record Linkage methods help us combine multiple data sets from different sources when a single data set with all necessary information is unavailable or when data collection on additional variables is time consuming and extremely costly. Linkage errors are inevitable in the linked data set because of the unavailability of error‐free unique identifiers. A small amount of linkage errors can lead to substantial bias and increased variability in estimating parameters of a statistical model. In this paper, we propose a unified theory for statistical analysis with linked data. Our proposed method, unlike the ones available for secondary data analysis of linked data, exploits record linkage process data as an alternative to taking a costly sample to evaluate error rates from the record linkage procedure. A jackknife method is introduced to estimate bias, covariance matrix and mean squared error of our proposed estimators. Simulation results are presented to evaluate the performance of the proposed estimators that account for linkage errors. more »

Award ID(s):: 1758808

PAR ID:: 10078219

Author(s) / Creator(s):: Han, Ying ; Lahiri, Partha

Publisher / Repository:: Wiley-Blackwell

Date Published:: 2018-10-17

Journal Name:: International Statistical Review

Volume:: 87

Issue:: S1

ISSN:: 0306-7734

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Journal Article:
https://doi.org/10.1111/insr.12295

More Like this