The emerging landscape of health research based on biobanks linked to electronic health records: Existing resources, statistical challenges, and potential opportunities

Beesley, Lauren J.  (ORCID:0000000237885944); Salvatore, Maxwell  (ORCID:0000000236591514); Fritsche, Lars G.; Pandit, Anita; Rao, Arvind; Brummett, Chad; Willer, Cristen J.; Lisabeth, Lynda D.; Mukherjee, Bhramar

doi:10.1002/sim.8445

Citation Details

The emerging landscape of health research based on biobanks linked to electronic health records: Existing resources, statistical challenges, and potential opportunities

Biobanks linked to electronic health records provide rich resources for health‐related research. With improvements in administrative and informatics infrastructure, the availability and utility of data from biobanks have dramatically increased. In this paper, we first aim to characterize the current landscape of available biobanks and to describe specific biobanks, including their place of origin, size, and data types. The development and accessibility of large‐scale biorepositories provide the opportunity to accelerate agnostic searches, expedite discoveries, and conduct hypothesis‐generating studies of disease‐treatment, disease‐exposure, and disease‐gene associations. Rather than designing and implementing a single study focused on a few targeted hypotheses, researchers can potentially use biobanks' existing resources to answer an expanded selection of exploratory questions as quickly as they can analyze them. However, there are many obvious and subtle challenges with the design and analysis of biobank‐based studies. Our second aim is to discuss statistical issues related to biobank research such as study design, sampling strategy, phenotype identification, and missing data. We focus our discussion on biobanks that are linked to electronic health records. Some of the analytic issues are illustrated using data from the Michigan Genomics Initiative and UK Biobank, two biobanks with two different recruitment mechanisms. We summarize the current body of literature for addressing these challenges and discuss some standing open problems. This work complements and extends recent reviews about biobank‐based research and serves as a resource catalog with analytical and practical guidance for statisticians, epidemiologists, and other medical researchers pursuing research using biobanks. more »

Award ID(s):: 1712933

PAR ID:: 10453676

Author(s) / Creator(s):: Beesley, Lauren J. ; Salvatore, Maxwell ; Fritsche, Lars G. ; Pandit, Anita ; Rao, Arvind ; Brummett, Chad ; Willer, Cristen J. ; Lisabeth, Lynda D. ; Mukherjee, Bhramar

Publisher / Repository:: Wiley Blackwell (John Wiley & Sons)

Date Published:: 2019-12-20

Journal Name:: Statistics in Medicine

Volume:: 39

Issue:: 6

ISSN:: 0277-6715

Page Range / eLocation ID:: p. 773-800

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Journal Article:
https://doi.org/10.1002/sim.8445

More Like this