Cross-study Learning for Generalist and Specialist Predictions

Ren, Boyu; Patil, Prasad; Dominici, Francesca; Parmigiani, Giovanni; Trippa, Lorenzo

Citation Details

Jointly using data from multiple similar sources for the training of prediction models is increasingly becoming an important task in many fields of science. In this paper, we propose a framework for {\it generalist and specialist} predictions that leverages multiple datasets, with potential heterogenity in the relationships between predictors and outcomes. Our framework uses ensembling with stacking, and includes three major components: 1) training of the ensemble members using one or more datasets, 2) a no-data-reuse technique for stacking weights estimation and 3) task-specific utility functions. We prove that under certain regularity conditions, our framework produces a stacked prediction function with oracle property. We also provide analytically the conditions under which the proposed no-data-reuse technique will increase the prediction accuracy of the stacked prediction function compared to using the full data. We perform a simulation study to numerically verify and illustrate these results and apply our framework to predicting mortality based on a collection of variables including long-term exposure to common air pollutants. more »

Award ID(s):: 1810829

PAR ID:: 10176140

Author(s) / Creator(s):: Ren, Boyu; Patil, Prasad; Dominici, Francesca; Parmigiani, Giovanni; Trippa, Lorenzo

Date Published:: 2020-07-24

Journal Name:: ArXivorg

ISSN:: 2331-8422

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
The DOI is not currently available.

More Like this