Title: Data and incentives
“Big data” gives markets access to previously unmeasured characteristics of individual agents. Policymakers must decide whether and how to regulate the use of this data. We study how new data affects incentives for agents to exert effort in settings such as the labor market, where an agent's quality is initially unknown but is forecast from an observable outcome. We show that measurement of a new covariate has a systematic effect on the average effort exerted by agents, with the direction of the effect determined by whether the covariate is informative about long‐run quality versus a shock to short‐run outcomes. For a class of covariates satisfying a statistical property that we callstrong homoskedasticity, this effect is uniform across agents. More generally, new measurements can impact agents unequally, and we show that these distributional effects have a first‐order impact on social welfare.  more » « less
Award ID(s):
1851629
PAR ID:
10561642
Author(s) / Creator(s):
;
Publisher / Repository:
Econometric Society
Date Published:
Journal Name:
Theoretical Economics
Volume:
19
Issue:
1
ISSN:
1933-6837
Page Range / eLocation ID:
407 to 448
Subject(s) / Keyword(s):
Big data, forecasting, effort incentives, career concerns
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Data sharing is central to various applications such as fraud detection, ad matching, and improving patient care. However, each solution to data sharing is bespoke and cost-intensive, hampering value generation. We identify the lack of abstractions to control data release as the culprit of the problem. For example, it is common to have constraints on whether to share data that depend on the result of sharing, and evaluating these constraints requires sharing in the first place, leading to a standstill. To help people build solutions to a wide variety of data sharing applications, we proposeprogrammable dataflows, which consist of two components. The first component is an abstraction, thecontract, which agents use to communicate the intent of a data sharing action and evaluate its consequences before the dataflow takes place. This helps agents control the release of their data. The second component is acontract programming model(CPM), which allows agents to program data sharing applications catered to each problem’s needs with the contract abstraction. We describe how to deploy those applications on a data escrow to ensure data remains protected from unintended data releases. Our evaluation shows 1) the contract abstraction permits representing a wide range of sharing problems, 2) CPM permits writing programs for complex data sharing problems and 3) quantitatively, our improvements to CPM make sharing programs run efficiently. 
    more » « less
  2. Dynamic max-min fair allocation (DMMF) is a simple and popular mechanism for the repeated allocation of a shared resource among competing agents: in each round, each agent can choose to request or not for the resource, which is then allocated to the requesting agent with the least number of allocations received till then. Recent work has shown that under DMMF, a simple threshold-based request policy enjoys surprisingly strong robustness properties, wherein each agent can realize a significant fraction of her optimal utility irrespective of how other agents' behave. While this goes some way in mitigating the possibility of a 'tragedy of the commons' outcome, the robust policies require that an agent defend against arbitrary (possibly adversarial) behavior by other agents. This however may be far from optimal compared to real world settings, where other agents are selfish optimizers rather than adversaries. Therefore, robust guarantees give no insight on how agents behave in an equilibrium, and whether outcomes are improved under one. Our work aims to bridge this gap by studying the existence and properties of equilibria under DMMF. To this end, we first show that despite the strong robustness guarantees of the threshold based strategies,no Nash equilibrium existswhen agents participate in DMMF, each using some fixed threshold-based policy. On the positive side, however, we show that for the symmetric case, a simple data-driven request policy guarantees that no agent benefits from deviating to a different fixed threshold policy. In our proposed policy agents aim to match the historical allocation rate with a vanishing drift towards the rate optimizing overall welfare for all users. Furthermore, the resulting equilibrium outcome can be significantly better compared to what follows from the robustness guarantees. Our results are built on a complete characterization of the steady-state distribution under DMMF, as well as new techniques for analyzing strategic agent outcomes under dynamic allocation mechanisms; we hope these may prove of independent interest in related problems. 
    more » « less
  3. null (Ed.)
    We exhibit a natural environment, social learning among heterogeneous agents, where even slight misperceptions can have a large negative impact on long‐run learning outcomes. We consider a population of agents who obtain information about the state of the world both from initial private signals and by observing a random sample of other agents' actions over time, where agents' actions depend not only on their beliefs about the state but also on their idiosyncratic types (e.g., tastes or risk attitudes). When agents are correct about the type distribution in the population, they learn the true state in the long run. By contrast, we show, first, that even arbitrarily small amounts of misperception about the type distribution can generate extreme breakdowns of information aggregation, where in the long run all agents incorrectly assign probability 1 to some fixed state of the world, regardless of the true underlying state. Second, any misperception of the type distribution leads long‐run beliefs and behavior to vary only coarsely with the state, and we provide systematic predictions for how the nature of misperception shapes these coarse long‐run outcomes. Third, we show that how fragile information aggregation is against misperception depends on the richness of agents' payoff‐relevant uncertainty; a design implication is that information aggregation can be improved by simplifying agents' learning environment. The key feature behind our findings is that agents' belief‐updating becomes “decoupled” from the true state over time. We point to other environments where this feature is present and leads to similar fragility results. 
    more » « less
  4. Summary We present machine learning estimators for causal and predictive parameters under covariate shift, where covariate distributions differ between training and target populations. One such parameter is the average effect of a policy that alters the covariate distribution, such as a treatment that modifies surrogate covariates used to predict long-term outcomes. Another example is the average treatment effect for a population with a shifted covariate distribution. We propose a debiased machine learning method to estimate a broad class of these parameters in a statistically reliable and automatic manner. Our method eliminates regularization bias arising from the use of machine learning tools in high-dimensional settings, and relies solely on the parameter’s defining formula. It employs data fusion by combining samples of target and training data to eliminate bias. We give asymptotic theory that allows the sample sizes of the training and target datato grow at different rates. Computational experiments and an empirical study of the impact ofminimum-wage increases on teen employment, using the difference-in-differences framework with unconfoundedness, demonstrate the effectiveness of our method. 
    more » « less
  5. Abstract An experiment was implemented in the 2023 wave of a US household panel study to assess the effects of a shortened field period on data collection outcomes. Following the recent adoption of sequential mixed-mode designs by panel studies worldwide, it has been observed that interview completion for respondents offered the initial mode of web is faster compared to those initially offered the telephone. This study describes an experiment designed to evaluate whether the new mixed-mode designs can support an accelerated field period and achieve cost savings while still meeting fieldwork goals. We assessed a shorter field period of 20 weeks against the standard 28-week field period and randomized study participants to each condition. The treatment group received accelerated fieldwork protocols over a 20-week data collection period, and a control group received the same protocols over the standard 28-week period. We compare the effect of the shortened duration on fieldwork outcomes, including response rates, sample composition, interviewer effort, time to interview completion, survey costs, and interview quality. We find that the accelerated protocol yields higher response rates, lower interviewer effort, and cost savings with no differences in sample composition or decrements to interview quality. We describe the strengths and limitations of the study and provide suggestions for future research on fieldwork duration. 
    more » « less