Background Many public health departments use record linkage between surveillance data and external data sources to inform public health interventions. However, little guidance is available to inform these activities, and many health departments rely on deterministic algorithms that may miss many true matches. In the context of public health action, these missed matches lead to missed opportunities to deliver interventions and may exacerbate existing health inequities. Objective This study aimed to compare the performance of record linkage algorithms commonly used in public health practice. Methods We compared five deterministic (exact, Stenger, Ocampo 1, Ocampo 2, and Bosh) and two probabilistic record linkage algorithms (fastLink and beta record linkage [BRL]) using simulations and a real-world scenario. We simulated pairs of datasets with varying numbers of errors per record and the number of matching records between the two datasets (ie, overlap). We matched the datasets using each algorithm and calculated their recall (ie, sensitivity, the proportion of true matches identified by the algorithm) and precision (ie, positive predictive value, the proportion of matches identified by the algorithm that were true matches). We estimated the average computation time by performing a match with each algorithm 20 times while varying the size of the datasets being matched. In a real-world scenario, HIV and sexually transmitted disease surveillance data from King County, Washington, were matched to identify people living with HIV who had a syphilis diagnosis in 2017. We calculated the recall and precision of each algorithm compared with a composite standard based on the agreement in matching decisions across all the algorithms and manual review. Results In simulations, BRL and fastLink maintained a high recall at nearly all data quality levels, while being comparable with deterministic algorithms in terms of precision. Deterministic algorithms typically failed to identify matches in scenarios with low data quality. All the deterministic algorithms had a shorter average computation time than the probabilistic algorithms. BRL had the slowest overall computation time (14 min when both datasets contained 2000 records). In the real-world scenario, BRL had the lowest trade-off between recall (309/309, 100.0%) and precision (309/312, 99.0%). Conclusions Probabilistic record linkage algorithms maximize the number of true matches identified, reducing gaps in the coverage of interventions and maximizing the reach of public health action.
more »
« less
Using Public Records to Scaffold Joint Sense Making
We share three suggestions for how teachers can more productively use board work to scaffold joint sense making: (1) make the public record precise; (2) purposefully organize the public record; and (3) take advantage of the public record by referencing it in meaningful ways.
more »
« less
- Award ID(s):
- 1720410
- PAR ID:
- 10433820
- Date Published:
- Journal Name:
- Mathematics Teacher
- ISSN:
- 0972-2440
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
null (Ed.)Abstract During 2013, multiple tornadoes occurred across Australia, leading to 147 injuries and considerable damage. This prompted speculation as to the frequency of these events in Australia, and whether 2013 constituted a record year. Leveraging media reports, public accounts, and the Bureau of Meteorology observational record, 69 tornadoes were identified for the year in comparison to the official count of 37 events. This identified set and the existing historical record were used to establish that, in terms of spatial distribution, 2013 was not abnormal relative to the existing climatology, but numerically exceeded any year in the bureau’s record. Evaluation of the environments in which these tornadoes formed illustrated that these conditions included tornado environments found elsewhere globally, but generally had a stronger dependence on shear magnitude than direction, and lower lifting condensation levels. Relative to local environment climatology, 2013 was also not anomalous. These results illustrate a range of tornadoes associated with cool season, tropical cyclone, east coast low, supercell tornado, and low shear/storm merger environments. Using this baseline, the spatial climatology from 1980 to 2019 as derived from the nonconditional frequency of favorable significant tornado parameter environments for the year is used to highlight that observations are likely an underestimation. Applying the results, discussion is made of the need to expand observing practices, climatology, forecasting guidelines for operational prediction, and improve the warning system. This highlights a need to ensure that the general public is appropriately informed of the tornado hazard in Australia, and provide them with the understanding to respond accordingly.more » « less
-
Much of the green stormwater infrastructure (GSI) in Baltimore, Maryland, USA, has been installed voluntarily by nonprofits and community groups, yet no comprehensive record of these installations previously existed. We worked with nonprofit stakeholders and Baltimore’s Department of Public Works to compile such a record, using both information provided by these agencies and publicly available data sources such as annual reports and newspaper articles. This dataset includes all voluntary green stormwater infrastructure projects that we were able to identify by the end of 2019, with the first known installation completed in 2001. The dataset includes two data tables, one with project-level information, and one with the locations of individual GSI facilities included in each project.more » « less
-
Differentially private (DP) machine learning often relies on the availability of public data for tasks like privacy-utility trade-off estimation, hyperparameter tuning, and pretraining. While public data assumptions may be reasonable in text and image data, they are less likely to hold for tabular data due to tabular data heterogeneity across domains. We propose leveraging powerful priors to address this limitation; specifically, we synthesize realistic tabular data directly from schema-level specifications — such as variable names, types, and permissible ranges — without ever accessing sensitive records. To that end, this work introduces the notion of "surrogate" public data — datasets generated independently of sensitive data, which consume no privacy loss budget and are constructed solely from publicly available schema or metadata. Surrogate public data are intended to encode plausible statistical assumptions (informed by publicly available information) into a dataset with many downstream uses in private mechanisms. We automate the process of generating surrogate public data with large language models (LLMs); in particular, we propose two methods: direct record generation as CSV files, and automated structural causal model (SCM) construction for sampling records. Through extensive experiments, we demonstrate that surrogate public tabular data can effectively replace traditional public data when pretraining differentially private tabular classifiers. To a lesser extent, surrogate public data are also useful for hyperparameter tuning of DP synthetic data generators, and for estimating the privacy-utility tradeoff.more » « less
-
City council meetings are vital sites for civic participation where the public can speak directly to their local government. By addressing city officials and calling on them to take action, public commenters can potentially influence policy decisions spanning a broad range of concerns, from housing, to sustainability, to social justice. Yet studies of these meetings have often been limited by the availability of large-scale, geographically-diverse data. Relying on local governments’ increasing use of YouTube and other technologies to archive their public meetings, we propose a framework that characterizes comments along two dimensions: local concerns (e.g., housing, election administration), and societal concerns (e.g., functional democracy, anti-racism). Based on a large record of public comments we collect from 15 cities in Michigan, we produce data-driven taxonomies of the local concerns and societal concerns that these comments cover, and employ machine learning methods to scalably apply our taxonomies across the entire dataset. We then demonstrate how our framework allows us to examine the salient local concerns and societal concerns that arise in our data, as well as how these aspects interact.more » « less
An official website of the United States government

