Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 11:00 PM ET on Thursday, August 13 until 12:00 AM ET on Friday, August 14 due to maintenance. We apologize for the inconvenience.


This content will become publicly available on December 1, 2027

Title: Inferring fine-grained migration patterns across the United States
Abstract Fine-grained migration data illuminate demographic, environmental, and health phenomena. However, United States migration data have serious drawbacks: public data lack spatial granularity, and higher-resolution proprietary data suffer from multiple biases. To address this, we develop a method that fuses high-resolution proprietary data with coarse Census data to create MIGRATE: annual migration matrices capturing flows between 47.4 billion US Census Block Group pairs—approximately four thousand times the spatial resolution of current public data. Our estimates are highly correlated with external ground-truth datasets and improve accuracy relative to raw proprietary data. We use MIGRATE to analyze national and local migration patterns. Nationally, we document demographic and temporal variation in homophily, upward mobility, and moving distance—for example, rising moves into top-income-quartile block groups and racial disparities in upward mobility. Locally, MIGRATE reveals patterns such as wildfire-driven out-migration that are invisible in coarser previous data. We release MIGRATE as a resource for migration researchers.  more » « less
Award ID(s):
2339427 2516270
PAR ID:
10686390
Author(s) / Creator(s):
; ; ; ;
Publisher / Repository:
Nature Communications
Date Published:
Journal Name:
Nature Communications
Volume:
17
Issue:
1
ISSN:
2041-1723
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Abstract Lifestyle recovery captures the collective effects of population activities as well as the restoration of infrastructure and business services. This study uses a novel approach to leverage privacy-enhanced location intelligence data, which is anonymized and aggregated, to characterize distinctive lifestyle patterns and to unveil recovery trajectories after 2017 Hurricane Harvey in Harris County, Texas (USA). The analysis integrates multiple data sources to record the number of visits from home census block groups (CBGs) to different points of interest (POIs) in the county during the baseline and disaster periods. For the methodology, the research utilizes unsupervised machine learning and ANOVA statistical testing to characterize the recovery of lifestyles using privacy-enhanced location intelligence data. First, primary clustering using k-means characterized four distinct essential and non-essential lifestyle patterns. For each primary lifestyle cluster, the secondary clustering characterized the impact of the hurricane into four possible recovery trajectories based on the severity of maximum disruption and duration of recovery. The findings further reveal multiple recovery trajectories and durations within each lifestyle cluster, which imply differential recovery rates among similar lifestyles and different demographic groups. The impact of flooding on lifestyle recovery extends beyond the flooded regions, as 59% of CBGs with extreme recovery durations did not have at least 1% of direct flooding impacts. The findings offer a twofold theoretical significance: (1) lifestyle recovery is a critical milestone that needs to be examined, quantified, and monitored in the aftermath of disasters; (2) spatial structures of cities formed by human mobility and distribution of facilities extend the spatial reach of flood impacts on population lifestyles. These provide novel data-driven insights for public officials and emergency managers to examine, measure, and monitor a critical milestone in community recovery trajectory based on the return of lifestyles to normalcy. 
    more » « less
  2. null (Ed.)
    Abstract Understanding dynamic human mobility changes and spatial interaction patterns at different geographic scales is crucial for assessing the impacts of non-pharmaceutical interventions (such as stay-at-home orders) during the COVID-19 pandemic. In this data descriptor, we introduce a regularly-updated multiscale dynamic human mobility flow dataset across the United States, with data starting from March 1st, 2020. By analysing millions of anonymous mobile phone users’ visits to various places provided by SafeGraph, the daily and weekly dynamic origin-to-destination (O-D) population flows are computed, aggregated, and inferred at three geographic scales: census tract, county, and state. There is high correlation between our mobility flow dataset and openly available data sources, which shows the reliability of the produced data. Such a high spatiotemporal resolution human mobility flow dataset at different geographic scales over time may help monitor epidemic spreading dynamics, inform public health policy, and deepen our understanding of human behaviour changes under the unprecedented public health crisis. This up-to-date O-D flow open data can support many other social sensing and transportation applications. 
    more » « less
  3. Abstract Non-pharmacologic interventions (NPIs) promote protective actions to lessen exposure risk to COVID-19 by reducing mobility patterns. However, there is a limited understanding of the underlying mechanisms associated with reducing mobility patterns especially for socially vulnerable populations. The research examines two datasets at a granular scale for five urban locations. Through exploratory analysis of networks, statistics, and spatial clustering, the research extensively investigates the exposure risk reduction after the implementation of NPIs to socially vulnerable populations, specifically lower income and non-white populations. The mobility dataset tracks population movement across ZIP codes for an origin–destination (O–D) network analysis. The population activity dataset uses the visits from census block groups (cbg) to points-of-interest (POIs) for network analysis of population-facilities interactions. The mobility dataset originates from a collaboration with StreetLight Data, a company focusing on transportation analytics, whereas the population activity dataset originates from a collaboration with SafeGraph, a company focusing on POI data. Both datasets indicated that low-income and non-white populations faced higher exposure risk. These findings can assist emergency planners and public health officials in comprehending how different populations are able to implement protective actions and it can inform more equitable and data-driven NPI policies for future epidemics. 
    more » « less
  4. Abstract Preparedness for adverse events is critical to building urban resilience to climate-related risks. While most extant studies investigate preparedness patterns based on survey data, this study explores the potential of big digital footprint data (i.e. population visits to points of interest (POI)) to investigate preparedness patterns in the real case of Hurricane Ida (2021). We further investigate income and racial inequality in preparedness by combining the digital footprint data with demographic and socioeconomic data. A clear pattern of preparedness was seen in Louisiana with aggregated visits to grocery stores, gasoline stations, and construction supply dealers increasing by nearly 9%, 12%, and 10% respectively, representing three types of preparedness: survival, mobility planning, and hazard mitigation. Preparedness for Hurricane Ida was not seen in New York and New Jersey states. Inequality analyses for Louisiana across census block groups (CBGs) demonstrate that CBGs with higher income have more (nearly 8% greater) preparedness in visiting gasoline stations, while CBGs with a larger percentage of the white population have more preparedness in visiting grocery stores (nearly 12% more) in the lowest income groups. The results indicate that income and racial inequality differ across different preparedness in terms of visiting different POIs. 
    more » « less
  5. ABSTRACT Population ecology has amassed a significant volume of demographic data across the Tree of Life. Together, these data enable comparative analyses at unprecedented taxonomic and biogeographic scales to examine patterns of demographic performance and their mechanisms. However, macroecological analysis of heterogeneous data and models from diverse study systems comes with risks, and care must be taken to ensure that the patterns from comparative approaches are biologically meaningful, rather than driven by model‐specific artifacts. Recently, a balancing approach has been proposed as a solution to “distorted” population structure, particularly for evaluating transient (short‐term) population dynamics. We argue that some distortion is the result of true biological processes, and that balancing over‐corrects for distortion due to census timing (pre‐ vs. post‐breeding). We lay out the relationship between demographic census design and the issues purported to be solved by balancing. Using a large dataset of carefully‐selected matrix population models from plants and animals, we demonstrate that balancing changes biological interpretation of the relationship between reproductive traits and demographic resilience. We also highlight how application of balancing outside of its narrow original application to transient metrics can be problematic. We argue that meaningful comparisons require tailored approaches that respect the structure and context of demographic data. A more nuanced strategy–based on the biological realities of life cycles, census design, and reproductive strategies–will improve the robustness and interpretation of comparative demographic analyses. 
    more » « less