<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>In Silico Human Mobility Data Science: Leveraging Massive Simulated Mobility Data (Vision Paper)</title></titleStmt>
			<publicationStmt>
				<publisher>ACM</publisher>
				<date>06/30/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10582628</idno>
					<idno type="doi">10.1145/3672557</idno>
					<title level='j'>ACM Transactions on Spatial Algorithms and Systems</title>
<idno>2374-0353</idno>
<biblScope unit="volume">10</biblScope>
<biblScope unit="issue">2</biblScope>					

					<author>Andreas Züfle</author><author>Dieter Pfoser</author><author>Carola Wenk</author><author>Andrew Crooks</author><author>Hamdi Kavak</author><author>Taylor Anderson</author><author>Joon-Seok Kim</author><author>Nathan Holt</author><author>Andrew Diantonio</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Human mobility data science using trajectories or check-ins of individuals has many applications. Recently, we have seen a plethora of research efforts that tackle these applications. However, research progress in this field is limited by a lack of large and representative datasets. The largest and most commonly used dataset of individual human trajectories captures fewer than 200 individuals, while datasets of individual human check-ins capture fewer than 100 check-ins per city per day. Thus, it is not clear if findings from the human mobility data science community would generalize to large populations. Since obtaining massive, representative, and individual-level human mobility data is hard to come by due to privacy considerations, the vision of this work is to embrace the use of data generated by large-scale socially realistic microsimulations. Informed by both real data and leveraging social and behavioral theories, massive spatially explicit microsimulations may allow us to simulate entire megacities at the person level. The simulated worlds, which do not capture any identifiable personal information, allow us to perform “in silico” experiments using the simulated world as a sandbox in which we have perfect information and perfect control without jeopardizing the privacy of any actual individual. In silico experiments have become commonplace in other scientific domains such as chemistry and biology, permitting experiments that foster the understanding of concepts without any harm to individuals. This work describes challenges and opportunities for leveraging massive and realistic simulated alternate worlds for in silico human mobility data science.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">INTRODUCTION</head><p>Data sets capturing individual-level movements have been used for <ref type="bibr">(1)</ref> modeling and analyzing human mobility patterns (e.g., <ref type="bibr">[15,</ref><ref type="bibr">16,</ref><ref type="bibr">20,</ref><ref type="bibr">69,</ref><ref type="bibr">70,</ref><ref type="bibr">87,</ref><ref type="bibr">88,</ref><ref type="bibr">102,</ref><ref type="bibr">110,</ref><ref type="bibr">111]</ref>), <ref type="bibr">(2)</ref> recommending locations to users based on previously visited locations (for example, <ref type="bibr">[11, 12, 24-27, 35, 48-50, 52, 54-56, 58-61, 68, 86, 100, 101, 112, 115, 116, 119-121, 124, 125, 127-131, 133, 134]</ref>, surveyed in <ref type="bibr">[6]</ref>), <ref type="bibr">(3)</ref> predicting the next location to be visited by an individual (e.g., <ref type="bibr">[5,</ref><ref type="bibr">13,</ref><ref type="bibr">33,</ref><ref type="bibr">53,</ref><ref type="bibr">57,</ref><ref type="bibr">84,</ref><ref type="bibr">96,</ref><ref type="bibr">114]</ref>), and (4) suggests new friends to individuals based on similar interests observed by visiting similar locations <ref type="bibr">[14,</ref><ref type="bibr">34,</ref><ref type="bibr">67,</ref><ref type="bibr">80,</ref><ref type="bibr">89,</ref><ref type="bibr">99,</ref><ref type="bibr">104,</ref><ref type="bibr">108,</ref><ref type="bibr">113]</ref>), Publicly available real-world data sets have been the driving force for human mobility data science in recent years.</p><p>These data sets mainly comprise trajectory data and location-based social network (LBSN) data. Trajectory data captures the location of individuals at a relatively high frequency, such as at 1Hz (one location per second per individual). LBSN data captures both 1) so-called check-ins, that is, the location of individuals only when a point of interest (POI) such as a restaurant is visited as well as 2) a social network between individuals. However, publicly available real-world trajectory and LBSN data sets exhibit certain weaknesses:</p><p>&#8226; Data sparsity: In all existing data sets that capture check-ins of individuals, the vast majority of users have less than ten check-ins <ref type="bibr">[60]</ref>. This results in the density of the data used in experimental studies on LBSNs being only usually around 0.1% <ref type="bibr">[60]</ref> (of a complete dataset that would capture every individual). For trajectory data, the most commonly used data set is the GeoLife dataset <ref type="bibr">[136,</ref><ref type="bibr">137]</ref> which captures the locations of 178 individuals in Beijing. Given the population of 21.58 million (2018) in Beijing, this means that the proportion of captured individuals compared to the real population is less than 0.001%. It becomes very challenging to infer patterns of human behavior from such a small sample. In addition, it is not clear whether the individuals included in real-world datasets are a representative sample of the population, as the demographics of anonymous users are unknown. Furthermore, the study in <ref type="bibr">[51]</ref> has shown that the lower bound of predictability of human spatio-temporal behavior (defined in <ref type="bibr">[51]</ref>) is as low as 27%. They conclude that "Researchers working with LBSN data sets are often confronted with doubts regarding the quality or potential of their data sets. " and that "it is reasonable to be skeptical" <ref type="bibr">[51]</ref>.</p><p>But imagine the research possibilities in a world where we had complete trajectory data of 100% of individuals of a large population such as Bejing, China.</p><p>&#8226; Privacy Concerns: Individual human mobility data is considered Personal Identifiable Information (PII) as it allows one to trace an individual's identity. Acquiring, storing, and publishing of individual human mobility data requires the consent of individuals. Even if such consent is given, users may later revoke this consent, for instance, by deleting their LBSN account. This limits, for good reasons, our ability to acquire additional individual-level human mobility data.</p><p>But imagine the research possibilities stemming from collecting individual-level trajectory data that does not jeopardize the privacy of any real human individual.</p><p>&#8226; No Ground-Truth Behavior: There is no way to assess, in existing human mobility data, whether location updates or check-ins are accurate and complete or if updates are missing. It is also difficult to derive the underlying behaviors that lead to an individual's decision to visit a particular POI. For example, did an individual visit a restaurant to eat by themselves? To meet a friend? To have a business meal? Or to work in the restaurant? What preferences lead an individual to choose one grocery store over another? In real-world individual mobility data, the link between users in the data and the corresponding individual in the real world is lost (intentionally Fig. <ref type="figure">1</ref>. The envisioned in silico mobility data science process -(left:) A massive microsimulation is created to simulate realistic human behavior specified by a user through an AI-supported builder tool. (middle:) The microsimulation generates massive datasets, including high-fidelity trajectories of all individuals over years of simulation time. This data, which is 100% accurate and complete (in the simulated world) is then sampled to generate realistic datasets. (right:) These datasets are then used to perform mobility data science tasks in the simulated in silico world as if it was the real world. The results of these tasks can then be compared to the ground truth data (of the simulated in silico world) for validation.</p><p>for privacy). Without knowing the underlying human behavior that led to the observed mobility, it is difficult to infer patterns of human behavior and to predict future mobility.</p><p>But imagine a world where we can go back in time to ask people about the purpose of their mobility to understand why an individual visited a place of interest.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1">In Silico Human Mobility Data Science</head><p>Our vision is to create such a world as illustrated in Figure <ref type="figure">1</ref>. A digital world or a "sandbox" that 1) provides massive, complete, and high-fidelity trajectory data of 100% of the (simulated) population that 2) does not jeopardize the privacy of any real human, and 3) allows insights into the underlying behavior that led individuals to make the trips they made.</p><p>We envision creating such a world "in silico" through massive microsimulation of socially plausible human behavior and mobility. To clarify this notion of "in silico" science, let us first define the underlying notion of in vitro science, which, instead of performing experiments on living organisms (in vivo) performs experiments inside a test tube: Definition 1.1 (In Vitro; Latin: "In Glass"). In vitro refers to a process or studies that are performed with microorganisms, cells, or biological molecules outside their normal biological context [...] performed or taking place in a test tube, culture dish, or elsewhere outside a living organism <ref type="bibr">[19,</ref><ref type="bibr">64]</ref>.</p><p>In vitro scientific experiments are commonplace outside of computer science. In biology, in vitro experiments are commonly performed to understand the behavior of parts of organisms (such as cells/tissue of humans). For example, the effects of a certain chemical or biotoxin can be evaluated in vitro on a small tissue sample rather than in vivo on a living human person to avoid unethical health risks for the human. Similarly, in human mobility data science, working directly with the trajectories of individual humans without their explicit consent is also unethical, as human trajectories are considered Personal Identifiable Information (PII) and using such data may put individuals at serious risk. For example, such trajectory data could enable an attacker (robber, stalker) to identify where a vulnerable person (for example, a child) may be found alone and without protection.</p><p>For human mobility data science, it is difficult to separate data from an individual in a safe way to enable in vitro science, which is the goal of efforts in the field of location privacy (e.g., <ref type="bibr">[29,</ref><ref type="bibr">47,</ref><ref type="bibr">105,</ref><ref type="bibr">107]</ref>. This task is difficult because trajectories are very unique to individual humans. It has been shown to be sufficient to use four spatial points to uniquely identify most individuals, even among a large population of people <ref type="bibr">[17,</ref><ref type="bibr">90]</ref>. Thus, despite efforts to anonymize the data, privacy attacks remain a risk.</p><p>In addition to in vitro science, the concept of in silico science has been used recently as a paradigm of scientific discovery.</p><p>Definition 1.2 (In Silicio; Latin: "In silicon"). In silico refers to a process or studies that are performed entirely via computer simulation <ref type="bibr">[93]</ref>.</p><p>In biology, in silico science simulates parts of the human body (such as individual cells) and allows the study of the interaction of cells with chemicals without any risk to any (real, not simulated) living being <ref type="bibr">[77]</ref>. In the case of human mobility data science, the organism that we want to study is an entire population such as that of a city or another universe of discourse (the world that we want to study). We can take advantage of massive microsimulations of individual humans to create a digital twin of this world. In this in silico world, we have complete access to all of the atoms (the individual agents). Furthermore, we can use the simulation as a sandbox that allows us to make changes to the simulation to see how policy interventions may affect the population over time. Just like in biology and other domains, results obtained from in silico simulation may not fully or accurately predict the effects in the real world. However, in cases where an in vivo (directly in the living body) analysis is too intrusive or even impossible (in the case of individual human mobility data due to privacy), and an in vitro analysis may be difficult on a large scall (in the case of individual human mobility data due to the risk of deindentification even using only a small part of human mobility traces) an in silico analysis allows us to evaluate hypotheses and gain an understanding that may likely be reflected in the real world.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.2">Challenges for In Silico Human Mobility Data Science</head><p>For in silico human mobility data science to be effective, we require a simulated in silico world that mimics the real world. This is challenging, as it requires the microsimulation of realistic individual behavior. As Nobel Prize laureate Murray Gell-Mann famously said, "Think how hard physics would be if particles could think" <ref type="bibr">[76]</ref> and as Richard Feynman added, "Imagine how much harder physics would be if electrons had feelings". In our in silico world, the particles are the simulated individual humans, which may act irrationally in the real world. Creating a simulation that captures realistic human behavior (and thus, resulting individual human mobility) requires close collaborations with experts in the social sciences. Towards such a socially realistic massive microsimulation of individual human mobility, Section 2 surveys existing data sets of individual human mobility and their limitations in terms of size and representativeness. Then, Section 3 surveys existing simulation frameworks to generate individual human mobility data and the limitations of state-of-the-art to generate individual human mobility data that captures all of the following: 1) realistic human behavior, defined as having the locations visited by individuals grounded in a realistic purpose that leads to visiting locations such as going to work and visiting friends (rather than individuals following random walks), 2) realistic human movement, defined as a realistic motion that takes an individual from one location to another, following Then, Section 5 describes only a small sample of applications and research directions that would be enabled by massive individual human mobility datasets if our vision came true, and Section 6 concludes this vision paper.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">LIMITATIONS OF STATE-OF-THE-ART DATASETS</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1">Existing Trajectory Datasets</head><p>Real-world trajectory data sets are a scarce resource due to the privacy implications of making such data public. Also, service providers consider such data sets invaluable when it comes to providing a competitive product and are thus somewhat unwilling to give researchers even sizable data sets. Given these considerations, the mobility data science community is grateful for two datasets that have been made available publicly and that have been used widely within the community and that are summarized in Table <ref type="table">1</ref>. The first dataset is the Geolife GPS trajectory dataset <ref type="bibr">[135,</ref><ref type="bibr">136]</ref>,</p><p>which was collected and shared by Microsoft Research Asia. This dataset captures detailed trajectories of 178 users in Beijing over a period of over four years (from April 2007 to October 2011). This dataset recorded a broad range of users' movements, including not only life routines like going home and going to work but also some entertainment and sports activities, such as shopping, sightseeing, dining, hiking, and cycling. Although this data set is excellent in terms of quality and fidelity, it is unfortunately very small. It is difficult to infer broad mobility patterns from a set of only 178 users, especially in a large city such as China. The small sample size makes tasks such as event detection, friend recommendation, or contact tracing difficult, as more than 99.99% of the population of Beijing is missing from the data.</p><p>It is also not clear if this sample is representative, which allows us to infer patterns learned from the sample to the entire population of Beijing.</p><p>The second dataset is the T-Drive trajectory data sample <ref type="bibr">[122,</ref><ref type="bibr">123]</ref> which was (also) collected and shared by Microsoft</p><p>Research Asia and captures one-week trajectories of 10,357 taxis in Beijing. While the number of individuals captured in this dataset is much larger than in GeoLife, T-Drive captures the trajectories of taxis, not individuals. Thus, consecutive trajectories of the same taxi may not correspond to the same passenger. While useful for applications such as traffic prediction, this dataset is very limited in terms of providing insights into individual human mobility and behavior as it is impossible to understand the sequence of places that individuals have visited.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Existing LBSN data sets</head><p>Table 2 summarizes the main characteristics of publicly available data sets that are used extensively by the LBSN research community. Woodstock '18, June 03-05, 2018, Woodstock, NY Z&#252;fle, Pfoser, Wenk et al. Foursquare: In terms of the number of users and check-ins, the largest publicly available LBSN data set was collected from Foursquare <ref type="bibr">[109]</ref>. However, this data set provides no social network information.</p><p>Yelp: A large dataset is published by Yelp as part of the Yelp data set Challenge <ref type="bibr">[117]</ref>. This data set provides additional information, such as user location ratings, user comments, user information, and location information. Again, this data set does not provide social network information.</p><p>Synthetic Check-In Data: The problem of using sparse and noisy real-world LBSN data has already been identified in previous work (e.g., <ref type="bibr">[3,</ref><ref type="bibr">4,</ref><ref type="bibr">48]</ref>). However, none of these works have proposed a way to obtain plausible check-in data.</p><p>For example, <ref type="bibr">[3,</ref><ref type="bibr">4,</ref><ref type="bibr">48]</ref> generated user location check-ins at random using parametric distributions without considering the semantics of the movement. While <ref type="bibr">[85]</ref> created additional check-ins by replicating Gowalla and Brightkite data, thus creating more data for run-time evaluation purposes but without creating more information.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3">Mobility Flow Data</head><p>There exist many other sources of mobility data that are not at individual level but aggregated to regions for which the pairwise flow between regions is reported. Such data includes the Longitudinal Employer-Household Dynamics (LEHD)</p><p>Origin-Destination Employment Statistics (LODES) dataset for the United States <ref type="bibr">[31]</ref>, which contains home-work communiting flows aggregated to census tract level. Safegraph Inc <ref type="bibr">[82]</ref> made available free access to high-resolution foot traffic data [83] as part of the SafeGraph COVID-19 Data Consortium. This dataset includes check-ins of 35 million anonymized mobile devices in the US, more than 10% of the US population. However, in this dataset, individuals are aggregated to census block groups (CBGs), which are statistical divisions of the U.S. population containing between 600 and 3,000 individuals. As there are no unique identifiers for individuals, it is not possible to combine flows (from one census block group to another) into trajectories, as there is no way to link multiple flows onto the same individual. Thus, SafeGraph does not provide individual check-in data, but only aggregated data that cannot be used to infer individual check-in sequences of individuals. In addition, there are also cell phone trace datasets, which capture the locations of individuals but aggregated to their nearest cell tower and have been used to help understand aggregated human mobility <ref type="bibr">[97]</ref>. However, the problem with cell phone traces is that locations are aggregated to nearest cell towers, which is an aggregation even more coarse than using census block groups covering multiple square kilometers even in cities.</p><p>Therefore, it is impossible to assess individual visited locations based on cell phone trace data.</p><p>Summarizing, all data types mentioned in this chapter do provide information on individual human mobility flow, but due to aggregation to large spatial unites or due to a loss of association of individual users to visited locations, these datasets allow inferring specific locations visited by individuals as possible using trajectory data (see Section 2.1) and location-based social network check-in data (see Section 2.2).</p><p>Table <ref type="table">3</ref>. Existing Simulators for Individual Trajectory Data. By realistic movement, we refer to realistic of how individuals move from one location to another, which may include waiting at traffic lights or delays due to traffic; by realistic behavior, we refer to the underlying causality of why individuals move between locations, including behaviors such as commuting, meeting friends, and walking their dog. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Patterns of</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Traditional Individual Human Trajectory Simulators</head><p>Individual-level microsimulators have been used for decades to represent and analyze individual human mobility. Early work on massive trajectory generation <ref type="bibr">[79]</ref> considers the semantics of movement as well as infrastructure information (buildings and obstacles) in the generation process. Brinkhoff's Moving Objects Generator <ref type="bibr">[9]</ref> uses a road network file as input and simulates the shortest paths between random vertices in the network. Moving objects then move at a constant velocity along their shortest path. Once objects reach their destination, a new destination vertex is selected at random. However, the simplified movement at constant speeds along the shortest paths and the selection of random locations as destinations is not a realistic model reflective of real human mobility. A similar simulator is ViewNet <ref type="bibr">[45,</ref><ref type="bibr">46]</ref> which imposes a simple traffic model in which people cannot overtake each other on roads. This model allows for the simulation of traffic jams. However, the interactions between individuals lowers the scalability of this simulator to at most a few thousand simulated individuals. The BerlinMod <ref type="bibr">[22]</ref> simulator is the first to simulate some realistic behavior, including the commuting of individuals between home and work locations and visiting random places within their neighborhood outside of work. In contrast, the Hermoupolis <ref type="bibr">[78]</ref> simulator simulates individual movement to a random destination but simulates coordinated behavior among individuals, such as groups of individuals moving in flocks thus forming clusters of trajectories. Note that in all of the aforementioned simulators, the goal was not to create a realistic world. Their goal was to create benchmark datasets that could be used to evaluate the efficiency of index structures or to cluster flocks of trajectories.</p><p>An important extension of the simulators was proposed by the MNTG <ref type="bibr">[66]</ref> simulator. While not a simulator itself, MNTG is a wrapper that enriches any of the aforementioned simulators with a web-interface that allows for the selection of a study region, generates a road network and POIs from OpenStreetMap, feeds these datasets to one of the existing simulators and sends the generated datasets for the selected study region via email. The advantage of MNTG is that it allows for the extension of existing simulators to any study region of a user's choice, rather than having to use the network dataset(s) provided by the simulators (such as the Berlin network for BerlinMod).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Realistic Traffic Movement Microsimulation</head><p>In addition, there are also microscopic traffic simulation toolkits such as SUMO <ref type="bibr">[44]</ref> and MATSim <ref type="bibr">[98]</ref>  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">Human Mobility Microsimulation</head><p>To simulate realistic individual human behavior, a simulation was recently published called "Patterns of Life" simulation [39, 139] using Maslow's Hierarchy of Needs <ref type="bibr">[62]</ref> as a driver of simulated individual human behavior. In this simulation, driven by physiological, safety, love, and esteem needs, agents perform activities to satisfy these needs by moving across a spatial road network to find places to eat, work, take shelter, engage in recreational activities, meet friends, and return home to be with family and sleep. By meeting others, agents make new friends and strengthen their friendships. When meeting at places, agents form social links which, in turn, affect the places agents visit. The purpose of this model is to act as a digital sandbox environment for social scientists to assess different methods and tools for analyzing complex social phenomena in an agent-based simulated world. The simulation was used in DARPA's Ground</p><p>Truth Program as a "plausible alternate world" in which we have perfect knowledge of the ground truth of how the world works, how individuals make their decisions, and how social ties are formed. Figure <ref type="figure">2</ref> shows a screenshot of this simulation and provides a link to a video of a simulation run (the spatial network populated with agents on the left and the social network between agents on the right). As agents go to work, restaurants, and recreational sites, they meet with new friends and create a social fabric. A detailed description of this simulation can be found in <ref type="bibr">[139]</ref>. The Patterns of Life model used synthetic environments and synthetic population <ref type="bibr">[40]</ref>. This choice was deliberate to prevent other performers <ref type="foot">1</ref> , who were tasked to explain, to predict, and to prescribe changes to this world Figure <ref type="figure">3</ref> shows a screenshot and provides a link to a video of a simulation of the French Quarter, New Orleans, LA, using real-world road and building data from OpenStreetMap <ref type="bibr">[75]</ref> and U.S. Census information to initialize the agent population. Again, the location of agents in the geographical space (located at places of interest and across the road network) is shown on the left and their location in the social network is shown on the right. Using this model, we simulated the outbreak of an infectious disease. The color of agents (both in geographical and social spaces) corresponds to the disease status of an agent using the SEIR (Susceptible, Exposed, Infectious, Recovered) model. This simulation allows tracking the spread of diseases simultaneously across the geographical map and the social network. Other graphs in Figure <ref type="figure">3</ref> show the time series of infected agents (center), the time series of the average number of friends per agent (bottom, center), and the current distribution of the number of friends per agent (bottom, right). Note that this is not a COVID-19 simulation. This simulation was created and published <ref type="bibr">[42]</ref> in 2019 -before COVID-19 was discovered. The contagion used in this simulation was a generic flu-like disease.</p><p>Although the goal of this simulation was to represent socially plausible human behavior, agents in this simulation move along the shortest paths (at constant speed) to reach their intended destinations without any realistic movement or traffic model. Also, due to the simulated complex human behavior based on agent needs on agents' social networks, this simulation can only scale to a few thousand individual agents. In addition, while realistically modeling an individual's needs to visit places, this simulated world may not reflect the real world, as "all models are wrong, but some are useful" <ref type="bibr">[8]</ref>. For example, it may not reflect that some individuals may have a favorite restaurant or a hairdresser they have been going to for years. But informed by real-world data about the environment (road network and buildings of a city), we hope that such a simulation, while not being a true representation of the real world, can be representative of some aspects of a real population to generate realistc trajectories.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">A VISION OF SCALABLE SIMULATION OF REALISTIC MOVEMENT AND BEHAVIOR</head><p>As we observe from Table <ref type="table">3</ref>, there exist individual-level trajectory simulations that can be scaled to millions of individuals.</p><p>However, these simulations lack realistic behavior (as in what, when, and why to visit certain places) and lack realistic movement (as in how to move between locations). We also see that there are traffic microsimulations that capture realistic mobility, including traffic jams and traffic lights. But in these simulations, the simulated origin and destination pairs of individuals are chosen at random or extrapolated from aggregated flow datasets as surveyed in Section 2.3.</p><p>There are also simulations that have agents exhibit realistic human behavior, but these simulations do not simulate realistic movement and are not yet scalable to millions of agents.</p><p>Our vision is a simulation that satisfies all three requirements, as follows: 1) scalability to millions of agents, 2) realistic human behavior, and 3) realistic human movement. From this simulation, individual-level trajectory mobility datasets can be generated that could be leveraged for human mobility data science research without risk to privacy for real human individuals. We envision that this could be achieved through a framework depicted in Figure <ref type="figure">4</ref> that combines current capabilities for realistic individual behavior simulation (such as in <ref type="bibr">[39,</ref><ref type="bibr">139]</ref>) with current capabilities for realistic individual movement simulation (such as in <ref type="bibr">[44]</ref>). For the purpose of scalability, we propose iterating between a closely coupled realistic behavior simulation and a realistic movement simulation. At the beginning of each</p><p>In Silico Human Mobility Data Science: Leveraging Massive Simulated Mobility Data (Vision Paper) Woodstock '18, June 03-05, 2018, Woodstock, NY simulation day, the individual human behavior simulation creates a daily plan for an agent prioritized by the needs of an agent. For example, such a daily plan of an agent may consist of the following: 1) Wake up at 6:00am, 2) Go to work at 8:00am, 3) Go to a restaurant for lunch at 12:00pm, 4) Go to a bar to meet friends at 6:00pm, 5) Go home at 10pm, with a list of priorities such as: I) work for eight hours, II) eat food, III) meet friends. This agent plan and set of priorities are then passed on to the traffic simulation generate movement data. During the execution of the traffic simulation, delays may occur. For example, the individual may arrive at work one hour late. The prioritization of the behavior simulation will then be used to decide that, in order to work for the full eight hours (which is the agent's first priority obtained from the behavior simulation, for example, due to a pressing financial safety need), the agent has to skip lunch or skip meeting friends. Prioritization will cause this agent to go to the restaurant, but skip going to the bar as a consequence of arriving late at work. The movement simulation will yield trajectories and check-ins of the current day. These are then passed back to behavior simulation to inform decision-making for the next day. The idea here is to inform the behavioral simulation with information about what has actually happened. For example, in the example above, the simulation will be informed that the individual was able to satisfy their financial safety and food needs, but did not meet any of their friends, such that the social links that would be reinforced through the planned meeting are not actually reinforced, and the individual will also not meet any new friends they would have met if they had gone to the bar instead of working longer. Once the behavior simulation is updated giving the actual movement of the previous day, the loop in Figure <ref type="figure">4</ref> restarts by planning the next day for each agent.</p><p>To achieve realistic human behavior, an in-silico world also needs to consider a realistic environment including daily weather changes and seasonal patterns. We know that weather conditions significantly affect traffic <ref type="bibr">[122]</ref>. General human behavior such as going for a walk, going to build a snowman, or having ice cream also depends on weather conditions. To the best of our knowledge, there is no existing large-scale agent-based simulation of human behavior that simulates weather conditions and seasonal weather and temperature patterns. But weather patterns can be learned and simulated for any city (using history weather patterns) and we hypothesize that including weather will be paramount in obtaining realistic human mobility data.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Simpifying Model Building</head><p>A necessary step in white-box model-based approaches is to build models and configure their parameters. However, this often requires expertise in modeling and can be very time-consuming for even expert modelers. To accelerate this process, we envision an advanced user interface beyond traditional programming languages or scripts. Human vision is the most efficient sensory system to comprehend a large amount of information in a short period of time. As such, well-defined diagrams are an effective means to conceive model concepts. Taking this into account, a model builder can provide a graphical user interface (GUI) that visually encodes entities, relationships, and constraints <ref type="bibr">[41]</ref>.</p><p>Although a GUI is intuitive, editing models can still be overwhelming, especially when the models are complex and require many template constructs. The modeling and configuration process requires multiple iterations. To advance the iterative process, we envision a multimodal large language model (MLLM) <ref type="bibr">[21]</ref> to create an initial graphical model and update the model via chat and voice interfaces similar to ChatGPT <ref type="bibr">[71,</ref><ref type="bibr">72]</ref>. Large Language Models (LLM) (e.g., ChatGPT, Llama2 <ref type="bibr">[63]</ref>) pre-trained from a vast corpus of text data (e.g., 500B tokens of code for Llama2) can capture complex context from user queries and generate comprehensive embeddings (e.g., 70B parameters for Llama2).</p><p>LLMs are also at the heart of AI-assisted programming tools such as Github Copilot <ref type="bibr">[30]</ref> and Amazon Codewhisperer <ref type="bibr">[1]</ref>, which, when provided with a programming problem in natural language, are capable of generating solution code. Yes. Save all trajectories in GeoParquet format.</p><p>That's perfect. I want Fairfax, Virginia, USA.</p><p>All set! Do you want me to run a simulation?</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Realtime interaction</head><p>Fig. <ref type="figure">5</ref>. Interactive model builder powered by multimodal large language model with chat interface .</p><p>Amazon Codewhisperer even introduced Amazon Q, an interactive generative AI-powered assistant that provides guidance such as code explanations, transformations, and suggestions through a simple conversational interface.</p><p>In envisioning such a conversational approach, Figure <ref type="figure">5</ref> illustrates how a modeler can interact with an AI assistant to build a model and configure simulations through the chat interface. Compared to earlier attempts to integrate natural language processing (NLP) <ref type="bibr">[91]</ref> or natural language understanding (NLU) <ref type="bibr">[92]</ref>, our vision is to take advantage of a prompt incorporated into the GUI model builder. The AI assistant can be considered a mediator between the GUI model builder and the user. State-of-the-art MLLMs <ref type="bibr">[106]</ref> that interact with modelers will dramatically improve processes and productivity. The more context we can provide in this process, the better and more specific the final model will be.</p><p>Recent ChatGPT improvements, such as a code interpreter or Advanced Data Analysis <ref type="bibr">[73]</ref> enable downloading and visualizing data within a prompt, as well as browsing available datasets. Key capabilities of the AI assistant for model building can be summarized as follows:</p><p>&#8226; Manipulate (load/save/add/update/remove) model components.</p><p>&#8226; Recommend/suggest datasets or components.</p><p>&#8226; Download/catalog datasets.</p><p>&#8226; Configure parameters and settings.</p><p>To the best of our knowledge, the idea of AI-assisted world building has never been used in the model and simulation community and has the potential to disrupt the field of simulation and modeling. Specifically, such an approach promises the bridge the gap between domain experts (who know what to model) and system builders (who know how to implement it).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2">Simulation vs. (Deep) Generation</head><p>This work is motivated by the fact that since large trajectory datasets are sorely missing, realistic synthetic trajectory datasets are of great importance. The authors of <ref type="bibr">[132]</ref> propose a new framework called "End-to-End Trajectory Generation with spatiotemporal validity constraints" (EETG). The EETG framework employs factorized latent sequential deep generative models to disentangle global and local semantics while learning trajectory representations in an end-toend fashion. Global semantics capture trip reasons, e.g., commuting, a stroll downtown, airport pickup. Local semantics are related to a specific location and time, and capture spatiotemporal autoregressive patterns, effectively modeling the dependency of subsequent trajectory samples. To ensure that this generative process also produces trajectories that resemble real-world trajectories, i.e., obey motion physics, the work in <ref type="bibr">[132]</ref> also introduces a novel constrained optimization solution that reduces the probability of generating invalid and irrational vehicle trajectories, such as speed limits in combination with turn angles.</p><p>Figure <ref type="figure">6</ref> provides a heatmap representation (stacked semi-opaque points representing trajectory position samples) of input/original trajectory datasets (Figures <ref type="figure">6a</ref> and <ref type="figure">6c</ref>) and the corresponding generated datasets (Figures <ref type="figure">6b</ref> and <ref type="figure">6d</ref>) for</p><p>Beijing (T-drive data <ref type="bibr">[123]</ref>) and Porto, Portugal <ref type="bibr">[38]</ref>.</p><p>The results of Figure <ref type="figure">6</ref> and other examples presented in <ref type="bibr">[132]</ref> show that end-to-end trajectory generation without any knowledge of the underlying road network can produce sets of trajectories that resemble network-constrained movement. As such, end-to-end generation is a viable alternative to simulation-based generation.</p><p>But one of the challenges in deep learning generative model is hallucination, that is, generating responses that are either factually incorrect, nonsensical, or disconnected from the input prompt. To ensure that this generative process also produces trajectories that resemble real-world trajectories, i.e., obey motion physics, the work in <ref type="bibr">[132]</ref> also introduces a novel constrained optimization solution that reduces the probability of generating invalid and irrational vehicle trajectories, such as speed limits in combination with turn angles. As such, end-to-end generation is a viable alternative to simulation-based generation in scenarios where it is not critical to explain underlying motivation and decision-making process of human behaviors. The more context we can provide in this process, the better and more specific the final model will be. The AI assistant can be beneficial for non-domain experts as well as domain experts.</p><p>Model components developed by other users can be integrated using the model builder, and the AI assistant can suggest model components regardless of one's expertise. Even if users adopt model components outside of their expertise, their model is explainable because such simulation models consist of white box components unlike black box models.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.3">Bias and Ethical Considerations of AI-Generated In-Silco Worlds</head><p>Our main idea of using an AI assistant to support the building of an agent-based model is to abstract the implementation aspect and allow non-technical users, such as social scientists and epidemiologists to implement aspects of human behavior without having to write code. But one can take this vision one step ahead, and use an AI to generate the rules of human behavior for the AI itself to implement. Such an approach would take the human out of the look and let an AI generate an in-silico world that reflects the AI's understanding of human behavior.</p><p>Aside from the computational challenges of a model of human mobility that not only reflects behavior relevant to a research area (such as the spread of an infectious disease) but all human behavior that an AI can "think" of, we also see potential challenges concerning data bias and privacy.</p><p>Data Bias: Abstractly, an AI captures human knowledge that was used to train it. This human knowledge is often taken from public textual data sources such as Wikipedia or news sources. But such sources are often biased to high entropy (or interesting) events. For example, news articles may report criminal activity in one city, but won't report the lack of a such activity in another city. Using such data, that is biased to high entropy events may cause the AI to overestimate the frequency of such events. Data generated from such an AI-modelled (and implemented) simulation may thus overestimate the frequency and magnitude of rare events. Using such a dataset for research may confirm </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Data</head><p>Representative Unbiased Spatial Temporal Demographic Application Available? Data? Data? Information? Information? Information?</p><p>Table 4. Properties of real-world datasets available in different applications described in Section 5.</p><p>spurious hypothesis relating to such rare events. For example, a study modeling and simulating the number of crime events may yield an "interesting" result only because the simulation it was trained on overrepresents interesting data.</p><p>Data Privacy: Another problem with having the AI not only guide the implementation but also the modeling is the problem of data privacy. A big advantage of our envisioned in-silico simulation is that simulated individual agents to not related to any real-world humans, as all data is based on aggregate information such as census data. But by using an AI (and thus, the knowledge stored in the AI using a large language model) to generate worlds, we may indeed capture information about real-world individuals that may have been captured in the data used for training the large language model. Thus, any claims that research using such an in-silico world does not include any personally identifiable information may not longer hold.</p><p>Due to these considerations, we would like to advice caution in using a large language model-based AI in the modeling of an in-silico world for human mobility data science.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">APPLICATIONS AND CHALLENGES FOR HUMAN MOBILITY DATA SCIENCE USING SCALABLE AND REALISTIC IN SILICO WORLDS</head><p>Consider an in silico world that represents the realistic behavior and movement of a large number of individuals and generates large sets of individual-level mobility data from this world. This section describes some possible applications using such data for broad impacts in science and everyday life. To illustrate how in silico data collection may support research in each of the following application domains, Table <ref type="table">4</ref> shows the limitations of publicly available real-world datasets for each application which can be addressed by using data generated in silico. For example, in all of the following applications, real-world datasets pertain to a small subset of the population selected by participation in a corresponding data collection application (such as a location-based social network or a contract tracing application). We know that such data will overrepresent the technically savvy and underrepresent old and vulnerable populations. By using data simulated in-silico, we can evaluate how such lack of representative data collection may affect applications such as community detection and contact tracing. For each application, details are described in the following.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1">In Silico Human Mobility Data Science for Social Link Prediction</head><p>This use case envisions a social link prediction data science challenge using existing simulation capabilities detailed in <ref type="bibr">[39]</ref>. The idea is to generate a very large set of individual trajectories and check-ins (many orders of magnitude larger than existing real-world datasets) along with the dynamic social network of friendship between the simulated individuals. Using this dataset as ground truth, we can publish a sample of "observed" trajectories, check-ins, and social connections to challenge the community to solve the problem of estimating missing social connections that are not captured in this sample.</p><p>Example ground-truth data sets for this use case can be found at OSF (<ref type="url">https://osf.io/e24th/</ref>). Due to the excessive size of some of these LBSN data sets introduced in the following, we recommend that researchers interested in using the data re-run the simulation locally instead of downloading the data directly. Parameterized executables are available for download ( <ref type="url">https://github.com/gmuggs/pol</ref>), and our simulation is fully serialized and deterministic, such that the data generated locally is guaranteed to be identical to the downloadable data. Details on the underlying simulation to model realistic human behavior can be found in <ref type="bibr">[139]</ref> with specific settings (including simulation initialization) to generate realistic datasets detailed in <ref type="bibr">[39]</ref>.</p><p>The goal of the generated data sets discussed in the following is to act as benchmark data sets for the LBSN community, where publicly available datasets (as surveyed in Table <ref type="table">2</ref>) are very small and biased to a small subset of the population using the respective data collection phone applications. We generated a mix of real and synthetic urban settings.</p><p>Real road network and point-of-interest data were obtained from OpenStreetMap (OSM) <ref type="bibr">[74]</ref>, where we downloaded data for the greater New Orleans, LA (NOLA) metropolitan area, and the George Mason University (GMU) office for Geo-Information Systems (GIS) Facilities Archives [28] provided us with data for the Fairfax, VA campus of GMU. We also generated two synthetic urban datasets of different size and layout denoted as Small (TownS) and Large (TownL).</p><p>These synthetic urban components were created using a spatial network and a place generator based on a generative grammar similar to L-systems described in <ref type="bibr">[40]</ref>.</p><p>Within NOLA, the area we concentrated on was the historic French Quarter (FQ). The GMU area captures the main campus in Fairfax, VA. Both areas were prepared using QGIS desktop GIS software <ref type="bibr">[81]</ref> and JOSM <ref type="bibr">[37]</ref>. Data preparation involved editing the data sets to produce three separate shapefiles <ref type="bibr">[23]</ref>: (i) building footprints, (ii) transportation networks (road and sidewalk layers), and (iii) building purpose (i.e., residential, commercial, etc.).</p><p>Table <ref type="table">5</ref> provides an overview of the generated output data from the location-based social network simulation. It</p><p>shows the number of agent check-ins and the number of social links attributed to each of the scenarios. We observe that the number of check-ins increases for all study areas linearly with the number of agents. This is plausible, as the number of hours per day that agents can spend to satisfy their needs and visit sites is independent of other agents.</p><p>However, we do see that the number of social links increases super-linearly with the number of agents. This can be explained by more agents leading to larger co-locations of agents, creating chances for each pair of agents in the same co-location to become friends. We note that the generated temporal social network may have more edges than we have agent pairs. This is due to the temporal nature of the network. It reports changes over time and as such a single pair of agents can have multiple friends and unfriend events. The reported number corresponds to the number of new edges added to the temporal social network, regardless of the duration of these events. The super-linear growth of the social network also explains the super-linear run-time to create each data set, ranging from less to one hour for the 1000 agent instances to 10.5h for 5000 agents. Besides (i) the number of check-ins and (ii) social links, we also report (iii) the run-time of each simulation and (iv) the resulting data size in Table <ref type="table">5</ref>. It is interesting to see that even small simulations can create sizable results with a longer duration. For example, the smaller synthetic urban component that had the longest (221mo) simulation period produced a 5.5GB result data set. However, the actual run-time of this simulation is shorter than a simulation with more agents that had a shorter simulation period, e.g., NOLA-5K -15mo produced only a 2.3GB data set and took 1233min to run vs. NOLA-1k -221mo produced a 5.5GB data set with a run-time of 774min.</p><p>Increasing the number of agents results in more complex data structures, e.g., social networks, which in turn increases the run-time of the algorithms to process them at each step of the simulation. Figure <ref type="figure">7a</ref> shows the average number of friendships per simulation time measured in 5-minute steps for all the 1K networks. In all cases there is a three-month (one month is equal to 12 * 24 * 30 = 8640 steps) settling time during which agents establish friendships (starting with an empty social network). After this phase, the friendship degrees fluctuate around the mean values for each simulation.</p><p>We observe that the two real networks exhibit a denser social network, due to a more uneven distribution of agents, leading to large groups of agents to co-locate at sites to become friends.</p><p>For a more detailed view of the resulting social network, Figure <ref type="figure">8</ref> shows two visualizations of the social networks of 1K agents exemplary for the large synthetic network and NOLA at the end of the 15mo simulation. These visuals show different types of network structures, such as two to three large social communities for the synthetic TownL, and one large community for NOLA. Since it is hard to describe the evolution of a social network over time, we have created a video for each of the four spatial areas showing the social network evolution over the 129, 600 steps within the 15mo simulation time. These videos can be found at <ref type="url">https://mdm2020.joonseok.org</ref> and show how social networks evolve from small isolated cliques into a large and complex network showing different sub-structures. The video can also help explain the patterns observed in Figure <ref type="figure">7a</ref> in which agents make friends during the day, while losing some at midnight at which time we periodically lower the weights of the social network.</p><p>These videos also show, at each step of the simulation, the distribution of the number of friends per agent, an exemplary one is shown in Figure <ref type="figure">7b</ref> for the small synthetic network (TownS-1K). We observe that all case studies exhibit realistic long-tail distributions of the number of friends: While most agents have 5-25 friends, there are outliers having 50+ friends, but also agents having only three or fewer friends. This observation agrees well with a limit on the number of people with whom one can maintain stable social relationships (cf. the Dunbar number <ref type="bibr">[138]</ref>). In our simulated world, we observe an emerging limit of about 35 friends resulting from agents striving to maximize their number of friends (to satisfy their love need) while having to prioritize lower needs in Maslow's hierarchy (sleep, food, money). More details on how agents become friends through collocation and and how friendship ties decay over time if not maintained can be found in <ref type="bibr">[139]</ref>.</p><p>Woodstock '18, June 03-05, 2018, Woodstock, NY Z&#252;fle, Pfoser, Wenk et al.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2">Emergency Response</head><p>Emergency response is driven by planning in relation to different crisis scenarios and informed by considerable amounts of information and data. No unifying model is available that integrates the available information and goes beyond a static view of a community and its infrastructure. We can use in silico mobility science in emergency preparedness exercises to better gauge situations and to simulate potential response scenarios, e.g., experimentally evaluating evacuation scenarios due to extreme weather or due to an anthropogenic disaster. By capturing the complex dynamics of human behavior during emergencies, simulations can help identify potential bottlenecks, evacuation challenges, and areas that require additional resources. This information can be used to refine hazard mitigation plans and allocate resources more effectively, ultimately reducing the potential impact of disasters.</p><p>Simulating emergency events in our in silico world allows us to study the behavior of simulated individuals given different emergency response strategies. Many people may immediately leave the city using a shortest-path approach, but others may first want to ensure the safety of their dependents, such as children and the elderly. Other individuals may simply stay at home and forego evacuation due to their beliefs, or due to a lack of communication of the situation (e.g., due to language or technological barriers which can be simulated). Such a result would provide us with a very large trajectory dataset of all individuals capturing their behavior and movements. Data may show bottlenecks (for example, in traffic or communication) that may be difficult to predict without a simulation, or using existing evacuation simulations that assume perfect (yet unrealistic) behavior in which all individuals immediately seek the shortest path out of the evacuation area <ref type="bibr">[7,</ref><ref type="bibr">94]</ref>.</p><p>In addition to understanding "what will happen?" in a disaster scenario, our in silico world also allows us to run experiments directly in the simulated world by using prescriptive analytics focusing on answering "how to make X happen?" questions. For example, we can evaluate the effectiveness of different traffic management policies and different strategies to effectively communicate evacuation notifications to individuals and optimize optimal strategies.</p><p>In this specific context, an in silico simulation can become the computational foundation of a digital twin, which can serve as a powerful tool to raise public awareness of hazards and mitigation strategies. By visualizing the potential impacts of disasters and the benefits of mitigation measures, the digital twin can engage and educate citizens, encouraging their participation in community preparedness efforts.</p><p>Improved emergency preparedness and response strategies can lead to reduced response times, more efficient resource allocation, and minimized damage. Raising public awareness encourages citizens to actively participate in related initiatives, building more resilient communities. A better prepared community will experience fewer casualties and reduced property damage during emergencies, leading to significant savings, especially considering the long-term impacts of human losses and the high costs associated with property damage and recovery efforts. Furthermore, the intangible benefits of an in silico "sandbox" would include enhanced public safety, improved inter-agency collaboration, and increased public awareness.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3">Infectious Disease Spread Prediction and Digital Contract Tracing</head><p>During the outbreak of an infectious disease, contact tracing allows health officials to warn individuals who may have been exposed to an infectious disease. Digital contact tracing applications allow to systematically detect such contacts, even contacts with strangers that an individual may not be able to remember. Unfortunately, due to the COVID-19 pandemic, it has been shown that contact tracing applications were ineffective due to an insufficient adoption of the population <ref type="bibr">[10]</ref>. In an in-silico world, we can learn how underrepresented groups in contact tracing applications will affect the spread of an infectious disease by giving insights on what we don't know in the real world.</p><p>We can simulate the spread of an emerging infectious disease and evaluate optimal policies to mitigate the spread.</p><p>By having a ground truth of all infection states of all agents (in our in silico world) we can evaluate how much sampled data we need to be able to successfully trace and mitigate a disease. For example, we know that most contact tracing apps have an uptake rate of less than 10%. That means that for any disease transmission event, we only have a chance of 0.1 2 = 1% of capturing both the infector and the infected in the contact tracing app <ref type="bibr">[65]</ref>. To answer the question of what degree of participation of a contact tracing app is needed to be effective has been tackled in recent work <ref type="bibr">[126]</ref> using up-sampling of real-world data (which only captures a small portion of the ground truth population). Such up-sampling however, assumes that the data is a representative example of the population, whereas we know that real-world human mobility datasets (see Section 2) and real-world infectious disease datasets are a biased sample of the population <ref type="bibr">[32]</ref>.</p><p>Having a ground truth of all individual human mobility data as well as a ground truth of all infections would allow us to gain a deeper understanding of the capabilities of different infectious disease spread prediction and contact tracing models.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.4">Human Mobility Modeling</head><p>High fidelity, fine-grained, modeling of human mobility is of critical interest to the Intelligence Community, in particular for establishing models of "normal" movement capable of encoding the diversity of human movement present across times, locations, and people <ref type="bibr">[18]</ref>. Having realistic models of normal models may allow for the identification of anomalous movements such as loitering or circling around an area or event of interest, unusual movement between locations, unusual volume of movement to a point of interest, or unusual congregations of individuals. But real-world human mobility datasets are very sparsely sampled and biased to users for specific apps.</p><p>Current research for understanding human mobility only provides a high-level understanding. The key limitation in achieving this goal is the lack of ground-truthed mobility datasets. But without full knowledge of the underlying activities, it is impossible to characterize what movement is practically detectable as anomalous. In silico individual mobility data science will allow us to understand what types of activities can be modeled and provide datasets that are grounded in these activities, thus allowing the generation of mobility datasets where each trajectory is labeled with the ground truth activity that caused the individual to perform the observed mobility. Having such a ground truth will allow for testing our capabilities of modeling such activities and of identifying outliers. For example, we may simulate the trajectories of a pickpocket circling near a landmark, and test our abilities of identifying such behavior directly from trajectory data.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.5">Trajectory Benchmark Data.</head><p>For a fair comparison of different research methods, it is paramount to compare solutions on the same data sets. As discussed in Section 2, publicly available data sets lack volume, temporal information, and ground truth to reliably generalize knowledge that can be mined from them. As an alternative to using these existing data sets, the mobility data science community has proposed solutions to efficiently crawl data from LBSN data providers <ref type="bibr">[36]</ref>. However, this data is the intellectual property of the respective LBSN providers, and publishing their data for research benchmarks will violate their license agreements. As it is not possible to crawl data from the past, researchers will ultimately find themselves comparing their solutions on similar, but not identical data sets crawled at different times. Generated data from our in silicio simulations can fill this gap. They allow different research groups, at different times, to evaluate their algorithms for different problems on the same data sets. Furthermore, simulated benchmark data is extensible. If researchers choose to use a simulation to generate a new data set for their particular application, then the corresponding parameter file can be added to a repository. However, for very large trajectory and LBSN data sets, which may exceed 10TB of filesize, instead of downloading the data directly, researchers can share their data by only providing the self-executable simulation for local re-generation of data due to bandwidth constraints. Towards this vision of large-scale simulated trajectory benchmark data, a first Data and Resource paper has recently been published <ref type="bibr">[2]</ref> providing terabytes of trajectory data (more than a million times larger than existing trajectory data sets), simulating four different regions of interest for tens of years of simulation time. The simulation can easily be adapted to new study regions as described in <ref type="bibr">[43]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.6">Location Recommendation.</head><p>While there exists a large body of work on the problem of location recommendation (e.g., <ref type="bibr">[6,</ref><ref type="bibr">13,</ref><ref type="bibr">48,</ref><ref type="bibr">100,</ref><ref type="bibr">115]</ref>) these works use the publicly available datasets surveyed in Section 2. Due to the small sampling size of these datasets, it becomes difficult to understand how the proposed solutions would generalize to a full population. Using an in-silico world, we are able to bridge this gap.</p><p>To recommend locations, a simulation may allow agents to rate sites on a one to five-star scale. This rating could be determined by a deterministic function of the agents' preferences and the locations' attributes. To leverage this simulation for location recommendation, our simulation can be extended to obfuscate ratings by random noise (of parameterizable degree). This obfuscation can be deliberately biased, such as giving low ratings a higher chance to appear, with medium ratings more likely to be omitted. Such data would allow researchers to experimentally compare existing methods and evaluate the effect of bias between observed and ground truth ratings. Such comparison enables us to answer the question of recommendation systems' generalizability to the whole population, or if they overfit their models towards a sub-population of individuals that use the recommendation service.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.7">Community Detection.</head><p>For the task of Community Detection (often referred to as social network clustering), existing work (e.g., <ref type="bibr">[103,</ref><ref type="bibr">118]</ref>) also uses the relatively small datasets surveyed in Section 2. This makes it difficult to understand if the communities found on these small datasets are representative of the entire population. To leverage in-silico human mobility simulation for community detection and social network clustering, we can extend the simulation to impose circles of friends (i.e., strongly connected groups) in our social network. Then, by observing co-locations from the data, we can see which existing solutions can best approximate the imposed ground truth social networks. This data generation provides the ground truth for communities which can be used to evaluate the accuracy of community detection algorithms.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6">CONCLUSION</head><p>Human mobility data science using individual-level data has the potential to improve our understanding of human behavior and has many applications with broad impacts on society. However, existing datasets of human mobility are too limited to trust results and generalize to broad populations and existing simulation frameworks do not capture realistic human behavior. The vision presented herein is to bridge this gap and reignite research on mobility data science by breaking the chains of small data sets and embracing in silico human mobility data science. In silico human mobility data science can leverage socially realistic and scalable simulation to generate massive datasets of individual-level trajectories. We envision a world builder which allows lay users to rapidly create an in silico world that captures human behavior of interest to be studied, and which allows learning from and adapting to samples of real-world trajectory data.</p><p>Based on such in silico worlds for which we have limitless human mobility data, we describe applications that transcend the boundaries of current human mobility data science, such as broadly identifying friendship from (co-)locations, recommending locations to users, and identifying trajectories that correspond to malicious behavior.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>Woodstock '18, June 03-05, 2018, Woodstock, NY Z&#252;fle, Pfoser, Wenk et al.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_1"><p>The simulation was designed for DARPA's Ground Truth Program. In Ground Truth, Technical Area 1 performers developed simulators for artificial but socially-plausible worlds. Technical Area</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_2"><p>performers that did not know how simulations worked had to explain how the world works, predict what would happen in the future, and prescribe some remedy to achieve certain goals (e.g., minimizing the number of infected agents). Here, 'other performers' indicate TA2 teams.</p></note>
		</body>
		</text>
</TEI>
