<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Probabilistic simulation supports generalizable intuitive physics</title></titleStmt>
			<publicationStmt>
				<publisher>Cognitive Science Society</publisher>
				<date>07/24/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10542746</idno>
					<idno type="doi"></idno>
					
					<author>Haoliang Wang</author><author>Khalid Jedoui</author><author>Rahul Venkatesh</author><author>Felix Binder</author><author>Joshua B Tenenbaum</author><author>Judith Fan</author><author>Daniel Yamins</author><author>Kevin A Smith</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[How do people perform general-purpose physical reasoning across a variety of scenarios in everyday life? Across two stud ies with seven different physical scenarios, we asked participants to predict whether or where two objects will make contact. People achieved high accuracy and were highly consistent with each other in their predictions. We hypothesize that this robust generalization is a consequence of mental simulations of noisy physics. We designed an “intuitive physics engine” model to capture this generalizable simulation. We find that this model generalized in human-like ways to unseen stimuli and to a different query of predictions. We evaluated several state-of-the-art deep learning and scene feature models on the same task and found that they could not explain human predictions as well. This study provides evidence that human’s robust generalization in physics predictions are supported by a probabilistic simulation model, and suggests the need for structure in learned dynamics models.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Every day, we interact with the physical world in a variety of ways. We might start the morning by pouring cereal into a bowl, and later stably stacking that bowl with the rest of the dishes in the sink. Around the house, we might use a book to stabilize a wobbly chair, or throw trash into the trash can. Later with our kids we might build towers with blocks, or determine the right shot in a game of billiards. All of these scenarios require knowledge of many different principles of physics: from containment to stability to ballistic motion to collision dynamics. Yet we handle each of these tasks naturally, and often with little effort. But how are people able to do such general-purpose physical reasoning?</p><p>One hypothesis that has grown in prominence over the past decade is that we have a cognitive module that can perform general purpose, probabilistic physics simulation, often termed the Intuitive Physics Engine <ref type="bibr">(Battaglia, Hamrick, &amp; Tenenbaum, 2013;</ref><ref type="bibr">Ullman, Spelke, Battaglia, &amp; Tenenbaum, 2017;</ref><ref type="bibr">Smith et al., in press)</ref>. Under this hypothesis, general physics understanding arises because the simulation engine contains more primitive components for modeling the world -representations of objects and the forces they exert on each other, latent properties such as mass or elasticity that constrain how objects respond to forces, and key dynamic quantities such as momentum and events such as collisions -and combines them with uncertainty about the state of the world to reason probabilistically about a wide range of scenarios we might expect to encounter in day-to-day life. Thus, much like the physics engines that underlie many computer simulations, these building blocks of knowledge can be combined to model much more complex and combinatorial situations.</p><p>While this hypothesis has received quantitative support from many studies, a crucial aspect of it has never been explicitly tested. Prior studies that model human intuitive physics have typically focused on just one scenario at a time: e.g., how or whether a stack of objects might fall <ref type="bibr">(Battaglia et al., 2013;</ref><ref type="bibr">Hamrick, Battaglia, Griffiths, &amp; Tenenbaum, 2016;</ref><ref type="bibr">Zhou, Smith, Tenenbaum, &amp; Gerstenberg, 2023)</ref>, how moving objects will bounce off each other and fixed obstacles <ref type="bibr">(Smith &amp; Vul, 2013;</ref><ref type="bibr">Ullman, Stuhlm&#252;ller, Goodman, &amp; Tenenbaum, 2018;</ref><ref type="bibr">Gerstenberg, Goodman, Lagnado, &amp; Tenenbaum, 2021;</ref><ref type="bibr">Neup&#228;rtl, Tatai, &amp; Rothkopf, 2021)</ref>, or how liquid will pour <ref type="bibr">(Bates, Yildirim, Tenenbaum, &amp; Battaglia, 2019;</ref><ref type="bibr">Kubricht et al., 2016)</ref>. While these models all use a physics engine at their core, across this research, modelers make different assumptions about the particulars of the physics engine, and fit different parameters to capture uncertainty about the state of the scene or how physical events resolve. This modeling approach risks overfitting to specific scenarios, and thus cannot answer the question of whether people have a general purpose physics simulator, or use different systems for different physical principles. Indeed, another theory of human physical reasoning is that our judgments are based on inferences from past experience. This idea was first manifested in exemplar-based models and simple heuristics <ref type="bibr">(Gilden &amp; Proffitt, 1989;</ref><ref type="bibr">Nusseck, Lagarde, Bardy, Fleming, &amp; B&#252;lthoff, 2007;</ref><ref type="bibr">Proffitt, Kaiser, &amp; Whelan, 1990;</ref><ref type="bibr">Sanborn, Mansinghka, &amp; Griffiths, 2013)</ref>: that people might base their judgments exclusively on combinations of features of the initial scene configuration without explicit reference to physical dynamics. Similar ideas have also been expressed by recent neural network models that learn to predict dynamics by watching videos. Proponents of this approach suggest that learning physics from raw data provides two benefits: these models can extract generalizable physical principles more flexibly than if the models were to rely on a fixed simulator, and can work directly from visual inputs in a way that physical simulation models on their own do not. A range of models have been proposed that express a spectrum of assumptions about what parts of physics should be learned, from those that attempt to jointly learn a scene representation and dynamics with few assumptions about the structure <ref type="bibr">(Babaeizadeh et al., 2020)</ref>, to models that assume the scene structure is known and try to learn only how objects interact <ref type="bibr">(Han et al., 2022)</ref>, and many in between. While these neural networks are often intended purely to advance an AI system's understanding of the physical world, they have been proposed as hypotheses for how infants learn physics <ref type="bibr">(Piloto, Weinstein, Battaglia, &amp; Botvinick, 2022)</ref>, and have been used</p><p>Dominoes Support Collide Contain Drop Link Roll Cue 500 ms Stimulus 450 ms A B Exp. 1 Exp. 2 Yes / No</p><p>Is the red object going to hit the yellow area?</p><p>Figure <ref type="figure">1</ref>: A: The seven different scenarios testing different physics principles. In each image, the object colored in red is the target object and the yellow area on the ground is the zone. B: Participants are cued with the target and zone objects, observe a short video, and predict either whether or where the target will contact the zone.</p><p>to predict both behavior and neural activity in monkeys performing physics prediciton tasks <ref type="bibr">(Nayebi, Rajalingham, Jazayeri, &amp; Yang, 2023)</ref>.</p><p>In this paper, we test whether human physical predictions can be explained by approximate probabilistic inference in a single, general physics simulator across a wide range of everyday settings. We use an adapted version of the Physion dataset <ref type="bibr">(Bear et al., 2021)</ref> which was designed to test arbitrary models' physics understanding in a variety of scenarios against both ground truth and human beliefs. We specifically test the generalizability of models: how well models can explain human predictions in scenarios that they have not been fitted or trained on. We show that an intuitive physics engine generalizes to these unseen scenarios in a human-like way, explaining human behavior only slightly worse than expected by the noise ceiling. We compare a variety of state-of-the-art deep learning models that encompass a range of assumptions about what is learned. Some jointly learn representations and dynamics with little structure, testing whether physics can be learned directly from video. Others learn to parse images into scene representations in a variety of ways -from few assumptions about scene structure to strong assumptions about 3D world structure -and learn dynamics on top of that representation with a recurrent network, in order to test how well these models produce representations that support learning physics. We also compare models that make heuristic predictions based on initial scene features. We find that an intuitive physics engine model captures human judgments and generalizes to unseen scenarios and novel tasks remarkably well, and far better than the deep learning and feature-based models tested. These results both support the mental simulation hypothesis as a generalizable mechanism for intuitive physical reasoning and point to the value of including stronger and more structured inductive biases into neural network models of intuitive physics.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Human experiments</head><p>To evaluate people's physical predictions across a wide range of scenarios, we adapted seven rigid body scenes from the Physion (Bear et al., 2021) dataset. These seven scenarios test a variety of physical concepts (Fig. <ref type="figure">1A</ref>): chains of collisions (dominoes), the stability of a stack of objects (support), the particular ways collisions resolve (collide), whether one object can contain another (contain), how individual objects fall (drop), how a collection of objects can be knocked over (link), and how objects roll or slide down a slope (roll).</p><p>In each trial, there is a target object and zone. The goal is to predict either whether (Exp. 1) or where (Exp. 2) the target object will contact the zone at some point in the future.</p><p>Each scenario consists of 150 trials (1050 total), varying in scenario-specific configurations (see <ref type="bibr">Bear et al. (2021)</ref> for details on their construction). Each trial consisted of a 450ms video in which the target object does not yet touch the zone. The trials were designed so that if the video had continued, in half of them the target would touch the zone within the next 2 seconds, but would never touch in the other half.</p><p>Experiment 1: Will it collide?</p><p>We first asked participants to make a binary judgement of whether they think the target object will contact the zone after watching a short video clip.</p><p>Participants 350 participants (50 per scenario; 198 female; all native English speakers) recruited from Prolific completed the experiment. Each participant was shown all 150 stimuli from a single scenario. Data from 33 participants were excluded for failing our preregistered inclusion attention checks. The experiment lasted approximately 15 minutes and participants were paid $3.50.</p><p>Task procedure The structure of our task is shown in Fig. <ref type="figure">1B</ref>. Each trial began with a 500-1500ms fixation cross. Participants then saw the first frame of the video for 500ms with the target object and zone flashing a red and yellow overlay respectively, followed by the stimulus video for 450ms. Participants then saw a screen with buttons to indicate "YES" (the target would contact the zone) or "NO" (it would not). Before the main task, participants observed 10 familiarization trials for which the full movie was shown post-prediction.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head><p>We first examined how often participants' predictions of contact agreed with the simulation outcome from the</p><p>... ... ... ... . . . ... physics engine used to create the stimuli. We found that people achieved high accuracy (proportion correct = 0.81, 95% CI=[0.77, 0.84]) and their performance was substantially above chance across all seven scenarios (t(6) =16.49, p &lt; 10 -5 ). Participants' data also demonstrated variation in performance across scenarios, achieving the highest accuracy in Dominoes (0.84) and lowest in Roll (0.74).</p><p>Even though people made errors on some trials, these errors were consistent across participants (cross-trial bootstrapped split-half reliability=0.94, 95% CI=[0.92, 0.96], Fig. <ref type="figure">3A</ref>). This pattern of high but imperfect accuracy and reliable errors is especially useful when comparing models with humans: to be a good explanation of how people make physical predictions across scenarios, a model should not only achieve high accuracy, but also err in the same ways as people do.</p><p>Experiment 2: Where will they touch?</p><p>Here we investigate more fine grained predictions by asking participants to indicate where they believe the target object will first contact the zone.</p><p>Participants A separate group of 245 participants (35 per scenario; 157 female; all native English speakers) recruited from Prolific completed the experiment. The experiment lasted approximately 16 minutes and paid $3.75.</p><p>Task procedure The task procedure was identical to the "Will it collide" task except that participants were asked to place a circular disk where they believed the target would first contact the zone (Fig. <ref type="figure">1B</ref>). For each scenario, the stimuli were the same as the previous experiment except that we filtered the 150 trials to only include trials where (a) the target object contacted the zone, and (b) this collision happened at a location that was unoccluded by other objects (Collide: 50 trials, Contain: 44, Dominoes: 68, Drop: 63, Link: 63, Roll: 65, Support: 48). After showing the first 450ms of the stimulus, the video froze on the final frame and participants used their cursor to position a disk on the target zone. Only the part of the disk overlapping the zone was displayed. When participants had placed the disk at their desired location, they clicked a "NEXT" button to register their prediction.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head><p>For each trial, we measured the center point of participants' disk placement positions as the 3D location in world coordinates. We excluded participants' placements that were off the zone by the disk radius (i.e. the disk had no overlap with the zone at all during the experiment, indicating that participants were not following the instructions or misclicked), accounting for about 5% of the data.</p><p>To assess how far off people are from the contact points given by the ground truth stimulus, we first calculated the Euclidean distance between the mean human predictions and the ground truth contact point for each trial, and averaged across trials. Because this metric is sensitive to the area of the target zone, for each trial we divided the distance by the standard deviation of participants' placements on that trial (analogous to d &#8242; in signal detection theory). We found similar patterns to the "will it" task: participants' predictions are significantly closer to ground truth than expected by chance, and this is true for every physical scenario (mean normalized dis-tance=1.39, 95% CI=[1.23, 1.51], t(6) = 43.25, p &lt; 10 -8 ). Furthermore, we calculated the split half distance between participants (the average distance between the mean predictions of evenly splitting participants into two random groups) and found that they are highly consistent with each other (mean distance=0.50, 95% CI=[0.38, 0.71], Fig. <ref type="figure">3B</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Intuitive Physics Engine</head><p>The two experiments described previously demonstrate that across a wide range of physical scenarios, people make good predictions but are also biased in systematic ways. We argue that these predictions can be characterized using a noisy physics engine that runs probabilistic simulations -that the characteristic patterns of errors and biases we observe in participants' data can be mostly explained by uncertainty about the state of the scene after watching the video plus a noisy, approximately correct simulator that transforms those initial states into a distribution over outcomes. We formalize this hypothesis in an Intuitive Physics Engine (IPE) model.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The architecture of the IPE</head><p>To model people's prediction in naturalistic 3D environments, we used Unity3D as the underlying physics engine and customized to add noise to model sources of uncertainty in humans. Following <ref type="bibr">Smith and Vul (2013)</ref>, we considered uncertainty along three different axes (see Fig. <ref type="figure">2</ref>): . Across all seven scenarios, the IPE better captures human response patterns than ground truth, and loses almost no predictive power across scenarios in the "all-but" fitting regime.</p><p>Error bars are 95% CIs.</p><p>Perceptual uncertainty: We modeled people's uncertainty in visual perception by adding noise to the initial positions and rotations of the objects. The starting position of each of the objects was perturbed around the true position by two-dimensional Gaussian noise parameterized by standard deviation &#963; x , &#963; y ; and the rotation by von Mises noise around the z axis parameterized by concentration &#954; z . <ref type="foot">1</ref>Physical property uncertainty: We capture people's uncertainty about physical properties that vary across objects but are not directly observable -the mass of different objects -by adding Gaussian noise to the true mass, parameterized by standard deviation &#963; m , truncated at zero.</p><p>Dynamic uncertainty: We considered people's uncertainty about how collision will resolve, by perturbing the resultant collision impulse force's magnitude by Gaussian noise around its true value with the standard deviation &#963; F , and direction by a spherical von Mises distribution centered on the true angle of the impulse with a concentration parameter &#954; &#952; .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Fitting model parameters</head><p>In order to determine the set of noise parameters that best describes human behavior, we fit the six parameters defined above to participants' data on the "will it" task. Because the "where" task requires no additional modeling assumptions, we can use the same model to compare to participants' data on this task, and thus can treat model performance on that task as generalization to a separate task.</p><p>We fit parameters by simulating the scenes for 2.5s with the noisy IPE 20 times for each scene. We measured the RMSE between the proportion of IPE runs that predict contact for each trial, and the proportion of participants that do. We minimized this RMSE using the HyperOpt package <ref type="bibr">(Bergstra, Yamins, &amp; Cox, 2013)</ref>.</p><p>We used two regimes for fitting. In the "all scenarios" regime, we fit the IPE to 20% of trials from each scenario (210 trials total). We then assessed performance on the 80% of trials the model had not been fit on (840 trials), testing generalization to new trials. In the "all-but-one" regime, we fit seven separate IPE models: each fit on the trials from six of the seven scenarios, then assessed performance of each of those models on the unseen scenario. Overall model performance was calculated by averaging over the performance of each model on its held-out scenario. This regime tests even stronger generalization: whether the uncertainty measured in separate scenarios can explain human predictions in scenarios uninvolved in the fitting <ref type="bibr">(Wang, Allen, Vul, &amp; Fan, 2022)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Model Results</head><p>We first test whether a single noisy simulator can explain a range of human judgments by evaluating the IPE's predictions against humans' on both the "will it" and the "where" task. We then assess how well a set of state-of-the-art deep learning networks explain human predictions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The IPE is physical-domain-general</head><p>Will it contact? We find that a single parameterization of the IPE can explain human judgments across scenarios. Using the "all" fitting regime, the IPE achieves human-level performance across all seven scenarios on the test set (mean accuracy=0.83, 95% CI=[0.79, 0.86], Fig. <ref type="figure">4A</ref>), and more importantly also has high correlation with human responses (mean correlation=0.87, 95% CI=[0.83, 0.90], Fig. <ref type="figure">3A</ref>), only slightly worse than could be expected by the human noise ceiling. We also compare against how well participants would be fit by assuming perfectly accurate predictions, and find that the IPE correlates with human predictions better (t(13)=2.42, p = 0.01, Fig. <ref type="figure">3A</ref>). This indicates that the IPE not only captures overall human performance, but also makes similar predictions on individual trials. Importantly, the IPE explains predictions better than the ground truth answers, suggesting that the uncertainty inherent in the model leads to uncertainty in outcomes that produce human-like errors across scenarios.</p><p>As a stronger test of generalization, we assess how well the IPE explains human data in the "all-but-one" fitting regime, where the model is assessed on scenarios it has not observed during parameter fitting. The IPE maintained high accuracy (mean accuracy=0.81, 95% CI=[0.76, 0.86]) and correlation with human responses (mean correlation=0.87, 95% CI=[0.84, 0.89]), and was nearly the same when it had access to trials from all scenarios -a pattern that held across all scenarios (Fig. <ref type="figure">3A</ref>). Thus uncertainty about physical properties can be assessed in one set of scenarios and extrapolated to separate scenarios without noticeably affecting performance.</p><p>Where will it contact? To evaluate the IPE against humans on precise location predictions, we reused the noise parameters fit on the "will it" task and extracted the contact location information between the target object and zone. The pattern of results is similar to those for the "will it" task: the</p><p>Training Protocol all scenarios all-but-one Model Class 2D static scene model 3D aware static scene model end-to-end dynamics model chance level IPE ground truth 0 0.4 0.8 0.2 1 0.6 -0.2 scene features A B D C Normalized distance to human response Normalized distance to true contact point 0.6 0.7 0.8 0.5 3 4 5 2 1 3 4 5 2 1 H u m a All deep learning and feature-based models perform worse than the IPE, and their performance suffers when generalizing in the "all-but-one" training protocol.</p><p>IPE's distance to the true contact remained the same across fitting regimes (all: 1.03, 95% CI=[0.88, 1.21], all-but-one: 1.02, 95% CI=[0.88, 1.18]; Fig. <ref type="figure">4C</ref>). The average normalized distance between the model's predictions and human placements is 1.07, 95% CI=[0.89, 1.21], significantly less than the distance between people and ground truth (t(13)=2.64, p = 0.01) and at the same time did not change when generalizing across scenarios (normalized distance=1.10, 95% CI=[0.91, 1.25], Fig. <ref type="figure">3B</ref>). Note that the "all-but-one" fitting regime is an extremely strong test of generalization: the model must generalize across participants, physical scenarios, and even the type of prediction. However, the IPE does capture human performance slightly worse on the "where" task than the "will it" task. This could be due to the aforementioned strong generalization hindering performance, because, e.g., one set of participants have different amounts of uncertainty than the other, or because simply judging whether contact will occur is a coarser measure than judging where contact would occur, and so parameter estimates should be less precise.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Comparing to deep learning and feature models</head><p>While we have shown that the IPE can explain human predictions across scenarios, another theory suggests human-like physics understanding can arise from less structured learning. Thus, in this section, we aim to test the generalizability of state-of-the-art deep learning models as well as models that make physical predictions based on scene features, and compare their predictions on the same stimuli to humans and IPE.</p><p>We selected state-of-the-art models from three representative model architecture classes. These models were either pretrained or finetuned on the Physion dataset. (1) We assessed a set of 2D static scene understanding models with an LSTM trained on Physion scenes to predict next-frame dynamics from scene representations with various amounts of structure <ref type="bibr">(DINO, Oquab et al., 2023;</ref><ref type="bibr">MAE, He et al., 2021;</ref><ref type="bibr">R3M, Nair, Rajeswaran, Kumar, Finn, &amp; Gupta, 2022;</ref><ref type="bibr">and ResNet He, Zhang, Ren, &amp; Sun, 2016)</ref>. This assesses whether the scene representations learned by these models support efficient learning of dynamics. (2) We assessed 3D scene understanding models with the same LSTM training on their scene representations (DFM, <ref type="bibr">Tewari et al., 2023, and</ref><ref type="bibr">PixelNERF, Yu, Ye, Tancik, &amp;</ref><ref type="bibr">Kanazawa, 2021)</ref>. This assesses whether richer, 3D-aware scene representations might support better prediction. (3) We assess end-to-end dynamics models pretrained on Physion scenes (MCVD, <ref type="bibr">Voleti, Jolicoeur-Martineau, &amp; Pal, 2022;</ref><ref type="bibr">FitVid, Babaeizadeh et al., 2020;</ref><ref type="bibr">and TECO, Yan, Hafner, James, &amp; Abbeel, 2022)</ref>. These models asses whether human-like physics knowledge could be learned in an unstructured manner from video.</p><p>In order to compare deep learning models to humans on the "will it" task, for each model, we first extracted features by showing the human stimulus (450ms) and concatenated them with the "simulated" features output by the model's dynamics predictor. To get a binary output from the models, we then froze the parameters of the model and fit a logistic regression on the features. The parameters for the logistic regression were fit on a separate set of stimulus provided by the Physion dataset, with the ground truth object contact labels acting as supervision. We evaluated these models on the same unseen 840 experimental trials that the IPE was evaluated on. As seen in Fig. <ref type="figure">4</ref> AB, none of the deep learning models reached human levels of accuracy, and they did not correlate with human predictions as well as the IPE. In the "all-but-one" regime we trained the deep models on six out of the seven scenarios and tested them on the held-out scenario, and found that performance dropped noticeably across the two tasks (p &lt; 10 -3 for all models for both accuracy and correlation), with many of the models only marginally exceeding chance levels. Next, we evaluated the deep models' predictions of where it believed contact would occur. Unlike the IPE, all the deep models compute on 2D images rather than 3D world coordinates, so we used logistic regression on the same model features as before to output a prediction probability distribution on a 16&#215;16 grid over the image and then projected the center of each cell in the grid to 3D world coordinates (see Fig. <ref type="figure">5</ref>). <ref type="foot">2</ref>In order to compare between humans, the IPE, and deep learning models, we needed to align their predictions. First, we transformed human and IPE predictions by translating their predictions on world coordinates back into 2D image coordinates, similarly binning them into 16 &#215; 16 grids, approximating the prediction using center of the cell and then projecting the center back to 3D world coordinates. Second, because participants and the IPE were only allowed to make predictions on the zone area, we re-normalized the prediction probability distribution from the deep learning models to be only on the zone. We then measured probability-weighted distance between the grid center points on world coordinates as a metric for prediction distance. <ref type="foot">3</ref> As shown in Fig. <ref type="figure">4</ref> CD, the deep learning models did not capture human performance as well as the IPE, and always had decreased predictivity in the "all-but-one" regime, though here most performed above chance, providing evidence that they had learned something about the physics of these scenes as a whole.</p><p>We also evaluated feature-based models on the same tasks, using objects' position, rotation, size, shape and velocity at the stopping frame (i.e. 450ms) individually as features, as well as a combination of all five, and then fit a linear readout model in the same way we did for deep learning models for each feature (including fitting to the 16x16 grid and normalizing for the "where" task). The combination of all scene features could predict whether contact would occur relatively well in the "all" training regime, but these features could not generalize to unseen scenarios in the "all-but-one" regime (Fig. <ref type="figure">4 AB</ref>). The scene features also poorly predicted "where" the objects would contact, with performance falling to close to chance in the "all-but-one" scenario (Fig. <ref type="figure">4 CD</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>In this paper, we tested the hypothesis that human physical predictions can be explained by approximate probabilistic inference in a single, general physics simulator across a wide range of everyday settings. Across two experiments, we found that a physics engine that runs probabilistic simulations generalized to unseen stimuli in human-like ways, but a set of state-of-the-art deep learning models and feature-based models do not yet reach that level.</p><p>One major point of difference between the IPE and the deep learning models we tested here is the input encoding: the IPE takes 3D information of the scene as inputs whereas the deep learning models compute on pixels. Learning directly from pixels can allow for greater flexibility in the representation of scenes and dynamics, but imposes the challenge of learning to extract scene information rather than that information being provided. In this case, however, it appears these representations do not support longer term predictions in untrained scenarios that require understanding the physics of the world. In future work, we will evaluate a broader set of deep learning models that impose greater structure on the learning of physics -e.g., graph neural networks that work from scene representations and explicitly parse the world into objects and their relations <ref type="bibr">(Mrowca et al., 2018;</ref><ref type="bibr">Li, Wu, Tedrake, Tenenbaum, &amp; Torralba, 2018;</ref><ref type="bibr">Han et al., 2022;</ref><ref type="bibr">Battaglia, Pascanu, Lai, Jimenez Rezende, &amp; Kavukcuoglu, 2016;</ref><ref type="bibr">Allen et al., 2022)</ref> -as well as noisy simulation models that rely on scene parsing models to provide information about the world <ref type="bibr">(Wu, Lu, Kohli, Freeman, &amp; Tenenbaum, 2017)</ref>. Systematic testing of broader sets of models can help inform us what additional structure is required to develop more human-like understandings of the physical world.</p><p>While the IPE model predicts human response patterns well, it is still below the noise ceiling. This is a pattern found in many studies going back to <ref type="bibr">Battaglia et al. (2013)</ref>, and is likely due to the fact that humans cognitive simulations are not exactly the same as computer physics engines, but instead have different implementations and additional simplifications <ref type="bibr">(Bass, Smith, Bonawitz, &amp; Ullman, 2021;</ref><ref type="bibr">Chen, Allen, Cheyette, Tenenbaum, &amp; Smith, 2023;</ref><ref type="bibr">Li et al., 2023)</ref>. Further research into the structure of human physical representations and simulations will be required to close this gap.</p><p>The physical world is complex and open-ended, yet we easily reason about a wide range of scenarios that we might encounter in everyday life. The current study suggests that this robust generalization behavior often comes from having a generalizable mental model of the physical world and the ability to continuously simulate forward about how the world will unfold. In the long term, such studies will help us to understand and implement the computational mechanisms needed for the deep learning models to be more human-like.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0"><p>Following<ref type="bibr">Battaglia et al. (2013)</ref>, we only consider position and rotation uncertainty along a plane because most objects are resting on the ground or another object, and so uncertainty along the z-axis would cause objects to either float or interpenetrate.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1"><p>We found empirically that structuring our problem as a 16&#215;16 grid classification task improved the readout training performance compared to a position regression task, while also guaranteeing finegrained predictions.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2"><p>We also considered Wasserstein distance over grid distributions, and found qualitatively similar patterns of results.</p></note>
		</body>
		</text>
</TEI>
