<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Quantifying Observed Prior Impact</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>01/01/2021</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10310167</idno>
					<idno type="doi">10.1214/21-BA1271</idno>
					<title level='j'>Bayesian Analysis</title>
<idno>1936-0975</idno>
<biblScope unit="volume">-1</biblScope>
<biblScope unit="issue">-1</biblScope>					

					<author>David E. Jones</author><author>Robert N. Trangucci</author><author>Yang Chen</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[When summarizing a Bayesian analysis, it is important to quantify the contribution of the prior distribution to the final posterior inference because this informs other researchers whether the prior information needs to be carefully scrutinized, and whether alternative priors are likely to substantially alter the conclusions drawn. One appealing and interpretable way to do this is to report an effective prior sample size (EPSS), which captures how many observations the information in the prior distribution corresponds to. However, typically the most important aspect of the prior distribution is its location relative to the data, and therefore traditional information measures are somewhat deficit for the purpose of quantifying EPSS, because they concentrate on the variance or spread of the prior distribution (in isolation from the data). To partially address this difficulty, Reimherr et al. ( 2014) introduced a class of EPSS measures based on prior-likelihood discordance. In this paper, we take this idea further by proposing a new measure of EPSS that not only incorporates the general mathematical form of the likelihood (as proposed by Reimherr et al., 2014) but also the specific data at hand. Thus, our measure considers the location of the prior relative to the current observed data, rather than relative to the average of multiple datasets from the working model, the latter being the approach taken by Reimherr et al. (2014). Consequently, our measure can be highly variable, but we demonstrate that this is because the impact of a prior on a Bayesian analysis can intrinsically be highly variable. Our measure is called the (posterior) mean Observed Prior Effective Sample Size (mOPESS), and is a Bayes estimate of a meaningful quantity. The mOPESS well communicates the extent to which inference is determined by the prior, or framed differently, the amount of sampling effort saved due to having relevant prior information. We illustrate our ideas through a number of examples including Gaussian conjugate and non-conjugate models (continuous observations), a Beta-Binomial model (discrete observations), and a linear regression model (two unknown parameters).]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>Prior knowledge and assumptions are central to many statistical problems, and in practice it is important to assess their impact on the final inference. When such an assessment is missing, it can be difficult to tell whether the results could be reproduced with a different prior, or whether similar studies could be made more efficient by incorporating existing information that was neglected. Of course, each researcher chooses whichever prior seems most appropriate to them, but reporting the impact of the prior allows others, and even the original researchers, to better interpret the results. These considerations are important in many scientific studies. For example, <ref type="bibr">Chen et al. (2019)</ref> propose a Bayesian analysis of the brightnesses of a large collection of stars based on a multi-telescope astronomical dataset, and highlight that scientific prior distributions provide key information about each of the specific instruments and play a substantial role in the final inference. For both the scientists directly involved in the study, and also others who rely on their work, it is important to understand the role of the prior distributions used, e.g., do the priors associated with one particular instrument have a much greater impact on the inference than those for other instruments?</p><p>One appealing and interpretable way to assess prior impact is to provide a measure of the effective prior sample size (EPSS), i.e., the approximate number of observations to which the information in the prior is equivalent. Gaussian conjugate models offer a canonical example: with observed data y i iid &#8764; N (&#956;, &#963; 2 ), for i = 1, . . . , n, and conjugate prior distribution &#956; &#8764; N (&#956; 0 , &#963; 2 /r), the posterior distribution of &#956; is N (w n &#563;n + (1w n )&#956; 0 , &#963; 2 /(n+r)), where w n = n/(n+r). Based on the denominator n+r in the expression for the posterior variance, the effect of the prior appears to be equivalent to that of r samples, so we say that the EPSS is r. However, this formulation faces two challenges: (a) it is not immediately clear how to generalize beyond conjugate models; and more importantly, (b) when &#956; 0 is arbitrarily different to &#563;n , the prior impact on the posterior mean is arbitrarily large, and is therefore clearly not equivalent to that of r samples.</p><p>EPSS measures have gained substantial attention in the literature, and a number of strategies have been proposed in response to the two challenges above, e.g., <ref type="bibr">Clarke (1996)</ref>, <ref type="bibr">Morita et al. (2008)</ref>, and <ref type="bibr">Morita et al. (2010)</ref>. Most of the strategies proposed rely on a comparison between the actual prior &#960; and a default or baseline prior &#960; b , e.g., the improper prior &#960; b (&#956;) &#8733; 1 would be a natural choice for the baseline prior in the Gaussian conjugate model above. This comparative information approach is necessary because there is no universal "non-informative" prior against which to measure prior impact, and Bayesian inference cannot be conducted without a prior. Early generalizations along these lines sought to match the prior &#960; to a hypothetical posterior distribution constructed using the baseline prior &#960; b and some hypothetical previous samples, that is, they interpreted the prior &#960; as the posterior from a previous analysis. The EPSS is then defined as the number of observations used in the hypothetical posterior distribution, e.g., <ref type="bibr">Clarke (1996)</ref> and <ref type="bibr">Morita et al. (2008)</ref>. These approaches successfully generalize the notion of EPSS, but do not address concern (b) regarding the real impact of the prior when the data mean and prior mean differ substantially. Indeed, these methods do not consider the observed data or the real posterior distribution at all. <ref type="bibr">Reimherr et al. (2014)</ref> instead suggested minimizing the discrepancy between two posterior distributions, one using the real prior &#960; and the other using the baseline prior &#960; b . In this case the EPSS is defined as the difference in the number of samples used by the two posteriors. Similar ideas have also been proposed in slightly different contexts, e.g., see <ref type="bibr">Lin et al. (2007)</ref> and <ref type="bibr">Wiesenfarth and Calderazzo (2019)</ref>. The <ref type="bibr">Reimherr et al. (2014)</ref> method offers many improvements over early approaches and goes beyond simply capturing the variance of the prior; it also partially quantifies the impact of the prior location. However, it averages over the data using the bootstrap, and therefore does not quantify the impact of the prior for the specific analysis carried out with the observed data at hand, which is of most interest in practice.</p><p>Another recent approach introduced by <ref type="bibr">Neuenschwander et al. (2020)</ref> defines the EPSS as the expected local-information-ratio (ELIR), i.e., the prior mean of the ratio of the prior information and Fisher information (of a single observation). This approach has the elegant property that a sample of size n from the posterior predictive distribution has an effective sample size of n plus the EPSS. However, similarly to the strategies already mentioned, this method does not take into account the observed data and therefore does not fully capture the impact of the prior on the Bayesian analysis at hand.</p><p>In this paper, we follow a similar approach to <ref type="bibr">Reimherr et al. (2014)</ref> but propose a new EPSS measure which addresses the above limitations by conditioning on the observed data, and thereby directly quantifies the prior impact for the actual analysis performed. Our measure is called the mean Observed Prior Effective Sample Size (mOPESS), where 'Observed' indicates that the observed data is treated as fixed, and 'mean' refers to an average over additional future samples drawn from the posterior predictive distribution (the purpose of which will become clear). This new measure was inspired by the work of <ref type="bibr">Efron and Hinkley (1978)</ref> which highlighted that observed Fisher information is sometimes more useful than expected Fisher information. By providing an explicit definition of the mOPESS in terms of future observations, we also identify the real-world estimand of interest, which we call the Observed Prior Effective Sample Size (OPESS), i.e., a quantity that would be realized if future samples were actually collected. The interpretation of the OPESS is essentially the number of additional samples that must be combined with the baseline prior &#960; b in order to obtain similar inference to that under our actual prior &#960;. In other words, the OPESS communicates how much sampling effort is saved by having access to the information in the prior &#960;, rather than only the default information captured by &#960; b . Further appealing properties of our mOPESS measure include a Bayes estimate interpretation and no lower limit on the observed data sample size n. The latter property is important because prior impact is often most pronounced, and therefore of most interest, when the sample size is small. In contrast, <ref type="bibr">Reimherr et al. (2014)</ref> require n to be large because their method relies on the bootstrap and an accurate estimate of the "true parameter" value, see Section 2.2 for a review. In summary, our approach represents a substantially improved method for quantifying prior impact in practice, and its real-world interpretation makes it a valuable tool for clearly reporting the contribution of priors in Bayesian analyses.</p><p>One possible limitation of our approach is the need for a baseline or default prior against which to compare the prior at hand. However, in our opinion, a key purpose of an EPSS measure is to communicate the impact of a prior, and for this objective comparing to a standard prior that many researchers are already familiar with is in fact a strength rather than a weakness. Of course, if necessary our mOPESS measure could be computed for several different baseline priors, but we suspect one version will usually provide a sufficient summary of prior impact. As mentioned above, most of the existing literature on EPSS measures has similarly concluded that comparing against a baseline prior is desirable or necessary or both. This paper is organized as follows. Section 2 briefly overviews some topics from the broader literature connected to EPSS measures and their uses, then summarizes the EPSS methods on which we build, and lastly provides a motivating Gaussian example to illustrate our mOPESS measure. Section 3 defines the mOPESS and discusses its computation, first in general and then in the specific case of the motivating Gaussian example introduced in Section 2.3. Section 4 provides intuition and theory supporting our method. Section 5 provides additional numerical results in the form of non-conjugate Gaussian, Beta-Binomial, and regression model examples. Section 6 provides a summary. Proofs are given in the appendices.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Connections with Existing Work and a Motivating Example</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1">EPSS and the Broader Literature on Prior Distributions</head><p>To provide greater context for the importance of EPSS measures, we now briefly discuss several concepts that have been studied in the literature on prior distributions, and highlight their connections to EPSS measures.</p><p>Prior-likelihood conflict, e.g., <ref type="bibr">Evans et al. (2006)</ref>; <ref type="bibr">Bousquet (2008)</ref>; <ref type="bibr">Walter and Augustin (2009)</ref>; <ref type="bibr">Nott et al. (2020</ref><ref type="bibr">Nott et al. ( , 2021))</ref>. The topic of prior likelihood-data conflict is very much related to EPSS measures in the sense that such conflict can indicate that the prior distribution has a large influence on the final analysis. On the other hand, measures of EPSS are not restricted to quantifying prior-likelihood conflict: in the case of prior-likelihood alignment or weakly informative priors, the mOPESS measure proposed here gives non-zero values which describe the extent to which the prior is facilitating the inference, e.g., by increasing posterior concentration.</p><p>Sample size determination, e.g., <ref type="bibr">Wang et al. (2002)</ref>; <ref type="bibr">Sahu and Smith (2006)</ref>; <ref type="bibr">Clarke et al. (2006)</ref>; <ref type="bibr">Gupta et al. (2016)</ref>. In clinical trials, a classical problem is to determine the sample size needed to achieve a certain power when performing a hypothesis test for the presence of a treatment effect, see <ref type="bibr">Sahu and Smith (2006)</ref> for more detailed discussion and examples. In its classical form this line of research does not usually emphasize prior impact and is not closely related to our work. However, the concepts of prior-likelihood conflict and EPSS are important for sample size determination in adaptive clinical trials, where priors are typically constructed from an interim analysis, e.g., <ref type="bibr">Hobbs et al. (2013)</ref> uses effective historical sample size (EHSS) to determine a randomization procedure for allocating patients. Similarly, <ref type="bibr">Wiesenfarth and Calderazzo (2020)</ref> compare EPSS measures that quantify prior information in terms of historical/external samples (e.g., <ref type="bibr">Morita et al., 2008)</ref> or current/new samples (e.g., <ref type="bibr">Reimherr et al., 2014)</ref>; and tailor the method in <ref type="bibr">Reimherr et al. (2014)</ref> to the adaptive design setting. <ref type="bibr">Wiesenfarth and Calderazzo (2020)</ref> specifically investigates priors that adaptively discard prior information in the case of prior-data conflict, e.g., robust mixture <ref type="bibr">(Berger et al., 1986)</ref>, power <ref type="bibr">(Ibrahim et al., 2015)</ref> and commensurate <ref type="bibr">(Hobbs et al., 2012)</ref> priors. Our mOPESS measure could be applied in such studies to identify other similarly "adaptive" priors and to better quantify their impact in Bayesian clinical trials.</p><p>Bayesian prior sensitivity analysis, e.g., <ref type="bibr">Berger (1990)</ref>; <ref type="bibr">Weiss (1996)</ref>; <ref type="bibr">Roos et al. (2011</ref><ref type="bibr">Roos et al. ( , 2015))</ref>. This line of work is aimed at assessing the impact on the posterior inference when the prior distribution is perturbed, thus possibly identifying parameters for which the posterior inference is highly sensitive to changes in the prior. In our work, we only assess the impact of the given prior on the current inference, as compared to a baseline prior, and do not consider whether this impact would be different for other similar priors. On the other hand, if posterior inference substantially varied across a group of priors, then our mOPESS measure would likely be high for some or all of the priors in question, and so our measure could be used to detect this type of sensitivity.</p><p>Prior construction for objective Bayes, e.g., <ref type="bibr">Kass and Wasserman (1996)</ref>; <ref type="bibr">Ghosh et al. (2011b)</ref>; <ref type="bibr">Berger et al. (2015)</ref>; <ref type="bibr">Consonni et al. (2018)</ref>; <ref type="bibr">Leisen et al. (2020)</ref>. For example, <ref type="bibr">Ghosh et al. (2011a)</ref> constructs "objective priors" by maximizing an approximate expression for the distance between the prior and the posterior under a general divergence criterion. In general this line of research focuses on constructing priors in such a way as to avoid the subjective nature of Bayesian inference, e.g., ensuring the prior has little influence in cases where the data contain substantial information. Our approach is about assessing the prior influence and giving an intuitively quantitative measure of the impact of an informative prior or subjective prior. Moreover, as opposed to approximate measures based on asymptotic expansions (e.g., <ref type="bibr">Ghosh et al., 2011a)</ref>, our method works on the exact posteriors and is particularly designed for finite sample settings, in which prior impact is potentially substantial. In our approach, reference priors (typically, but not necessarily, non-informative) are only used as a baseline to compare our working priors against. On the other hand, our mOPESS measure could likely be used to assess whether a prior is a suitable alternative to the current non-informative prior of choice.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2">Some Existing Methods of Measuring EPSS</head><p>Suppose that &#960; is our prior distribution for a collection of unknown parameters of interest &#952; &#8712; &#920;. Let &#960; b be a baseline prior and x = {x 1 , . . . , x n } be unknown hypothetical previous data with probability density f (x|&#952;). Imagine that our real prior &#960; is the posterior distribution q b (&#8226;|x) &#8733; f (x|&#8226;)&#960; b (&#8226;) computed using the unknown hypothetical dataset x. Under this formulation, <ref type="bibr">Clarke (1996)</ref> considers</p><p>where KL(g, h) denotes the Kullback-Leibler divergence (KL divergence) defined as &#920; log(g(&#952;)/h(&#952;))g(&#952;)d&#952;, and X is the support of f (for simplicity we assume X to be the same for all &#952; &#8712; &#920;). In words, the approach of <ref type="bibr">Clarke (1996)</ref> is to find the hypothetical dataset x * that, when combined with the baseline prior &#960; b , produces the posterior distribution q b (&#8226;|x * ) with minimum KL divergence from our true prior &#960;. The EPSS can then be quantified as the number of individual observations contained in x * . Note that the density f is a user specified hypothetical distribution for prior data, and is not necessarily the same as the model for any actual data.</p><p>The above approach is distinguished from most other methods (such as those mentioned below) in that it gives a specific dataset x * which represents the information in the prior. An advantage of this approach is that x * can potentially capture other aspects of the information contained in &#960; in addition to the EPSS. However, x * has no concrete relation to the likelihood or data at hand, which we consider to be a drawback, at least when the impact of the prior on a specific analysis is of primary interest. <ref type="bibr">Morita et al. (2008)</ref> adopt a similar approach but measure distance using the difference of the second derivative of the log densities rather than KL divergence. Furthermore, to avoid the peculiarity of reporting a specific dataset x * , and to take account of uncertainty regarding the hypothetical dataset; they take an expectation over x, i.e., they compute</p><p>. This treatment of the hypothetical previous data x may be preferable to that of <ref type="bibr">Clarke (1996)</ref>. But for the purpose of assessing prior impact on a specific analysis, the <ref type="bibr">Morita et al. (2008)</ref> method suffers from the same fundamental problem of not taking the likelihood of any actual data into account.</p><p>To address this limitation, <ref type="bibr">Reimherr et al. (2014)</ref> introduced the notion of priorlikelihood discordance and incorporated it in their measures of EPSS. The key change they proposed was to compare two posterior distributions rather than comparing a prior to a (hypothetical) posterior. To make the comparison, under each prior &#960;, they consider the expected mean squared error (MSE) when a draw from the posterior is used to estimate the true parameter &#952; T , i.e.,</p><p>where the "posterior MSE", as defined in <ref type="bibr">Reimherr et al. (2014)</ref>, is</p><p>Reimherr </p><p>where &#952; is the maximum likelihood estimate of &#952; T based on y, and &#219; is computed by averaging over datasets of size k drawn from the empirical distribution (hence the constraint that k n). By a slight abuse of terminology, we refer to their averaging method as bootstrapping (as they do). One of the novel aspects of this formulation is that their estimate of EPSS, r, is allowed to be negative. This is helpful when, for example, we are trying to assess if &#960; is a low-information prior and therefore might feasibly have less impact than &#960; b .</p><p>The approach of <ref type="bibr">Reimherr et al. (2014)</ref> described above has a number of advantages over earlier methods: (i) it focuses on the impact of the prior on posterior inference;</p><p>(ii) it incorporates the likelihood, although for reduced data size; and (iii) it proposes a potentially reasonable method for generating datasets to combine with &#960; and &#960; b (bootstrapping). There are however still some limitations of their approach. Firstly, their method averages over the data and therefore their measure of EPSS does not tell us what the impact of the prior is on the inference using the observed data y, which is of most interest in practice. Secondly, their approach relies on bootstrapping the data and estimating &#952; T which both require n to be large, but the impact of a prior is usually greatest and of most interest when n is small. Lastly, their use of MSE is not necessarily the best way of quantifying the difference between two posterior distributions and therefore the prior impact. Indeed, there is in fact no reason to introduce the notion of a true parameter value &#952; T in order to measure prior impact.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3">Motivating Gaussian Example</head><p>Suppose that we have observed data</p><p>&#8764; N(&#956;, &#963; 2 ), for i = 1, . . . , n. Assume that &#963; 2 is known, that our prior for &#956; is a conjugate prior, denoted &#960;(&#956;) &#8801; N (&#956; 0 , &#955; 2 0 ), and that the baseline prior is &#960; b (&#956;) &#8733; 1. Suppose that &#956; = &#956; 0 = 0, &#963; 2 = 1, and &#955; 2 0 = 0.1. We now use this example to illustrate the behavior of our EPSS measure, which is called the mean Observed Prior Effective Sample Size (mOPESS). All mathematical and computational details are deferred to Section 3.</p><p>The top left panel of Figure <ref type="figure">1</ref> shows the mOPESS plotted against the data mean &#563;n . It can be seen that the mOPESS increases with the difference |&#563; n&#956; 0 | between the prior mean and data mean, which is a key feature of our proposed measure of EPSS. The top right panel of Figure <ref type="figure">1</ref> shows the quantiles of the Observed Prior Effective Sample Size (OPESS), i.e., it summarizes the distribution of which the mOPESS is the mean for any given &#563;n . The OPESS represents a real-world quantity (defined in Section 3.1) which captures the impact of the prior, and can in principle be observed by collecting more samples. Variability in the OPESS for a fixed value of &#563;n indicates genuine uncertainty about the future observations. The mOPESS (i.e., the mean of the OPESS distribution) averages over this uncertainty and is our preferred single number summary of prior impact.</p><p>The bottom left panel of Figure <ref type="figure">1</ref> corresponds to the point indicated by a "+" symbol in the top left panel, i.e., one of two points with the largest value of |&#563; n&#956; 0 |. In particular, the bottom left panel shows the priors &#960; and &#960; b , as well as the corresponding posterior distributions denoted q n and q b n , respectively. Note that the posterior distribution q n is pulled towards zero by the informative conjugate prior &#960;. The top left panel shows that in this case the mOPESS is larger than for values of &#563;n that are closer to the prior mean &#956; 0 = 0. In particular, the mOPESS has a value of about 14.5, which has the interpretation that on average an investigator using q b n would need to collect 14.5 additional samples to obtain similar inference to an investigator using &#960;, for the specific value of &#563;n currently at hand, i.e., for &#563;n &#8776; -0.6.</p><p>The bottom right panel of Figure <ref type="figure">1</ref> illustrates the case where |&#563; n&#956; 0 | is smallest across the points plotted in the top left panel, i.e., the point indicated by a cross in the top left panel. From the bottom right panel we can see that in this case both Figure <ref type="figure">1</ref>: (Left top) mOPESS as a function of &#563;n . The solid curve shows a LOESS (LOcal polynomial regrESSion) fit to the points. (Right top) 95% quantile (long dash curve), median (short dash curve), and 5% quantile (dash-dot curve) of the OPESS as a function of &#563;n . The solid curve is the same as in the left panel. (Left and right bottom) posteriors q n and q b n (dashed lines) and priors &#960; and &#960; b (solid lines) for the dataset indicated by a green "+" symbol and a cross, respectively, in the top left plot.</p><p>posteriors are centered very close to zero. Specifically, the posteriors have substantial overlap because the conjugate prior &#960; is centered at &#956; = &#956; 0 = 0 &#8776; &#563;n and so does not cause q n to have a substantially different mean to q b n , only a smaller variance. Returning to the top left panel we can see that this case corresponds to a mOPESS value of around 10.5, which is one of the smallest among the points plotted.</p><p>In conclusion, we can see that the mOPESS is larger the further &#563;n is from &#956; 0 = 0 and that this is because the prior has more impact on the posterior in these cases. Thus, at least in this simple example, our mOPESS measure of EPSS seems to have an intuitive interpretation that well captures the way the prior impact changes with the observed data.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">mOPESS Definition and Computation</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">mOPESS Definition</head><p>Our guiding intuition is that we want to know how many extra samples are needed to obtain similar inference under the baseline prior as that under our informative prior. To that end, for each m = n + r, where r &#8712; Z &#8805;0 , we introduce a hypotheti-</p><p>The superscripts '(m)' are necessary because we do not assume that x (m+1) = x (m) &#8746; {x m+1 }, for reasons to be explained in Section 3.2. If x (m) was known for all m then intuitively we would choose the EPSS to be r = mn for the m that minimizes the distance between the two posterior distributions</p><p>, where &#960; is our real prior whose EPSS is to be measured, &#960; b is the baseline prior, and f is the model. We denote the distance for a given m by D(q b (&#8226;|x (m) ), q(&#8226;|y)). The collection of expanded datasets is denoted by x L = {x (n) , x (n+1) , . . . , x (L) }, where L is the maximum feasible value of m, or in other words Ln is the maximum feasible magnitude of the EPSS associated with &#960;.</p><p>We must account for the possibility that our prior &#960; is in fact less impactful than the baseline prior &#960; b . This happens when the prior &#960; is more diffuse than the baseline &#960; b or is similarly diffuse but has greater location agreement with the data than &#960; b . Thus, we also consider the alternative distance D(q b (&#8226;|y), q(&#8226;|x (m) )), where the extra hypothetical samples are combined with our real prior &#960; rather than with the baseline &#960; b . Again we do not assume that x(m) = x (m) , for reasons to be explained in Section 3.2. We write (n+1) , x(n+1) , . . . , x (L) , x(L) } to denote all the future samples combined, and for conciseness introduce the notation D(m) and D(m) as shorthand for the distances D q b (&#8226;|x (m) ), q(&#8226;|y) and D q b (&#8226;|y), q(&#8226;|x (m) ) , respectively. We can now define the underlying quantity of interest.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Definition 3.1. For a given realization of x all</head><p>L , the Observed Prior Effective Sample Size (OPESS) is</p><p>where</p><p>The OPESS is negative when S n (x all L ) = -1 because this suggests that &#960; is less informative than the baseline prior &#960; b . In practice, the future samples x all L are unknown and therefore the OPESS must be estimated. The mOPESS defined in Definition 3.2 below is simply the posterior mean of the OPESS, and as such provides a convenient estimate of the OPESS. Posterior quantiles of the OPESS distribution and other summaries could also be reported to provide a measure of uncertainty.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Definition 3.2. The theoretical (posterior) mean Observed Prior Effective Sample Size (mOPESS) is</head><p>where X is the domain of x all L and p(&#8226;|y, &#960;) is the corresponding posterior predictive distribution.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Discussion of the General mOPESS Definition and Computation</head><p>In practice, it makes sense for the posterior predictive distribution in Definition 3.2 to factorise, i.e.,</p><p>as we now explain. A researcher who does not want to use &#960; would not collect many additional samples and then attempt to find the m to minimize the distance between their posterior and q(&#8226;|y). Instead, the researcher would simply collect a fixed number of additional samples r = mn (fixed in the sense that no minimization of the distance to q(&#8226;|y) is performed). Therefore the correct question to ask when trying to estimate the OPESS of &#960; is as follows: if there are multiple independent researchers each of whom chooses a different value of r, then whose inference will most closely agree with our inference? This is why, for the mOPESS to correspond to normal scientific procedure, the hypothesized future samples x (m) (and x(m) ) need to be conditionally independent across values of m, given &#952;. We avoid assuming that xL = x L for essentially the same reason: for the interpretation of the mOPESS to correspond to normal scientific procedure, we cannot assume that each individual researcher computes both q b (&#8226;|x (m) ) and q(&#8226;|x (m) ) and then decides which to use depending on whether D(m) or D(m) is smaller.</p><p>We instead assume there are two researchers in the population for each value of m, one who computes q b (&#8226;|x (m) ) and one who computes q(&#8226;|x (m) ), and since the researchers will likely have different laboratories it is natural to assume that x (m) and x(m) are independent. Unconditionally, all the future samples are dependent, which corresponds to the real-world in that all additional samples collected would be generated using the same underlying, but unknown, value of &#952;.</p><p>In this paper we set the discrepancy measure D to be the 2-Wasserstein distance, and from hereon replace "D" by "W 2 " in our notation. For p &#8805; 1, let u and v be probability measures defined on M with finite p th moment. The p-Wasserstein distance between u and v is defined as</p><p>where &#915;(u, v) denotes the set of measures on M&#215;M with marginals u and v respectively, and d is a metric on M. Conveniently, in the case of multivariate Gaussian distributions the 2-Wasserstein distance can be computed in closed form, and more generally there are efficient software packages for approximating it given samples from the two distributions at hand, e.g., <ref type="bibr">Schuhmacher et al. (2019)</ref>. On the other hand, our framework is general and other measures of posterior discrepancy could also be used, e.g., Kullback-Leibler Algorithm 1: General procedure for computing the mOPESS.</p><p>Step 1: Compute q b n &#8801; q b (&#8226;|y) and q n &#8801; q(&#8226;|y).</p><p>Step 2: For j = 1, . . Step 3: Report the (estimated) mOPESS:</p><p>(KL) divergence or mean squared error as adopted by <ref type="bibr">Clarke (1996)</ref> and <ref type="bibr">Reimherr et al. (2014)</ref>, respectively. General f -divergences <ref type="bibr">(Ali and Silvey, 1966;</ref><ref type="bibr">Sason and Verd&#250;, 2016)</ref>, of which KL divergence is a special case, provide further options.</p><p>Algorithm 1 summarizes how to estimate the mOPESS in practice. In Step 2, S denotes the number of realizations of the OPESS simulated. The procedure is widely applicable and can be implemented for a large family of models beyond the specific cases considered in this paper. Naturally, we use analytical forms of the posterior distributions and the Wasserstein distances when available; otherwise, we use approximation strategies such as importance sampling or Markov chain Monte Carlo (MCMC) methods (e.g., <ref type="bibr">Marin and Robert, 2007;</ref><ref type="bibr">Liu, 2008;</ref><ref type="bibr">Brooks et al., 2011)</ref>. For convenience, in Algorithm 1 and the remainder of this paper, we use mOPESS to refer to the estimate M n = 1 </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">Implementation and Discussion for the Gaussian Example</head><p>Algorithm 2 provides a detailed version of Algorithm 1 for the case of the Gaussian example introduced in Section 2.3. We applied Algorithm 2 to 300 simulations of &#563;n , and set S = 10,000 in Step 2, i.e., for each value of &#563;n , the value of M n was computed via 10,000 Monte Carlo samples of x all L . The left panel of Figure <ref type="figure">2</ref> shows the posteriors q n and q b n (dashed lines) for a single example observed value of &#563;n , with n = 20. The priors &#960; and &#960; b are also plotted (solid lines). The right panel of Figure <ref type="figure">2</ref> shows the distribution of the mOPESS, i.e., of M n , across 300 datasets. Interestingly, in the current context M n is quite variable and is always higher than the nominal EPSS of 10 (vertical line). The nominal EPSS is 10 because of the following three information based analogies between the prior and data: (i) if n = 10 then the Fisher information is n/&#963; 2 = 1/&#955; 2 0 = 10, (ii) if n = 10 then &#563;n &#8764; &#960;, and (iii) for any n, the posterior distribution is N (0, &#963; 2 /(n + 10)).</p><p>In Section 4.1 we illustrate that there is a good explanation for this disagreement with the nominal EPSS: classical information measures consider the prior in isolation Algorithm 2: mOPESS computation for Gaussian example.</p><p>Step 1: Compute the initial posterior distributions</p><p>where</p><p>Step 2: For j = 1, . .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>. S:</head><p>Part a: Generate extra samples x all L = x L &#8746; xL by drawing &#956; * &#8764; q n and then drawing x all</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Part b: For</head><p>Part c: For m = n, . . . , L, compute the distances</p><p>n given by (3.1).</p><p>Step 3: Report the (estimated) mOPESS:</p><p>and can typically only correspond to the prior impact if there is no data. As soon as some data are collected there is always some disagreement between the prior and the data and therefore M n is usually greater than the nominal EPSS, at least in the current Gaussian conjugate model example. On the other hand, in the right panel of Figure <ref type="figure">1</ref>, the 5% quantile of M n (x all L ) (dash-dot curve) shows that there often exist Figure <ref type="figure">2</ref>: (Left) posteriors q n and q b n (dashed lines) and priors &#960; and &#960; b (solid lines) for a single simulated dataset y. (Right) distribution of the mOPESS (M n ) across 300 simulated datasets. The vertical line shows the nominal EPSS of 10. some realizations of x all L such that the value of the OPESS M n is less than the nominal EPSS. Indeed, q b m may by chance be closest to q n after mn &lt; 10 additional samples.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Method Justification and Theory for the Gaussian Example</head><p>In this section we focus on the Gaussian conjugate model introduced in Sections 2.3 and 3.3, for which we derive theoretical results to justify our proposed method. These results can be seen as general in the sense that the Bernstein von Mises Theorem ensures that the posterior distribution is asymptotically Gaussian under mild conditions, see <ref type="bibr">Van der Vaart (2000)</ref> and references therein. On the other hand, the main role of these results is to provide a foundation for understanding and developing the mOPESS, and we emphasize that our method does not rely on asymptotic posterior normality or consistency of the MLE (Maximum Likelihood Estimate). Indeed, our approach is distinguished from many existing methods in that it is designed for small or moderate sample size settings, in which the prior can substantially impact the inference.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Justification of Sampling Distribution for Extra Observations</head><p>Let r = mn and denote the additional samples collected by s </p><p>Recall that in our approach described in Section 3.1 the future samples are drawn from the posterior predictive distribution (3.3)-(3.4) under &#960;, meaning that</p><p>In contrast to our approach, <ref type="bibr">Morita et al. (2008)</ref> sample from the distribution of hypothetical previous data and <ref type="bibr">Reimherr et al. (2014)</ref> bootstrap the observed data. Our proposed sampling method is therefore not the only option and in order to provide justification for our choice it is instructive to consider the behavior of M n under several sampling methods. To investigate this, Proposition 4.1 below considers the case where the empirical mean of the future samples sr is exactly equal to its theoretical mean, denoted &#947;, e.g., &#947; = &#956; n (the posterior mean) under our approach of sampling the future observations from the posterior predictive distribution. If the behavior of M n for sr = &#947; does not make sense then there is little hope that the corresponding sampling method is useful, and if it does make sense then the investigation may offer valuable insights. The proof of Proposition 4.1 is given in Appendix A by <ref type="bibr">Jones et al. (2021)</ref>. Result (a) of Proposition 4.1 corresponds to our proposed method of sampling the future samples from the posterior predictive distribution (3.3)-(3.4). The result is consistent with the top right panel of Figure <ref type="figure">1</ref> in Section 2.3, which shows that the median (dashed curve) value of M n is always equal to or greater than z. To gain further intuition consider the distance W 2 (m) under the conditions of Proposition 4.1:</p><p>Inspecting <ref type="bibr">(4.4)</ref> reveals that the second term captures the nominal EPSS: setting m = n + z makes the standard deviation of the baseline posterior &#960; b m match that of the conjugate posterior q n , so the second term of (4.4) equals zero, and thus giving z extra samples to the baseline prior minimizes the second term. However, the first term of (4.4) reveals that there is an intuitive reason for the value of M n to often be larger than z: disagreements between the prior and the data as captured by (&#956; n&#563;n ) 2 = (z/(z + n)) 2 (&#956; 0&#563;n ) 2 mean that the two posteriors will not be centered in the same location, and the (n/m) 2 term in <ref type="bibr">(4.4)</ref> suggests that greater agreement is expected to be obtained by adding further samples to the baseline prior, i.e., by increasing m. Thus, our definition of M n correctly identifies that simply reporting the classical information content of the prior as determined by its standard deviation is not sufficient: we must also take into account the impact of the prior location relative to the data. Of course, the results in Proposition 4.1 also take account of W 2 (m), see Appendix A by <ref type="bibr">Jones et al. (2021)</ref> for details.</p><p>Bayesian methodology stipulates that the extra samples must be drawn from the posterior predictive distribution, as above, but results (b) and (c) of Proposition 4.1 additionally reveal that several other natural approaches do not work well. Result (b) supposes that &#947; = &#563;n which corresponds to the case where the future observations are sampled from the empirical distribution (bootstrap) or from the baseline posterior predictive distribution, i.e., the posterior predictive distribution under &#960; b and conditional on only y. The first part of result (b), M n = z for &#563;n &#8776; &#956; n , is similar to what is seen in Figure <ref type="figure">1</ref>. However, the second part of result (b), M n &lt; 0 for large |&#563; n&#956; n |, does not make sense in the current scenario firstly because z &gt; 0 secondly because intuitively the prior impact is large when &#563;n is far from &#956; n . <ref type="bibr">Reimherr et al. (2014)</ref> avoided this problem by defining the EPSS so that negative values convey a disagreement between the prior and the data, but there are limitations of their approach as discussed in Section 2.2. Furthermore, there is always some disagreement between the data and the prior so we find it conceptually more appealing to always have positive prior impact (unless our prior is less informative than the baseline).</p><p>Result (c) of Proposition 4.1 corresponds to the case where the additional samples are drawn from the conjugate prior distribution &#960;: just as some might argue that the future data sampling method should not be "contaminated" by the prior, others may argue that it should not be "contaminated" by the data! Under the scenario of the proposition, this sampling scheme yields M n = z, which is at least never negative. However, simply recovering the nominal EPSS regardless of the magnitude of |&#956; 0&#563;n | does not convey differences in the impact of the prior, which is the purpose of having a measure of prior impact. In summary, drawing the extra samples from the posterior predictive distribution (3.3)-(3.4) seems to yield the most reasonable behavior.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2">Theoretical Posterior Distribution of the OPESS</head><p>To study the variation in the OPESS for a given observed dataset, we now derive the theoretical distribution of the OPESS conditional on y for the Gaussian conjugate posterior example discussed in Sections 2.3 and 3.3. More generally, the distribution of the OPESS will typically be hard to derive, but it can be empirically approximated, see Algorithm 1 Step 2. Lemma 4.1 below gives the distribution of the distances W 2 (m) and W 2 (m) conditional on &#563;n and &#956; (drawn from q n , given by (3.7) in Algorithm 2). The proof is given in Appendix B by <ref type="bibr">Jones et al. (2021)</ref>. We condition on both &#563;n and &#956; because then the two distances are independent which facilitates derivation of the OPESS distribution. The distance distributions conditional on only &#563;n are given in Appendix D by <ref type="bibr">Jones et al. (2021)</ref>. Lemma 4.1 states that both distances follow shifted non-central &#967; 2 distributions, whose non-centrality parameters depend on &#956; through &#955; m and &#948; m (given in the lemma statement). &#8764; N(&#956;, &#963; 2 ), for i = n + 1, . . . , m. Then we have</p><p>where</p><p>and</p><p>where</p><p>conditional on &#563;n and &#956;, W 2 q b m , q n and W 2 q m , q b n are independent.</p><p>Theorem 4.1 below gives the posterior distribution of the OPESS conditional on &#563;n . The proof is given in Appendix C by <ref type="bibr">Jones et al. (2021)</ref>. In the theorem statement, v denotes a possible value of M n (e.g., in the notation P (M n = v|&#563; n )), and t is a dummy variable for the distance corresponding to M n = v, i.e., the distance W 2 (n + v), if v &#8805; 0, and the distance W 2 (n + |v|), otherwise. The result gives a separate expression for the case M n = 0 because when m = n the distance between the posteriors (i.e., W 2 (n)) is not random, meaning that an integral over the distance dummy variable t is not required. In the case v &#8712; Z/{0}, the integrands specified are tractable because the products are truncated at M (t) and M (t) (defined in Appendix C by <ref type="bibr">Jones et al. (2021)</ref>), which are finite for all values of t &#8804; &#963; 2 /(n + z). m , q n and W 2 (m) = W 2 q m , q b n , respectively, as given in <ref type="bibr">Lemma 4.1. Lastly,</ref><ref type="bibr">let g(t,</ref><ref type="bibr">&#956;,</ref><ref type="bibr">v</ref>, M, M ) denote the function that gives P (min</p><p>where t m = (tc 2 m )/&#964; m and tm = (t -c2 m )/&#954; m , and M and M are known functions (see Appendix C by <ref type="bibr">Jones et al. (2021)</ref>). Then</p><p>m otherwise, and M (t) and M (t) are finite integers for all values of t &#8804; &#963; 2 /(n + z).</p><p>Figure <ref type="figure">3</ref> shows two examples of the conditional posterior distribution of the OPESS given &#563;n and &#956;, where &#563;n = &#956; = 0 in the left panel and &#563;n = &#956; = 0.45 = 2&#963;/ &#8730; n in the right panel. We plot the conditional posterior distribution to gain intuition about how the particular draw of &#956; from &#960; impacts the conditional distribution of the OPESS. This is important because in reality the value of &#956; is fixed but unknown, and it is therefore valuable to understand how the distribution of the OPESS changes when we simulate the future samples based on different fixed choices of &#956;. In Figure <ref type="figure">3</ref>, the linedot density is a close Monte Carlo approximation to the theoretical conditional density of the OPESS given &#563;n and &#956;, and was obtained by simulating from the theoretical conditional distributions of W 2 (m) and W 2 (m) and averaging the resulting values of the integrand given in Theorem 4.1 (except in the case P (M n = 0) for which no Monte Carlo approximation is needed). The red crosses show the empirical distribution of the OPESS obtained by directly applying the first two steps of Algorithm 1, except with the modification that &#956; is fixed in Step 2(a). Figure <ref type="figure">3</ref> illustrates that for &#563;n (and &#956;) farther from the prior mean &#956; 0 = 0 (right panel) the conditional OPESS distribution has larger mean and is more right-skewed. This corroborates the numerical results seen in Figure <ref type="figure">1</ref>. For some values of &#563;n and &#956; the conditional posterior distribution of the OPESS is bi-modal, with one mode at positive values and one at negative values (not shown). For other &#563;n and &#956;, there is a mode at M n = 0, which is prominent in examples where the mOPESS is less than the nominal EPSS. Such scenarios are discussed further in Section 5.2. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Examples Beyond the Gaussian Conjugate Case</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1">Non-Conjugate Gaussian Model</head><p>We now investigate the properties of Algorithm 1 when using a non-conjugate prior distribution for &#956;. The setting is identical to that outlined in Section 2.3, except that we set the informative prior to be a t-distribution, &#960; t (&#956;) &#8801; T (&#956;|&#957;, &#956; 0 , &#955; 2 0 ). The degreesof-freedom parameter, &#957;, controls the heaviness of the tails, and &#956; 0 and &#955; 0 are location and scale parameters, respectively. We set &#956; 0 = 0 and &#955; 2 0 = 0.1, and investigate two choices of &#957;, namely, &#957; = 4, 100. As &#957; &#8594; &#8734; the t-distribution density converges to a Gaussian density, so for large &#957; we expect the relationship between the mOPESS and &#563;n to be similar to that seen in the top left-hand panel of Figure <ref type="figure">2</ref>. However, for relatively small values of &#957;, we expect the relationship between the mOPESS and &#563;n to be different, because in such cases both the posterior variance and the posterior skewness have nonnegligible dependence on &#563;n&#956; 0 (in contrast to the conjugate Gaussian example of Section 2.3).</p><p>The top-left panel of Figure <ref type="figure">4</ref> shows the relationship between the mOPESS and &#563;, for the t-distribution prior &#960; t . As in Section 2.3, the values of &#563; were chosen to be the quantiles (k -0.5)/100, for k = 1, . . . , 100, of its distribution, namely a zero mean Gaussian with variance 1 n . The relationship for &#957; = 100, indicated by a solid red line, mimics the quadratic curve shown in the top-left panel of Figure <ref type="figure">1</ref>, but we see a concave relationship for &#957; = 4, represented by a dashed black line. The latter pattern shows that, for &#957; = 4, the mOPESS is smaller when |&#563; n&#956; 0 | is larger. This result can be explained by the fact that the posterior variance of &#956; increases with |&#563; n&#956; 0 |, see the upper right-hand panel of Figure <ref type="figure">4</ref> (dashed black line). We computed the posterior variance for each (&#957;, &#563;n ) pair with Monte Carlo integration using 5&#215;10 5 posterior draws generated by RStan <ref type="bibr">(Stan Development Team, 2020b,a)</ref>. In the case of &#957; = 100 (solid red line), the posterior variance of &#956; is nearly constant for all values of |&#563; n&#956; 0 |, and To gain further insight, we define the equivalent nominal EPSS of the t-distribution prior by matching posterior variances, i.e., we define it to be the nominal EPSS of the conjugate Gaussian prior &#960; which yields the same posterior variance as under the tdistribution prior &#960; t . The bottom left panel of Figure <ref type="figure">4</ref> shows that when &#957; = 4 and &#563;n = -0.6, the equivalent nominal EPSS is 4 (dashed black line), whereas for &#563;n near 0 it is 10. In contrast, the equivalent nominal EPSS of a t-distribution prior with &#957; = 100 (solid red line) is 10 for all &#563;n &#8712; [-0.6, 0.6], as in the case of a conjugate Gaussian prior with z = 10 (shown as a dotted blue line overlapping with the solid red line).</p><p>The pattern described above suggests that subtracting the equivalent nominal EPSS from the mOPESS might shed more light on the comparison between the cases &#957; = 4 and &#957; = 100. We call the quantity resulting from this subtraction the excess mOPESS. The bottom right-hand panel in Figure <ref type="figure">4</ref> shows the excess mOPESS, and reveals that the excess increases with |&#563; n&#956; 0 |. Indeed, the relationship is now seen to be what we may have initially expected: a t-distribution prior with &#957; = 4 (dashed black line) has a qualitatively similar, but smaller, impact than a t-distribution prior with &#957; = 100 (solid red line).</p><p>A third factor that is unexplored in the tetraptych above is the posterior skewness of &#956; under the two values of &#957;. As |&#563; n&#956; 0 | increases, so too does the skewness of the posterior for &#956;. Furthermore, the sensitivity of the posterior skewness to the quantity &#563;n&#956; 0 is a function of the degrees of freedom parameter &#957;. Skewness of the posterior also impacts the OPESS. If we generate a &#956; from the tail of &#960; t n , then M n is nearly certain to be negative. For example, suppose that &#563;n &lt; &#956; 0 . Then our posterior &#960; t n will be rightskewed. Consequently, Algorithm 1 will tend to generate more posterior draws from the region &#956; E[&#956;|&#563; n , &#960; t ] than from the region &#956; E[&#956;|&#563; n , &#960; t ]. In the former region, P (W 2 (m) &lt; W 2 (m)|&#956;, &#563;n , &#960; t ) is small because the extra samples generated from the predictive distribution will decrease W 2 (q m , q b n ), but will increase W 2 (q b m , q n ), leading to a negative value of M n . Because the magnitude of the skewness decreases with |&#563; n&#956; 0 |, smaller values of |&#563; n&#956; 0 | yield fewer negative M n values.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2">Beta-Binomial Model</head><p>Suppose {y i , 1 &#8804; i &#8804; n} are independent observations taking values in the set {0, 1}. The unknown parameter &#952; is the probability that y i = 1. We set the informative prior to be &#960;(&#952;) &#8801; Beta(&#945;, &#946;), where &#945;, &#946; are known hyperparameters, and the baseline prior to be</p><p>&#8764; Bernoulli(&#952;) for i = n+1, . . . , m (but since &#952; is unknown it is drawn from its posterior distribution when computing the mOPESS, see Algorithm 1 Step 2(a)). Let q A n and q B m denote the posterior distribution using the original data y = (y 1 , . . . , y n ) T and the expanded dataset x (m) , respectively, under priors A and B. Also define (F A n ) -1 and (F B m ) -1 to be the quantile functions associated with these posterior densities. Then it can be shown that the 2-Wasserstein distance between q A n and q B m is</p><p>, see</p><p>Theorem 2 in <ref type="bibr">Cambanis et al. (1976)</ref>. Unfortunately, this distance cannot be expressed in closed form in the case of Beta distributions, but it can be approximated to high precision using numerical integration, which is the approach we take.</p><p>In our simulations we set &#945; = &#946; = 5, which corresponds to a nominal EPSS of &#945; + &#946; -2 = 8. The subtraction of 2 highlights that the standard nominal EPSS is relative to the prior sample size of the flat prior Beta(1, 1), which is also our baseline prior &#960; b . To investigate the prior impact across different datasets, we sample 1,000 datasets of size n = 20 with replacement from the sex ratio dataset presented in Section 2.4 of <ref type="bibr">Gelman et al. (2013)</ref>. The sex ratio dataset consists of the biological sexes of 980 babies born to mothers with placenta previa: 437 of the babies are female, a proportion of 0.446 of the total.</p><p>The top left panel of Figure <ref type="figure">5</ref> shows the mOPESS values obtained across the 1,000 simulated datasets. The M n estimates have a similar pattern as in the Gaussian conjugate model of Section 2.3, except that M n is less than the nominal EPSS for datasets with &#563;n = 0.5. The top right panel of Figure <ref type="figure">5</ref> shows that the median of the posterior distribution of M n is similar to the mean. It also shows that the posterior distribution is much wider for datasets with means near 0.1. This is due to the fact that the posterior distribution of &#952; is strongly right skewed when &#563;n is much less than the prior mean of Figure <ref type="figure">5</ref>: (Top left) mOPESS as a function of &#563;n . Each point corresponds to one of the 1,000 simulated datasets. (Top right) Quantiles of the posterior distribution of M n as a function of &#563;n including the median (dashed black curve), 95% quantile (long dash green curve), and 5% quantile (dash-dot blue curve). The horizontal solid line shows the nominal EPSS of 8. (Bottom left and right) Comparison of the posteriors q n and q b n and their respective priors in the case where &#563;n = 0.1 and &#563;n = 0.5, indicated in the top left panel by a green "+" and cross, respectively. 0.5. The skewness is in turn reflected in the future observations x all L , which are simulated conditional on a draw of &#952;, and this results in large posterior uncertainty for M n . In particular, the right-skewness of q b n means that a large number of future samples can be needed to move the distribution to the right, resulting in some large OPESS values, and the right-skewness of q n means that future samples to the left can substantially shift it, which results in some negative OPESS values. That posterior skewness is reflected in the resulting OPESS distribution is a strength of our approach. Indeed, the large spread in the posterior distribution of M n corresponds to genuine uncertainty about the number of extra samples that need to be collected in order to minimize the distance between q n and q n (or q m and q b n ). Next, we examine the relationship between the mOPESS and the nominal EPSS in two cases, namely, those where the prior mean and &#563;n are highly discrepant and perfectly aligned, respectively. The top left panel of Figure <ref type="figure">5</ref> indicates a simulated dataset for which &#563;n = 0.1 (green "+"), and the bottom left panel shows the corresponding initial posterior distributions q n and q b n (as well as the prior distributions). In this case, the mOPESS value is high (around 12) because the initial posteriors are very different. The bottom right panel of Figure <ref type="figure">5</ref> shows an analogous plot for a dataset with &#563;n = 0.5, i.e., that labeled with green cross in the top left panel. In this case, the initial posteriors are very similar, which is why the mOPESS value is low (approximately 7.5, the variation across datasets is due to Monte Carlo error). In particular, the mOPESS value is less than the nominal EPSS of 8, a phenomenon that did not occur in the Gaussian conjugate model example of Section 2.3 for any of the datasets considered.</p><p>The low mOPESS value occurs here due to circumstances that arise, in this case, due to the discreteness of the data and the future data, as we now explain. The data mean &#563;n exactly matches the mean of the prior &#960;, which in turn makes the means of q n and q b n exactly equal. Thus, q n and q b n are very similar to begin with, and it is unclear whether adding more samples to one of these posteriors will further reduce the distance between them. Adding more samples to the baseline posterior q b n could reduce the width of the distribution, and therefore may lead to greater agreement with q n . However, the discrete nature of the data means that one additional sample with value one or zero will necessarily move the posterior mean away from 0.5, therefore potentially increasing the 2-Wasserstein distance between the two posteriors. Of course, if we draw an even number of extra samples then their average may be close to 0.5, so a reduction in the width of the baseline posterior q b n may be achieved without any substantial change in the mean. However, based on the nominal EPSS value, the approximate number of extra samples needed for matching the posterior widths is 8, but the probability of achieving an average of 0.5 (or very close to this) when drawing around 8 samples is not sufficiently high, and consequently the distance W (n) is often smaller than W 2 (m) and W 2 (m) for all m &gt; n. Thus, for many simulations of x all L , we have M n = 0, meaning that the mOPESS M n is shrunk towards zero.</p><p>In summary, the mOPESS may be less than the nominal EPSS when the means of the initial posteriors q n and q b n are very similar relative to the size of z. In particular, adding extra samples to one posterior may not reduce the 2-Wasserstein distance between the two posteriors because: (i) if few extra samples are added then the variability in their mean can introduce discrepancies between the posterior means, and (ii) adding many extra samples will introduce discrepancies in the spreads since the initial discrepancy will be over-corrected. Thus, the smallest distance between the posteriors may often be achieved when M n = 0, in which case M n will be small in magnitude. Note that this phenomenon can occur in the Gaussian conjugate model example if n z (whereas in Section 2.3 we set n = 2z).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3">Simple Linear Regression Model</head><p>We now consider the setting of a simple linear regression model:</p><p>, &#969; i &#8764; N (0, 1), <ref type="bibr">(5.1)</ref> for i = 1, . . . , n, where &#963; 2 is known, and &#946; = (&#946; 1 , &#946; 2 ) are the unknown model parameters. We note that in simple linear regression models, distributional assumptions on covariates are typically not made. We assume that the distribution of the covariates &#969; i is known in order to simplify the algorithm for generating hypothetical samples. Let right panel in Figure <ref type="figure">7</ref>, which shows the mOPESS versus ( &#946;n ) [2]&#947; 0 , we see that the maximum mOPESS occurs at the maximum observed value of ( &#946;n ) [2]&#947; 0 (because the maximum value of this discrepancy is large compared to that for ( &#946;n ) [1]&#956; 0 ). In summary, the mOPESS generalizes to two dimensions as we would expect it to, with joint dependence on &#946;n&#951; 0 .</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="6" xml:id="foot_0"><p>SummaryIn this paper, we have proposed the mean Observed Prior Effective Sample Size (mOPESS) as a measure of the impact of the prior distribution on the Bayesian analysis at hand. Our measure is different from other methods proposed in literature in that we condition on the observed data, instead of averaging over the data. Furthermore, we do not rely on asymptotic results, meaning that our method can be applied to small or moderate sample size settings, where the prior impact is largest and of most interest in practice.</p></note>
		</body>
		</text>
</TEI>
