<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>&lt;i&gt;Colloquium&lt;/i&gt; : Gene expression in growing cells: A biophysical primer</title></titleStmt>
			<publicationStmt>
				<publisher>Reviews of Modern Physics</publisher>
				<date>10/01/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10556527</idno>
					<idno type="doi">10.1103/RevModPhys.96.041001</idno>
					<title level='j'>Reviews of Modern Physics</title>
<idno>0034-6861</idno>
<biblScope unit="volume">96</biblScope>
<biblScope unit="issue">4</biblScope>					

					<author>Ido Golding</author><author>Ariel Amir</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Cell growth and gene expression, two essential elements of all living systems, have long been the focus of biophysical interrogation. Advances in experimental single-cell methods have invigorated theoretical studies into these processes. However, until recently, there was little dialog between the two areas of study. In particular, most theoretical models for gene regulation assumed gene activity to be oblivious to the progression of the cell cycle between birth and division. But, in fact, there are numerous ways in which the periodic character of all cellular observables can modulate gene expression. The molecular factors required for transcription and translation-RNA polymerase, transcription factors, ribosomes-increase in number during the cell cycle, but are also diluted due to the continuous increase in cell volume. The replication of the genome changes the dosage of those same cellular players but also provides competing targets for regulatory binding. Finally, cell division reduces their number again, and so forth. Stochasticity is inherent to all these biological processes, manifested in fluctuations in the synthesis and degradation of new cellular components as well as the random partitioning of molecules at each cell division event. The notion of gene expression as stationary is thus hard to justify. In this review, we survey the emerging paradigm of cell-cycle coupled gene expression, with an emphasis on the global expression patterns rather than gene-specific regulation. We discuss recent experimental reports where cell growth and gene expression were simultaneously measured in individual cells, providing first glimpses into the coupling between the two, and motivating several questions. How do the levels of gene expression products -mRNA and protein -scale with the cell volume and cell-cycle progression? What are the molecular origins of the observed scaling laws, and when do they break down to yield non-canonical behavior? What are the consequences of cell-cycle dependence for the heterogeneity ("noise") in gene expression within a cell population? While the experimental findings, not surprisingly, differ among genes, organisms, and environmental conditions, several theoretical models have emerged that attempt to reconcile these differences and form a unifying framework for understanding gene expression in growing cells.
CONTENTSPrologue: Physical models of cellular processes I. Growth-agnostic models of gene expression A. Gene expression models with constant rates B. Comparison to experimental data and the two-state model for gene expression C. Continuous, non-Markovian models of gene expression II. Growth-dependent models of gene expressions: The effect of DNA replication A. Beyond the static picture B. Replication of the gene of interest C. Cell-cycle dependent transcription: Experimental observations D. Non-monotonic transcription patterns: observations and possible mechanisms III. Growth-dependent models of gene expression: The effects of global transcription and translation A. Correcting for cell growth by normalization by cell volume B. Constraints on the rates of transcription and translation 1. The bacterial "growth law" 2. The rate of protein production 3. The rate of mRNA production C. Solving the model D. Revisiting the assumptions -what limits transcription and translation? E. The limiting resource for transcription and translation: Experimental evidence 1. The scaling of protein levels with time and cell volume 2. The scaling of mRNA levels with time and cell volume 3. Experiments observing the slowdown of exponential growth IV. Outlook A. Extensions and generalizations of the continuous growth model B. Concentration homeostasis in growing cells C. Space: The final frontier]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Prologue: Physical models of cellular processes</head><p>In this Colloquium, we discuss biophysical models for the process of gene expression, and for the way this process is coupled to the progression of the cell cycle (these terms will be elaborated below). Before embarking on this discussion, we should say a few words about the general nature of physical models for cellular processes. This preamble is required, we believe, because those models are somewhat different, in terms of how they are constructed and used, from physical models of inanimate matter, with which some readers are perhaps more familiar. These differences, in turn, reflect the vast gap in knowledge and experimental amenability between the physics of living and non-living systems.</p><p>One category of biophysical models aims for a molecular, or even atomic, level of description. Such models have been extremely successful in elucidating the function of biological molecules <ref type="bibr">(Dill and Bromberg, 2011;</ref><ref type="bibr">Nelson et al., 2020)</ref>. However, even with advances in computational power, these models are typically limited to depicting just one or several molecules, over very short time scales (less than a millisecond). This makes the approach inadequate for capturing all but the simplest processes in the living cell, since these processes-such as the expression of genetic information, discussed below-typically involve numerous molecular players and take place over minutes and hours.</p><p>Making a molecularly detailed model of these processes is thus currently impractical 1 . In truth, even if such "full" cellular models were possible, to many physicists such a model would be unsatisfactory in its complexity -Jorge Luis Borges's story comes to mind, of "a Map of the Empire whose size was that of the Empire" <ref type="bibr">(Borges, 1999)</ref>.</p><p>For these reasons, cellular processes are typically conceptualized using simplified theoretical models, where a small number of molecules and interactions are considered explicitly, while many others are ignored or coarse-grained into other observables <ref type="bibr">(Amir and Balaban, 2018;</ref><ref type="bibr">Bialek, 2012;</ref><ref type="bibr">Bintu et al., 2005)</ref>. One need not be apologetic about the use of such phenomenological models. They tend to be more robust to the model details and lend themselves better to analytical approaches, which in turn can provide deeper insights into the physical principles at play <ref type="bibr">(Amir, 2020;</ref><ref type="bibr">Machta et al., 2013;</ref><ref type="bibr">Phillips, 2005)</ref>.</p><p>In some ways, this approach follows the traditional physics attitude as applied, e.g., to the description of fluids or elastic materials, where countless microscopic constituents are left out <ref type="bibr">(Taylor, 2005)</ref>. This omission is both a necessity-the positions and interactions of all atoms in the material are unknowable to us-and a choice, since it allows us to obtain a simple yet predictive depiction of the system. Biophysical models, too, reflect the combined constraints of ignorance and parsimony. However, much more than in the physics of non-living matter, the level of abstraction in biophysical models-what molecules, processes and interactions are included-reflects our ignorance of the underlying details. This ignorance, in terms of what is known or is even experimentally knowable, is overwhelming to a degree that is unfamiliar, and would perhaps be unacceptable, to modelers of nonliving systems <ref type="bibr">(Golding, 2022)</ref>.</p><p>Even for the best characterized cellular processes, such as the regulation of gene expression in bacteria, a major focus of this Colloquium, what we currently have is a partial list of the molecular players involved and the interactions between them, but little or no knowledge of the biophysical parameters characterizing these interactions, such as the rates of diffusion, binding, and assembly into molecular complexes. Similarly, what we can experimentally measure is, too, highly limited in terms of number of molecular species simultaneously detected (typically, only a few), the precision (typically, relative rather than absolute levels, averaged over many individual cells), and the temporal resolution, which typically under-samples much of the relevant kinetics.</p><p>The simplicity of biophysical models often reflects this limited knowledge regarding the systems under study, rather than an informed choice of which features to include and which ones to leave out <ref type="bibr">(Golding, 2022)</ref>. In other words, the "coarse-graining" process is driven by the need to remain anchored in known facts and make model predictions experimentally testable 2 . Thus, while there are select examples where a simple biophysical model may be argued to reflect an underlying simplicity of behavior, a-la Occam's razor (that of bacterial "growth laws" <ref type="bibr">(Scott et al., 2010)</ref>, relating ribosome levels to growth rates, is discussed later), in most instances model simplicity instead implies that we are ignorant of many details, which are swept under the proverbial "Occam's rug" <ref type="bibr">(Brenner, 1997;</ref><ref type="bibr">Golding, 2011)</ref>.</p><p>In addition to the many "known unknowns", for example, the rate constants of the regulatory interaction under study, even more worrisome are the "unknown unknowns", e.g., the presence of additional unrecognized interactions in the system. Making models more elaborate may make them appear more realistic, but typically only achieves the opposite, since the added details are inevitably less grounded in knowledge. Model elaboration can only be justified as a means to explore specific hypotheses that can be experimentally tested.</p><p>In this Colloquium, we will survey how the constraints discussed above have driven the development of models describing gene expression and its coupling to cell growth (with the accompanying changes in the amount of molecules 1 There have been recent attempts at constructing models for cellular function, which explicitly consider thousands of molecular players <ref type="bibr">(Karr et al., 2012;</ref><ref type="bibr">Thornburg et al., 2022)</ref>. But while these models use statistical inference approaches to address the challenge of parameterization, they leave unresolved the problem of additional unrecognized regulatory interactions in the system. 2 Tellingly, even the term "coarse-graining" possesses a different meaning here compared with statistical or condensed matter physics.</p><p>Whereas in those contexts it refers to a well-defined mathematical procedure, in the physics of living systems it is used more loosely, indicating an attempt to describe the system at some lower resolution without accounting for all of the underlying processes -even if there is no rigorous procedure for identifying the appropriate level of description or the necessary parameters.</p><p>driving gene expression-DNA, RNA polymerase, ribosomes). We will focus primarily on bacterial systems, where the knowledge infrastructure and the experimental tractability are significantly superior to the more complex eukaryotic and multicellular systems. But, as we have emphasized already, even for bacteria our ignorance is-by physics standards-overwhelming, and this ignorance has strong consequences for constructing theoretical models. Throughout, we discuss two classes of models: In some cases, we utilize stochastic models, which enable us to track (analytically or computationally) the relevant random variables, e.g., the discrete copy numbers of a particular protein. In other instances, we write expressions for the mean behavior, which therefore take the form of deterministic differential equations (either ordinary or partial). The latter approach typically approximates copy numbers as continuous variables, and also neglects the fluctuations in the temporal dynamics. This approach can therefore be described as a continuous-deterministic approximation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>I. GROWTH-AGNOSTIC MODELS OF GENE EXPRESSION</head><p>A. Gene expression models with constant rates</p><p>In the process of gene expression, a segment of the cell's genome (the gene) is transcribed repeatedly into a complementary, short lived messenger RNA <ref type="bibr">(mRNA)</ref>. Each mRNA molecule is then translated, again repeatedly, into the protein encoded by that gene <ref type="bibr">(Alberts et al., 2002)</ref>. Since each cell's identity, shape, and function are largely determined by which proteins it contains, gene expression and its regulation can be seen as the prime mover in the living cell 3 .</p><p>In constructing a biophysical model for gene expression, we must first note that the production of even a single protein involves thousands of stochastic molecular events. On the transcription side alone, these events include RNA Polymerase (RNAP) and transcription factors (TFs) searching the whole cell and genome to find the regulatory region of the gene (called promoter) and binding to it; changes in molecular conformation of RNAP that enable the initiation of transcription; and the basepair-by-basepair synthesis of mRNA by RNAP, until the gene termination site is reached. The synthesis of protein from mRNA, and the degradation of both mRNA and proteins, are likewise molecularly elaborate <ref type="bibr">(Alberts et al., 2002)</ref>. Moreover, these different molecular events are regulated by multiple cellular factors, and subject to feedback from downstream steps in the gene expression processes, in ways that are only partly understood <ref type="bibr">(Berry and Pelkmans, 2022)</ref>. Despite this complexity, gene regulation is often modeled using a mere four rates, corresponding to the production and degradation of mRNA and proteins <ref type="bibr">(Nelson et al., 2015;</ref><ref type="bibr">Ozbudak et al., 2002;</ref><ref type="bibr">Paulsson, 2005;</ref><ref type="bibr">Swain et al., 2002)</ref>(Fig. <ref type="figure">1</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>FIG. 1 A minimal stochastic model for gene expression.</head><p>3 The preceding statement, as any in biology, has important exceptions, such as the cell's use of post-translational protein modifications to encode additional information <ref type="bibr">(Alberts et al., 2002)</ref>. This subject is outside the scope of the current Colloquium.</p><p>This stochastic model can be succinctly summarized as:</p><p>Here, M and P denote the mRNA and protein copy numbers, respectively. k m is the transcription rate, k p the translation rate (per mRNA), and &#947; m , &#947; p the degradation rates for mRNAs and proteins, respectively 4 .</p><p>As can be surmised from the earlier discussion, coarse graining the molecular complexity of gene expression into this simple model does not so much reflect informed choices as an intuitive attempt at parsimony, whose legitimacy depended on the limited resolution of experimental data available until about two decades ago. Indeed, once experimental methods improved to allow measuring gene expression at finer resolution, the inadequacy of this simple model was revealed, leading to necessary modifications, as we shall discuss in Section I.B.</p><p>But first, let us examine the model in some detail. We will first consider the ensemble means of the observables, for which we may write down a set of readily solvable ODEs:</p><p>The "ensemble" here can be interpreted as consisting of the individual cells within a population. Since traditional biochemical methods for measuring mRNA and protein levels are performed in bulk, using millions of cells to obtain a single reading (Fig. <ref type="figure">2</ref>), the ensemble average is the natural observable to be calculated. Performing these experimental measurements, one finds that Eqs. ( <ref type="formula">5</ref>)-( <ref type="formula">6</ref>) neatly capture mRNA and protein kinetics during gene induction, i.e., when the gene is turned "on" (Fig. <ref type="figure">2</ref>). In other words, the standard model for gene expression appears to be consistent with the experimental data.</p><p>More recently, however, it has become possible to measure mRNA and protein numbers in an individual cell <ref type="bibr">(Cai et al., 2006;</ref><ref type="bibr">Skinner et al., 2013;</ref><ref type="bibr">Taniguchi et al., 2010;</ref><ref type="bibr">Yu et al., 2006)</ref> (Fig. <ref type="figure">3</ref>), thus allowing us to go beyond the population mean and examine the copy number distribution. Characterizing the statistics -rather than the mean alone -of expression level is significant for two reasons. First, cell-to-cell differences in protein levels may result in variations in phenotype, such that genetically identical cells, within a uniform environment, diverge in their behavior.</p><p>This cellular individuality plays a crucial role throughout biology, from the emergence of antibiotic persistence among bacteria, to cell differentiation in the early mammalian embryo, and numerous other examples <ref type="bibr">(Bal&#225;zsi et al., 2011;</ref><ref type="bibr">Eldar and Elowitz, 2010)</ref>. Thus, describing gene expression in individual cells is arguably more important than capturing it in the hypothetical "average cell".</p><p>4 A note on notation: Throughout the Colloquium, the distinction between molecule copy number and its concentration will be important.</p><p>We will use M , P to denote mRNA and protein copy numbers, respectively, and m, p for their concentrations.</p><p>Table I summarizes the notations used in the manuscript But there is also a second, biophysical reason to examine single-cell expression, and that is to provide stronger empirical challenge to the theoretical picture we presented in Fig. <ref type="figure">1</ref>. To do so, we will interrogate the model further by considering the stochastic fluctuations associated with it, and derive theoretical results regarding the copy-number statistics, which can then be compared to experimental data. As we shortly see, this exercise will prove insightful.</p><p>We begin by finding the steady-state distribution of mRNA copy number. To this end, we can write the master equation for the temporal dynamics of this distribution, P M (n, t) (the probability to have precisely n copies of mRNA as time t)<ref type="foot">foot_0</ref> :</p><p>Here, the first and second terms on the RHS correspond, respectively, to the incoming fluxes from production of an mRNA when there were n -1 copies in the cell, or degradation of a molecule when there were n + 1 copies in the cell. The last term on the RHS corresponds to degradation and production events that change n copies in a cell to n -1 and n + 1, respectively.</p><p>At steady-state, dP M (n,t) dt = 0, and we can therefore write:</p><p>What should we expect the solution for this equation to look like? If every mRNA molecule did not decay stochastically, but instead lived for precisely a time 1/&#947; m , then the number of mRNA molecules within the cell at a given time would equal the total number of molecules produced within a time-window 1/&#947; m -which is, of course, Poisson distributed with parameter k m /&#947; m . One can verify by direct substitution that this is also the exact solution to Eq. ( <ref type="formula">8</ref>). In fact, if the initial mRNA copy number is Poisson distributed, it can be readily shown that the distribution will remain</p><p>Poissonian at all times. To see this, we may substitute into Eq. ( <ref type="formula">7</ref>) a form that is Poisson distributed at any time, albeit with a time-dependent parameter &#951;(t), i.e. P M (n, t) = e -&#951;(t) &#951;(t) n /n!. This leads to the following solvable ODE for &#951;(t):</p><p>Next, we turn to the steady-state distribution of protein numbers, which is a little trickier to handle. Before delving into the equations, let us consider the stochastic kinetics of mRNA and protein numbers. Once an mRNA molecule is transcribed (a process occurring with rate k m ), proteins will start being produced at a rate k p . This will happen until the mRNA is degraded, the timing of which is exponentially distributed (and is typically much shorter than protein lifetime, hence we can ignore protein degradation for the purpose of this calculation). We will refer to the event where multiple proteins are translated from a single mRNA copy as a burst. Since protein production from each mRNA occurs at a constant rate, and since the mRNA lifetime distribution is exponential, the distribution &#957;(b) of the number of proteins b produced in a single burst will be given by:</p><p>with P oiss(x, &#955;) the Poisson distribution with parameter &#955;. This integral can be readily evaluated, resulting in a geometric distribution for the protein copy number produced within a burst:</p><p>The average burst size, b, is found to be, as expected:</p><p>To proceed and find the protein copy number distribution, it will be convenient to work with a continuous (i.e., Fokker-Planck) equation <ref type="bibr">(Friedman et al., 2006)</ref> rather than a discrete one, as we did in the case of mRNA above.</p><p>Clearly, for the existence of a stationary protein distribution we will need to consider a finite protein degradation rate (which was inconsequential for the previous calculation) -otherwise proteins will accumulate indefinitely. We will also assume that b &#8811; 1 such that the burst size distribution may be approximated as continuous. The continuous approach will be inaccurate when the protein level in a cell is low, but in practice these levels are often sufficiently high to make the results we will obtain a useful approximation (we will comment on the exact solution of the discrete equations shortly). The Fokker-Planck equation reads:</p><p>where x is the (now continuous) protein copy-number and the probability distribution &#957;(b) is given by Eq. ( <ref type="formula">11</ref>).</p><p>Equating the time derivative to zero leads us to an equation for the the steady state. One may verify by direct substitution that the (normalized) solution is given by the gamma distribution <ref type="bibr">(Friedman et al., 2006)</ref>):</p><p>x a-1 e -x/ b, (</p><p>with a = k m /&#947; p , assumed to be a large number, and thus &#915;(a) &#8776; [a -1]!, with [] indicating the nearest integer.</p><p>In fact, this form may be intuited by noting that a sum of n independent variables, each drawn from an exponential distribution, is gamma distributed with a shape parameter n (see chapter 6 of <ref type="bibr">(Amir, 2020)</ref>). In our case, each exponentially distributed variable corresponds to a single burst of proteins. Since the protein lifetime is 1/&#947; p , we expect n to equal the number of bursts in this time window, namely n = k m /&#947; p = a -which turns out to be the precise result<ref type="foot">foot_1</ref> . We note that the discrete case can also be solved, and in the limit of short mRNA lifetime yields a negative binomial distribution <ref type="bibr">(Shahrezaei and Swain, 2008)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Comparison to experimental data and the two-state model for gene expression</head><p>Now that we have found theoretical predictions for the distributions of mRNA and protein copy numbers, we turn to the experimental data <ref type="bibr">(Golding et al., 2005;</ref><ref type="bibr">So et al., 2011)</ref>. As can be seen in Fig. <ref type="figure">3</ref>, the measured mRNA distribution is, alas, very poorly fit by the Poisson distribution we predicted above. Intriguingly, the mRNA data is found to be well-fitted by a gamma (or a negative binomial) distribution -which was the (approximate) result we expected for the protein number distribution. What do we learn from this conundrum?</p><p>The insight lies in realizing that the gamma distribution arose from a model where one molecular species follows a birth-death process (i.e., it is produced and decays at constant rates) and a second species is made at a rate proportional to the copy number of the first one. We may, effectively, obtain the same result for the mRNA distribution if we postulate that the gene from which mRNA is produced can be either "on" or "off", and that the switching between the two states occurs at constant rates k ON and k OFF -see Fig. <ref type="figure">4</ref>. This model, commonly referred to as the "twostate" (or "telegraph") model <ref type="bibr">(Paulsson, 2005)</ref>, will produce bursts of mRNA, that -since the stochastic dynamics is formally identical to that of protein production in the simpler, one-state, model analyzed earlier -will be exponentially distributed. A Fokker-Planck equation, analogous to Eq. ( <ref type="formula">13</ref>), can then be set up for the steady-state mRNA copy number distribution, leading to the gamma distribution (or, if we treat mRNA numbers as discrete rather than continuous, a negative binomial distribution).</p><p>FIG. <ref type="figure">4</ref> The two-state model for stochastic gene expression.</p><p>The two-state model for transcription is able to capture mRNA statistics both at steady-state and during gene induction, and was further validated by following the stochastic kinetics of mRNA production in live cells (Fig. <ref type="figure">5</ref>), which exhibits the exponentially distributed transcription "bursts" predicted by the model <ref type="bibr">(Golding et al., 2005)</ref>.</p><p>Note that for the ensemble-average behavior, once we coarse-grain over the timescale of gene switching, the model reduces to the naive model we started with -and thus the agreement with the bulk experiments is retained <ref type="bibr">(Golding et al., 2005)</ref>.</p><p>Beyond bacteria, the two-state model has been shown to reproduce mRNA statistics in higher organisms, from yeast to mammalian tissues <ref type="bibr">(Sanchez and Golding, 2013;</ref><ref type="bibr">Skinner et al., 2016)</ref>. Notwithstanding this success, the mechanistic basis of gene on/off switching is still debated <ref type="bibr">(Jones and Elf, 2018;</ref><ref type="bibr">Sanchez and Golding, 2013)</ref>. In bacteria, both the kinetics of transcription-factor binding/unbinding <ref type="bibr">(Jones et al., 2014)</ref> and temporal changes in DNA supercoiling <ref type="bibr">(Sevier and Levine, 2017)</ref> have been put forward as plausible hypotheses. However, a caveat is in place: The success of the two-state model may be overstated. Often, the only available data is the measured distribution of mRNA (or protein) copy-number, and the two-state model has enough parameters to capture a large spectrum of such (non-negative) distributions, both unimodal and bimodal <ref type="bibr">(Munsky et al., 2012)</ref>. Consequently, two-state kinetics may be erroneously invoked when, in fact, other factors underlie the variability in mRNA numbers <ref type="bibr">(Sanchez and Golding, 2013;</ref><ref type="bibr">Zopf et al., 2013)</ref>. As will become evident later in this colloquium, cell growth and reproduction result in exactly such variability and must therefore be considered when interpreting mRNA and protein statistics. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Continuous, non-Markovian models of gene expression</head><p>Another improvement in experimental resolution, which necessitated a theoretical revision, was the ability to measure mRNA levels at an accuracy finer than a whole molecule, i.e., quantify the amounts of different parts of the same polymeric mRNA <ref type="bibr">(Chen et al., 2015;</ref><ref type="bibr">Wang et al., 2019)</ref>. While the models discussed above depict mRNA creation and elimination as point processes, this is, in fact, a rather poor approximation, especially in bacteria, where the timescale of synthesizing a full-length mRNA is comparable to the lifetime of these molecules, both typically on the order of several minutes <ref type="bibr">(Chen et al., 2015)</ref>. Consequently, much of the mRNA in the bacterial cell is expected to consist of partial rather than full-length molecules. mRNA number is thus better approximated as a continuous, rather than discrete, variable. In addition, its stochastic kinetics cannot be assumed to be Markovian, but instead exhibit finite memory of transcription initiation events. Writing the master equation for mRNA dynamics becomes more challenging, but still possible, using various heuristic approaches <ref type="bibr">(Jiang et al., 2021;</ref><ref type="bibr">Xu et al., 2016)</ref>. Furthermore, once mRNA kinetics becomes tractable at sub-molecular resolution, other intricacies of the gene expression process, which were ignorable earlier, reveal themselves. These include the coupling between multiple co-transcribing RNAPs, between them and the ribosomes translating the same transcript, and between those ribosomes and the enzymes degrading the mRNA molecule <ref type="bibr">(Chen et al., 2015;</ref><ref type="bibr">Iyer et al., 2018;</ref><ref type="bibr">Kim et al., 2019)</ref>. While the theoretical consideration of mRNA kinetics at sub-molecule resolution is an important direction for future work, it is outside the focus of this Colloquium.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>II. GROWTH-DEPENDENT MODELS OF GENE EXPRESSIONS: THE EFFECT OF DNA REPLICATION</head><p>A. Beyond the static picture "The dream of a bacterium is to become two bacteria", said Francois Jacob <ref type="bibr">(Jacob, 1965)</ref>. In other words, cells beget more cells, and, at least in the realm of unicellular organisms, they typically do so as rapidly as they can. One consequence of this is that (ignoring stochastic effects) The constant-rates models introduced in Chapter I ignore all these features 7 These models are stationary -the number of gene copies is held constant at one, reaction rates are unchanging, and the attractors, too, are stationary.</p><p>Evidently, these models must be revised to reliably capture gene expression in growing, proliferating cells. We now embark on the construction of these revised models, doing so in a gradual manner. In this chapter, we consider the impact of a single, discrete event: the replication of the encoding gene. As we will see, this doubling (referred to as a change in "gene dosage") creates a time varying gene expression pattern along the cell cycle. In the next chapter, we will shift our focus to the continuous aspects of cell growth and examine the modifications that this growth imposes on gene expression.</p><p>In discussing the kinetic effects of gene replication, We focus on mRNA, rather than protein, levels. The reason is that, as the step immediately downstream of the DNA, transcription responds first, and more dramatically, to the discontinuous change in dosage. The effect on protein levels is delayed and, owing to the longer lifetime of proteins,</p><p>7 But not out of ignorance of their significance. Some of the earliest stochastic models of gene expression considered the impact of gene replication, volume growth, and cell division (reviewed in <ref type="bibr">(Paulsson, 2005)</ref>). However, these models did not become part of the standard framework that was used for interpreting experimental data. The reason, again, was that the data was of insufficient resolution to constrain, and therefore inform, these cell cycle features..</p><p>temporally smoothed (recall Fig. <ref type="figure">2</ref>). In addition to these differences in kinetics, our experimental ability to follow transcription along the cell cycle currently outpaces that for protein kinetics (we return to this point in Section II.C).</p><p>While earlier single-cell measurements were still ignorant of the cell cycle phase of individual cells, and hence had to contend with mapping dosage changes to expression "noise" (see Box II.B), more recently it has become possible to measure how mRNA numbers vary with cell-cycle progression <ref type="bibr">(Pountain et al., 2024;</ref><ref type="bibr">Wang et al., 2019)</ref>. As we will see, these new studies reveal diverse patterns, some involving non-monotonous changes along the cell cycle. We discuss the possible interpretations of these empirical findings.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Replication of the gene of interest</head><p>Single-cell measurements of gene expression became prevalent at the beginning of the new millennium <ref type="bibr">(Elowitz et al., 2002;</ref><ref type="bibr">Ozbudak et al., 2002)</ref>, and drove extensive utilization of stochastic models for the process. These models typically followed the constant-rates formulation of Chapter I. This was because the single-cell measurements at the time were ignorant of the age of the individual cells (i.e., its cell cycle phase), thus "legitimizing" the exclusion of this critical observable from the models. Nevertheless, because the data was typically acquired from asynchronous populations of growing cells, gene dosage was expected to vary twofold within the population (Fig. <ref type="figure">7</ref>). To incorporate this feature into the models of gene expression, several studies <ref type="bibr">(Jones et al., 2014;</ref><ref type="bibr">So et al., 2011)</ref> used the assumption that mRNA levels will rapidly equilibrate to reflect the new gene dosage resulting from replication. In that case, a population of growing cells is considered to be composed of two subpopulations, of cells before and after gene replication. The size of each subpopulation will depend on the growth rate and the genomic position of the gene (recall Fig. <ref type="figure">6</ref>). Once double that of newborn cells ("short"). Figure adapted from <ref type="bibr">(Sep&#250;lveda et al., 2016)</ref>. Reprinted with permission from AAAS.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Box II.B Extrinsic versus Intrinsic noise</head><p>As we discussed in Chapter I, beyond the ensemble-averaged dynamics one is often interested in predicting the cell-tocell variance in expression. Cell growth and division will contribute to this variance. For example, the rate of mRNA production in Eq. ( <ref type="formula">5</ref>), k m , will vary between cells depending on their age (cell cycle phase) due to the difference in gene copy-number before and and after DNA replication, as well as changes in the copy number of RNAP and other molecular players involved in gene expression. Beyond the changes in production rate, asymmetric partitioning at cell division will also contribute to cell-to-cell variability. One way of representing all these different contributions to heterogeneity is by treating them as sources of "extrinsic noise", in addition to the "intrinsic noise" associated with the stochasticity of the kinetic scheme itself model <ref type="bibr">(Huh and Paulsson, 2011;</ref><ref type="bibr">Jones et al., 2014;</ref><ref type="bibr">Peterson et al., 2015;</ref><ref type="bibr">Swain et al., 2002)</ref>, as we shall now explain.</p><p>Consider the copy-number of a given mRNA or protein in the cell, within the class of stochastic models introduced in Chapter I, albeit when the transcription and translation rates k m , k p are time-dependent, reflecting their potential change along the cell cycle. The law of total variance enables one to decompose the variance of a random variable Y (e.g., M or P ) into two components, by conditioning on another variable or set of variables X <ref type="bibr">(Blitzstein and Hwang, 2015;</ref><ref type="bibr">Fu and Pachter, 2016;</ref><ref type="bibr">Hilfinger and Paulsson, 2011)</ref>:</p><p>The first term on the RHS corresponds to taking the variance of the variable (say, protein number) when conditioning on the parameters (in our case, k m and k p ), and then taking the expectation value over the probability distribution of these parameters. This term corresponds to the noise we expect in the simpler models with fixed rates. This is known as intrinsic noise. In the case of the two-state model, for example, it can be shown to scale as ( b + 1) P , where b is the burst size of Eq. ( <ref type="formula">12</ref>) and P is the mean expression level of the protein in question (as expected, the standard deviation, normalized by the mean expression level, will be smaller for highly expressed proteins). The second term in the decomposition accounts for the fluctuations in the reaction rates, and is known as extrinsic noise; For any given set of parameters we are conditioning on, consider the mean (the expectation value) of the variable of interest (e.g., protein level). The extrinsic noise is simply the variance of that expectation value over the distribution of the varying parameters.</p><p>The expected contribution to the extrinsic noise in gene expression from several cell-cycle features-e.g., gene doubling, cell growth and division-has been calculated <ref type="bibr">(Huh and Paulsson, 2011;</ref><ref type="bibr">Jones et al., 2014;</ref><ref type="bibr">Peterson et al., 2015;</ref><ref type="bibr">Swain et al., 2002)</ref>. The challenge, however, is that these noise contributions may eventually dominate over the intrinsic component reflecting the kinetic scheme-which is often our main interest. This limitation of noise-based analysis thus motivates directly measuring the cell-cycle phase of individual cells, and explicitly considering factors such as cell volume and gene dosage in the analysis of gene expression.</p><p>One assumption underlying the treatment above is that transcription rate is proportional to gene dosage, hence would double once the gene replicates. However, experimental data indicates that this simple proportionality is sometimes violated. Specifically, mRNA production rate may be a sublinear function of gene copy number, a feature referred to as "dosage compensation". Dosage compensation may be advantageous to the organism as a means of buffering expression against the unavoidable change of gene copy-number during cell growth. The subject is outside the premise of this Colloquium (reviewed in <ref type="bibr">(Bar-Ziv et al., 2016)</ref>).</p><p>Notwithstanding the possibility of dosage compensation, the approximation that mRNA levels instantaneously track gene dosage relies on the assumption that mRNA lifetime (which determines the adaptation time to dosage doubling) is negligible compared to the cell generation time. This, however, is again a poor assumption in rapidly growing bacteria, where the two timescales may be within a few-fold of each other. In that case, the temporal kinetics of mRNA (as described, e.g., by Eq. ( <ref type="formula">5</ref>)) must be solved, while matching the two boundary conditions: before and after replication of the gene (at time t r ), and before and after cell division (at time t d ). One arrives at the following expression for the population averaged mRNA number over time, M (t) <ref type="bibr">(Peterson et al., 2015)</ref>:</p><p>As expected, the finite mRNA lifetime results in a smoothed (filtered) response to the discontinuous change in gene dosage (see Fig. <ref type="figure">10</ref> below). Beyond the population mean, the contribution to extrinsic noise in mRNA numbers can also be calculated for this case, using the approach described in Box II.B <ref type="bibr">(Peterson et al., 2015)</ref> (see also <ref type="bibr">(Beentjes et al., 2020;</ref><ref type="bibr">Jia and Grima, 2023)</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Cell-cycle dependent transcription: Experimental observations</head><p>Bacterial cells can nowadays be grown and tracked under the microscope for many generations <ref type="bibr">(Wang et al., 2010)</ref>. By combining bright field and fluorescence microscopy, key events during the cell cycle-e.g., the initiation or termination of genome replication-can be detected in each cell and their timing recorded (Fig. <ref type="figure">8</ref>). But while this approach has allowed an empirical characterization of the bacterial cell cycle, the real-time measurement of mRNA and proteins production along the cell cycle remains largely outside the current experimental capability. Multiple challenges of live-cell microscopy contribute to this problem. For one, long-term measurement of gene expression relies on the detection of fluorescent proteins. To emit their signal, these proteins must undergo a slow and stochastic process of fluorophore maturation, which lowers the temporal resolution of the measurements to multiple minutes at best <ref type="bibr">(Balleza et al., 2018)</ref>. Consequently, the inference of cell-cycle expression patterns has remained a challenge (but see <ref type="bibr">(Rosenfeld et al., 2005;</ref><ref type="bibr">Walker et al., 2016;</ref><ref type="bibr">Zopf et al., 2013)</ref> for exceptions). In some instances, the dependence on fluorescent protein maturation can be circumvented by devising a detection scheme that relies on the change in localization of preexisting proteins rather than production of new ones. This approach has been used successfully to detect and count both mRNA (Fig. <ref type="figure">5</ref> above) and gene loci (Fig. <ref type="figure">7</ref> above) in live bacteria, and analogous schemes have been proposed for detecting translation <ref type="bibr">(Wu et al., 2016)</ref>. However, issues of detection sensitivity, spatiotemporal resolution, and the perturbation to cell physiology must be resolved before robust, long-term measurement of gene expression along the bacterial cell cycle becomes possible.</p><p>As an alternative to tracking gene activity in live cells, the details of cell-cycle dependent transcription can be revealed by analyzing snapshots of chemically fixed cells<ref type="foot">foot_2</ref> . This approach leverages the fact that for exponentially growing cells, cell size (in rod-shaped bacteria like E. coli, its length) can be approximately mapped to its age (cell cycle phase).This is demonstrated by the distribution of measured cell lengths, which reflects the expected statistics of cell ages during exponential growth <ref type="bibr">(Pountain et al., 2024)</ref>, as well as the measured gene copy-number, which exhibits a step-like change as a function of cell length, see Fig. 9 <ref type="bibr">(Wang et al., 2019)</ref>. An additional advantage offered by size-based analysis is that, in E. coli, the initiation of genome replication is triggered, on average, at a given cell volume rather than age <ref type="bibr">(Ho et al., 2018;</ref><ref type="bibr">Wallden et al., 2016;</ref><ref type="bibr">Zheng et al., 2016)</ref>, thus making size a natural axis along which to examine the effect of this event.</p><p>FIG. <ref type="figure">8</ref> Tracking the progression of the bacterial cell cycle in real time. Top, the "Mother Machine" microfluidic device enables high-throughput observation of mother cells (adapted from <ref type="bibr">(Wang et al., 2010)</ref>, with permission from Elsevier). Bottom, the progression of genome replication in E. coli is followed by fluorescently labeling the cellular replication machinery, or "replisome" (adapted from <ref type="bibr">(Kn&#246;ppel et al., 2023)</ref>, reproduced under Creative Commons Attribution License 4.0 (CC BY)).</p><p>FIG. <ref type="figure">9</ref> Cell length approximates cell-cycle progression in E. coli. Bacteria were chemically fixed and imaged to measure cell length and the copy number of a given genomic locus. The gene dosage exhibits a step-like change as a function of cell length, reflecting the progression of the cell cycle. Data by Tianyou Yao (unpublished). For the experimental method, see <ref type="bibr">(Wang et al., 2019)</ref>.</p><p>Using this approach to measure the mRNA level along the cell cycle for several strongly expressed promoters in E.</p><p>coli revealed good agreement with the model of Eq. ( <ref type="formula">16</ref>), where mRNA levels track gene dosage with a finite adaptation period (Fig. <ref type="figure">10</ref>). While the imaging-based method is limited to quantifying only a few promoters at a time, a recent study introduced an algorithm for sorting the full transcriptome of individual cells (obtained using single-cell RNA sequencing, scRNA-seq) along the cell cycle, thus opening the door to identifying the cell-cycle expression pattern across the whole genome ( <ref type="bibr">(Pountain et al., 2024)</ref>; <ref type="bibr">(Riba et al., 2022)</ref> apply an analogous approach in mammalian cells). The sequencing-based results agree well with the imaging data, and, for many E. coli genes, reveal a similar mRNA-follows-dosage pattern along the cell cycle, again captured well by Eq. ( <ref type="formula">16</ref>) (Fig. <ref type="figure">11</ref>). FIG. 11 Single-cell RNA sequencing (scRNA-seq) analysis reveals an mRNA-follows-dosage pattern across multiple E. coli genes. scRNA-seq expression, converted to mRNA copy-number, is plotted against cell age. Horizontal dotted lines indicate the inferred steady-states levels before and after gene replication. Red line, fit to Equation 16. Data from (Pountain et al., 2024). Additional analysis by Kevin McDonald and Tianyou Yao. D. Non-monotonic transcription patterns: observations and possible mechanisms</p><p>While multiple promoters exhibit the aforementioned step-like pattern of expression, other promoters show nonmonotonic changes in mRNA level along the cell cycle, where the expected increase accompanying gene replication is both preceded and followed by a decrease in expression (Fig. <ref type="figure">12</ref>) <ref type="bibr">(Wang et al., 2019)</ref>. The anecdotal observations using microscopy are again reflected in the RNA sequencing analysis, which suggests that many E. coli genes exhibit this behavior (Fig. <ref type="figure">13</ref>). These non-monotonic expression patterns, whose origin is currently unknown, provide an opportunity for testing some of the current ideas regarding the drivers of gene expression in growing cells. Here we briefly discuss two classes of hypotheses. 1. Replication-triggered transcription. An idea long discussed in the bacterial literature is that transcription of low-expression proteins takes place preferentially around the time of gene replication rather than with a uniform probability along the cell cycle <ref type="bibr">(Golding, 2019;</ref><ref type="bibr">Guptasarma, 1995)</ref>. The idea is premised on the many conceivable ways in which passing of the DNA replication machinery through the gene may transiently increase transcription, beyond the obvious change in dosage discussed above. The hypothesized effects include changes to DNA topology (supercoiling) ahead and behind the replicated gene <ref type="bibr">(Sevier and Levine, 2017)</ref>; changes in the spatial position of the gene in the cell during replication, making it more accessible to RNAP and ribosomes; and the displacement of bound transcription factors acting as repressors by the replication process-their removal resulting in transient transcriptional activity until they rebind <ref type="bibr">(Golding, 2019;</ref><ref type="bibr">Guptasarma, 1995)</ref>. Consistent with the latter hypothesis, the activity patterns of the repressor-controlled lac promoter appears to gradually shift, from the canonical dosagetracking at high expression to a pulsatile replication-associated one, as the repression by the transcription factor LacI tightens (Fig. <ref type="figure">12</ref>). Beyond the various biophysical effects of DNA replication, it is conceivable that the doubling of dosage itself may elicit a non-monotonic transcriptional response. Gene replication creates a step-like change in dosage, and the resulting effect on transcription can be seen as the "step response" of the genetic circuit of which the replicated gene is part. Depending on the topology of that circuit, e.g., the presence of one or more negative feedback loops <ref type="bibr">(Milo et al., 2002)</ref>, this response may be non-monotonic or even oscillatory <ref type="bibr">(Stricker et al., 2008)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">RNAP competition during genome replication.</head><p>In Chapter III we discuss the idea that transcription of a given gene is limited by RNAP availability, and thus that the different genes compete for this limited resource. The competition could conceivably result in non-monotonic transcription along the cell cycle, reminiscent of the empirical data of Figs. 12 and 13: rising when the gene-copy doubles, but diminishing at other times while the rest of the genome replicates, since this replication produces competing targets for RNAP. The expected activity pattern is complicated by the elaborate scheme of genome replication in rapidly growing bacteria, where multiple, nested replication events run simultaneously and the rate of new-genome (hence, new RNAP targets) production varies along the cell cycle <ref type="bibr">(Neidhardt et al., 1990;</ref><ref type="bibr">Wallden et al., 2016)</ref>(recall Fig. <ref type="figure">6</ref>).</p><p>Identifying the mechanistic origins of cell-cycle dependent transcription requires characterizing mRNA numbers across the genome and as a function of multiple parameters, including the cell's growth rate and each promoter's expression level and regulatory topology, interrogation that is now becoming possible thanks to emerging single-cell transcriptomics approaches based on imaging <ref type="bibr">(Dar et al., 2021)</ref> and sequencing <ref type="bibr">(Pountain et al., 2024)</ref>.</p><p>III. GROWTH-DEPENDENT MODELS OF GENE EXPRESSION: THE EFFECTS OF GLOBAL TRANSCRIPTION AND</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>TRANSLATION</head><p>In considering cell growth and reproduction, we focused above on a single element, the gene of interest, and examined the consequences of its replication for the expression of the corresponding mRNA. But this discrete replication event takes place as part of the doubling of cell volume and all cellular components, including the entire genome and the gene expression machinery. What is the effect of these global changes? To answer this question, we consider a simplified picture in which the cell undergoes continuous growth, while putting aside for now genome replication and cell division. As we will see, this level of description enables us to identify important constraints on the rates of transcription and translation in growing cells. An important outcome of the analysis is that the expression of an individual gene is coupled to that of all other genes in the same cell.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Correcting for cell growth by normalization by cell volume</head><p>As noted above, early studies of gene expression in single cells utilized a constant-rates formulation (Chapter I).</p><p>The change in gene dosage during cell growth, rather than modeled explicitly, was mapped to an added noise term (BOX II.B). But, of course, not only the gene of interest but all genes, as well as mRNAs and proteins, are expected to vary twofold within a population of growing cells (even ignoring the additional stochastic effects). A common way of addressing this heterogeneity became to normalize protein numbers (mRNA measurements were still uncommon at the time) by the cell volume, i.e., consider concentrations rather than molecular numbers, and interpret these numbers in a constant-rates framework <ref type="bibr">(Elowitz et al., 2002;</ref><ref type="bibr">Jia et al., 2022;</ref><ref type="bibr">Thomas and Shahrezaei, 2021)</ref>. Intuitively, this normalization should-at least to some extent-take care of the doubling of all cellular components during the cell cycle: for cells undergoing binary fission (such as E. coli) cellular concentrations of genes, mRNAs, and proteins are expected to be identical before and after cell division<ref type="foot">foot_3</ref> . To render this argument more rigorous, let us use Eqs. ( <ref type="formula">5</ref>)-( <ref type="formula">6</ref>)</p><p>to derive the temporal dynamics of the cellular concentrations of mRNA and protein, m and p. This requires us to explicitly consider cell growth, since volume increase inherently leads to the dilution of both species. Denoting by P the protein copy-number, we have p = P/V . Therefore, the dynamics of the ensemble-averaged concentration</p><p>Let us assume that, at a given point along the cell cycle, proteins are produced at a rate &#954;, which is potentially time-dependent (within the models of Chapter I, &#954; will be proportional to the instantaneous mRNA copy number).</p><p>Assuming exponential growth of cell volume, with rate &#955; = 1 V dV dt , we find:</p><p>We thus need to account for the effect of dilution by adding &#955; to the protein degradation rate &#947; p . The possible time dependence of &#954; allows for the possibility of a time-independent protein concentration as a stationary solution. In particular, if &#954;(t) &#8733; V (t), we see that the additional term on the RHS will allow for a fixed point of the dynamics, even in the absence of protein degradation: a homeostatic concentration level, at which dilution balances production.</p><p>As we will discuss in Section III.E, this is consistent with experimental observations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Constraints on the rates of transcription and translation</head><p>The preceding discussion left undetermined &#954;(t), the time-dependent rate of protein production in Eq. ( <ref type="formula">18</ref>). To determine it, we will need to develop a more complete model, which considers both transcription and translation simultaneously in the context of continuous cell growth. This section introduces such a model 11 .</p><p>1. The bacterial "growth law"</p><p>We first digress from the preceding discussion of stochastic gene expression, and briefly recap a celebrated "growthlaw" observed experimentally in growing microbes, including E. coli and budding yeast: as nutrient conditions are varied, one finds that the fraction of total protein mass in the cell taken by ribosomes increases linearly with the cellular growth rate <ref type="bibr">(Metzl-Raz et al., 2017;</ref><ref type="bibr">Scott et al., 2010)</ref>. There has been considerable work recently related to this growth law, outside the scope of the current Colloquium (see <ref type="bibr">(Scott and Hwa, 2023)</ref>). Here, we will ignore many of the subtleties and highlight the simple qualitative rationale for the observed behavior, which will become pertinent to our discussion of gene expression.</p><p>Since ribosomes produce all proteins within the cell -including other ribosomes -a coarse-grained model for their auto-catalytic production (neglecting degradation) would suggest:</p><p>where R(t) is the total number of ribosomes in the cell 12 , k tr is the translation rate (per ribosome), and &#934;r is fraction of ribosomes that are actively translating ribosomal proteins. Note that although each ribosome is composed of tens of smaller proteins, in this coarse-grained, simplified equation the ribosome is treated as a single, self-replicating entity.</p><p>Furthermore, we have neglected the fact that each ribosome has a large RNA component (in addition to proteins).</p><p>However, one may argue that producing ribosomal RNA is much "cheaper" than producing ribosomal proteins, as evidenced by the fact that in E. coli the fraction of RNAP in the proteome (a few percents, typically) is considerably lower than fraction of ribosomes. A more systematic discussion of these features is presented in <ref type="bibr">(Reuveni et al., 2017)</ref>.</p><p>The solution of Eq. ( <ref type="formula">19</ref>) is exponential growth of the ribosomal copy number with a rate proportional to &#934;r , which we also expect to equal the growth rate of the cell. Since ribosomes produce all other proteins as well, those proteins' numbers will also increase exponentially, and the fraction of ribosomes in the cell will equal &#934;r . We conclude that the growth rate should be proportional to the fraction of ribosomes in the proteome, thus providing a possible explanation for the experimentally observed "growth-law".</p><p>11 The model we propose is by no means "complete" in the sense of capturing all of the relevant cellular features -it is still a coarse-grained description which ignores many processes. Ref. <ref type="bibr">(Thomas et al., 2018)</ref> provides an example for a more comprehensive model, which also considers nutrient transport and metabolism. 12 We typically think of the rates of biochemical reactions as determined by the cellular concentrations of molecules. Why have we switched back to working with copy numbers? The scenario to have in mind is one where ribosomes are the limiting cellular resource for protein production, and once a ribosome completes translation of a given mRNA, it is only idle for a brief moment before encountering another mRNA and beginning translating again. We can thus assume that all ribosomes are producing proteins at any given time, arriving at Eq. ( <ref type="formula">19</ref>). Of course, the equation can be rewritten in terms of concentrations, but the derivation and interpretation are simpler when working with copy numbers. This will also be true for the subsequent derivations in this chapter. (In reality, only a finite fraction of ribosomes are active at any moment (Metzl-Raz et al., 2017). However, this will only introduce a prefactor of order unity into the equations, reflecting this fraction.)</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">The rate of protein production</head><p>At the heart of the simple model above lies the assumption that ribosomes are limiting for translation, with the protein production rate proportional to their copy number. This contrasts with the models of Chapter I, where changes in the rates of protein production were only associated with changes in mRNA copy numbers, while the ribosomal levels played no explicit role. Furthermore, since ribosomal numbers are expected to increase as the cell grows, the protein production rate would not be constant in time, in contrast to the assumptions of the earlier models.</p><p>What should the constant rates in, e.g., Eqs. ( <ref type="formula">5</ref>)-( <ref type="formula">6</ref>), be replaced with to be consistent with the ribosomal growth law?</p><p>To answer this question, recall that the ribosome-centric model we introduced corresponds to a scenario where ribosomes are always "hungry", and are actively translating some mRNA at any moment in time <ref type="bibr">(Lin and Amir, 2018)</ref>. Within this picture, the mRNAs corresponding to different genes compete for ribosomes' attention, and the protein production rate for gene i will depend not on the absolute level of the corresponding mRNA, but rather on its relative abundance in the pool of mRNAs, which determines the chance that the next ribosome to become available will encounter it. Under this scenario, the production rate of protein i reads:</p><p>with Mj the effective mRNA copy number of gene j -also accounting for its affinity for ribosome binding -and the summation is over all genes in the genome. For simplicity, below we neglect the heterogeneity in ribosomal binding affinity, and hence associate Mj with the actual mRNA copy number, M j . The total protein production rate (i.e.</p><p>the production rate summed over all proteins) will be, by construction, limited only by the ribosome number and independent of the mRNA levels. Those mRNA levels, however, dictate the relative rates of protein production for different genes. To determine these rates, we thus turn next to the laws governing transcription within the ribosome-centric model of cell growth.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">The rate of mRNA production</head><p>In analogy to the preceding discussion regarding translation, let us assume that the transcription rate of each gene is limited by the (time dependent) cellular number of RNA polymerases (RNAPs), which we denote by N<ref type="foot">foot_5</ref> . Similarly to what we previously assumed of ribosomes, we envision that RNAPs (or, as before, a finite fraction of them) are always busy transcribing, with the rate of mRNA production from a given gene dictated by the fraction of RNAPs actively transcribing that gene. Which specific gene is transcribed by the next available RNAP will depend on the particular gene's copy number and the affinity of RNAP to the promoter region of the gene. The propensity to transcribe can be further modulated by the action of transcription factors, proteins that bind the genome and affect gene expression by, e.g., sterically blocking the site RNAP should bind to <ref type="bibr">(Ptashne and Gann, 2002)</ref>. To incorporate these combined effects, we write an expression analogous to Eq. ( <ref type="formula">20</ref>), where the fraction of mRNAs corresponding to a particular gene determined its translation rate. Here, the corresponding quantity is one we term the gene allocation fraction &#934; i :</p><p>where g i is a coarse-grained quantity that reflects the copy number of a given gene as well as the regulatory features above -in other words, g i determines how competitive a given gene is in capturing RNAP's attention. The mRNA copy number then obeys:</p><p>with k tx the transcription rate (per RNAP) 14 , N the number of RNAPs, and &#947; m the mRNA degradation rate (which, for simplicity, we will assume to be identical for all genes).</p><p>Before proceeding to solve equations ( <ref type="formula">20</ref>) and ( <ref type="formula">22</ref>), it is important to note that, in both of these equations, one of the indices i corresponds to the genes encoding ribosomes, and another to those encoding RNAP. Specifically, we will use the notation &#934; r for the ribosome allocation fraction and &#934; n for the RNA polymerase allocation fraction. Also note that the translation rates in Eq. ( <ref type="formula">20</ref>) depend explicitly on the mRNA levels, which in turn are given by Eq. ( <ref type="formula">22</ref>)</p><p>-provided that we know the time-dependent RNAP level N (t).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Solving the model</head><p>Eqs. ( <ref type="formula">20</ref>) and ( <ref type="formula">22</ref>) comprise a closed set of equations for the production of both mRNA and proteins -including ribosomes and RNAPs themselves. These equations were written under the explicit assumptions that ribosomes (rather than mRNA copy numbers) limit the overall rate of protein production in the cell, and, similarly, RNAPs (rather than the gene copy number) limit the overall rate of RNA synthesis. We will later relax these assumptions, and, in doing so, Eqs. ( <ref type="formula">20</ref>) and ( <ref type="formula">22</ref>) will come to describe one regime (later denoted as regime I ) out of several possibilities described by the continuous growth model.</p><p>To proceed, consider Eq. ( <ref type="formula">20</ref>) for the case of genes encoding ribosomal proteins. It reads <ref type="bibr">(Lin and Amir, 2018)</ref>:</p><p>with M R the copy number of ribosomal mRNA. Next, we note that the solution of Eq. ( <ref type="formula">22</ref>) for the mRNA of gene i is given by:</p><p>At long times compared with the mRNA lifetime, the initial conditions for the mRNA copy number will not matter, and we conclude that:</p><p>Plugging this into Eq. ( <ref type="formula">23</ref>) leads to a closed equation for the ribosome number:</p><p>reproducing the functional form of the "growth law" of Eq. ( <ref type="formula">19</ref>). Previously, &#934;r was defined as the fraction of active ribosomes translating ribosomal proteins. Here, &#934; r is the gene allocation fraction -"hardcoded" into the DNA since it depends on the gene copy number and promoter strength (but also, potentially, on the modulation by transcription 14 Not to be confused with km of Eq. ( <ref type="formula">1</ref>).</p><p>factors). To see why these two quantities are identical, note that, according to Eq. ( <ref type="formula">25</ref>) the gene allocation fraction results in an identical mRNA fraction, which, in turn, implies (according to Eq. ( <ref type="formula">23</ref>)) the same ribosomal fraction.</p><p>Eq. ( <ref type="formula">26</ref>) for the ribosome number is closed, hence R(t) is now known, and predicted to be exponential in time.</p><p>This allows us to revisit Eq. ( <ref type="formula">20</ref>), but consider the expression of other proteins within the cell, finding along a similar vein:</p><p>the solution of which gives us:</p><p>with &#955; = k tr &#934; r the growth rate obtained from Eq. ( <ref type="formula">26</ref>).</p><p>It is useful to recast these equations in terms of concentrations, which will help reveal that the behavior we obtain indeed represents a steady-state (i.e., homeostasis) in terms of those cellular concentrations. This follows the logic of our discussion around Eq. ( <ref type="formula">18</ref>), but now the protein production rate &#954;(t) is obtained explicitly. Let us assume that the total protein concentration in the cell is fixed to a value c, independent of the changes to cell volume during growth. In Chapter IV we will discuss how such homeostasis may be achieved, but for now we can take it as an empirical observation <ref type="bibr">(Crissman and Steinkamp, 1973;</ref><ref type="bibr">Kubitschek et al., 1983;</ref><ref type="bibr">Rollin et al., 2023)</ref>. From Eq. ( <ref type="formula">27</ref>) we find that:</p><p>with p i denoting protein concentration and r the ribosome concentration.</p><p>According to this equation, the concentration of each protein is analogous to the position of an overdamped particle in a harmonic potential -it is subject to a linear restoring force, attracting it to the steady-state concentration of c&#934; i , where c is the total protein concentration. Note that since we wrote an ODE for the ensemble-average, stochasticity has been neglected; adding it would lead to small fluctuations around the steady-state solution, as shown in Fig. <ref type="figure">14</ref>.</p><p>With stochasticity present, the mechanical analogy with a particle in a harmonic potential essentially maps the dynamics of the protein concentration to the well-known Ornstein-Uhlenbeck process, corresponding to a confined particle subject to Brownian noise, and described by a simple Langevin equation <ref type="bibr">(Amir, 2020)</ref>. Without the restoring force, the particle would diffuse to infinity. Without the noise, the overdamped particle would reside in a particular "coordinate" (corresponding to the homeostatic concentration c&#934; i ). With both features present, there exists an equilibrium solution -in our problem, a stationary distribution for the concentration, centered around its fixed point in the absence of noise. In fact, from observing the fluctuations involved in gene expression, we may infer the relative contributions of intrinsic and extrinsic noise (discussed in box II.B) acting on a particular gene: In the absence of any extrinsic noise (and assuming no additional regulation), Eq. ( <ref type="formula">29</ref>) predicts that for stable proteins the strength of the "restoring force" is the growth rate &#955;. Extrinsic noise can be shown to weaken this restoring force <ref type="bibr">(Lin and Amir, 2021)</ref>. The experimental data for E. coli, corroborating the prediction of an effective, linear restoring force, is shown in Fig. <ref type="figure">15</ref>. Note that the fluctuations due to the stochastic term, which will supplement Eq. ( <ref type="formula">29</ref>), are what enables us to measure this restoring force -since without these fluctuations the cellular concentration would have been perfectly constant. This is reminiscent of the way in which natural variability between cells enables one to draw conclusions regarding cell size control <ref type="bibr">(Amir and Balaban, 2018;</ref><ref type="bibr">Ho et al., 2018;</ref><ref type="bibr">Kar et al., 2023</ref>).</p><p>An analogous equation can be written for the mRNA concentrations:  <ref type="formula">20</ref>) and ( <ref type="formula">22</ref>), are shown. The protein and mRNA levels increase exponentially with time, with strong fluctuations in the mRNA and much weaker ones for the proteins.</p><p>The background shows three individual trajectories of the stochastic dynamics, while the circles show the mean of 130 cell cycles (with the colored bands representing the standard deviation). The black lines are the theoretical predictions of exponential growth. The corresponding concentrations are presented in <ref type="bibr">(Lin and Amir, 2018)</ref>. Cells were assumed to divide symmetrically, based on the "adder" model (reviewed in Ref. <ref type="bibr">(Ho et al., 2018)</ref>; this particular choice for implementing divisions has little effect on the results). Adapted from <ref type="bibr">(Lin and Amir, 2018)</ref>, and reproduced with permission.</p><p>with &#955; the growth rate and c the total protein concentration. Considering again numbers instead of concentrations, Eqs. ( <ref type="formula">29</ref>) and ( <ref type="formula">30</ref>) tell us that at long times compared with the cell's doubling time, protein and mRNA copy numbers will increase exponentially (since concentrations are stationary and volume increases exponentially) <ref type="foot">15</ref> . This conclusion -namely, exponential growth, with identical rates, of protein and mRNA numbers -is robust to the introduction of stochasticity and cell divisions into the continuous growth model, as is illustrated in Fig. <ref type="figure">14</ref>. Here, Gillespie simulations <ref type="bibr">(Gillespie, 1976)</ref> were used to simulate the stochastic processes depicted in Fig. <ref type="figure">16</ref>, regime I.</p><p>The main simplifying assumption in the derivation above is the constancy of the effective fraction of gene copy number for each gene, &#934; i . This implied that the relative fraction of RNAPs transcribing any two genes does not change in time. Consequently, the relative mRNA levels corresponding to these genes (at times long compared with the mRNA lifetime) are also given by &#934; i , as are the resulting relative protein levels. In reality, of course, the DNA encoding all genes is replicated during cell growth. How would this affect the above calculation? If the duration of replicating the entire genome is small compared to the cell's doubling time, as is sometime the case for eukaryotes, the relevant fractions &#934; i would remain the same before and after the replication of the entire genome, and, assuming also that mRNA lifetime is short, the predictions above would hold. In fast growing bacteria, on the other hand, we are in a very different regime -replicating the genome typically takes a considerable fraction of, or even longer than, the birth-to-division time, and the resulting existence of multiple replications simultaneously further complicates the picture above <ref type="bibr">(Neidhardt et al., 1990</ref>) (Fig. <ref type="figure">6</ref>). This scenario has not yet been studied in depth within the class of models presented above. It may, conceivably, give rise to the non-monotonic gene expression patters discussed in chapter II, since after a particular gene is replicated (leading to a fast rise in its mRNA levels), the subsequent replication of additional genes, together with the effects of competition for the limiting RNAPs (as reflected in Eq. ( <ref type="formula">22</ref>)) is expected to result in a decrease in the mRNA levels of the gene in question.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>D. Revisiting the assumptions -what limits transcription and translation?</head><p>In constructing the model for gene expression in growing cells, we made specific assumptions about the factors limiting gene expression: RNAP for transcription, ribosomes for translation 16 . We now explore the consequences of relieving these assumptions. As noted, doing so leads to different dynamics for mRNA, protein, and cell growth, defining other regimes of the continuous growth model. Later, in Section III.E, we refer to various experimental studies and attempt to assess in which regime of the continuous growth model cells of various organisms reside.</p><p>First, we modify the premise of the original model by assuming that transcription is limited by the DNA amount in the cell, rather than RNAP availability as we posited initially (protein production is still assumed to be ribosomelimited). As motivation for studying this case, consider a gedanken experiment-we will later discuss an actual experiment of this sort-where cell volume continually increases as before, but the amount of DNA remains fixed.</p><p>The limiting resource for transcription is initially assumed, as before, to be RNAP, resulting, in accordance with Eq.</p><p>(30), in a constant cellular concentration of proteins. Cellular DNA, on the other hand, is gradually diluted. It is evident that, at some stage, DNA rather than RNAP will become limiting for transcription: a single DNA template would be insufficient to support transcription in an enormous cell.</p><p>To understand how this would come about mechanistically, we may consider the limit where the volume/DNA ratio is sufficiently large such that RNAPs are packed to their limit on the DNA; clearly, there is a physical limit to the number of RNAPs that can fit on any particular region of the DNA, in turn limiting transcription. Reaching the 16 A priori, one may imagine that protein production and cell growth will also be limited by the rate of transporting nutrients into the cell, which would depend on the surface area to volume ratio, hence on cell size and cell-cycle progression. Empirically, however, this does not appear to be the case: Studies in both mammalian cells <ref type="bibr">(Mu et al., 2020</ref>) and E. coli <ref type="bibr">(Zheng et al., 2016)</ref> found that even a dramatic perturbation to cell size did not alter cells' growth rate.</p><p>physical limit of RNAP occupancy is, however, not the only possibility. Ref. <ref type="bibr">(Lin and Amir, 2018)</ref> arrives at similar results by considering, instead, stochastic RNAP kinetics: binding/unbinding at the promoter, and the initiation of transcription when bound. The authors show that when the free RNAP concentration is low, the model reduces to that of Section III.C (i.e., RNAPs are limiting), but that this inevitably breaks down for large volume/DNA ratio, at which the amount of DNA becomes limiting for transcription.</p><p>Regardless of the underlying mechanism, when DNA becomes limiting for transcription, mRNA production follows:</p><p>see Table <ref type="table">I</ref> for a reminder of the variable definitions. The total amount of cellular mRNA depends on the DNA level in the cell, and will saturate to a constant rather than increase exponentially. However, importantly, so long as gene dosage stays unchanged the relative amounts of mRNAs between different genes will still obey:</p><p>with g i the effective gene copy numbers and &#934; i the gene-allocation fraction, defined in Eq. ( <ref type="formula">21</ref>). Since protein production is still described by Eq. ( <ref type="formula">20</ref>), and depends on the relative amounts of mRNAs, the previous predictions of the ribosome-centric model for protein production remain intact. In particular, protein levels still increase exponentially in time, and the ribosome growth law of Eq. ( <ref type="formula">19</ref>) remains valid. We refer to this regime, where transcription is DNA limited and translation is ribosome limited, as "regime II" of the continuous growth model, see Fig. <ref type="figure">16</ref>.</p><p>What happens if we further relax the assumption that ribosomes are limiting for protein production? In that case, the model becomes identical to the constant-rates model (aside from the effects of gene dosage due to DNA replication, as discussed in Chapter II), and protein accumulation becomes linear rather than exponential in time (in the absence of protein degradation, which would lead to its saturation at a finite value, and assuming a short mRNA lifetime).</p><p>We refer to this as "regime III" of the continuous growth model. The three regimes are summarized schematically in the diagram of Fig. <ref type="figure">16</ref>. regime I is the regime analyzed in Section III.C, where RNAP is limiting for transcription and ribosomes for translation. regime II is the regime where ribosomes (rather than mRNAs) still limit translation, but transcription is limited by the DNA template rather than RNAP. In regime III, as in the constant-rates model of Chapter I, DNA limits transcription while mRNAs limit translation 17 .</p><p>E. The limiting resource for transcription and translation: Experimental evidence</p><p>The analysis above indicated that the specific identity of the limiting factors for transcription and translation will result in different temporal dynamics of mRNA and protein levels. What does the experimental data suggest for different organisms? We begin by reviewing results for protein levels, then proceed to discuss mRNA.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">The scaling of protein levels with time and cell volume</head><p>The question of how protein levels scale with time or cell volume is a long-standing one. Already in the 1970's, work based on radioactive labeling showed that, in certain mammalian cells, protein numbers are proportional to cell volume <ref type="bibr">(Crissman and Steinkamp, 1973)</ref>. More recently, by flowing cells through a microfluidic device embedded in 17 For realistic parameter values, as the volume/DNA ratio increases, the transition to the regime where transcription is limited by DNA rather than RNAP occurs before the regime where mRNAs become limiting for translation -thus preventing a 4th regime where ribosomes are not limiting but RNAPs are <ref type="bibr">(Lin and Amir, 2018)</ref>. a cantilever, and measuring the latter's resonance frequency <ref type="bibr">(Cermak et al., 2016;</ref><ref type="bibr">Godin et al., 2010)</ref>, the buoyant mass of growing cells (which is typically dominated by proteins, as we know from other studies <ref type="bibr">(Hosios et al., 2016;</ref><ref type="bibr">Neidhardt et al., 1990</ref>)) was measured to the remarkable precision of a picogram -1 % of the mass of a typical E. coli cell (for mammalian cells, the relative accuracy is an order of magnitude higher). The signal is precise enough that the time-derivative of the mass can be evaluated. For linear growth of the biomass in time, this derivative is expected to be constant, while for exponential growth it will be proportional to the instantaneous cell mass. Data from four different system -the bacteria E. coli and B. subtilis, the budding yeast S. cerevisiae, and mammalian cancer cellswas inconsistent with linear growth but consistent with exponential growth <ref type="bibr">(Cermak et al., 2016;</ref><ref type="bibr">Godin et al., 2010)</ref>.</p><p>In fission yeast, recent experiments (using alternative methods to the one above, including Click Chemistry to detect incorporation of a methionine analogue) found that the rate of protein production was approximately proportional to cell volume, as would be expected if mass and volume grow exponentially <ref type="bibr">(Basier and Nurse, 2023)</ref>. The observation of approximately exponential growth hints that cells are in either regime I or regime II of the continuous growth model, in which ribosomes are limiting for translation. As noted in Section III.B.1 above, in the case of bacteria, the aforementioned "growth law" relating ribosome concentration to growth rate is, too, suggestive that ribosomes are limiting for protein production, thus reinforcing this conclusion.</p><p>Note that, in regimes I and II of the continuous growth model, if the duration of DNA replication is short compared to the cell cycle, total protein production will be exponential (since the allocation of each protein -including ribosomes -will be identical before and after DNA replication). While the gene dosage effects discussed in chapter II impact the expression levels of a given gene (in terms of both transcription and translation), in regimes I and II the total ribosome copy number is the only determinant of protein production. As long as an approximately constant fraction of ribosomes is devoted to ribosome production, their fraction in the proteome will remain constant and total protein production will be exponential. This is a reasonable approximation for many eukaryotes, since, as we mentioned, the duration of DNA replication may constitute a small fraction of the cell cycle, which within the continuum growth model implies constant protein allocations. But why this would be valid for bacteria such as E. coli, where the duration of DNA replication is comparable to the cell cycle duration, is not obvious. How can we explain then the approximately exponential growth of biomass measured in this case? Could there be deviations from exponential growth of mass that cannot be revealed using the current experimental setups (analogous to those observed for mammalian cells <ref type="bibr">(Mu et al., 2020)</ref>)?. Alternatively, the tight control exerted over ribosomal levels <ref type="bibr">(Neidhardt et al., 1990)</ref> might enable bacterial cells to maintain a constant ribosome fraction in the proteome throughout the cell cycle, leading to true exponential growth of biomass.</p><p>2. The scaling of mRNA levels with time and cell volume</p><p>In a similar manner, we may consider the change in mRNA levels during cell cycle progression, a question we began examining in Chapter II. In the current context, it is convenient to consider the scaling with cell volume. In regime I of the continuous growth model, mRNA level is proportional to cell volume (since both are exponential in time).</p><p>Importantly, the linear dependence between the two is agnostic as to the level of DNA, thus the same scaling will exist before and after DNA replication. This picture contrasts with that obtained in regimes II/III of the model, where DNA replication is reflected in a twofold jump in mRNA levels, preceded and followed by a plateau -absent in regime I, but consistent with the patterns we saw in Chapter II for many E. coli genes.</p><p>In contrast to the bacterial behavior, experimental data for mammalian cells <ref type="bibr">(Padovan-Merhar et al., 2015)</ref> shows a clear linear dependence between mRNA copy number of a given gene and cell volume, consistent with the expected regime I behavior under the assumption that RNAP limits transcription. In fission yeast, the experimental evidence supports the same conclusion: by studying mutants with differing cell size, it was found that global mRNA levels correlated with the RNAP occupancy (the fraction of RNAP bound to the promoter region) in a manner consistent with the above picture, where the polymerases form the limiting factor for transcription <ref type="bibr">(Zhurinsky et al., 2010)</ref>. Recent work revealed that transcription rates in fission yeast scale approximately linearly with cell volume, also consistent with this interpretation <ref type="bibr">(Basier and Nurse, 2023)</ref>. Alternative evidence, also supporting the RNAP-limiting picture in the same organism, was recently provided by experiments using single-molecule mRNA counting <ref type="bibr">(Sun et al., 2020)</ref>, which found a linear relation between mRNA number and cell size. Similar behavior was reported in budding yeast <ref type="bibr">(Swaffer et al., 2023)</ref>. However, this latter work suggested that the rate of mRNA degradation -which we so far assumed to be constant in time-also changes throughout the cell cycle, compensating for the sublinear dependence of the RNAP occupancy with cell size, and together leading to the linear relation.</p><p>The linear scaling between mRNA copy number and volume, reported for the evolutionarily distant mammalian cells, fission, and budding yeast, hints at the possibility of universal behavior. Nonetheless, as we saw in Chapter II, similar results were not reported for bacteria, in which diverse behavior is observed, and where matters are potentially complicated by the fact that the duration of DNA replication is comparable to the birth-to-division time, such that the gene-allocation fraction changes throughout the cell cycle.</p><p>A recent study measured both transcription and translation levels as a function of growth-rates in E. coli, albeit in bulk measurements rather than the single-cell level <ref type="bibr">(Balakrishnan et al., 2022)</ref>. While that study was thus unable to identify in which regime of the continuous growth model bacteria reside, the results nevertheless confirm several of the general predictions of the model: across the genome, there was a strong, linear correlation between the mRNA and protein levels of different genes, as is expected from Eqs. ( <ref type="formula">25</ref>) and ( <ref type="formula">28</ref>). This behavior is expected in all regimes Ref. <ref type="bibr">(Neurohr et al., 2019)</ref>.</p><p>experimental findings in budding yeast <ref type="bibr">(Chen et al., 2020)</ref>. Note that this would imply that the proteome composition would change with cell size: the super-linear genes would be over-represented in larger cells, and vice-versa. This is indeed observed experimentally in human cell lines <ref type="bibr">(Lanz et al., 2022)</ref>.</p><p>Another recent work extended the analysis discussed in the previous chapter, finding a new regime where the limiting factor for translation is neither ribosomes nor mRNA transcripts alone, but rather the formation of ribosome-mRNA complexes <ref type="bibr">(Calabrese et al., 2023)</ref>. One motivation for developing this model were experiments where yeast cells transcribed mRNAs that were quickly degraded before being translated. As the amount of such mRNA was increased, the cellular growth rate declined <ref type="bibr">(Kafri et al., 2016)</ref>. This result is surprising in light of the model presented earlier, where it was assumed that transcription is "cheap" and the burden on growth rate rises from protein production rather than transcription. If, instead, one considers a regime where ribosome-mRNA complexes are limiting, this behavior arises naturally: there would be a decrease in growth when producing more mRNAs, even when they are not successfully translated. Ref. <ref type="bibr">(Calabrese et al., 2023)</ref> proposes that the increased cost of ribosome-mRNA complexes could also serve as a regulatory checkpoint along the process of gene expression, such that anomalies at the transcription level induce a slow down in the downstream steps of protein production.</p><p>Finally, in Chapter III, the "gene allocation fraction" played a crucial role in determining the transcription rate (and through it, the protein production rate), but we neglected the possible effects of transcription factors (TFs), the proteins that mediate a direct effect by one gene on another and serve as building blocks for the regulatory networks that drive cell function <ref type="bibr">(Alon, 2006)</ref>. In principle, it is not difficult to include those in the model, as was done in Ref. <ref type="bibr">(Guo and Amir, 2021)</ref>. Conceptually, since the model considers explicitly all proteins in the cell, one needs to account for the TFs, and the term k tx in Eq. ( <ref type="formula">22</ref>) depends (non-linearly) on the concentration of these TFs. Ref. <ref type="bibr">(Guo and Amir, 2021)</ref> shows that adding such regulatory elements would, at some critical level of TFs, lead to a destabilization of the gene regulatory network, destroying the concentration homeostasis we derived in Chapter III (and leading to chaotic behavior instead). The broader point, however, is that TF activity, which is traditionally considered in a pairwise gene-on-gene manner, must instead be evaluated in the context of the full genome within the growing cell.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Concentration homeostasis in growing cells</head><p>In previous chapters, we discussed models where transcription and translation occur at a constant rate (independent of volume), we well as and ones where different genes and mRNAs "compete" with each other for the attention of the relevant molecular machinery. The latter class of models led naturally to relative concentrations which are maintained over time, fluctuating around a well-defined value. However, an explicit assumption we made was that the global concentration within the cell is maintained: we assumed, based on previous experimental evidence in both bacteria <ref type="bibr">(Basan et al., 2015;</ref><ref type="bibr">Kubitschek et al., 1983)</ref>, yeast <ref type="bibr">(Neurohr and Amon, 2020)</ref> and mammalian cells <ref type="bibr">(Crissman and Steinkamp, 1973)</ref>, that the total volume is proportional to the total amount of proteins within the cell. In cyanobacteria, it was shown that even when the number of chromosomes in each cell varied significantly, the protein concentration remained narrowly distributed <ref type="bibr">(Zheng and O'Shea, 2017)</ref>.</p><p>Indeed, within the models we discussed, fixing the global concentration would immediately imply that the concentration of every protein within the cell would be maintained -since as mentioned above the competition for the ribosomes naturally leads to the control of the relative concentrations<ref type="foot">foot_7</ref> . How the global concentration is maintained is an important problem that is the subject of current research, both theoretically and experimentally, as we briefly review.</p><p>First, to appreciate the problem at hand, let us consider a model where cell envelope components (membranes, and in some microbes, the cell wall) are produced constitutively by the ribosomes -without any feedback on the cellular density. How should we expect the density to behave within this model?</p><p>The answer depends on the cell geometry. Let us consider first the example of rod-shaped cells such as E. coli. The continuous growth model of chapter III led to exponential growth of the ribosome numbers. Under our assumption that a finite fraction of ribosomes are devoted to producing cell envelope components, we will find that, ignoring fluctuations, the cell surface area should also increase exponentially. This conclusion would also be true when considering multiple generations and, correspondingly, the total surface area of the entire progeny of a given cell. Since both total mass and surface area grow exponentially with the same rate, and assuming that the cell geometry is fixed, density is expected to be confined to a relatively narrow range even within this simplified model.</p><p>In fact, experimental data from various bacteria supports variants of this model: in such experiments, cell volume (rather than the ribosomal content) and surface area are concurrently quantified in bacterial cells during exponential growth. Ref. <ref type="bibr">(Harris and Theriot, 2016)</ref> used data of this type to suggest a model where:</p><p>This phenomenological model predicts that when perturbing the surface/volume ratio, S/V , it will relax exponentially to its steady-state value, with a relaxation rate equal to the growth rate -consistent with the experimental results for both E. coli and C. crescentus. Ref. <ref type="bibr">(Shi et al., 2021)</ref> proposed a similar picture, albeit with a time-delay between surface and volume production.</p><p>Recently, Ref. <ref type="bibr">(Oldewurtel et al., 2021)</ref> suggested that, in fact, the dynamics of surface area growth should be described by:</p><p>with M the cell mass. For exponential growth of volume and mass, Eqs. ( <ref type="formula">33</ref>) and ( <ref type="formula">34</ref>) would be identical, but the two models make distinct predictions for time-varying growth conditions, and the experiments of Ref. <ref type="bibr">(Oldewurtel et al., 2021</ref>) on E. coli agreed better with Eq. (34). A similar picture was also suggested to hold for B. subtilis <ref type="bibr">(Kitahara et al., 2022)</ref>. Note that for rod-shape cells with a well-defined diameter, the constancy of the surface/mass ratio (up to small deviations, due to the change in the volume/surface area ratio during cell growth) is equivalent to a constant cellular density &#961;, since we may write:</p><p>with the quantity V /S governed solely by the cell's geometry, and Eq. ( <ref type="formula">34</ref>) leading to a constant S/M ratio. The same arguments applied to a spherical cell would also lead to density homeostasis, albeit with non-negligible density fluctuations throughout the cell cycle, since the S/V ratio changes more significantly in this case compared to that of rod-shaped cells. We are currently unaware of experimental evidence suggesting that these scaling laws, relating surface-area, volume, and mass, are applicable in mammalian cells. This could be due to the experimental challenges involved, as the complex and changing shapes of mammalian cells make quantification of surface area extremely difficult.</p><p>One may view all of the aforementioned models coupling surface, volume, and mass growth as "open circuits"the cells do not "measure" concentration, and there is no direct feedback on it. However, effects such as molecular crowding may alter the diffusion of molecules and hence the rates of chemical reactions in the cell, providing precisely such feedback. For instance, Ref. <ref type="bibr">(Alric et al., 2022)</ref> demonstrated the effects of crowding on growth rate in budding yeast. However, the experiments were done using significant perturbations to the wild-type behavior, and whether crowding effects provide a biophysical regulatory cue for the small fluctuations around the typical density in organisms such as E. coli remains to be seen.</p><p>A more subtle mechanism has recently been proposed to control cellular concentration <ref type="bibr">(Rollin et al., 2023)</ref>. This work utilized the "pump-and-leak" model, commonly used to describe the transport of ions through the cell membrane.</p><p>In addition to highlighting the potential role of amino-acids in concentration homeostasis, their model explains the increase in volume and the concurrent decrease in cellular concentration upon entry into mitosis (the cell cycle phase where the chromosomes are segregated), known as "mitotic swelling", observed in mammalian cells <ref type="bibr">(Miettinen et al., 2022;</ref><ref type="bibr">Son et al., 2015;</ref><ref type="bibr">Zlotek-Zlotkiewicz et al., 2015)</ref>. An alternative mechanism to the aforementioned electrophysiological model could rely on the mechanical stress within the cell wall or membrane, as a means to "measure" and respond to changes in the intracellular concentration <ref type="bibr">(Mukherjee et al., 2023)</ref> (that would result in excess osmotic pressure and hence mechanical stresses in the cell envelope). Such mechanics-based feedback is resonant with the theoretical and experimental results of Refs. <ref type="bibr">(Amir et al., 2014;</ref><ref type="bibr">Amir and Nelson, 2012;</ref><ref type="bibr">Wong et al., 2017)</ref>.</p><p>Another unresolved puzzle regards the precise dependence of volume and biomass growth on time. In Chapter III we discussed experimental evidence for the approximately exponential growth of biomass across multiple organisms, including the bacteria E. coli and B. subtilis. Recent analysis of single-cell microscopy data suggested that the volume growth of these bacteria deviates from exponential growth, and is, in fact, faster than exponential ("superexponential") <ref type="bibr">(Kar et al., 2021;</ref><ref type="bibr">Nordholt et al., 2020)</ref>. If biomass growth is exponential but volume growth is not, cellular concentration is expected to be cell-cycle dependent -yet, as discussed above, other measurements have suggested that, at least in E. coli, it is not <ref type="bibr">(Kubitschek et al., 1983)</ref>. These apparently contradictory observations are yet to be reconciled. Adding another wrinkle to the story, it was also observed recently that the bacterium Mycobacterium tuberculosis grows approximately linearly at the single-cell level <ref type="bibr">(Chung et al., 2023)</ref> -hinting that, at least for this pathogen, the ribosome-centric framework of Chapter III might not be adequate, as it would lead to exponential growth.</p><p>In mammalian cells, experiments measuring buoyant mass revealed cell-cycle-dependent deviations from exponential growth <ref type="bibr">(Mu et al., 2020</ref>), yet such deviations were not observed in microscopy measurements of cell volume <ref type="bibr">(Cadart et al., 2018)</ref> (though it should be noted that the cell types were not identical). Thus, here again, obtaining a consistent picture from mass and volume measurements remains to be achieved -or, alternatively, establishing that cellular concentration is cell-cycle-dependent. Indeed, recent work in fission yeast suggested that cell density is not constant throughout the cell cycle <ref type="bibr">(Odermatt et al., 2021)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Space: The final frontier</head><p>A common feature of all models discussed in this Colloquium is that they do not explicitly consider the intracellular space: Ensemble means are described using ordinary differential equations in time, and fluctuations using the master equation, but nowhere do the spatial coordinates appear. Implicitly, the molecular encounters that underlie all cellular reactions are assumed to be diffusion driven, and to take place in a homogeneous, well-mixed environment.</p><p>In the model equations, the rates of encounter are coarse-grained into the reaction rates of transcription, translation, degradation, etc.</p><p>This approach, long the practice in modeling gene regulation <ref type="bibr">(Bintu et al., 2005;</ref><ref type="bibr">Paulsson, 2005)</ref>, is premised on the notion that the bacterial cell is sufficiently small, and internally uniform (lacking membranal compartments), such that the time it takes a molecule to diffuse across it (&lt; 1 sec <ref type="bibr">(Elowitz et al., 1999</ref>)) 20 is considerably shorter than the time scales for other processes under consideration, e.g., producing an mRNA or protein (&#8764; minutes), replicating the genome or the cell (&#8764; hour <ref type="bibr">(Cooper and Helmstetter, 1968;</ref><ref type="bibr">Neidhardt et al., 1990)</ref>). Under this premise, an explicit description of cellular spatiality is not required, and would unduly complicate models 21 .</p><p>However, there are reasons to suspect that the approximation of a homogeneous, diffusion-driven cell may be a poor one, even for bacteria. Bacterial cells have been revealed, in recent decades, to exhibit elements of spatial heterogeneity and organization, including, critically, in the machinery of gene expression. Considering E. coli, we now know that the key players in the gene expression process are localized preferentially to different regions of the cell, with only partial overlap between them: RNAP is enriched in the chromosomal region (the nucleoid), ribosomes are largely found at the periphery of the nucleoid, and the protein complex that degrades mRNA (the degradosome) is 20 An intriguing caveat is that, in contrast to the normal (Fickian) diffusion of individual proteins, large molecular complexes have been shown to exhibit anomalous diffusion (specifically, sub-diffusion), with possible consequences for molecular encounter kinetics <ref type="bibr">(Golding and Cox, 2006</ref>). 21 A similar argument is commonly applied to eukaryotic cells, with the exception that the membrane-separated nucleus and cytoplasm are considered as different compartments, with molecules shuttling between them <ref type="bibr">(Hansen et al., 2018)</ref>. The oversimplification of this depiction is demonstrated by the recent report of concentration gradients inside yeast cells, gradients which could lead to intracellular differences in molecular diffusivity and reaction rates <ref type="bibr">(Odermatt et al., 2021)</ref>.</p><p>anchored to the cell membrane <ref type="bibr">(Campos and Jacobs-Wagner, 2013)</ref>. Thus, the steps of gene expression conceivably proceed with spatial directionality, but this feature has only begun to receive theoretical attention <ref type="bibr">(Castellana et al., 2016)</ref>. Beyond this regional preference, it was found that, as in eukaryotes <ref type="bibr">(Cho et al., 2018)</ref>, bacterial RNAP further exhibits areas of high localized concentration (reported to form through liquid/liquid phase separation <ref type="bibr">(Ladouceur et al., 2020)</ref>), which colocalize with genomic loci that encode ribosomes <ref type="bibr">(Fan et al., 2023;</ref><ref type="bibr">Weng et al., 2019)</ref>. It is plausible that these RNAP clusters represent "transcription factories" <ref type="bibr">(Cook, 2010)</ref>, regions of increased transcriptional activity, functioning to ensure sufficient production of ribosomes in the cell. If true, this would provide ribosomal genes with competitive advantage over other genes in securing the limiting resource of RNAP, under the scenario discussed in Section III.B.3 above.</p><p>Similarly, although the evidence in this regard is more limited, several studies suggest the local spatial accumulation of bacterial transcription factors (the proteins that modulate expression level), both at their genomic site of production <ref type="bibr">(Kuhlman and Cox, 2013)</ref> and of targeted binding <ref type="bibr">(Sarkar-Banerjee et al., 2018)</ref>. The experimental observations (again mirroring analogous ones in eukaryotic cells <ref type="bibr">(Mir et al., 2017)</ref>), have been accompanied by several theoretical efforts to elucidate how spatial gradients of transcription factors form, and what the consequences are for their regulatory activity <ref type="bibr">(Kolesov et al., 2007;</ref><ref type="bibr">Kuhlman and Cox, 2013)</ref>.</p><p>The examples above highlight the limitations of the spaceless picture, where models of gene expression still, for the most part, reside. The inadequacy of this modeling approach is further suggested by the fact that, whereas current models can typically reproduce the ensemble-averaged expression, they fail to correctly predict some aspects of the associated fluctuations, in particular, the near-universal scaling relation between the mean and noise, observed across different genes and even different organisms <ref type="bibr">(Sanchez and Golding, 2013;</ref><ref type="bibr">So et al., 2011)</ref>. It may very well be that these failures reflect models' ignorance of intracellular space. Incorporating this feature thus remains a promising, and critical, direction for future research.</p><p>Table I: Notations used in the various gene expression models Variable Definition Equation where first appears M mRNA copy number 1 k m Transcription rate (when not limited by RNA polymerase) 1 &#947; m mRNA degradation rate 2 P Protein copy number 3 k p Translation rate (when not limited by ribosomes) 3 &#947; p Protein degradation rate 4 M Population-averaged mRNA copy number 5 P Population-averaged protein copy number 6 P M (n, t) mRNA copy number distribution 7 &#957;(b) Distribution of number of proteins b in a burst 10 P(x, t) Protein copy number distribution 13 t d Cell doubling time 16 t r Time of gene replication 16 p Protein concentration 17 V Cell volume 17 &#955; Cell growth rate 18 &#954; Protein production rate 18 R Ribosome copy number 19 k tr Rate of translation, per ribosome 19 N RNA polymerase copy number 22 &#934;r Fraction of ribosomes translating ribosomes 19 Mi Effective mRNA copy number (corresponding to gene i) 20 g i Effective gene copy number 21 &#934; i g i / j g j , gene allocation fraction 21 M i mRNA copy number (corresponding to gene i) 22 k tx Rate of transcription, per RNA polymerase 22 M R Ribosomal mRNA copy number 23 &#934; r Ribosome allocation fraction 26 &#934; n RNA polymerase allocation fraction 30 c Total cellular protein concentration 30</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="5" xml:id="foot_0"><p>We will use P to denote both discrete and continuous probability distributions (i.e., probability density functions in the mathematics notation).</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="6" xml:id="foot_1"><p>The fact that the sum of n independent, exponentially-distributed variables is gamma distributed is helpful in evaluating the convolution which appears in Eq. (13), when verifying that Eq. (14) is its stationary solution.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="8" xml:id="foot_2"><p>A caveat is that measuring the statistics of a particular bacterial trait on single lineages (as in Fig.8) is not mathematically equivalent to the doing so within an exponentially growing culture, see Ref.<ref type="bibr">(Levien et al., 2021)</ref> for a recent review, and Ref.<ref type="bibr">(Thomas, 2019)</ref> for results in the context of gene expression.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="9" xml:id="foot_3"><p>This could break down for proteins that are partitioned not according to the ratio of volumes<ref type="bibr">(Lin et al., 2019;</ref><ref type="bibr">Min and Amir, 2021)</ref>.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="10" xml:id="foot_4"><p>For simplicity, we omit the x notation used earlier for ensemble means.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="13" xml:id="foot_5"><p>Note that this need not be the case. In Chapter II we considered an alternative picture where the amount of DNA, rather than RNAP, was limiting. Later in this chapter, we will show how these two limits can be interpolated.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="15" xml:id="foot_6"><p>Note that since multiple cell divisions occur during this time, we should consider the total protein numbers in all the progeny of the initial cell considered.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="19" xml:id="foot_7"><p>It is not immediately clear how to reconcile the picture of a constant protein concentration with the observed mRNA kinetics in bacteria, discussed in Chapter II, which exhibit an increase by twofold or more following gene replication. It is possible that this discontinuity is filtered out at the protein level, or is averaged out at the whole-genome level.</p></note>
		</body>
		</text>
</TEI>
