<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Structural and Practical Identifiability of Phenomenological Growth Models for Epidemic Forecasting</title></titleStmt>
			<publicationStmt>
				<publisher>MDPI</publisher>
				<date>04/01/2025</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10661590</idno>
					<idno type="doi">10.3390/v17040496</idno>
					<title level='j'>Viruses</title>
<idno>1999-4915</idno>
<biblScope unit="volume">17</biblScope>
<biblScope unit="issue">4</biblScope>					

					<author>Yuganthi R Liyanage</author><author>Gerardo Chowell</author><author>Gleb Pogudin</author><author>Necibe Tuncer</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[<p>Phenomenological models are highly effective tools for forecasting disease dynamics using real-world data, particularly in scenarios where detailed knowledge of disease mechanisms is limited. However, their reliability depends on the model parameters’ structural and practical identifiability. In this study, we systematically analyze the identifiability of six commonly used growth models in epidemiology: the generalized growth model (GGM), the generalized logistic model (GLM), the Richards model, the generalized Richards model (GRM), the Gompertz model, and a modified SEIR model with inhomogeneous mixing. To address challenges posed by non-integer power exponents in these models, we reformulate them by introducing additional state variables. This enables rigorous structural identifiability analysis using the StructuralIdentifiability.jl package in JULIA. We validated the structural identifiability results by performing parameter estimation and forecasting using the GrowthPredict MATLAB Toolbox. This toolbox is designed to fit and forecast time series trajectories based on phenomenological growth models. We applied it to three epidemiological datasets: weekly incidence data for monkeypox, COVID-19, and Ebola. Additionally, we assessed practical identifiability through Monte Carlo simulations to evaluate parameter estimation robustness under varying levels of observational noise. Our results confirm that all six models are structurally identifiable under the proposed reformulation. Furthermore, practical identifiability analyses demonstrate that parameter estimates remain robust across different noise levels, though sensitivity varies by model and dataset. These findings provide critical insights into the strengths and limitations of phenomenological models to characterize epidemic trajectories, emphasizing their adaptability to real-world challenges and their role in informing public health interventions.</p>]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Phenomenological growth models are widely used to describe infectious disease dynamics, offering critical insights into the parameters governing their spread and control. A critical aspect for ensuring the reliability of epidemiological modeling is the identifiability of the model parameters <ref type="bibr">[1]</ref><ref type="bibr">[2]</ref><ref type="bibr">[3]</ref><ref type="bibr">[4]</ref><ref type="bibr">[5]</ref>, as it determines whether the values of the parameters can be accurately and uniquely estimated from available data. Structural identifiability analysis</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Structural Identifiability of Phenomenological Models</head><p>Structural identifiability analysis is a theoretical framework used to determine whether the parameters of a model can be uniquely inferred from unlimited, error-free observations. This analysis is essential for ensuring the mathematical integrity of a model before applying it to real-world data, as it confirms that parameter estimates are both unique and reliable. Various techniques, such as differential algebra methods, the Taylor series expansion method, and input-output approaches, have been developed to evaluate structural identifiability <ref type="bibr">[10,</ref><ref type="bibr">12]</ref>.</p><p>In this study, we employ the differential algebra method, a powerful approach that systematically eliminates unobserved state variables and derives differential algebraic polynomials involving the observed variables and model parameters <ref type="bibr">[12,</ref><ref type="bibr">22,</ref><ref type="bibr">23]</ref>. To achieve this, we utilize the StructuralIdentifiability.jl package in JULIA, which enables efficient derivation of these polynomials and facilitates rigorous identifiability analysis <ref type="bibr">[7]</ref>.</p><p>We reformulate the model structure for models with non-integer power exponents by introducing additional state variables. This reformulation not only preserves the fundamental structure of the model but also extends the applicability of identifiability analysis to previously intractable cases. By applying this approach, we ensure that parameters and state variables remain identifiable across models (3), ( <ref type="formula">6</ref>), ( <ref type="formula">9</ref>), <ref type="bibr">(11)</ref>, <ref type="bibr">(13)</ref>, <ref type="bibr">(15)</ref>, broadening the utility of these models in practical epidemic forecasting.</p><p>Structural identifiability theory: To analyze the structural identifiability of models, we rewrite them in the following compact form:</p><p>x &#8242; (t) = f (x, p),</p><p>where x(t) denotes the vector of state variables, y(t) represents the observations, and p denotes the vector of parameters. A parameter p in model ( <ref type="formula">1</ref>) is called structurally identifiable if its value can be uniquely determined from the observable y(t) (assuming noise-free measurements). A similar property for states is typically referred to as observability, but in the context of the present paper, it will be convenient for us to use the term identifiability in both cases. In fact, we will define identifiability for an arbitrary function of parameters and states. We will give a simplified version of the definition to avoid unnecessary technicalities; a formalized version can be found in <ref type="bibr">[24]</ref> (Definition 2.5).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Definition 1.</head><p>A function h(x, p) in the states and parameters of model (1) is said to be structurally globally identifiable if, for generic (x, p) and any ( x, p), the following implication holds: g(x, p) = g( x, p) implies h(x, p) = h( x, p).</p><p>If all the parameters are identifiable, then we will say that the model is identifiable.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Definition 2.</head><p>A function h(x, p) in the states and parameters of model ( <ref type="formula">1</ref>) is said to be structurally locally identifiable if, for generic (x, p), there exists a neighborhood such that for any ( x, p), the following implication holds: g(x, p) = g( x, p) implies h(x, p) = h( x, p).</p><p>While the definitions above reflect the intuitive notion of identifiability, when it comes to checking identifiability for specific models, they are not very convenient. One standard approach to assessing structural identifiability is via input-output equations (also referred to as the differential algebra approach). We will not describe input-output equations in full generality, only for the single-output model, since all the models considered in this paper belong to this class. For a single-output model, the input-output equation is the irreducible equation of minimal order satisfied by the output. We will normalize the equation by dividing by the leading coefficient (considered a polynomial in y and its derivatives) to obtain a monic polynomial. Let c 1 (p), . . . , c &#8467; (p) be the coefficients of this equation. Then, under a certain assumption on the Wronskian of the input-output equations (see, e.g., <ref type="bibr">[25]</ref> (Lemma 4.6)), the identifiable functions of the parameters are precisely the ones expressible in terms of c 1 , . . . , c &#8467; . In this study, we used the software StructuralIdentifiability.jl version v1.10.2 to assess identifiability; the software automatically checks if the aforementioned condition on Wronskians is fulfilled.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the generalized growth model</head><p>We derive input-output equations by eliminating unobserved state variables using the differential algebra method. This process generates algebraic relationships involving the observed variables and model parameters, enabling identifiability analysis. For example, the generalized growth model is defined by the following differential equation:</p><p>where C &#8242; (t) denotes incidences at time t, C(t) denotes the cumulative number of cases at time t, and &#945; denotes the growth rate of the infectious diseases such that 0 &#8804; &#945; &#8804; 1. Since epidemic data are typically collected as incident case counts over discrete time intervals, we assume that the incidence, C &#8242; (t), corresponds to the observed data. This assumption is made for methodological consistency and computational feasibility, as it simplifies the identifiability analysis by ensuring that our model structure directly aligns with real-world epidemic datasets. It also allows us to use a standardized observation function across different growth models, facilitating direct comparisons in both structural and practical identifiability assessments.</p><p>To facilitate this process, we introduce an additional state variable, which allows us to reformulate models with non-integer power exponents into a structure that can be analyzed using the differential algebra approach. For this purpose, we introduce an additional state variable x(t) to eliminate the non-integer power exponent, resulting in the extended version of the model by letting x(t) = rC &#945; (t); then, GGM becomes</p><p>In this study, we examine the scenario in which the observations are equal to the incidences; that is, y(t) = x(t). The monic input-output equation obtained from the StructuralIdentifiability.jl package in JULIA, with the observation y(t), is given by 0</p><p>In the context of structural identifiability, the input-output equation indicates that 2 -1/&#945; can be identified with the observation y(t), so &#945; is identifiable as well. Now, we can express C(t) = &#945; x 2 (t)</p><p>x &#8242; (t) in terms of identifiable parameters and state variables. As a result, C(t) becomes identifiable. This implies that the parameter r is identifiable. Therefore, model ( <ref type="formula">2</ref>) is structurally identifiable with the observation y(t). We state the following proposition: Proposition 1. The generalized growth model GGM is structurally identifiable from the observation of incidences y = x(t).</p><p>We would like to stress that the way the lifting to the extended model ( <ref type="formula">3</ref>) is performed is important to obtain correct identifiability results for the original model. Suppose that, instead of x(t) = rC &#945; (t), we would have introduced x(t) = C &#945; (t). Then, we would obtain a different extended system:</p><p>In this model, neither r nor x are identifiable because there exists an output-preserving transformation r &#8594; &#955;r, x &#8594; x &#955; for any nonzero number &#955;. The discrepancy between the identifiability results for different extended models can be explained as follows. The trajectories of the original versions of model (2) and of (3) are parametrized by three numbers: &#945;, r, C(0) for the former, and &#945;, C(0), x(0) for the latter. On the other hand, the space of trajectories of ( <ref type="formula">4</ref>) is four-dimensional, parametrized by &#945;, r, C(0), x(0). The images of the trajectories of (2) among the trajectories of (4) are constrained to a manifold x(t) -C &#945; (t) = 0. Thus, this is the additional trajectories that (4) possesses, which turn r into nonidentifiable status.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the generalized logistic growth model</head><p>The generalized logistic growth model is given by the following equation:</p><p>where r is the generalized growth rate, k is the final epidemic size, C &#8242; (t) denotes incidences at time t, C(t) denotes the cumulative number of cases at time t, and the parameter &#945; &#8712; [0, 1] denotes the different growth scenarios; the constant incidents is &#945; = 0, sub-exponential growth is 0 &lt; &#945; &lt; 1, and exponential growth is &#945; = 1. To obtain the extended version of the model without a non-integer exponent, we substitute x(t) = rC &#945; (t), where GLM becomes</p><p>Here, we consider the case where the observations y(t) = x(t) 1 -C(t) k correspond to the incidences. The input-output equation obtained from JULIA is normalized, with the observation y(t) given by 0 = y &#8242;&#8242;2 y 2 + y &#8242;&#8242; y &#8242;2 y -3&#945;</p><p>We will check that all the parameters can be expressed in terms of the coefficients of this equations, thus showing that they are identifiable. First, &#945; can be expressed from 2&#945;-1 &#945; , so it is identifiable. Next, k can be expressed from &#945; and the coefficient -&#945; 2 +1 &#945;k . This concludes that only parameters k and &#945; are identifiable.</p><p>By using the StructuralIdentifiability.jl package in JULIA, we obtained the state variable x(t), and C(t) is identifiable (observable) from the given observation. Then, r = x(t) C &#945; (t) can be written as a combination of identifiable parameters and the state variables. Therefore, the parameter r is also identifiable. We assert the following proposition: Proposition 2. The generalized logistic growth model (5) is structurally identifiable from the observation of incidences, y(t) = x(t) 1 -C(t) k .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the Richards model</head><p>The Richards model is given by the following equation:</p><p>We let x(t) = ( C(t) k ) a ; the Richards model becomes</p><p>Here, we consider the case where the observations, y(t) = rC(t)(1x(t)), correspond to the incidences. The normalized input-output equation of the model ( <ref type="formula">9</ref>) is given by</p><p>The parameter a can be expressed from the coefficient 2a+1 a+1 = 2 -1 a+1 . Next, parameter r can be expressed from a and coefficient ar(a-1) a+1 . Thus, both a and r are identifiable from the observation y(t). Now, we check the identifiability of the state variables; both x(t) and C(t) are identifiable from the observation. Therefore, parameter a can be written as a combination of the identifiable parameters and state variables. Thus, a is identifiable. We state the following proposition: Proposition 3. The Richards model (8) is structurally identifiable from the observation of incidences, y(t) = rC(t)(1x(t)).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the generalized Richards model</head><p>The model is represented by the following equation:</p><p>where r is the generalized growth rate, k is the final epidemic size, C &#8242; (t) denotes incidences at time t, C(t) denotes the cumulative number of cases at time t, and the parameter &#945; &#8712; [0, 1] denotes the different growth scenarios; the constant incidents is &#945; = 0, sub-exponential growth is 0 &lt; &#945; &lt; 1, and exponential growth is &#945; = 1, and the exponent a denotes the deviation from the symmetric s-shaped dynamics of the simple logistic curve.</p><p>To obtain the extended version without any non-integer power exponent, we let x(t) = rC p (t) and z(t) = (C/k) a (t), where GRM becomes</p><p>By using StructuralIdentifiability.jl in JULIA, we determined that the parameters a and &#945; are locally identifiable, and the state variables x(t) and z(t) are also locally identifiable. Furthermore, the product a 2 and the summation a + 2&#945; are globally identifiable. Since a is positive, it follows that a is globally identifiable, ensuring that &#945;, x(t), and z(t) are globally identifiable as well. To obtain the identifiability of parameters r and k, we express</p><p>in terms of identifiable states and parameters. Therefore, the parameters r and k are identifiable. We conclude the following proposition:</p><p>Proposition 4. The generalized Richards model <ref type="bibr">(10)</ref> is structurally identifiable from the observation of incidences y(t) = x(t)(1z(t)) under the assumption of the positivity of a and k.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the Gompertz model</head><p>The Gompertz model is given by the following equation:</p><p>When letting x(t) = re -at , GOM becomes</p><p>Here, we consider the case where the observations, y(t), correspond to the incidences,</p><p>The input-output equation of model ( <ref type="formula">13</ref>) is given by</p><p>Since a appears among the coefficients, it is identifiable. The StructuralIdentifiability package in JULIA shows that both C(t) and x(t) are identifiable from the observation y(t). Thus, the parameter r = x(t)e at can be written as a product of identifiable parameters and state variables. Thus, r is identifiable.</p><p>Proposition 5. The model ( <ref type="formula">12</ref>) is structurally identifiable from the observation of incidences, y(t) = C(t)x(t).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Structural identifiability results for the SEIR model with inhomogeneous mixing</head><p>The SEIR model with inhomogeneous mixing is given by the following system of equations:</p><p>Here, we consider the case where the observations are y(t) = &#954;E(t). By using StructuralIdentifiability.jl in JULIA, we determined that the parameters &#945;, &#947;, and &#954; are globally identifiable, whereas N is not identifiable. However, since N represents the total population size of a closed population, its value is known, and</p><p>With N being a known quantity, we treat it as an additional observation in the model. We then perform the identifiability analysis again and obtain the fact the state variable x(t) is identifiable from observations y(t) and N. Thus, we can express &#946; = x(t) I &#945; as a combination of identifiable parameters and state variables. Thus, model ( <ref type="formula">14</ref>) is structurally identifiable. The summary of the structural identifiability results for models (3), ( <ref type="formula">6</ref>), ( <ref type="formula">9</ref>), ( <ref type="formula">11</ref>), <ref type="bibr">(13)</ref>, and <ref type="bibr">(15)</ref>, obtained from StructuralIdentifiability.jl in JULIA, is given in Table <ref type="table">1</ref>.</p><p>Cases with time-varying N(t) introduce additional degrees of freedom into the model, potentially leading to the unidentifiability of key parameters unless additional constraints or external data sources (e.g., demographic data) are available. Future work could explore methods such as including auxiliary equations for N(t) or assuming known functional forms to restore identifiability.</p><p>Another critical challenge in real-world applications is the uncertainty in initial conditions. Since initial values for S(0), E(0), I(0), and R(0) are often unknown or estimated from limited data, they can significantly impact parameter identifiability and estimation accuracy. Given these considerations, we consider analyses with and without knowing the initial conditions of the systems. Proposition 6. The model ( <ref type="formula">14</ref>) is structurally identifiable from the observation of incidences, y(t) = &#954;E(t), with known initial conditions. Table <ref type="table">1</ref>. Summary of structural identifiability results. This table summarizes the structural identifiability results for the model parameters and the observability of state variables across six models: the generalized growth model (GGM), the generalized logistic model (GLM), the Richards model, the generalized Richards model (GRM), the Gompertz model (GOM), and the SEIR model with inhomogeneous mixing. The results are presented for cases with and without known initial conditions. The parameters marked as "globally identifiable" can be uniquely determined from perfect data, whereas "locally identifiable" parameters require specific constraints for unique estimation. The observability of state variables, which indicates whether they can be inferred from the data, is also included. These findings were obtained using the StructuralIdentifiability.jl package in JULIA. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Model</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Fitting Models to Real Epidemic Data</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Data</head><p>We performed parameter estimation using the GrowthPredict MATLAB Toolbox, a robust platform developed explicitly for fitting and forecasting time series trajectories based on phenomenological growth models using ordinary differential equations <ref type="bibr">[13]</ref>. The models outlined in Equations ( <ref type="formula">5</ref>) to <ref type="bibr">(14)</ref> were fitted to three sets of data: the weekly incidence curve of monkeypox data, COVID-19 data, and Ebola data, as described in the Data section. The computation of prediction intervals assumes that observed values are generated from the deterministic model with added normal error (mean zero and estimated variance). This assumption allows us to quantify uncertainty while maintaining consistency with the Monte Carlo framework used for practical identifiability analysis.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Parameter Estimation Method</head><p>We used the least squares (LSQ) method to perform the parameter estimation problem by minimizing the following objective function.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>pi = argmin</head><p>Here, f (t j , p i ) is the predicted epidemic trajectory given by model i, where i = {2, 3, 4, 5, 6}, and y t j is the time series data, where t j , j = 1, 2, 3, . . . , n are the discrete time points for the time series data. This method was chosen for its computational efficiency and ability to handle nonlinear parameter spaces effectively.</p><p>To evaluate model performance, we used the GrowthPredict MATLAB Toolbox, which provides several goodness-of-fit metrics, including the Akaike Information Criterion (AIC), mean absolute error (MAE), mean squared error (MSE), and weighted interval score (WIS) <ref type="bibr">[13]</ref>. The corresponding formulations are as follows.</p><p>Akaike information criterion (AIC) is computed as</p><p>where m is the number of model parameters, and n is the number of data points (calibration period). The sum of squared errors SSE is defined as</p><p>Mean absolute error (MAE) and mean squared error (MSE) are given by</p><p>The weighted interval score (WIS) is computed as</p><p>where K is the number of prediction intervals (PIs), &#7929; is the predictive median, and w k = &#945; k 2 for k = 1, 2, . . . , K, with w 0 = 1/2. The interval scoring function is defined as</p><p>where 1 (y &lt; l) is the indicator function that equals 1 if y &lt; l, and 0 otherwise. The terms l and u represent the &#945; 2 and 1 -&#945; 2 quantiles of the forecast F, respectively. The coverage of the 95% prediction interval is given by</p><p>where L i and U i are the lower and upper bounds of the 95% PIs, respectively, y t i represents the data, and 1 is an indicator variable that equals 1 if y t i is in the specified interval and is 0 otherwise. The estimated parameter values from fitting three sets of time series data (monkeypox, COVID-19, and Ebola) to models (5)-( <ref type="formula">14</ref>) are given in Tables <ref type="table">2</ref>, <ref type="table">3</ref>, and 4, respectively.    The figure illustrates the alignment of the GLM with the observed data, showing its ability to capture the epidemic dynamics effectively. Key features of the GLM, such as its flexibility to model sub-exponential growth, are evident in the close agreement between the predicted trajectory and the observed data points. The model parameters were estimated using the least squares method, and the fit was evaluated based on the Akaike Information Criterion (AIC), mean absolute error (MAE), and mean squared error (MSE). In the bottom panel, the solid red line is the median model fit and the dashed lines correspond to the 95% prediction intervals. The blue dots indicate the observed data points. The gray lines correspond to the mean of the model fits obtained from the parametric bootstrapping with 1000 bootstrap realizations, and the shaded region indicates the 95% prediction intervals (PIs), highlighting the model's ability to quantify uncertainty in its projections.   </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Practical Identifiability</head><p>Structural identifiability examines the feasibility of estimating parameters in a model with unlimited observation without measurement errors. However, real-world data are often collected discretely and may include significant measurement errors. Thus, the structural identifiability of a model does not imply the practically identifiable parameters. So, it becomes crucial to assess whether structurally identifiable parameters can be accurately estimated from data containing noise. Various methods are available to examine practical identifiability <ref type="bibr">[1,</ref><ref type="bibr">10,</ref><ref type="bibr">12,</ref><ref type="bibr">23,</ref><ref type="bibr">[26]</ref><ref type="bibr">[27]</ref><ref type="bibr">[28]</ref>. In this study, we employed the Monte Carlo simulation (MCS) method, described below <ref type="bibr">[1,</ref><ref type="bibr">10,</ref><ref type="bibr">12,</ref><ref type="bibr">23]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>&#8226;</head><p>We solved the models' Equations ( <ref type="formula">5</ref>), ( <ref type="formula">8</ref>), ( <ref type="formula">10</ref>), <ref type="bibr">(12)</ref>, and ( <ref type="formula">14</ref>), numerically using the true parameter values p obtained from the fitting of the model to experimental data from the data fitting section. Then, we obtained the observations (output) at the experimental time points.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>&#8226;</head><p>We generated M = 1000 datasets at each of the experimental time points, with four different levels of increasing measurement error, &#963; = 1%, 5%, 10%, 20%, by adding measurement errors to each given experimental data point. Each measurement error is assumed to be distributed in a normal distribution with zero mean and standard deviation &#963;.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>&#8226;</head><p>We fit the model to each of the 1000 datasets to estimate the parameter set pi using the objective function Equation ( <ref type="formula">16</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>&#8226;</head><p>We calculated the average relative estimation error (ARE) for each parameter in the model according to</p><p>where p(k) is the kth element of the true parameter set p, and</p><p>We used AREs values to determine whether each of the model's parameters is practically identifiable, using Definition 3 given below.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Definition 3.</head><p>Let ARE be the average relative estimation error of the parameter p (k) . Let &#963; be the measurement error.</p><p>A model is said to be practically identifiable if the parameters p (k) are practically identifiable for all k.</p><p>In the data fitting section, we fitted the model equations (Equations ( <ref type="formula">5</ref>) through ( <ref type="formula">14</ref>)) to three different datasets: monkeypox data, COVID-19 data, and Ebola data. Here, we focus on assessing the practical identifiability of each model using the data set that produced the best fit based on the lowest AIC values. Thus, our investigation focuses on determining whether the Richards model, GLM, and GRM demonstrate practical identifiability when applied to the monkeypox, Ebola, and COVID-19 datasets, respectively. Additionally, we evaluated the GRM, GOM, and SEIR models on the datasets where they demonstrated the best relative performance to further assess their practical identifiability.</p><p>We found that all models, Richards, GRM, GOM, GLM, and SEIR, are practically identifiable from the given datasets. Each model's parameters were reliably estimated, with well-constrained confidence intervals, indicating that the data provided sufficient information to determine the parameters uniquely. Moreover, these results are consistent with the structural identifiability findings, further validating the models' capacity to capture the dynamics of the epidemics studied. The results of our analysis are presented in Table <ref type="table">5</ref> for the Richard model, Tables <ref type="table">6</ref> and <ref type="table">7</ref> for GRM, Table <ref type="table">8</ref> for GLM, Table <ref type="table">9</ref> for GOM, and Table <ref type="table">10</ref> for SEIR.</p><p>Table 5. Results of Monte Carlo simulations assessing the practical identifiability of the Richards model (Equation ( <ref type="formula">8</ref>)) using virtual datasets generated for discrete monkeypox experimental data points. The table presents the average relative estimation errors (AREs) for each model parameter (r, a, and k) under varying noise levels (&#963; = 1%, 5%, 10%, and 20%) alongside the corresponding confidence intervals (CIs). These results demonstrate the model's robustness in parameter estimation across different levels of observational noise.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Parameter</head><p>r a k ARE &#963; = 1% 2.8867 &#215; 10 -5 3.4759 &#215; 10 -5 4.1286 &#215; 10 -7 CI [0.9200, 0.9200] [0.3600, 0.3600] [29,000.0000, 29,000.0000] ARE &#963; = 5% 8.5510 &#215; 10 -4 0.0015 2.7180 &#215; 10 -4 CI [0.9200, 0.9200] [0.3600, 0.3600] [28,999.6373, 29,000.3469] ARE &#963; = 10% 0.0020 0.0036 9.1805 &#215; 10 -4 CI [0.9200, 0.9200] [0.3600, 0.3600] [28,999.2199, 29,000.7637] ARE &#963; = 20% 0.0042 0.0074 0.0021 CI [0.9199, 0.9201] [0.3599, 0.3601] [28,998.3931, 29,001.5577]     </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Discussion</head><p>This study evaluates the structural and practical identifiability of six commonly used phenomenological growth models in epidemiology. It demonstrates their robustness in parameter estimation and applicability across diverse datasets. To assess structural identifiability, we reformulated the models to address challenges posed by non-integer power exponents; we created an extended structure with fewer parameters and additional equations while preserving the original degrees of freedom. Our findings confirm that all reformulated models are structurally identifiable, even with unknown initial conditions. The original model formats were applied for data fitting and practical identifiability analyses, revealing that all models remained practically identifiable under varying noise levels, with performance differing across datasets. The first step in validating these models involved assessing whether all unknown parameters were structurally identifiable, ensuring they could theoretically be uniquely determined from perfect, unlimited data. This step establishes a necessary foundation for reliable parameter estimation. We used the differential algebra method for this evaluation <ref type="bibr">[12,</ref><ref type="bibr">22,</ref><ref type="bibr">23]</ref>. Given the lack of dedicated software tools for analyzing systems of ordinary differential equations with non-integer exponents, we reformulated the models by introducing additional state variables. This reformulation enabled the use of the StructuralIdentifiability.jl package in Julia <ref type="bibr">[7]</ref>. Our findings demonstrated that all parameters in the reformulated models were structurally identifiable, even with unknown initial conditions. By assessing the observability of state variables, we inferred the structural identifiability of the original models, bridging the gap between the reformulated and original structures <ref type="bibr">[8,</ref><ref type="bibr">9]</ref>.</p><p>Model validation was conducted using the GrowthPredict MATLAB Toolbox, a specialized tool for fitting and forecasting time series trajectories based on phenomenological growth models. This toolbox facilitated comparisons across datasets and enabled the calculation of performance metrics such as AIC, MAE, MSE, and WIS. This toolbox was applied to three epidemiological datasets: weekly incidence data for monkeypox, COVID-19, and Ebola <ref type="bibr">[13]</ref>. To compare performance, metrics such as AIC, MAE, MSE, WIS, and coverage were calculated for each model. The selection of models based on AIC values revealed that certain models are better suited to specific epidemic contexts. For example, the Richards model provided the best fit for monkeypox data. In contrast, the GRM model excelled at fitting the COVID-19 data. The GLM model was most effective for the Ebola data.</p><p>In the next phase, we assessed practical identifiability by applying the models to datasets that produced the best fits, as indicated by the AIC values. We focused on the Richards, SEIR, and Gompertz models with the monkeypox dataset, the GLM with the Ebola dataset, and the GRM with the COVID-19 dataset. Practical identifiability was evaluated using Monte Carlo simulations, highlighting the robustness of parameter estimation under real-world conditions, where data are often noisy and limited. Other approaches, such as the Fisher Information Matrix (FIM), Profile Likelihood Method, and Bayesian methods, could complement this analysis <ref type="bibr">[1,</ref><ref type="bibr">10,</ref><ref type="bibr">12,</ref><ref type="bibr">23]</ref>. Our results showed that all the models were practically identifiable for their respective datasets. However, parameter estimation accuracy varied across models and datasets, emphasizing the sensitivity of these models to data quality. These findings underscore the importance of including practical identifiability analysis in epidemiological studies to ensure robust model validation.</p><p>By addressing the challenge of non-integer power exponents and validating the models under practical conditions, this study broadens the applicability of phenomenological models. It strengthens confidence in their use for epidemic forecasting. However, several limitations must be acknowledged. First, while introducing additional state variables facilitates structural identifiability analysis, further research is needed to assess its impact on computational efficiency and model interpretability. Second, the dependency of parameter estimation accuracy on data quality remains a challenge, emphasizing the need for high-quality, well-calibrated datasets. Third, alternative approaches such as Bayesian methods and Fisher Information Matrix analysis should be explored to complement the Monte Carlo simulations performed here. Finally, future research should examine the role of time-varying parameters and unknown initial conditions in shaping identifiability outcomes.</p><p>Our analysis highlights the structural and practical identifiability of six phenomenological models across various epidemiological datasets. The results underscore the utility of these models in forecasting disease dynamics and their adaptability to diverse epidemic contexts. However, the accuracy of parameter estimates is highly dependent on data quality, emphasizing the critical need for careful data collection and the integration of practical identifiability evaluations in the model selection process.</p></div></body>
		</text>
</TEI>
