<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Extended Convex Lifting for Policy Optimization of Optimal and Robust Control</title></titleStmt>
			<publicationStmt>
				<publisher>Proceedings of Machine Learning Research</publisher>
				<date>06/04/2025</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10631415</idno>
					<idno type="doi"></idno>
					
					<author>Yang Zheng</author><author>Chih-Fan Pai</author><author>Yujie Tang</author><author>Necmiye Ozay</author><author>Laura Balzano</author><author>Dimitra Panagou</author><author>Alessandro Abate</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Many optimal and robust control problems are nonconvex and potentially nonsmooth in their policy optimization forms. In this paper, we introduce the Extended Convex Lifting (ECL) framework, which reveals hidden convexity in classical optimal and robust control problems from a modern optimization perspective. Our ECL framework offers a bridge between nonconvex policy optimization and convex reformulations. Despite non-convexity and non-smoothness, the existence of an ECL for policy optimization not only reveals that the policy optimization problem is equivalent to a convex problem, but also certifies a class of first-order non-degenerate stationary points to be globally optimal. We further show that this ECL framework encompasses many benchmark control problems, including LQR, state-feedback and output-feedback H-infinity robust control. We believe that ECL will also be of independent interest for analyzing nonconvex problems beyond control.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>The classical optimal and robust control problems, including linear quadratic regulator (LQR), linear quadratic Gaussian (LQG) control, and H &#8734; control, have been extensively studied <ref type="bibr">(Kalman, 1963;</ref><ref type="bibr">Levine and Athans, 1970)</ref>. It is well known that almost all of these problems are nonconvex in the space of controller (i.e., policy) parameters. Nevertheless, classical techniques based on controller reparameterizations <ref type="bibr">(Scherer and Weiland, 2015;</ref><ref type="bibr">Boyd et al., 1994)</ref> or Riccati equations <ref type="bibr">(Zhou et al., 1996)</ref> have been established to characterize optimal or suboptimal controllers. These classical techniques do not optimize over the policy parameters directly, and often require an explicit system model. In contrast, the optimization landscapes of optimal and robust control problems offer a fruitful alternative perspective. In this approach, the control costs are viewed as functions of the policy parameters, and their analytical and geometrical properties are studied <ref type="bibr">(Lewis, 2007;</ref><ref type="bibr">Hu et al., 2023;</ref><ref type="bibr">Talebi et al., 2024)</ref>. This perspective is naturally amenable to data-driven design paradigms such as reinforcement learning and learning-based control <ref type="bibr">(Recht, 2019)</ref>.</p><p>However, this policy optimization perspective for control generally leads to nonconvex and potentially nonsmooth problems. For example, the set of feedback gains K that stabilize the system &#7819; = Ax + Bu via u = Kx is already nonconvex; if we consider output-feedback controller synthesis such as LQG, then the parameterized set of dynamic policies can even be disconnected <ref type="bibr">(Tang et al., 2023)</ref>. Furthermore, the LQG cost function may have spurious stationary points <ref type="bibr">(Zheng et al., 2022)</ref>, and there may exist uncountably many globally optimal policies lying on a manifold induced by similarity transformations <ref type="bibr">(Zheng et al., 2022;</ref><ref type="bibr">Tang et al., 2023;</ref><ref type="bibr">Kraisler and Mesbahi, 2024)</ref>.</p><p>In addition to non-convexity, non-smoothness may also arise when considering robust control problems. A typical performance measure for robust control is the H &#8734; norm of a certain closed-loop transfer function <ref type="bibr">(Zhou et al., 1996)</ref>, which is known to be both nonconvex and nonsmooth in the policy space <ref type="bibr">(Apkarian and Noll, 2006;</ref><ref type="bibr">Lewis, 2007)</ref>.</p><p>For nonconvex and nonsmooth optimization, it is generally very hard to establish theoretical guarantees for local search algorithms. Nevertheless, recent findings have revealed benign nonconvex landscape properties in several benchmark control problems, including LQR <ref type="bibr">(Fazel et al., 2018;</ref><ref type="bibr">Mohammadi et al., 2022;</ref><ref type="bibr">Fatkhullin and Polyak, 2021)</ref>, risk-sensitive control <ref type="bibr">(Zhang et al., 2021)</ref>, LQG <ref type="bibr">(Tang et al., 2023;</ref><ref type="bibr">Zheng et al., 2022;</ref><ref type="bibr">Duan et al., 2024)</ref>, dynamic filtering <ref type="bibr">(Umenberger et al., 2022;</ref><ref type="bibr">Zhang et al., 2023)</ref>, H &#8734; control <ref type="bibr">(Hu and Zheng, 2022;</ref><ref type="bibr">Guo and Hu, 2022;</ref><ref type="bibr">Tang and Zheng, 2023)</ref>, and distributed control <ref type="bibr">(Furieri et al., 2020a)</ref>. Many of these works leveraged the idea that the control problems under investigation admit suitable convex reformulations. However, these existing works are mostly on a case-by-case basis. Our work aims to provide a unified framework that explains the benign nonconvex landscape properties of these iconic control problems.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Our Contributions -Extended Convex Lifting (ECL)</head><p>We introduce a unified framework, called Extended Convex Lifting (ECL), to reveal hidden convexity in classical optimal and robust control problems from a modern optimization perspective. The core idea behind ECL stems from the existing results in control theory that, via a suitable change of variables, many optimal and robust control problems admit "convex reformulations" <ref type="bibr">(Scherer and Weiland, 2015;</ref><ref type="bibr">Boyd et al., 1994)</ref>. By fitting the change of variables into the ECL framework, we can analyze the nonconvex landscape of the corresponding policy optimization problem by exploiting its hidden convexity. Figure <ref type="figure">1</ref> provides a schematic illustration of ECL. Specifically, we begin with the epigraph of the objective J(K), and lift a certain subset of its closure to a higher-dimensional set L lft by incorporating Lyapunov variables. Then, we construct a smooth bijection &#934; that maps L lft to a product space F cvx &#215; G aux where F cvx is convex. This bijection encodes the change of variables associated with the control problem. In many control problems, F cvx can be represented by linear matrix inequalities (LMIs), while G aux accounts for similarity transformations of output-feedback policies.</p><p>Despite non-convexity and non-smoothness, the existence of an ECL not only shows that the policy optimization problem is equivalent to a convex problem (Theorem 2.1) but also identifies a class of non-degenerate stationary points that are globally optimal (Theorem 3.1). Our ECL framework covers many iconic control problems; many recent results on the global optimality of (nondegenerate) stationary points, such as LQR <ref type="bibr">(Fazel et al., 2018;</ref><ref type="bibr">Mohammadi et al., 2022)</ref>, LQG <ref type="bibr">(Tang et al., 2023)</ref>, state-feedback H &#8734; control <ref type="bibr">(Guo and Hu, 2022)</ref>, output-feedback H &#8734; control <ref type="bibr">(Tang and Zheng, 2023)</ref>, are special cases once the corresponding ECL is constructed.</p><p>We point out that our ECL framework is more general than <ref type="bibr">Tang et al. (2023)</ref>; <ref type="bibr">Umenberger et al. (2022)</ref>; <ref type="bibr">Guo and Hu (2022)</ref>; <ref type="bibr">Sun and Fazel (2021)</ref>; <ref type="bibr">Mohammadi et al. (2022)</ref>, as it can directly handle both state-feedback and output-feedback policies, as well as smooth and nonsmooth cost functions. More importantly, the ECL framework naturally distinguishes between degenerate and non-degenerate policies, reflecting the subtleties between strict and non-strict LMIs in control. Due to space limitations, we omit most proofs in this paper; detailed proofs, additional discussions, and illustrative examples can be found in our extended reports <ref type="bibr">(Zheng et al., 2023</ref><ref type="bibr">(Zheng et al., , 2024))</ref>.</p><p>2. Extended Convex Lifting (ECL) for Benign Non-convexity</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">A Motivating Example</head><p>We first provide a motivating example. Consider an LTI system &#7819;(t) = Ax(t)+Bu(t)+w(t) with</p><p>, where w(t) is a white Gaussian noise with E[w(t)w(&#964; )] = 4&#948;(t-&#964; )I. We aim to design a state-feedback policy</p><p>Standard calculations show that the cost is</p><p>Following the classical change of variables Y = KX where X solves the Lyapunov equation (A + BK)X + X(A + BK) T + 4I 2 = 0, we obtain the nonlinear mapping</p><p>One can check that g is invertible, and the cost function after applying this mapping becomes</p><p>By standard techniques in convex analysis, one can show that h(Y ) is convex. Thanks to the smooth bijection g, minimizing J(K) is now equivalent to minimizing the convex function h(Y ). As a result, any stationary point of J(K) is globally optimal, and in fact, can be shown to be unique. This motivating example demonstrates how one can utilize a proper change of variables to certify the global optimality of stationary points via convex analysis for policy optimization. In the next subsection, we propose a general framework for nonconvex policy optimization that can accommodate a much wider range of benchmark problems admitting convex reformulations.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">The Extended Convex Lifting (ECL) Framework</head><p>Consider a policy optimization problem where the objective is f : D &#8594; R, with D &#8838; R d being its domain. To study the landscape of f , we resort to its strict and non-strict epigraphs defined by</p><p>, the mapping &#934; directly outputs &#947; in the first component).</p><p>The notion of ECL gives extensive flexibility by 1) adding an extra variable to lift the epigraph to a higher dimension, 2) relaxing the projection of the lifted set to sit between the strict epigraph and the closure of the non-strict epigraph (chain of inclusion), and 3) introducing an auxiliary set G aux to extend the convex image under a diffeomorphism. The extra variable &#958; often corresponds to Lyapunov variables, and G aux is often related to similarity transformations of dynamic policies.</p><p>The interested reader may wonder why we need such a peculiar chain of inclusion in (2a). A simpler and more straightforward requirement for the lifting process might be</p><p>(3)</p><p>Evidently, (2a) includes (3) as a special case. In Section 4, we will present an ECL for outputfeedback H &#8734; control where the more general (2a) is necessary, which is largely due to the intricacy between strict and non-strict LMIs in the convex reformulations of control problems. This intricacy is important for global optimality, but has been less emphasized before since classical results often focused on suboptimal controller design; see <ref type="bibr">(Zheng et al., 2023</ref><ref type="bibr">(Zheng et al., , 2024) )</ref> for details. An immediate benefit of ECL is that we can reformulate the minimization of f (x) over x &#8712; D as a convex problem.</p><p>Theorem 2.1 Let f : D &#8594; R be continuous and equipped with an</p><p>Thanks to the diffeomorphism &#934;, the proof is straightforward but requires careful reasoning about (non-)strict epigraphs; see <ref type="bibr">Zheng et al. (2024, Theorem 3</ref>.1) for a detailed proof. In Theorem 2.1, the function f can be nonsmooth and nonconvex, but the existence of an ECL reveals its hidden convexity in the sense that optimizing f (x) over x &#8712; D is equivalent to a convex problem. Theorem 2.1 provides the rationale behind convex re-parameterizations of many control problems <ref type="bibr">(Scherer and Weiland, 2015;</ref><ref type="bibr">Boyd et al., 1994)</ref>. Note that we only guarantee an infimum instead of a minimum in Theorem 2.1 (the infimum may not always be achieved in H &#8734; control).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Non-degenerate Policies and Global Optimality 2</head><p>In addition to convex reformulation as shown in Theorem 2.1, the existence of an ECL can further reveal global optimality of certain first-order stationary points for the potentially nonconvex and nonsmooth function f . This allows us to optimize f (x) by direct local search without knowing the particular form of the ECL, which is particularly important for learning-based model-free control.</p><p>Before proceeding, we re-emphasize that the chain of inclusion (2a) is critical to the construction of ECL in many control problems. This chain of inclusion allows existence of points (x, f (x)) that are not covered by &#960; x,&#947; (L lft ), as well as members of &#960; x,&#947; (L lft ) that are only accumulation points of epi &#8805; (f ). We introduce the notion of (non-)degeneracy to characterize the former type of points.</p><p>The set of non-degenerate points in D will be denoted by D nd .</p><p>In Section 4.1, we will show that all stabilizing state-feedback policies for LQR and H &#8734; control are non-degenerate using standard ECL constructions. We will also see that for output-feedback H &#8734; control, the set of non-degenerate policies defined in our previous work <ref type="bibr">(Tang and Zheng, 2023)</ref> corresponds to the set of non-degenerate points per Definition 3.1 under the ECL framework.</p><p>We now present one main technical result of this paper, which provides global optimality certificates for stationary points that are non-degenerate. </p><p>The proof of Theorem 3.1 has a strong geometric intuition, but the details are technically involved, which are given in <ref type="bibr">Zheng et al. (2024, Section 3.3</ref>). Theorem 3.1 guarantees that nondegenerate stationarity implies global optimality for any subdifferentially regular function with an ECL. Subdifferentially regular functions are a very large class of functions, covering all optimal and robust control problems discussed in Section 4.</p><p>We now provide a corollary for the case where (3) holds for the ECL. As mentioned before, many state-feedback and full-order output-feedback controller synthesis problems are nonconvex in their natural forms but admit "convex reformulations" in terms of LMIs using a suitable change of variables. We argue that our notion of ECL presents a unified treatment for many of these convex reformulations. In Section 4, we will present some ECL construction details for benchmark optimal and robust control problems. We point out that, despite the wide use of convex reformulations in control, exact constructions of ECL require special care, especially for output-feedback control problems (due to the subtleties between strict and non-strict LMIs).</p><p>Remark 3.1 (Degenerate points and saddles) By (2a), &#960; x,&#947; (L lft ) may not cover the whole nonstrict epigraph. As a result, Theorem 3.1 does not provide global optimality guarantees for degenerate stationary points x &#8712; D\D nd since they cannot be covered by the convex reparameterization. 4 Suboptimal saddle points for f might exist even when equipped with an ECL. Indeed, it has been revealed that LQG policy optimization has strictly sub-optimal saddle points <ref type="bibr">(Tang et al., 2023, Thereom 5</ref>) <ref type="bibr">(Zheng et al., 2022, Theorem 2)</ref>, which are all degenerate per Definition 3.1.</p><p>3. Subdifferential regularity allows us to relate the Clarke subdifferential with ordinary directional derivatives. 4. In this sense, some classical LMI reformulations are not "equivalent" to the original control problems, especially for output-feedback cases. </p><p>Notations: A K,X := (A + BK)X + P (A + BK) T ; A t K,P := (A + BK) T P + P (A + BK); A X,Y := AX + BY + (AX + BY ) T .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Applications in Optimal and Robust Control</head><p>In this section, we present the ECL constructions for several benchmark optimal and robust control problems. Theorems 2.1 and 3.1 can then be directly applied to their policy optimization formulations. Due to space limitations, we omit the mathematical justifications of these ECL constructions, and refer interested readers to <ref type="bibr">Zheng et al. (2023</ref><ref type="bibr">Zheng et al. ( , 2024) )</ref> for detailed proofs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">State-Feedback Policy Optimization</head><p>Consider a continuous-time LTI system</p><p>where x(t) &#8712; R n is the state variable, u(t) &#8712; R m is the control input, and w(t) &#8712; R n is the disturbance on the system process. We introduce B w &#8712; R n&#215;n for a unified treatment of LQR and H &#8734; control in this section. z(t) represents the performance signal, where Q &#10928; 0 and R &#8827; 0. We assume (A, B) is controllable and (Q 1/2 , A) is observable.</p><p>For both LQR and H &#8734; control, their cost values depend on B w only via B w B T w . We thus define W := B w B T w , and assume B w = W 1/2 without loss of generality. We also assume that W &#8827; 0. We consider the class of state-feedback policies of the form u(t) = Kx(t), with K &#8712; R m&#215;n , to regulate z(t) under the influence of w(t). The set of stabilizing state-feedback policies, parameterized by K, is then K := K &#8712; R m&#215;n max i Re &#955; i (A + BK) &lt; 0 .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.1.">LINEAR QUADRATIC REGULATOR (LQR)</head><p>In the LQR problem, w(t) is assumed to be a white Gaussian noise with unit intensity. The policy optimization formulation for LQR is then (see, e.g., <ref type="bibr">Mohammadi et al. (2022)</ref>),</p><p>where X K is the positive semidefinite solution to the Lyapunov equation (A + BK)X K + X K (A + BK) + W = 0. It is known that J LQR (K) is smooth and nonconvex, but has a unique 5 stationary point (which is globally optimal) and is gradient dominated on any sublevel set <ref type="bibr">(Mohammadi et al., 2022)</ref>. These nice landscape properties are closely related to the hidden convexity of ( <ref type="formula">5</ref>).</p><p>The ECL construction for LQR policy optimization is presented in the second column of Table <ref type="table">1</ref>. Specifically, the construction consists of three steps:</p><p>Step 1: Lifting. We first define the lifted set L LQR , as delineated in the second row of Table <ref type="table">1</ref>. By the theory of linear quadratic control, we have K &#8712; K and &#947; &#8805; J LQR (K) if and only if there exists X such that (K, &#947;, X) &#8712; L LQR . This further implies that &#960; K,&#947; (L LQR ) = epi &#8805; (J LQR ).</p><p>Step 2: Convex set. We define the convex set F LQR as given in the third row of Table <ref type="table">1</ref>, and let the auxiliary set be G LQR = {0}. The first three constraints in the definition of F LQR are obviously convex, and the convexity of the last inequality follows by the joint convexity of tr(QX + X -1 Y T RY ) with respect to (X, Y ).</p><p>Step 3: Diffeomorphism. We employ the classical change of variables Y = KX and define &#934; LQR (K, &#947;, X) = (&#947;, KX, X), where (KX, X) represents the variable &#950; 1 in ECL. This mapping naturally satisfies (2b). We do not need &#950; 2 here as no similarity transformation exists. One can check by standard calculation that F LQR = &#934; LQR (L LQR ), and that &#934; LQR admits an inverse on F LQR given by <ref type="formula">5</ref>). The ECL framework then immediately implies the following well-known results: 1. Theorem 2.1 justifies that the LQR problem ( <ref type="formula">5</ref>) can be reformulated as a convex program: 6  min</p><p>and their optimal solutions K * and (&#947; * , Y * , X * ) are related by K * = Y * X * -1 and &#947; * = J LQR (K * ). The policy optimization for LQR can be viewed as a convex problem in disguise. 2. Any stationary point K &#8902; of J LQR is globally optimal, confirmed by &#960; K,&#947; (L LQR ) = epi &#8805; (J LQR )</p><p>and Corollary 3.1.</p><p>We mention that existing literature has further proved that J LQR (K) has a unique stationary point, is coercive, and is L-smooth and gradient dominated over any sublevel set <ref type="bibr">(Mohammadi et al., 2022;</ref><ref type="bibr">Watanabe and Zheng, 2025)</ref>. These properties are fundamental to establishing global convergence of direct policy search and their model-free extensions for solving LQR <ref type="bibr">(Malik et al., 2020;</ref><ref type="bibr">Mohammadi et al., 2022)</ref>. It would be interesting to investigate how to refine our ECL framework to cover these properties.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.2.">STATE-FEEDBACK H &#8734; CONTROL</head><p>In state-feedback H &#8734; control, we consider w(t) as adversarial disturbance with bounded energy, and the goal is to minimize the maximum energy gain from the disturbance w(t) to the performance signal z(t). It is a standard result in robust control that the state-feedback H &#8734; control problem can be formulated as inf</p><p>5. This uniqueness is due to W &#8827; 0; see <ref type="bibr">Watanabe and Zheng (2025)</ref> for details on non-unique optimal LQR gains. 6. The infima for both the original policy optimization and the convex reformulation can be achieved, due to the coerciveness of JLQR(K) and the compactness of (Y, X) (&#947;, Y, X) &#8712; FLQR for any given &#947; &gt; 0.</p><p>where T zw (K, s) is the transfer matrix from w(t) to z(t) when the policy u(t) = Kx(t) is applied, and &#8741; &#8226; &#8741; H&#8734; denotes the H &#8734; norm. Note that the infimum of (6) may not be attained. The H &#8734; policy optimization problem ( <ref type="formula">6</ref>) is nonconvex and nonsmooth, but it admits a convex reformulation <ref type="bibr">(Scherer and Weiland, 2015)</ref>. It has been recently revealed in <ref type="bibr">Guo and Hu (2022)</ref> that for the discrete-time version of (6), any Clarke stationary point is globally optimal. Our aim here is to construct an ECL for (6), which then allows us to draw conclusions from Theorems 2.1 and 3.1 directly. The construction process is very similar to the LQR case, and is delineated in the third column of Table <ref type="table">1</ref>. One key difference is that we employ the bounded real lemmas to handle the H &#8734; case.</p><p>Specifically, our ECL construction consists of the following steps:</p><p>Step 1: Lifting. We define the lifted set L &#8734; = (K, &#947;, P ) L &#8734; &#10927; 0, P &#8827; 0 , where P is an extra Lyapunov variable, and the matrix L &#8734; is given in the second row of Table <ref type="table">1</ref>.</p><p>Step 2: Convex set. We define a convex set</p><p>Step 3: Diffeomorphism. We employ the classical change of variables Y = KP -1 , X = P -1 and introduce the mapping &#934; &#8734; (K, &#947;, P ) = (&#947;, KP -1 , P -1 ), where (KP -1 , P -1 ) represents the variable &#950; 1 . Similar to the LQR case, no auxiliary variable &#950; 2 is needed.</p><p>For the construction above, we have the following results.</p><p>Proposition 4.1 Consider the state-feedback H &#8734; policy optimization problem (6), where (A, B) is controllable, B w has full row rank, and Q &#8827; 0, R &#8827; 0.</p><p>1. For any K &#8712; R m&#215;n and &#947; &#8712; R, we have K &#8712; K and &#947; &#8805; J &#8734; (K) if and only if there exists P such that (K, &#947;, P ) &#8712; L &#8734; . This further implies &#960; K,&#947; (L &#8734; ) = epi &#8805; (J &#8734; ).</p><p>2. The mapping &#934; &#8734; is a C &#8734; diffeomorphism between the lifted set L &#8734; and the convex set F &#8734; .</p><p>The proof is not very difficult, but one needs to be careful about some technical subtleties in (non)-strict Riccati inequalities. The details are provided in <ref type="bibr">Zheng et al. (2024, Appendix C.4</ref>). Proposition 4.1 guarantees that (L &#8734; , F &#8734; , {0}, &#934; &#8734; ) is an ECL of J &#8734; (K) in (6). For this ECL, we further have the following nice results, which are consequences of Theorem 2.1 and Corollary 3.1, and the fact that J &#8734; (K) is sudifferentially regular.</p><p>Theorem 4.1 Under the conditions of Proposition 4.1, the following statements hold. 1. Problem (6) is equivalent to the convex problem inf (&#947;,Y,X)&#8712;F&#8734; &#947;. 2. All stabilizing policies K &#8712; K are non-degenerate with respect to the ECL (L &#8734; , F &#8734; , {0}, &#934; &#8734; ). 3. Any Clarke stationary point of (6) is globally optimal.</p><p>We note that the global optimality of Clarke stationary points for (6) has not been reported before. This is the continuous-time counterpart of the discrete-time result in <ref type="bibr">Guo and Hu (2022)</ref>. We also note that the infimum of (6) may not be achieved, in which case the Clarke stationary point does not exist. An explicit SISO example is provided in <ref type="bibr">Zheng et al. (2024, Appendix C.2)</ref>.</p><p>Remark 4.1 The diffeomorphisms &#934; LQR and &#934; &#8734; are essentially in the same form and follow from the classical change of variable K = Y X -1 <ref type="bibr">(Khargonekar and Rotea, 1991;</ref><ref type="bibr">Boyd et al., 1994;</ref><ref type="bibr">Bernussou et al., 1989)</ref>. This change of variable is able to linearize many bilinear matrix inequalities that appear in state-feedback control problems, most of which are related to the Lyapunov inequality (A + BK)X + X(A + BK) T &#8826; 0 (see <ref type="bibr">Boyd et al., 1994, Chapter 7</ref> for a historical perspective). As we will see in the next subsection, the linearization for dynamic output-feedback policies turns out to be much more complicated, and we will utilize the techniques in <ref type="bibr">Scherer et al. (1997)</ref>; <ref type="bibr">Scherer and Weiland (2015)</ref> to construct ECLs for output-feedback H &#8734; control. &#9633;</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Output-Feedback Policy Optimization</head><p>In this subsection, we show that ECL is also applicable to policy optimization with dynamic outputfeedback policies. Due to space limitations, we only treat the output-feedback H &#8734; control problem in this paper; see <ref type="bibr">Zheng et al. (2023</ref><ref type="bibr">Zheng et al. ( , 2024) )</ref> for the case of LQG control. Consider an LTI system with partial observations,</p><p>where y(t) &#8712; R p is the vector of measured outputs available for feedback control, and w(t) &#8712; R n , v(t) &#8712; R p are the disturbances on the system process and measurement at time t. We define</p><p>and, without loss of generality, assume B w = W 1/2 and D v = V 1/2 . We adopt the same performance signal z(t) in (4). The following assumption is standard.</p><p>Assumption 4.1 (A, B) and (A, W 1/2 ) are controllable, and (C, A) and (Q 1/2 , A) are observable. Moreover, Q &#10928; 0, W &#10928; 0 and R &#8827; 0, V &#8827; 0.</p><p>To properly regulate z(t), we consider full-order dynamic output-feedback policies of the form</p><p>where &#958;(t) &#8712; R n is the internal state, and A K , B K , C K and D K are matrices of proper dimensions that specify the policy dynamics. The matrix K is used for parameterizing the dynamic policies. By closing the feedback loop, we can write the transfer matrix from the disturbance d(t) = w T (t) v T (t) T to the performance signal z(t) in the following form:</p><p>where A cl (K), B cl (K), C cl (K) and D cl (K) are certain matrix-valued functions that characterize the state-space model of the closed-loop system (see <ref type="bibr">Zheng et al., 2023, Appendix A.5</ref> for details). Now consider a standard output-feedback H &#8734; policy optimization problem,</p><p>with C n denoting the set of internally stabilizing dynamic policies C n = K A cl (K) is stable . It is known that the policy optimization problem ( <ref type="formula">8</ref>) is nonsmooth and nonconvex, which has complicated landscape properties. Our ECL construction for J &#8734;,n (K) is based on the change of variables given in <ref type="bibr">Scherer et al. (1997)</ref>; <ref type="bibr">Scherer and Weiland (2015)</ref>, and the details are as follows:</p><p>Step 1: Lifting. We first introduce the lifted set L &#8734; by</p><p>where P 12 denotes the n &#215; n submatrix of P corresponding to the first n rows and last n columns. The extra variable P plays the role of the lifting variable in ECL.</p><p>Step 2: Convex and auxiliary sets. We let the convex set be</p><p>where (&#923;, X, Y ) corresponds to &#950; 1 , and M (&#947;, &#923;, X, Y ) is an affine operator whose definition is omitted here due to space limitations. The auxiliary set is GL n = T &#8712; R n&#215;n det T &#824; = 0 .</p><p>Step 3: Diffeomorphism. We define the mapping &#934; &#8734; by</p><p>, (P -1 ) 11 , P 11 , P 12 , (K, &#947;, P ) &#8712; L &#8734; ,</p><p>where &#934; M = P 12 B K C(P -1 ) 11 + P 11 BC K (P -1 ) 21 + P 11 (A + BD K C)(P -1 ) 11 + P 12 A K (P -1 ) 21 , &#934; H = P 11 BD K + P 12 B K , and &#934; F = D K C(P -1 ) 11 + C K (P -1 ) 21 .</p><p>The following proposition justifies that (L &#8734; , F &#8734; , GL n , &#934; &#8734; ) is an ECL for J &#8734;,n (K). The proof is technically involved and is given in the report <ref type="bibr">Zheng et al. (2024, Appendix D)</ref>; the main difficulty lies in that (3) does not hold anymore, and we need to establish the chain of inclusion (2a).</p><p>Proposition 4.2 Under Assumption 4.1, we have i</p><p>Then, by Theorems 2.1 and 3.1, we get the following corollary.</p><p>Corollary 4.1 Under Assumption 4.1, the output-feedback H &#8734; policy optimization problem (8) is equivalent to a convex problem in the sense that inf K&#8712;Cn J &#8734;,n (K) = inf (&#947;,&#923;,X,Y )&#8712;F &#8734; &#947;. Furthermore, for a Clarke stationary point K &#8712; C n (i.e., 0 &#8712; &#8706;J &#8734;,n (K)), if K is non-degenerate in the sense of Definition 3.1, then it is globally optimal for (8).</p><p>We can now see that our ECL framework covers the output-feedback H &#8734; policy optimization problem as a special case, and provides global optimality certificates for non-degenerate Clarke stationary points. Note that the equivalence to the convex reformulation is essentially the same as <ref type="bibr">(Scherer and Weiland, 2015, Chapter 4.2.3</ref>), but the global optimality of non-degenerate H &#8734; policies cannot be derived from <ref type="bibr">(Scherer and Weiland, 2015, Chapter 4.2.</ref>3) due to its use of strict LMIs. Finally, we remark that our ECL can also cover a class of distributed control problems under the condition of quadratic invariance <ref type="bibr">(Furieri et al., 2020a,b)</ref>; see <ref type="bibr">Zheng et al. (2024, Section 4.4)</ref> for details.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Conclusion</head><p>This paper introduced the ECL framework to reveal hidden convexity in nonconvex and potentially nonsmooth policy optimization for optimal and robust control. Particularly, we have shown that, with the existence of an ECL for nonconvex policy optimization, all non-degenerate stationary policies are globally optimal. We have built explicit ECLs for LQR, state feedback H &#8734; control, and dynamic output-feedback H &#8734; control. We hope the ECL framework will be useful for analyzing nonconvex problems in other areas beyond control. We are also interested in developing principled first-order policy optimization algorithms that leverage both convex and nonconvex information within the ECL framework.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>&#169; 2025 Y. Zheng, C.-F. Pai &amp; Y. Tang.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_1"><p>We allow d2 = 0, in which case we adopt the convention Gaux = {0} and identify Fcvx &#215; {0} with Fcvx.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_2"><p>Some materials rely on the notion of Clarke subdifferential, which extends subdifferential to nonconvex nonsmooth functions. We refer the readers to Clarke (1990) for details, orZheng et al. (2023, Appendix B)  for a brief review.</p></note>
		</body>
		</text>
</TEI>
