If you see this, something is wrong
First published on Monday, Sep 21, 2026 and last modified on Monday, Sep 21, 2026 by François Chaplais.
College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China Email
College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China Email
Policy iteration, data-driven control, off-policy reinforcement learning, optimal control.
This article investigates data-driven policy iteration (PI) for continuous-time linear systems without requiring an initially stabilizing policy. Standard infinite-horizon PI is not self-starting because its policy-evaluation step is well posed only when the feedback gain is stabilizing. However, verifying this property is difficult when the system matrices are unknown. To remove this requirement, we develop a finite-horizon bootstrap method. The key idea is to perform policy evaluation over a compact interval for a shifted system, where the evaluation equation is well defined for arbitrary bounded time-varying policies. We show that, for a sufficiently long horizon, the initial-time optimal gain of the shifted finite-horizon problem, when applied as a constant feedback gain, achieves a prescribed stability margin for the original system. We then derive a data-driven implementation from an off-policy identity evaluated along trajectories of the original plant. We use basis-function approximations to reconstruct the finite-horizon value matrix and policy, and we characterize the resulting error through a perturbed policy-improvement recursion. A data-driven Lyapunov certificate is further introduced to verify admissibility of the candidate gain before it is used to initialize infinite-horizon PI. Numerical studies of a batch reactor and a two-mass-spring system demonstrate the effectiveness of the proposed bootstrap method.
Policy iteration (PI) is a widely used Newton-type method for continuous-time linear quadratic regulator (LQR) problems [1, 2, 3]. PI is initialized with a stabilizing feedback gain and proceeds by repeatedly evaluating the current policy and updating the control law. The iteration preserves closed-loop stability, generates a monotonically nonincreasing sequence of value matrices, and converges locally at the Newton rate [4, 5]. Integral and off-policy reinforcement learning (RL) formulations implement both steps using measured state–input data without explicitly identifying the system matrices [6, 7, 8, 9, 10, 11]. Together, these properties make PI attractive for data-driven optimal control. Yet standard infinite-horizon PI is not self-starting [12, 13, 14]. The policy-evaluation step in infinite-horizon PI requires the current feedback gain to stabilize the closed-loop system. If the gain is not stabilizing, closed-loop trajectories need not decay, the infinite-horizon cost is generally not finite, and the associated Lyapunov equation no longer defines a valid value matrix. Consequently, the set of admissible policies defines the domain on which infinite-horizon policy evaluation admits a valid value-function interpretation. When the system matrices are unknown, constructing and certifying an admissible gain may require precisely the model information that data-driven PI seeks to avoid. This circular dependence poses a structural obstacle to implementing infinite-horizon PI.
Value iteration (VI) avoids the need for an initial stabilizing feedback gain by initializing a value matrix instead [15, 16, 17]. In this setting, existing continuous-time VI schemes typically rely on diminishing or small step sizes and may use bounded-set projections or reset rules [18, 19]. Their first-order updates may also converge more slowly than the local Newton iteration underlying PI. Hybrid iteration (HI) uses VI to reach the admissible-policy region and then switches to PI to exploit its fast local convergence [20, 21, 22]. The HI framework has also been applied to Markov jump systems [23, 24], networked control systems [25], and autonomous vehicles [26]. Although HI reduces reliance on an initial stabilizing feedback gain, its VI stage still requires step-size selection, boundedness verification, and an explicit VI-to-PI switching criterion.
| References |
Initial stabilizing policy |
Stabilizing intermediate policy |
Infinite-horizon subproblem |
Step-size mechanism |
Bias mechanism |
| Standard PI [8] | Required | Required | Required | N/A | N/A |
| VI [18] | Not required | N/A | Not required | Required | N/A |
| HI [20, 21, 22] | Not required |
Required after switching |
Required after switching |
Required | N/A |
| Bias PI [27] | Relaxed |
Required for auxiliary system |
Required | N/A | Required |
| Homotopy PI [28, 29, 30] | Not required | Required | Required | N/A | Required |
| This paper | Not required | Not required | Not required | Not required | Not required |
A second class of methods modifies the policy-evaluation operator to enlarge the initialization region. Bias PI introduces a bias parameter into policy evaluation and transfers the stability requirement from the original closed-loop matrix to a shifted auxiliary matrix [27]. A smaller bias yields PI-like behavior, whereas a larger bias can expand the initialization region at the cost of making the iteration more VI-like. Related ideas have been explored for discrete-time systems [31] and zero-sum games [32, 33]. Bias PI thus relaxes the initial admissibility requirement while retaining linear matrix equations at each iteration. Implementing the method requires selecting a sufficiently large bias to ensure that auxiliary policy evaluation is well posed. Global variants additionally use boundedness tests and bias-adjustment rules.
Homotopy-based methods take a different approach. They start from an artificially stabilized system that admits a trivial or readily constructed stabilizing policy. They then gradually deform the auxiliary system into the original plant while maintaining a stabilizing policy for each intermediate problem. Homotopic PI methods move the artificial dynamics toward those of the original plant using a cumulative continuation factor and provide an explicit bound on the closed-loop pole location [28, 30]. A more recent predictor-corrector homotopy method constructs a path of stabilizing Riccati solutions and uses first-order differential information to predict the controller for the next auxiliary problem, thereby enabling faster continuation along the path [29]. Related homotopy-based methods have subsequently been extended to nonlinear systems [34], inverse RL [35], multiagent systems [36], and Markov jump systems [37]. These methods substantially relax the requirement for an initially stabilizing policy for the original plant. Nevertheless, each continuation stage still involves an infinite-horizon subproblem, and the corresponding intermediate policy must be stabilizing to ensure that policy evaluation is well posed. Thus, homotopy methods address the initialization difficulty by constructing a stability-preserving path through a sequence of auxiliary infinite-horizon problems. Table 1 summarizes the initialization requirements of the representative continuous-time data-driven RL methods discussed above and highlights the distinctions between these methods and the proposed finite-horizon bootstrap method.
The difficulty of evaluating a nonstabilizing policy stems from the unbounded evaluation interval. By contrast, over a compact interval, the terminal-value policy-evaluation equation is well posed for any bounded time-varying feedback policy, regardless of closed-loop stability. This distinction suggests using finite-horizon PI as a bootstrap stage rather than as a replacement for the original infinite-horizon problem. The finite-horizon stage generates the stabilizing policy needed to initialize infinite-horizon PI. However, implementing this bootstrap from data raises two main challenges. First, the time-varying finite-horizon value matrix and policy must be reconstructed using trajectories of the original plant alone, without collecting data from the auxiliary shifted system. Second, approximation and regression errors inevitably affect this reconstruction. The resulting constant gain must therefore be independently certified to stabilize the original plant before being used to initialize infinite-horizon PI. The main contributions are as follows:
Notation: Throughout this article, we use \( \mathbb{R}^{n}\) for the \( n\) -dimensional real Euclidean space and \( \mathbb{R}^{m\times n}\) for the space of real \( m\times n\) matrices. The identity matrix of order \( n\) is written as \( I_n\) , whereas \( \mathbf{0}\) represents a zero scalar, vector, or matrix whose dimensions are inferred from context. For a matrix \( \textbf{A}\) , we use \( \sigma_{\min}(\textbf{A})\) to represent its smallest singular value, while \( \textbf{A}^{\dagger}\) refers to its Moore–Penrose pseudoinverse. When \( \textbf{A}\) is square, \( \sigma(\textbf{A})\) represents its spectrum, and its spectral abscissa is given by \( \alpha(\textbf{A}) = \max_{\lambda\in\sigma(\textbf{A})} \operatorname{Re}(\lambda) \) , where \( \operatorname{Re}(\lambda)\) represents the real part of \( \lambda\) . Moreover, \( \mathbb{C}_{-} = \{z\in\mathbb{C}:\operatorname{Re}(z)<0\} \) denotes the open left-half complex plane. The minimum and maximum eigenvalues of a symmetric matrix \( \textbf{X}\) are denoted by \( \lambda_{\min}(\textbf{X})\) and \( \lambda_{\max}(\textbf{X})\) , respectively. For matrix-valued function \( \textbf{H}\) defined on \( [0,T]\) , let \( \|\textbf{H}\|_{\infty,T} = \sup_{t\in[0,T]}\|\textbf{H}(t)\|_F . \) Given a matrix \( \textbf{A}=[a_{ij}]\in\mathbb{R}^{m\times n}\) , \( \operatorname{vec}(\textbf{A})\) is formed by stacking the columns of \( \textbf{A}\) into a single vector. For a symmetric matrix \( \textbf{P}=[p_{ij}]\in\mathbb{R}^{n\times n}\) , define \( \operatorname{vecs}(\textbf{P})=[p_{11},2p_{12},…,2p_{1n},p_{22},2p_{23},…,p_{nn}]^T.\) For a vector \( x=[x_1,…,x_n]^T\in\mathbb{R}^{n}\) , define \( \operatorname{vecv}(x)=[x_1^2,x_1x_2,…,x_1x_n,x_2^2,x_2x_3,…,x_n^2]^T.\) For matrices \( X_1,…,X_q\) having the same number of columns, \( \operatorname{col}(X_1,…,X_q) = [X_1^{T},…,X_q^{T}]^{T}\) denotes their vertical concatenation.
Consider the continuous-time linear system
(1)
where \( x(t)\in \mathbb{R}^{n}\) is the system state and \( u(t)\in \mathbb{R}^{m}\) is the control signal. The system matrices \( \mathbf{A}\in\mathbb{R}^{n\times n}\) and \( \mathbf{B}\in\mathbb{R}^{n\times m}\) are constant but unavailable. The infinite-horizon quadratic cost is
(2)
where \( \mathbf{Q}=\mathbf{Q}^{T}\geq \mathbf{0}\) and \( \mathbf{R}=\mathbf{R}^{T}>\mathbf{0}\) are the state and control penalties, respectively. The pair \( (\mathbf{A},\mathbf{B})\) is assumed to be stabilizable, and \( (\mathbf{A},\mathbf{Q}^{1/2})\) is assumed to be detectable.
Definition 1
A constant feedback gain \( \mathbf{K}\in\mathbb{R}^{m\times n} \) is called admissible if \( \mathbf{A}-\mathbf{B}\mathbf{K}\) is Hurwitz. For a prescribed constant \( \varepsilon\geq 0\) , \( \mathbf{K}\) is called \( \varepsilon\) -admissible if
(3)
When \( \mathbf{A}\) and \( \mathbf{B}\) are available, the optimal control law is
(4)
where \( \mathbf{P}^{*}=(\mathbf{P}^{*})^{T}\geq\mathbf{0}\) is the unique stabilizing solution of the algebraic Riccati equation
(5)
The standard continuous-time PI for solving (5) is Kleinman’s iteration. The following standard result of Kleinman’s iteration is recalled.
Lemma 1
Let \( \mathbf{K}^{[0]}\) be any admissible gain matrix, i.e., \( \sigma (\mathbf{A}-\mathbf{B}\mathbf{K}^{[0]})\subset \mathbb{C}_{-}\) . For \( j=0,1,2,… \) , let \( \mathbf{P}^{[j]}=(\mathbf{P}^{[j]})^{T}\) solve
(6)
and update the control gain by
(7)
Then, the following properties hold:
Lemma 1 shows that Kleinman’s iteration is not self-starting, because its well-posedness and stability-preserving property both require an admissible initial gain. The problem addressed in this paper is therefore to construct, from measured data, a constant gain \( \mathbf{K}^{[0]} \) satisfying (3), which can serve as the initializer for infinite-horizon data-driven PI.
The central difficulty in removing the initial admissible controller is that the infinite-horizon policy evaluation operator is intrinsically defined only on stabilizing policies. For a fixed gain \( \mathbf{K}\) , the evaluation step requires the Lyapunov equation
When \( \mathbf{A}-\mathbf{B}\mathbf{K}\) is Hurwitz, its solution admits the value representation
If \( \mathbf{K}\) is not stabilizing, the state-transition matrix \( e^{(\mathbf{A}-\mathbf{B}\mathbf{K})t}\) does not decay exponentially. Since the cost is accumulated over the unbounded interval \( [0,\infty)\) , the integral is generally not finite, and the evaluation no longer defines a valid cost matrix.
As shown in Fig. 1, a finite-horizon evaluation avoids this obstruction because it is posed on a compact interval. For any bounded time-varying gain \( \mathbf{K}(t)\) on \( [0,T] \) , the terminal-value Lyapunov equation
(8)
has a unique solution on \( [0,T]\) with a terminal weight \( \textbf{M}=\textbf{M}^{T}\geq\textbf{0}\) . This well-posedness does not require the time-varying closed-loop system to be exponentially stable. Thus, unlike the infinite-horizon evaluation, the finite-horizon evaluation is well defined for arbitrary bounded policies and can therefore be used before an admissible gain is available.
The finite-horizon stage must still produce a gain that is admissible for the original infinite-horizon problem. To this end, the evaluation is performed for the shifted dynamics \( \dot x(t)=\mathbf{A}_{\rho }x(t)+\mathbf{B}u(t)\) with \( \mathbf{A}_{\rho }=\mathbf{A}+\rho \mathbf{I}\) and \( \rho>\varepsilon\) . Let \( \mathbf{K}_{\rho,T}^\ast(t)\) denote the optimal finite-horizon gain of this shifted problem. As \( T\) increases, the initial-time gain \( \mathbf{K}_{\rho,T}^\ast(0)\) approaches the stabilizing infinite-horizon gain \( \mathbf{K}_\rho^\ast=\mathbf{R}^{-1}\mathbf{B}^T\mathbf{P}_\rho^\ast , \) where \( \mathbf{P}_\rho^\ast\) solves the shifted algebraic Riccati equation. Since \( \mathbf{A}_\rho-\mathbf{B}\mathbf{K}_\rho^\ast\) is Hurwitz, applying the same gain to the original system gives
Therefore, by continuity of the spectral abscissa, \( \mathbf{K}_{\rho,T}^\ast(0)\) is also \( \varepsilon\) -admissible for all sufficiently large \( T\) .
This section presents the model-based version of the proposed finite-horizon bootstrap method. The model-based construction clarifies the role of the finite horizon and serves as the reference algorithm for the data-driven development.
Let \( \varepsilon \geq 0\) be the desired stability margin in (3). Choose a design parameter \( \rho >\varepsilon \) and define the shifted system matrix \( \mathbf{A}_{\rho }\) . The following assumption is imposed on the shifted pair.
Assumption 1. The pair \( (\mathbf{A}_{\rho},\mathbf{B})\) is stabilizable, and \( (\mathbf{A}_{\rho},\mathbf{Q}^{1/2})\) is detectable.
Assumption 1 guarantees the existence of the stabilizing solution \( \mathbf{P}_{\rho }^{\ast }\) to the shifted algebraic Riccati equation
(9)
The control gain is \( \mathbf{K}_{\rho }^{\ast }=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\rho }^{\ast }\) . Since \( \mathbf{A}_{\rho }-\mathbf{B}\mathbf{K}_{\rho }^{\ast }\) is Hurwitz, the matrix \( \mathbf{A}-\mathbf{B}\mathbf{K}_{\rho }^{\ast }=\mathbf{A}_{\rho }-\mathbf{B}\mathbf{K}_{\rho }^{\ast }-\rho \mathbf{I}\) is Hurwitz with a decay margin greater than \( \rho \) . The purpose of the finite-horizon stage is to approximate this gain without requiring an admissible initial policy.
For a horizon \( T>0\) and a terminal matrix \( \mathbf{M}=\mathbf{M}^{T}\geq \mathbf{0}\) , consider the shifted finite-horizon Riccati differential equation
(10)
with terminal condition \( \mathbf{P}_{\rho ,T}^{\ast }(T)=\mathbf{M}\) . The associated finite-horizon optimal gain is
(11)
The finite-horizon gain in (11) can be computed by a finite-horizon PI without requiring a stabilizing initial policy. Let \( \mathbf{K}_{\rho ,T}^{[0]}(t)\) be any bounded time-varying matrix on \( [0,T]\) . For \( j=0,1,2,… \) , solve
(12)
with \( \mathbf{P}_{\rho ,T}^{[j]}(T)=\mathbf{M}\) , and update
(13)
Because (12) is solved over a compact interval, the initial policy \( \mathbf{K}_{\rho ,T}^{[0]}(t)\) need not stabilize the shifted system. As will be shown below, the finite-horizon PI sequence converges to \( \mathbf{K}_{\rho ,T}^{\ast }(t)\) . After convergence, with \( \bar{j}\) denoting the terminal iteration index, the initializer is taken as \( \mathbf{K}^{[0]}=\mathbf{K}_{\rho ,T}^{[\bar{j}]}(0)\) . In the model-based setting, admissibility can be checked directly by verifying
(14)
If (14) holds, \( \mathbf{K}^{[0]}\) is used to initialize Kleinman’s infinite-horizon PI (6)–(7) for the system (1). Otherwise, after the finite-horizon PI has converged for the current horizon, the horizon is updated according to
(15)
Algorithm 1 summarizes the model-based finite-horizon bootstrap procedure. The following theorem justifies the algorithm by separating the fixed-horizon finite-horizon PI property from the admissibility guarantee obtained by enlarging the shifted horizon.
Theorem 1
Let \( \rho >\varepsilon \geq 0\) , \( T>0\) , and \( \mathbf{M}=\mathbf{M}^{T}\geq \mathbf{0}\) . For any bounded piecewise-continuous initial policy \( \mathbf{K}_{\rho ,T}^{[0]}(t)\) on \( [0,T]\) , generate \( \mathbf{P}_{\rho ,T}^{[j]}(t)\) and \( \mathbf{K}_{\rho ,T}^{[j+1]}(t)\) by (12)–(13). Then, the following statements hold.
The sequence satisfies
(16)
and \( \lim_{j\rightarrow \infty }\mathbf{P}_{\rho ,T}^{[j]}(t)=\mathbf{P}_{\rho ,T}^{\ast }(t),\) \( \lim_{j\rightarrow \infty }\mathbf{K}_{\rho ,T}^{[j]}(t)=\mathbf{K}_{\rho ,T}^{\ast }(t)\) uniformly on \( [0,T]\) .
If Assumption 1 holds and \( \mathbf{K}_{T}=\mathbf{K}_{\rho ,T}^{\ast }(0)\) , then \( \lim_{T\rightarrow \infty }\mathbf{K}_{T}=\mathbf{K}_{\rho }^{\ast }\) . Moreover, there exists \( T_{\varepsilon }>0\) such that, for all \( T\geq T_{\varepsilon }\) ,
(17)
If Assumption 1 holds, then for every \( T\geq T_{\varepsilon }\) , there exists an integer \( j_{\varepsilon }(T)\) such that
(18)
Proof
(1) For each \( j\geq 0\) , define \( \mathbf{A}^{[j]}(t)=\mathbf{A}_{\rho }-\mathbf{B}\mathbf{K}_{\rho ,T}^{[j]}(t)\) and \( \mathbf{G}^{[j]}(t)=\mathbf{Q}+(\mathbf{K}_{\rho ,T}^{[j]}(t))^{T}\mathbf{R}\mathbf{K}_{\rho ,T}^{[j]}(t)\) . Since \( \mathbf{K}_{\rho ,T}^{[0]}\) is bounded and piecewise continuous on \( [0,T]\) , \( \mathbf{A}^{[0]}(t)\) and \( \mathbf{G}^{[0]}(t)\) are bounded on \( [0,T] \) . Therefore, the terminal-value Lyapunov equation (12) has a unique symmetric solution on \( [0,T]\) . Let \( \boldsymbol{\Phi }^{[j]}(\tau ,t)\) be the transition matrix generated by \( \dot{z}(\tau )=\mathbf{A}^{[j]}(\tau )z(\tau )\) . The solution is
(19)
The right-hand side is finite for every \( t\in \lbrack 0,T]\) . Moreover, once \( \mathbf{P}_{\rho ,T}^{[j]}\) is obtained, it is continuous on the compact interval \( [0,T]\) and hence bounded. Thus \( \mathbf{K}_{\rho ,T}^{[j+1]}=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\rho ,T}^{[j]}\) is bounded. By induction, the recursion is well defined for every \( j\geq 0\) .
(2) Consider the trajectory generated by the improved policy \( u(t)=-\mathbf{K}_{\rho ,T}^{[j+1]}(t)x(t)\) for the shifted system. Along this trajectory,
(20)
where \( \Delta \mathbf{K}^{[j]}(t)=\mathbf{K}_{\rho ,T}^{[j]}(t)-\mathbf{K}_{\rho ,T}^{[j+1]}(t)\) . Subtracting the Lyapunov equation satisfied by \( \mathbf{P}_{\rho ,T}^{[j+1]}\) gives
(21)
where \( \mathbf{E}^{[j]}(t)=\mathbf{P}_{\rho ,T}^{[j]}(t)-\mathbf{P}_{\rho ,T}^{[j+1]}(t)\) . Therefore,
Thus \( \mathbf{P}_{\rho ,T}^{[j+1]}(t)\leq \mathbf{P}_{\rho ,T}^{[j]}(t)\) for \( t\in \lbrack 0,T].\) Next, \( \mathbf{P}_{\rho ,T}^{[j]}(t)\) is the finite-horizon value matrix associated with the policy \( u(t)=-\mathbf{K}_{\rho ,T}^{[j]}(t)x(t)\) for the shifted system. Since \( \mathbf{P}_{\rho ,T}^{\ast }(t)\) is the optimal finite-horizon value matrix, we have \( \mathbf{P}_{\rho ,T}^{\ast }(t)\leq \mathbf{P}_{\rho ,T}^{[j]}(t)\) for \( t\in \lbrack 0,T]\) . This proves the monotonicity relation. It remains to prove convergence. Since \( \mathbf{P}_{\rho ,T}^{\ast }(t)\leq \mathbf{P}_{\rho ,T}^{[j]}(t)\leq \mathbf{P}_{\rho ,T}^{[0]}(t)\) , the sequence \( \{\mathbf{P}_{\rho ,T}^{[j]}\}\) is uniformly bounded on \( [0,T]\) . Moreover, \( \mathbf{K}_{\rho ,T}^{[j+1]}(t)=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\rho ,T}^{[j]}(t)\) implies that \( \{\mathbf{K}_{\rho ,T}^{[j]}\}_{j\geq 1}\) is also uniformly bounded on \( [0,T]\) . From (12), the derivatives \( \{\dot{\mathbf{P}}_{\rho ,T}^{[j]}\}_{j\geq 1}\) are uniformly bounded. Hence \( \{\mathbf{P}_{\rho ,T}^{[j]}\}\) is equicontinuous on \( [0,T]\) . For each fixed \( t\in \lbrack 0,T]\) , the monotone bounded sequence \( \{\mathbf{P}_{\rho ,T}^{[j]}(t)\}\) has a limit in the Loewner order \( \mathbf{P}_{\infty }(t)\) . The equicontinuity on the compact interval \( [0,T]\) , together with the pointwise convergence, implies that the convergence to \( \mathbf{P}_{\infty }\) is uniform. Moreover, \( \mathbf{K}_{\rho ,T}^{[j]}\) converges uniformly to \( \mathbf{K}_{\infty }(t)=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\infty }(t)\) . Passing to the limit in the integral form of (12) yields
with \( \mathbf{P}_{\infty }(T)=\mathbf{M}\) , where \( \mathbf{P}_{\infty }\) satisfies the finite-horizon Riccati equation (10). By uniqueness of (10), \( \mathbf{P}_{\infty }(t)=\mathbf{P}_{\rho ,T}^{\ast }(t)\) , \( t\in \lbrack 0,T]\) . Therefore, \( \mathbf{P}_{\rho ,T}^{[j]}\) and \( \mathbf{K}_{\rho ,T}^{[j]}\) converge uniformly to \( \mathbf{P}_{\rho ,T}^{\ast }\) and \( \mathbf{K}_{\rho ,T}^{\ast }\) , respectively, on \( [0,T]\) .
(3) Let \( \mathbf{S}(\tau )\) be the solution of the forward Riccati flow
(22)
with \( \mathbf{S}(0)=\mathbf{M}\) . By the time transformation \( \tau =T-t\) , the finite-horizon Riccati solution satisfies \( \mathbf{P}_{\rho ,T}^{\ast }(t)=\mathbf{S}(T-t)\) , and in particular \( \mathbf{P}_{\rho ,T}^{\ast }(0)=\mathbf{S}(T)\) . Under Assumption 1, the continuous-time Riccati flow converges to the stabilizing solution of the shifted algebraic Riccati equation (9), that is, \( \lim_{\tau \rightarrow \infty }\mathbf{S}(\tau )=\mathbf{P}_{\rho }^{\ast }.\) Therefore, \( \lim_{T\rightarrow \infty }\mathbf{P}_{\rho ,T}^{\ast }(0)=\mathbf{P}_{\rho }^{\ast }\) and \( \lim_{T\rightarrow \infty }\mathbf{K}_{\rho ,T}^{\ast }(0)=\mathbf{K}_{\rho }^{\ast }\) . Since \( \mathbf{A}_{\rho }-\mathbf{B}\mathbf{K}_{\rho }^{\ast }\) is Hurwitz,
Because the spectral abscissa is continuous with respect to the matrix entries, and continuous with respect to \( \mathbf{K}\) , there exists a neighborhood \( \mathcal{N}\) of \( \mathbf{K}_{\rho }^{\ast }\) such that \( \alpha (\mathbf{A}-\mathbf{B}\mathbf{K})<-\varepsilon \) , \( \forall \mathbf{K}\in \mathcal{N}\) . Since \( \lim_{T\rightarrow \infty }\mathbf{K}_{\rho ,T}^{\ast }(0)=\mathbf{K}_{\rho }^{\ast }\) , there exists \( T_{\varepsilon }>0\) such that (17) holds.
(4) Fix any \( T\geq T_{\varepsilon }\) . From statement 3), \( \alpha (\mathbf{A}-\mathbf{B}\mathbf{K}_{T})<-\varepsilon \) . By continuity of the spectral abscissa, there exists a neighborhood \( \mathcal{N}_{T}\) of \( \mathbf{K}_{T}\) such that \( \alpha (\mathbf{A}-\mathbf{B}\mathbf{K})<-\varepsilon\) , \( \forall \mathbf{K}\in \mathcal{N}_{T}\) . From statement 2), \( \lim_{j\rightarrow \infty }\mathbf{K}_{\rho ,T}^{[j]}(0)=\mathbf{K}_{\rho ,T}^{\ast }(0)=\mathbf{K}_{T}\) . Therefore, there exists an integer \( j_{\varepsilon }(T)\) such that \( \mathbf{K}_{\rho ,T}^{[j]}(0)\in \mathcal{N}_{T}\) , \( \forall j\geq j_{\varepsilon }(T)\) . This proves the theorem.
The multiplicative update (15) gives \( T_{s}=\gamma ^{s}T_{0}\) . Since \( \gamma >1\) , \( T_{s}\rightarrow \infty \) as \( s\rightarrow \infty \) . Theorem 2 implies that the horizon is sufficiently large if
(23)
then \( T_{s}\geq T_{\varepsilon }\) . Therefore, for every \( s\geq s_{\varepsilon }\) , there exists \( j_{\varepsilon }(T_{s})\) such that
(24)
Thus, the finite-horizon bootstrap mechanism ensures that an \( \varepsilon\) -admissible initializer can be obtained after a sufficiently large horizon and sufficiently many finite-horizon PI iterations. In the model-based algorithm, this admissibility is verified directly by the spectral test (14).
This section develops a data-driven realization of the finite-horizon PI stage in Algorithm 1. The data are collected from the original system \( \dot{x}(t)=\mathbf{A}x(t)+\mathbf{B}u_{b}(t)\) , where \( u_{b}(t)\) is a behavior input.
For the current policy \( u(t)=-\mathbf{K}_{\rho ,T}^{[j]}(t)x(t)\) , define \( \mathbf{Q}_{\rho ,T}^{[j]}(t)=\mathbf{Q}+(\mathbf{K}_{\rho ,T}^{[j]}(t))^{T}\mathbf{R}\mathbf{K}_{\rho ,T}^{[j]}(t)\) and \( \mathbf{z}^{[j]}(t)=u_{b}(t)+\mathbf{K}_{\rho ,T}^{[j]}(t)x(t).\) The signal \( \mathbf{z}^{[j]}(t)\) is the off-policy correction term. Let \( V^{[j]}(x,t)=x^{T}\mathbf{P}_{\rho ,T}^{[j]}(t)x\) . Along the behavior trajectory, one obtains
(25)
Integrating (25) over each stored interval \( I_k=[t_k,t_{k+1}]\subset[0,T]\) , \( k=1,…,N_d\) , gives
(26)
Remark 1
Although the finite-horizon PI is formulated for the shifted matrix \( \mathbf{A}_\rho\) , no trajectory of the shifted system is required. The off-policy identity (26) is evaluated along the behavior trajectory of the original plant (1), and the effect of the shift is represented explicitly by the term involving \( \rho\) . Therefore, when the horizon is enlarged from \( T_s\) to \( T_{s+1}\) , one only needs behavior data of the original system over the enlarged interval.
Choose differentiable basis functions \( \phi _{\ell }(t)\) , \( \ell =1,… ,N_{p}\) with \( \phi _{\ell }(T)=0\) and basis functions \( \psi _{\ell }(t)\) , \( \ell =1,… ,N_{k}\) . Approximate \( \mathbf{P}_{\rho,T}^{[j]}(t)\) and \( \mathbf{K}_{\rho,T}^{[j+1]}(t)\) as
(27)
where \( \mathbf{P}_{\ell }^{[j]}=(\mathbf{P}_{\ell }^{[j]})^{T}\) , and
(28)
where \( \Delta _{P}^{[j]}(t)\) and \( \Delta _{K}^{[j+1]}(t)\) denote the approximation residuals induced by the chosen basis functions.
Remark 2
The expansions (27) and (28) are introduced to turn the time-varying unknowns \( \mathbf{P}_{\rho,T}^{[j]}(t)\) and \( \mathbf{K}_{\rho,T}^{[j+1]}(t)\) in (26) into a finite set of coefficient matrices, so that the off-policy identity (26) can be solved from data as the linear regression. In (27), the term \( \mathbf{M}\) and the condition \( \phi_\ell(T)=0\) enforce the terminal constraint \( \mathbf{P}_{\rho ,T}^{[j]}(T)=\mathbf{M}\) by construction.
To facilitate the solution of (26) using the representations in (27)-(28), we define the following interval-wise quantities for each stored interval \( I_{k}\)
Then (26) is equivalent to
(29)
where
Stacking all stored intervals gives
(30)
where
where \( \boldsymbol{\Omega}_{\rho,T}^{[j]} \in\mathbb{R}^{N_{d}\times N_{\theta}}\) , \( \theta^{[j]}\in\mathbb{R}^{N_{\theta}}\) , and \( \mathbf{b}_{\rho,T}^{[j]}\in\mathbb{R}^{N_{d}}\) , with \( N_{\theta} = N_{p}\frac{n(n+1)}{2}+N_{k}mn\) . The residual \( \mathbf{r}_{\rho,T}^{[j]}\) is induced by the approximation errors associated with the selected basis functions. The following lemma provides an explicit upper bound on this term.
Lemma 2
Fix \( T>0\) and \( j\geq 0\) . Define \( \delta _{P}^{[j]}=\sup_{t\in \lbrack 0,T]}\Vert \Delta _{P}^{[j]}(t)\Vert _{F}\) , \( \delta _{K}^{[j+1]}=\sup_{t\in \lbrack 0,T]}\Vert \Delta _{K}^{[j+1]}(t)\Vert _{F}\) , \( X_{T}=\sup_{t\in \lbrack 0,T]}\Vert x(t)\Vert _{2}\) , \( Z_{T}^{[j]}=\sup_{t\in \lbrack 0,T]}\Vert \mathbf{z}^{[j]}(t)\Vert _{2}\) , and let \( h_{\max }=\max_{1\leq k\leq N_{d}}(t_{k+1}-t_{k})\) . Then, for \( k=1,… ,N_{d}\) , \( |r_{k}^{[j]}|\leq \eta _{T}^{[j]}\) and
(31)
where \( \eta _{T}^{[j]}=\eta _{P}^{[j]}+\eta _{K}^{[j+1]}\) with
Proof
From the definition of \( r_{k}^{[j]}\) , the endpoint term satisfies
(32)
Similarly, the integral terms containing \( \Delta _{P}^{[j]}(t)\) and \( \Delta _{K}^{[j+1]}(t)\) are bounded, respectively, by \( 2\rho h_{\max }X_{T}^{2}\delta _{P}^{[j]}\) and \( 2h_{\max }Z_{T}^{[j]}\Vert \mathbf{R}\Vert _{2}X_{T}\delta _{K}^{[j+1]}\) . The proof is complete.
Lemma 3
Fix a horizon \( T>0\) and an iteration index \( j\geq 0\) . Let \( \mathbf{P}_{\rho ,T}^{[j]}(t)\) be the exact policy-evaluation solution associated with the current data-driven policy \( \widehat{\mathbf{K}}_{\rho ,T}^{[j]}(t)\) , and define \( \mathbf{K}_{\rho ,T}^{[j+1]}(t)=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\rho ,T}^{[j]}(t)\) . Let \( \theta _{0}^{[j]}\) collect the coefficient matrices \( \mathbf{P}_{\ell }^{[j]}\) and \( \mathbf{K}_{\ell }^{[j+1]}\) in the decompositions (27)-(28). Let \( \widehat{\theta }^{[j]}=\left( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\right) ^{\dagger }\mathbf{b}_{\rho ,T}^{[j]}\) , and let \( \widehat{\mathbf{K}}_{\rho ,T}^{[j+1]}(t)\) be reconstructed from the gain-coefficient blocks of \( \widehat{\theta }^{[j]}\) . If \( \rm {rank} \left( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\right) =N_{\theta }\) , then
(33)
where
(34)
where \( \mathbf{D}_{K,\ell }^{[j]}\in \mathbb{R}^{m\times n}\) are the gain coefficient blocks of \( \left( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\right) ^{\dagger }\mathbf{r}_{\rho ,T}^{[j]}.\) Moreover,
(35)
where \( c_{\psi ,T}=\sup_{t\in \lbrack 0,T]}\left( \sum_{\ell =1}^{N_{k}}\psi _{\ell }^{2}(t)\right) ^{1/2}\) .
Proof
By the definition of \( \theta _{0}^{[j]}\) , the exact coefficients satisfy \( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\theta _{0}^{[j]}=\mathbf{b}_{\rho ,T}^{[j]}+\mathbf{r}_{\rho ,T}^{[j]}.\) Since \( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\) has full column rank, we have \( \left( \boldsymbol{\Omega }_{\rho ,T}^{[j]}\right) ^{\dagger }\boldsymbol{\Omega }_{\rho ,T}^{[j]}=I_{N_{\theta }}\) . Therefore,
(36)
Let \( \widehat{\mathbf{K}}_{\ell }^{[j+1]}\) and \( \mathbf{K}_{\ell }^{[j+1]}\) denote the gain-coefficient blocks of \( \widehat{\theta }^{[j]}\) and \( \theta _{0}^{[j]}\) , respectively. Then
(37)
Using \( \widehat{\mathbf{K}}_{\rho ,T}^{[j+1]}(t)=\sum_{\ell =1}^{N_{k}}\psi _{\ell }(t)\widehat{\mathbf{K}}_{\ell }^{[j+1]}\) and (28), we obtain (33) and (34). For every \( t\in \lbrack 0,T]\) , Cauchy’s inequality gives
(38)
Thus,
(39)
Since \( \Vert (\boldsymbol{\Omega }_{\rho ,T}^{[j]})^{\dagger }\Vert _{2}=1/\sigma _{\min }(\boldsymbol{\Omega }_{\rho ,T}^{[j]})\) , (35) follows directly from Lemma 3.
Lemma 4 allows the data-driven update to be viewed as an inexact finite-horizon PI step. For a fixed horizon \( T>0\) , define the finite-horizon policy-improvement mapping as follows. For any bounded matrix-valued function \(\mathbf K:[0,T]\to\mathbb R^{m\times n}\), let \( \mathbf{P}_{\rho ,T}(t;\mathbf{K})\) be the unique solution on \([0,T]\) of
(40)
with terminal condition \( \mathbf{P}_{\rho ,T}(T;\mathbf{K})=\mathbf{M}\) . Define
(41)
Then the exact finite-horizon PI recursion can be written as \( \mathbf{K}_{\rho ,T}^{[j+1]}=\mathcal{F}_{\rho ,T}(\mathbf{K}_{\rho ,T}^{[j]})\) . Moreover, by Lemma 4, the data-driven finite-horizon recursion has the inexact form
(42)
Theorem 2
Fix \( T>0\) , and let \( \mathbf{K}_{\rho ,T}^{\ast }(t)\) be defined by (11). Suppose that the data-driven finite-horizon recursion satisfies (42). Then the following statements hold.
There exist \( a_{T}>0\) and \( \bar{\delta}_{T}>0\) such that, if \( \Vert \mathbf{K}-\mathbf{K}_{\rho ,T}^{\ast }\Vert _{\infty ,T}\leq \bar{\delta}_{T}\) , then
(43)
Choose \( \delta _{T}\in (0,\bar{\delta}_{T}]\) with \( \chi _{T}=a_{T}\delta _{T}<1\) . Assume that, for some \( j_{0}\geq 0\) , \( \left\Vert \widehat{\mathbf{K}}_{\rho ,T}^{[j_{0}]}-\mathbf{K}_{\rho ,T}^{\ast }\right\Vert _{\infty ,T}\leq \delta _{T}\) and \( \left\Vert \mathcal{E}_{T}^{[j]}\right\Vert _{\infty ,T}\leq \bar{\eta}_{T}\) , \( \bar{\eta}_{T}\leq (1-\chi _{T})\delta _{T}\) , \( j\geq j_{0}\) . Then, for all \( q\geq 0\) ,
(44)
Consequently,
(45)
Proof
Let \( \mathbf{K}\) lie in a sufficiently small bounded neighborhood of \( \mathbf{K}_{\rho ,T}^{\ast }\) . Let \( \mathbf{S}(t)=\mathbf{P}_{\rho ,T}(t;\mathbf{K})-\mathbf{P}_{\rho ,T}^{\ast }(t)\) and \( \widetilde{\mathbf{K}}(t)=\mathbf{K}(t)-\mathbf{K}_{\rho ,T}^{\ast }(t)\) . Since \( \mathbf{K}_{\rho ,T}^{\ast }(t)=\mathbf{R}^{-1}\mathbf{B}^{T}\mathbf{P}_{\rho ,T}^{\ast }(t)\) , the optimal finite-horizon Riccati equation can be rewritten as
(46)
Subtracting this equation from the policy-evaluation equation for \( \mathbf{P}_{\rho ,T}(t;\mathbf{K})\) gives
(47)
with \( \mathbf{S}(T)=\mathbf{0}\) , where \( \mathbf{A}_{K}(t)=\mathbf{A}_{\rho }-\mathbf{B}\mathbf{K}(t)\) . Let \( \boldsymbol{\Phi }_{K}(\tau ,t)\) be the transition matrix of \( \mathbf{A}_{K}(t)\) . On the compact interval \( [0,T]\) and for \( \mathbf{K}\) in a bounded neighborhood of \( \mathbf{K}_{\rho ,T}^{\ast }\) , there exists \( m_{T}>0\) such that
(48)
Therefore,
(49)
It follows that
(50)
Thus (43) holds with \( a_{T}=\left\Vert \mathbf{R}^{-1}\mathbf{B}^{T}\right\Vert _{2}Tm_{T}^{2}\Vert \mathbf{R}\Vert _{2}.\) Whenever \( \Vert \widehat{\mathbf{K}}_{\rho ,T}^{[\bar{q}]}-\mathbf{K}_{\rho ,T}^{\ast }\Vert _{\infty ,T}\leq \delta _{T}\) with \( \bar{q}=j_{0}+q\) , (43) implies
(51)
The inequality above, together with the assumptions \( \Vert \widehat{\mathbf{K}}_{\rho,T}^{[j_0]} -\mathbf{K}_{\rho,T}^{\ast}\Vert_{\infty,T} \leq \delta_T\) and \( \bar{\eta}_T\leq(1-\chi_T)\delta_T\) , implies by induction that \( \left\| \widehat{\mathbf{K}}_{\rho,T}^{[j_0+q]} -\mathbf{K}_{\rho,T}^{\ast} \right\|_{\infty,T} \leq \delta_T\) for all \( q\geq 0. \) Thus, we have
where \( \bar q=j_0+q\) . This implies (44). Since \( 0<\chi_T<1\) , taking \( q\rightarrow\infty\) yields
(52)
If \( \Vert \mathcal{E}_{T}^{[j]}\Vert _{\infty ,T}\rightarrow 0\) , the same scalar recursion with \( 0<\chi _{T}<1\) gives \( \Vert \widehat{\mathbf{K}}_{\rho ,T}^{[j]}-\mathbf{K}_{\rho ,T}^{\ast }\Vert _{\infty ,T}\rightarrow 0\) , which is uniform convergence on \( [0,T]\) .
Remark 3
Equation (42) identifies the data-driven recursion as a perturbed finite-horizon policy-improvement iteration. Theorem 5 shows that the local quadratic behavior of the nominal map \( \mathcal F_{\rho,T}\) is retained modulo the perturbation \( \mathcal{E}_{T}^{[j]}\) . Thus, the effect of basis truncation and data conditioning is quantified through \( \mathcal{E}_{T}^{[j]}\) : uniformly small perturbations yield practical convergence, whereas vanishing perturbations recover convergence to \( \mathbf{K}_{\rho,T}^{\ast}\) .
Corollary 1
Suppose that Assumption 1 holds. Let \( \rho >\varepsilon \geq 0\) , and choose \( T\geq T_{\varepsilon }\) , where \( T_{\varepsilon }\) is given in Theorem 2. Suppose further that the hypotheses of Theorem 5(ii) hold. Then there exists \( r_{T}>0\) such that
(53)
If \( \frac{\bar{\eta}_{T}}{1-\chi _{T}}<r_{T}\) , then \( \widehat{\mathbf{K}}_{\rho ,T}^{[j]}(0)\) is \( \varepsilon \) -admissible for the original system (1) for all sufficiently large \( j\) .
Proof
Since \( T\geq T_{\varepsilon }\) , we have \( \alpha \left( \mathbf{A}-\mathbf{B}\mathbf{K}_{\rho ,T}^{\ast }(0)\right) <-\varepsilon .\) By continuity of the spectral abscissa with respect to the feedback gain, there exists \( r_{T}>0\) such that (53) holds. On the other hand, Theorem 5(ii) yields
(54)
which implies that, for all sufficiently large \( j\) ,
(55)
Therefore, by (53), \( \alpha \left( \mathbf{A}-\mathbf{B}\widehat{\mathbf{K}}_{\rho ,T}^{[j]}(0)\right) <-\varepsilon .\) Hence, \( \widehat{\mathbf{K}}_{\rho ,T}^{[j]}(0)\) is \( \varepsilon \) -admissible for all sufficiently large \( j\) .
Corollary 6 gives a theoretical admissibility guarantee, but the neighborhood radius involved there is model dependent and cannot be checked directly from data. Therefore, we introduce a data-driven Lyapunov certificate for the candidate constant gain generated by the finite-horizon stage. Let \( \mathbf{K}_{c}=\widehat{\mathbf{K}}_{\rho ,T}^{[j+1]}(0)\) be a candidate constant gain generated by the data-driven finite-horizon recursion. Choose a matrix \( \mathbf{W}_{c}=\mathbf{W}_{c}^{T}>\mathbf{0}\) . The purpose of the following test is to determine, using the stored behavior data only, whether there exists a matrix \( \mathbf{P}_{c}=\mathbf{P}_{c}^{T}>\mathbf{0}\) such that \( (\mathbf{A}-\mathbf{B}\mathbf{K}_{c}+\varepsilon \mathbf{I})^{T}\mathbf{P}_{c}+\mathbf{P}_{c}(\mathbf{A}-\mathbf{B}\mathbf{K}_{c}+\varepsilon \mathbf{I})=-\mathbf{W}_{c}\) . If such a matrix \( \mathbf{P}_{c}\) exists, then \( \mathbf{K}_{c}\) is \( \varepsilon \) -admissible for the original system.
Define \( \mathbf{L}_{c}=\mathbf{B}^{T}\mathbf{P}_{c} \) and \( z_{c}(t)=u_{b}(t)+\mathbf{K}_{c}x(t)\) . The desired Lyapunov equation is equivalent to
(56)
Along the trajectory \( \dot{x}(t)=\mathbf{A}x(t)+\mathbf{B}u_{b}(t),\) one has
(57)
Integrating (57) over \( I_{k}=[t_{k},t_{k+1}]\) gives
(58)
For each interval \( I_{k}\) , define
Then (58) can be written as
(59)
Stacking all stored intervals yields
(60)
where
with
The following result gives a data-driven admissibility certificate for the candidate gain.
Lemma 4
Let \( \mathbf{K}_{c}\) be a candidate constant gain and let \( \mathbf{W}_{c}=\mathbf{W}_{c}^{T}>\mathbf{0}\) be given. Suppose that the stored data satisfy the data-richness condition
(61)
where
If (60) admits a solution \( \theta _{c} \) whose component satisfies \( \mathbf{P}_{c} = \mathbf{P}_{c}^{T} > \mathbf{0},\) then \( \mathbf{K}_{c}\) is \( \varepsilon \) -admissible.
Proof
Let \( \theta _{c}\) be a solution of (60), and let the associated matrices be \( \mathbf{P}_{c}=\mathbf{P}_{c}^{T}\) . By the system dynamics, for any symmetric matrix \( \mathbf{P}_{c}\) ,
(62)
Combining this identity with (58) gives, for all stored intervals,
(63)
By the rank condition (61), it follows that
(64)
Since \( \mathbf{P}_{c}>\mathbf{0}\) and \( \mathbf{W}_{c}>\mathbf{0}\) , the Lyapunov theorem implies that \( \mathbf{A}-\mathbf{B}\mathbf{K}_{c}+\varepsilon \mathbf{I}\) is Hurwitz, which means that \( \mathbf{K}_{c}\) is \( \varepsilon \) -admissible.
The data-driven bootstrap scheme is summarized in Algorithm 2 and illustrated in Fig. 2. It consists of two nested loops: an inner finite-horizon PI loop that generates a candidate gain at a fixed horizon, and an outer horizon-expansion loop that is activated whenever the candidate fails the data-driven admissibility test. Once the certificate succeeds, the candidate gain is returned as an initializer for the subsequent infinite-horizon PI.
Remark 4
Since the quantities reconstructed from (27)–(28) satisfy the off-policy identity only up to \( \mathbf r_{\rho,T}^{[j]}\) , a stability test built on them would inherit basis-truncation errors and would not certify the constant gain applied to the original system. Therefore, after \( \mathbf K_c=\widehat{\mathbf K}_{\rho,T}^{[j+1]}(0)\) is generated, an independent Lyapunov variable \( \mathbf P_c\) is introduced to verify admissibility for \( \mathbf A-\mathbf B\mathbf K_c+\varepsilon \mathbf I\) .
Remark 5
The proposed data-driven Lyapunov test in Algorithm 2 is both sufficient and necessary in the exact-data setting. This is enabled by introducing an independent Lyapunov variable \( \mathbf{P}_{c}\) , rather than testing stability using a value matrix inherited from the learning iteration. In contrast, the stopping criteria used in the hybrid and homotopy-based initialization schemes considered in [28, 20] are sufficient stability tests and may therefore certify a gain only after it has already entered the admissible region.
In this section, a batch reactor system [38] is used to illustrate the proposed finite-horizon bootstrap method. The system matrices are given by
The open-loop poles are \( 1.9910\) , \( 0.0635\) , \( -8.6659\) , and \( -5.0566\) , which shows that the uncontrolled system is unstable. The weighting matrices are selected as \( \mathbf{Q}= \rm {diag} (1,1,0.1,0.01)\) , \( \mathbf{R}=I_{2}\) , and the terminal weight is chosen as \( \mathbf{M}=\mathbf{0}\) . We first validate the model-based finite-horizon bootstrap procedure. The design parameters are set as \( \varepsilon =0.2\) , \( \rho =0.6\) , \( T_{0}=0.1\) , \( \epsilon_{\mathrm{PI}}=10^{-7}\) and \( \gamma =1.25\) . The finite-horizon policy is initialized by \( \mathbf{K}_{\rho ,T_{s}}^{[0]}(t)=\mathbf{0}\) for \( t\in \lbrack 0,T_{s}]\) , and the horizon is updated according to \( T_{s+1}=\gamma T_{s}\) .
The numerical results are summarized in Table 2 and Fig. 3. It can be seen that, as the finite horizon \( T_s\) increases, the spectral abscissa of the closed-loop matrix \( \mathbf{A}-\mathbf{B}\mathbf{K}_{c}\) decreases monotonically in this example. When \( T_{s}=0.3052\) , the finite-horizon bootstrap produces a gain satisfying \( \alpha (\mathbf{A}-\mathbf{B}\mathbf{K}_{c})<-0.2. \) Therefore, the algorithm terminates and the obtained gain is accepted as an \( \varepsilon \) -admissible initializer.
| \( s\) | \( T_s\) | Iterations | \( \alpha(\mathbf{A}-\mathbf{B}\mathbf{K}_c)\) |
| 0 | 0.1000 | 4 | 1.8108 |
| 1 | 0.1250 | 4 | 1.6755 |
| 2 | 0.1563 | 4 | 1.4501 |
| 3 | 0.1953 | 5 | 1.0694 |
| 4 | 0.2441 | 5 | 0.3866 |
| 5 | 0.3052 | 5 | -0.5822 |
The resulting stabilizing initializer is
The corresponding closed-loop poles are \( -0.5822\pm 0.5852i\) , \( -8.3664\) , and \( -7.6089\) , where \( i\) represents the imaginary unit, which confirms that the proposed finite-horizon bootstrap successfully generates an admissible initial policy.
We next evaluate the data-driven bootstrap procedure in Algorithm 2. The certificate matrix is chosen as \( \mathbf{W}_{c}=I_{4}\) . The initial states are randomly generated with norms in \( [0.3,0.7]\) . The basis functions (27)–(28) are selected as
Thus \( N_{p}=N_{k}=8\) . The data-driven PI loop is initialized by \( \widehat{\mathbf{K}}_{\rho ,T_{s}}^{[0]}(t)=\mathbf{0}\) . The data-driven results are shown in Fig. 4. For all tested horizons, the data-richness condition for the Lyapunov certificate is satisfied with \( \rm {rank} =18\) , and the finite-horizon regression matrix has full column rank, i.e., \( \rm {rank} (\boldsymbol{\Omega}_{\rho,T_s}^{[j]})=144\) . For numerical implementation, the linear equation (60) is solved in the least-squares sense. The relative residual
(65)
is reported only as a numerical accuracy indicator, where \( \widetilde{\boldsymbol{\Psi}}_c\) and \( \widetilde{\mathbf{b}}_c\) denote the row-scaled data matrix and right-hand side. As \( T_s\) increases, the spectral abscissa of the candidate closed-loop matrix moves toward the \( \varepsilon\) -admissible region. The candidates generated for the first five horizons are rejected by the data-driven Lyapunov certificate. At \( T_s=0.3052\) , the data-driven finite-horizon PI converges in five iterations and produces
For this gain, the Lyapunov certificate (60) has relative residual \( 7.432\times 10^{-4}\) and \( \lambda_{\min}(\mathbf{P}_c)=5.938\times 10^{-2}>0\) . Therefore, the certificate succeeds and Algorithm 2 terminates. Thus the gain certified from data is an \( \varepsilon\) -admissible initializer for the subsequent infinite-horizon PI. Using \( \mathbf{K}_c\) as the initial policy, the subsequent off-policy RL converges to the optimal solution \( \mathbf{K}^\ast\) within a few iterations. As shown in Fig. 5, both \( \|\mathbf{P}^{[j]}-\mathbf{P}^\ast\|\) and \( \|\mathbf{K}^{[j]}-\mathbf{K}^\ast\|\) decrease rapidly, confirming that the bootstrap gain provides a valid initializer for the infinite-horizon data-driven PI stage.
We next consider the two-mass-spring system used in the homotopic PI study [28] to compare Algorithm 2 with existing data-driven RL methods under the same plant and cost. The system matrices are
The physical parameters are \( k_{1}=1.5 \mathrm{N/m}\) , \( k_{2}=1 \mathrm{N/m}\) , \( m_{1}=1.1 \mathrm{kg}\) , and \( m_{2}=0.9 \mathrm{kg}\) . The weights are selected as \( \mathbf{Q}= \rm {diag} (1,1,2,1)\) and \( \mathbf{R}=I_{1}\) .
For the existing data-driven homotopic PI method [28], the reported simulation uses \( \mathbf{K}^{[0]}=\mathbf{0}\) and \( \beta -\alpha ^{[0]}=0.1\) . The data are collected over \( [0,5]\) from the initial condition \( x(0)=[0.2,0.2,0.2,0.2]^{T} \) , with a probing input formed by sine and cosine signals of different frequencies. As shown in Fig. 6, the learned stabilizing gain is
which gives the closed-loop poles \( -1.51455\pm 1.66164i\) and \( -1.47887\pm 0.752358i\) .
For comparison, the hybrid PI method [20] uses \( \mathbf{P}^{[0]}=0.1I_{4}\) , \( \epsilon _{j}=1/(j+2)\) , and \( \mathcal{B}_{q}=\{\mathbf{P}>\mathbf{0}:\Vert \mathbf{P}\Vert \leq 5(q+1)\}\) , \( q=0,1,… \) . At each VI iteration, \( \mathbf{H}^{\left[ j\right] }\) and \( \mathbf{K}^{\left[ j\right] }_{v}\) are recovered from the data equation, and \( \widetilde{\mathbf{P}}^{\left[ j+1\right] }\) is generated by the VI update. If \( \widetilde{\mathbf{P}}^{\left[ j+1\right] }\notin \mathcal{B}_{q}\) , the iteration is reset to \( \mathbf{P}^{[0]}\) and \( q\) is increased by one. Otherwise, the VI stage is terminated once the data-driven Lyapunov certificate \( \mathbf{P}^{\left[ j\right] }>\mathbf{0}\) and \( \mathbf{H}^{\left[ j\right] }-2(\mathbf{K}^{[j]}_{v})^{T}\mathbf{R}\mathbf{K}^{\left[ j\right] }_{v}<\mathbf{0}\) is satisfied, where \( \mathbf{H}^{\left[ j\right] }=\mathbf{A}^{T}\mathbf{P}^{\left[ j\right] }+\mathbf{P}^{\left[ j\right] }\mathbf{A}\) . Fig. 7 shows the VI-stage learning process. After one reset caused by the bounded-set test, the spectral abscissa \( \alpha(\mathbf{A}-\mathbf{B}\mathbf{K}^{[j]}_{v})\) remains negative and decreases gradually, as shown in Fig. 7. Meanwhile, \( \lambda_{\max}(\mathbf{H}^{[j]}-2(\mathbf{K}^{[j]}_{v})^{T}\mathbf{R}\mathbf{K}^{\left[ j\right] }_{v})\) crosses zero at the \( 193\) rd VI iteration, which activates the data-driven termination certificate. The certified initializer is
and the corresponding closed-loop poles are \( -0.3376\pm1.6540i\) and \( -0.3607\pm0.7862i\) .
For Algorithm 2, we use \( \varepsilon =0\) , \( \rho =0.1\) , \( T_{0}=0.1\) , \( \gamma =1.25\) , \( \mathbf{M}=\mathbf{0}\) , and \( \mathbf{W}_{c}=I_{4}\) . The finite-horizon policy is initialized by \( \widehat{\mathbf{K}}_{\rho ,T_{s}}^{[0]}(t)=\mathbf{0}\) , and the same basis functions and stopping tolerance \( 10^{-7}\) as in the preceding Section 5.1 are used. As shown in Fig. 8, the data-driven bootstrap procedure succeeds at the first tested horizon \( T_s=0.1 \) . The resulting candidate gain is
For this gain, the data-driven Lyapunov certificate gives \( \lambda_{\min}(\mathbf{P}_c)=10.84354>0\) and a relative residual \( 4.732\times10^{-4}\) , which is below the prescribed tolerance. The closed-loop spectral abscissa is \( \alpha(\mathbf{A}-\mathbf{B}\mathbf{K}_{\mathrm{c}}) =-1.0445\times10^{-2}<0\) , confirming that the certified initializer is admissible. The results in Figs. 6-8 illustrate the distinct advantage of the proposed bootstrap mechanism. The homotopic PI [28] and hybrid PI [20] both succeed in generating admissible initializers, but their certification requires \( 101\) stabilizing iterations and \( 193\) iterations, respectively. By contrast, Algorithm 2 succeeds at the first tested horizon \( T_s=0.1\) and the finite-horizon PI converges in only three iterations.
This article develops a finite-horizon bootstrap method for data-driven infinite-horizon PI without requiring an initial stabilizing policy. By exploiting the well-posedness of finite-horizon policy evaluation and enlarging the horizon for a shifted system, the method generates an admissible candidate gain. An off-policy implementation reconstructs the time-varying value matrix and policy from measured trajectories, while an independent data-driven Lyapunov certificate verifies the candidate gain before it initializes infinite-horizon PI. Numerical studies demonstrate the effectiveness of the proposed method.
[1] Optimal Control John Wiley & Sons 2012
[2] Optimal Control Theory for Applications Cham, Switzerland: Springer, 2003.
[3] Adaptive dynamic programming-regulated extremum seeking for distributed feedback optimization iElectron. Lettėctr. Engėac 2025 70 11 7675–7682 Nov.
[4] On an iterative technique for Riccati equation computations iElectron. Lettėctr. Engėac 1968 13 1 114–115 Feb.
[5] Efficient off-policy Q-learning for data-based discrete-time LQR problems iElectron. Lettėctr. Engėac 2023 68 5 2922–2933 May
[6] Learning-based adaptive optimal control of linear time-delay systems: A value iteration approach AutoMechatronicsatica 2025 171 111944
[7] Adaptive optimal control for continuous-time linear systems based on policy iteration AutoMechatronicsatica 2009 45 2 477–484
[8] Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics AutoMechatronicsatica 2012 48 10 2699–2704
[9] Data-driven near optimization for fast sampling singularly perturbed systems iElectron. Lettėctr. Engėac 2024 69 7 4689–4694 Jul.
[10] Recent advances on off-policy reinforcement learning for optimization control iElectron. Lettėctr. Engėc in press, DOI: 10.1109/TCYB.2026.3683384
[11] Output feedback Q-learning for discrete-time linear zero-sum games with application to the H-infinity control AutoMechatronicsatica 2018 95 213–221
[12] Adaptive dynamic programming and adaptive optimal output regulation of linear systems iElectron. Lettėctr. Engėac 2016 61 12 4164–4169 Dec.
[13] Data-driven inverse reinforcement learning control for linear multiplayer games iElectron. Lettėctr. EngėNeural Netw.ls 2024 35 2 2028–2041 Feb.
[14] Off-Policy Reinforcement Learning for \( H_{\infty}\) Control of Linear Discrete-Time Systems with Network Induced Dropouts iElectron. Lettėctr. Engėac 2025 70 12 8000–8015 Dec.
[15] Value iteration for continuous-time linear time-invariant systems iElectron. Lettėctr. Engėac 2023 68 5 3070–3077 May
[16] Value iteration adaptive dynamic programming for optimal control of discrete-time nonlinear systems iElectron. Lettėctr. Engėc 2016 46 3 840–853 Mar.
[17] Balancing value iteration and policy iteration for discrete-time control iElectron. Lettėctr. EngėsMechatronicSoft Computṡ 2020 50 11 3948–3958 Nov.
[18] Value iteration and adaptive dynamic programming for data-driven adaptive optimal control design AutoMechatronicsatica 2016 71 348–360
[19] Adaptive optimal control of linear periodic systems: An off-policy value iteration approach iElectron. Lettėctr. Engėac 2021 66 2 888–894 Feb.
[20] Resilient reinforcement learning and robust output regulation under denial-of-service attacks AutoMechatronicsatica 2022 142 110366
[21] Distributed FilterNet reinforcement learning for achieving output consensus in heterogeneous multiplayer multiagent systems iElectron. Lettėctr. EngėNeural Netw.ls 2026 37 2 575–588 Feb.
[22] Experimental validation of data-driven adaptive optimal control for continuous-time systems via hybrid iteration: An application to rotary inverted pendulum iElectron. Lettėctr. Engėie 2024 71 6 6210–6220 Jun.
[23] Secure control for Markov jump cyber-physical systems subject to malicious attacks: A resilient hybrid learning scheme iElectron. Lettėctr. Engėc 2024 54 11 7068–7079 Nov.
[24] Secure optimal control of Itô stochastic Markov jump systems subject to DoS attacks: A hybrid learning algorithm AutoMechatronicsatica 2026 183 112681
[25] Prescribed-time observer-based HI-RL secure output tracking control for heterogeneous MASs under DoS attacks iElectron. Lettėctr. EngėsMechatronicSoft Computṡ 2026 56 1 709–723 Jan.
[26] Cooperative adaptive cruise control of connected and autonomous vehicles via hybrid iteration iElectron. Lettėctr. Engėvt 2026 75 6 8793–8804 Jun.
[27] Bias-policy iteration based adaptive dynamic programming for unknown continuous-time linear systems AutoMechatronicsatica 2022 136 110058
[28] Homotopic policy iteration-based learning design for unknown linear continuous-time systems AutoMechatronicsatica 2022 138 110153
[29] A homotopy method for continuous-time model-free LQR control based on policy iteration iElectron. Lettėctr. Engėcaa 2025 12 8 1673–1682 Aug.
[30] Adaptive dynamic programming for optimal control of unknown LTI system via interval excitation iElectron. Lettėctr. Engėac 2025 70 7 4896–4903 Jul.
[31] Model-free \( \lambda\) -policy iteration for discrete-time linear quadratic regulation iElectron. Lettėctr. EngėNeural Netw.ls 2023 34 2 635–649 Feb.
[32] Data-driven single-loop policy iteration control of uncertain singularly perturbed systems iElectron. Lettėctr. Engėac 2025 70 12 8314–8320 Dec.
[33] Novel single-loop policy iteration for linear zero-sum games AutoMechatronicsatica 2024 163 111551
[34] Adaptive optimal control of unknown nonlinear systems via homotopy-based policy iteration iElectron. Lettėctr. Engėac 2024 69 5 3396–3403 May
[35] Memory-Efficient Inverse Reinforcement Learning for Multiplayer Differential Games iElectron. Lettėctr. Engėc 2025 55 11 5545–5558 Nov.
[36] Cooperative Optimal Output Tracking for Discrete-Time Multiagent Systems: Stabilizing Policy Iteration Frameworks iElectron. Lettėctr. Engėac 2026 71 4 2746–2753 Apr.
[37] A Parallel Homotopic Optimized Control Scheme of Uncertain Nonlinear Markov Jump Systems and Its Applications iElectron. Lettėctr. Engėase 2025 22 19403–19414
[38] Stability analysis of networked control systems iElectron. Lettėctr. Engėcst 2002 10 3 438–446 May