
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39278944
71844
10.1038/s41598-024-71844-y
Article
Neural network optimal control for tripartite UAV confrontation systems based on fuzzy differential game
Fu Xingjian fxj@bistu.edu.cn

Yan Hang
https://ror.org/04xnqep60 grid.443248.d 0000 0004 0467 2584 School of Automation, Beijing Information Science and Technology University, Beijing, 100192 China
16 9 2024
16 9 2024
2024
14 215478 5 2024
31 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
The neural network optimal control strategy based on a fuzzy differential game is proposed for the tripartite UAV confrontation systems consisting of the attackers, defenders, and targets. Firstly, the tripartite UAV mutual confrontation model is constructed and a nonlinear differential control system is established. Secondly, combining the fuzzy evaluation method and differential game theory, the tripartite UAV are divided into two parts of the confrontation game: attackers-defenders and attackers-targets. The optimal control strategies for the attackers, defenders and targets parties are derived separately. Then, the tripartite UAV game model is considered to be difficult to solve directly. The evaluation neural network is introduced to approximate the optimal value function using an adaptive dynamic programming method. The convergence of the evaluation neural network weights and the stability of the nonlinear differential control system are proved by using Lyapunov stability theory. Finally, the effectiveness of the tripartite UAV confrontation game control strategy designed in this paper is verified by simulation.

Keywords

Differential game
Fuzzy evaluation methods
Tripartite UAVs
Confrontation games
Neural networks
Subject terms

Applied mathematics
Computer science
Information technology
Aerospace engineering
National Natural Science Foundation of China62103057 Fu Xingjian issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

With the continuous progress of aviation technology and the increasing military demand, the UAV-related technology has been fully developed, and its importance will become more and more prominent in various fields. In the civil field, UAVs play an important role in agriculture, logistics, environmental monitoring, etc., and the productivity and accuracy of data collection are improved. In the military field, the application of UAVs in reconnaissance, target striking, and battlefield monitoring has further strengthened the combat capability of the military1–4. UAVs are being used for more dangerous and high-risk missions than traditional manned fighters. Through advanced technologies such as remote operation and autonomous navigation, UAVs are able to realize collaborative operations and the overall combat capability is enhanced. This provides commanders with more options and tactical flexibility5–8.

For the problem of multi-UAV confrontation, there are usually complex conflicts of interest among all participants, leading to complex tactics and decision-making. To cope with this situation, game theory has been widely used to analyze and solve the optimal strategy problem under multi-party conflicts. In9, a new multi-UAV confrontation model is proposed. Information uncertainty is considered in this model. A posture matrix is introduced to express this uncertainty. To address the effects of uncertain information, an interval decision-making approach is used to solve the game model. In10, the problem of information uncertainty faced in the decision making of UAV attack and defense game is investigated. An interval number is used to represent the uncertain information. A quantum particle swarm optimization algorithm is used to solve the mixed-strategy Nash equilibrium solution of the two-party UAV adversarial game. In11, for the hostile combat environment at sea, the adversarial combat model and objective function between two UAV swarms are established separately. A heuristic SL-PSO algorithm is proposed to solve the optimal Nash strategy of the game between the two UAVs, so as to optimally allocate the strike missions between the UAVs.

Differential game is an important branch of game theory. It primarily involves making decisions in a dynamic environment and considering the reactions and behaviors of other participants. In UAV dynamic confrontation, differential game can be used to make strategic and tactical decisions to maximize self-interest and reduce the advantage of the adversary. In12, optimal control rules and differential game theory to calculate the optimal paths of UAVs on both sides of the escape chase is used. Artificial neural networks are used to predict the flight trajectories of the escaping party UAVs. In13, an innovative feedforward controller for the problem of missile interception of a target vehicle is designed. An adaptive dynamic programming algorithm is utilized to solve the Nash equilibrium strategy of the zero-sum differential game. An evaluation neural network is constructed to approximate the optimal solution of Hamilton–Jacobi-Isaacs (HJI) equation online. In14, the engagement problem of multiple offensive bombs intercepting a target vehicle is addressed. A differential game guidance method applied to incomplete target information is proposed. In15, the attack-defense game guidance problem for interceptor precision strikes against maneuvering targets is addressed. A differential game guidance law based on adaptive dynamic programming (ADP) is devised. A neural network approximation of the optimal cost function is introduced, and the optimal strategies of both sides of the attack and defense game are solved.

Most of the research results in the above references are about two parties playing a dynamic adversarial game. In real situations, there may also be multi-party games, which is also the focus of some scholars' research. In16, a tripartite confrontation game problem is addressed for the attacker, defender and target UAV. The scheme which the target and the defender cooperate to form a team against the attacker is designed to maximize the benefits to the target and the defender team. In17, the problem for an attacker against an actively defending aircraft in a tripartite confrontation is addressed. A novel strategy combining linear quadratic and differential game theory is proposed. Differential game theory and the solution of the Hamiltonian equation are utilized. A combined guidance strategy for each belligerent is obtained. In18, the active target defense problem is considered. The attacker missile pursues the target aircraft and the defender missile is responsible for intercepting the attacker. Based on the value functions of the tripartite confrontation, the optimal strategy for the tripartite confrontation game is devised. In19, the pursuit and escape problem with multiple attackers, multiple defenders and a single target is considered. A tripartite optimal control strategy based on a linear quadratic differential game was designed. Bipartite graph matching algorithm is utilized. The pursuit problem between multiple attackers and multiple defenders is transformed into the form of a multi-group two-player zero-sum differential game. The optimal strategies for both attackers and defenders are solved.

Most of the above studies on three-party adversarial games consider cooperation between two parties, and the three-party adversarial game is transformed into a two-party adversarial game problem for research. In this paper, a game control method based on fuzzy differential game is proposed for the three-party UAV adversarial game problem of attacker, defender, and target. The main contributions of this paper are denoted as follows:The confrontation model between the three-party UAVs is established, and the three-party UAVs are divided into two parts of attacker-defender and attacker-target to be studied.

A fuzzy evaluation method is used to determine the scaling coefficient of the attacking party in the two-part game, and differential game theory is utilized to derive the Nash equilibrium control strategies for the two-part game respectively.

Since the Hamilton–Jacobi–Isaacs (HJI) equation is difficult to be solved directly, two evaluation neural networks are introduced to approximate the optimal value function by utilizing the adaptive dynamic programming method.

Simulations are conducted to compare with the traditional proportional guidance control method, and the effectiveness and superiority of the designed tripartite UAV adversarial game control strategy are verified.

Tripartite UAVs modeling and problem description

The tripartite UAVs confrontation game problem for the attacker, defender, and target is considered in this paper. The attacker wants to pursue to the target, the defender intercepts the attacker, the target avoids the attacker, then the tripartite game situation is formed.

The confrontation relationship among the tripartite UAVs is shown in Fig. 1. A, D and T denote the attacker, defender, and target of the UAVs, respectively. VA, VD and VT are attacker speed, defender speed, and target speed, respectively; θAT and RAT denote the line-of-sight angle and relative distance between the attacking party A and the target party T, respectively; θAD and RAD are the line-of-sight angle and relative distance between the attacker A and the defender D, respectively; uA, vD and wT denote the control inputs for each of A, D and T, respectively, oriented perpendicular to the direction of the respective velocity vectors; α, β and γ are the track angles of A, D and T respectively.Fig. 1 The confrontation relationship between the tripartite UAVs.

It is assumed that the systems of the attacker A, the defender D and the target T are all first-order systems. Combining the tripartite UAVs motion relations represented in Fig. 1, the motion equations for A, D and T can be obtained20.

The motion equation of the attacker A is:1 x˙A=VAcosαy˙A=VAsinαα˙=aAVAa˙A=uA-aAτA

where, xA,yA is the position of the attacker A in the two-dimensional plane; aA is its lateral acceleration; τA is a time constant taking the value τA=0.1.

The motion equation for the defender D is:2 x˙D=VDcosβy˙D=VDsinββ˙=aDVDa˙D=vD-aDτD

where, xD,yD is the position of the defender D in the two-dimensional plane; aD is lateral acceleration; τD is a time constant taking the value τD=0.1.

The motion equation for the target T is:3 x˙T=VTcosγy˙T=VTsinγγ˙=aTVTa˙T=wT-aTτT

where, xT,yT is the position of the target T in the two-dimensional plane; aT is its lateral acceleration; τT is the time constant, which takes the value of τT=0.1.

Suppose that the speeds of the attacker A, the defender D and the target T remain constant in the end game phase. Then the relative motion equations between A and D, T can be established.

The relative motion equation between attacker A- defender D is21:4 R˙AD=VAD=VDcosβ-θAD-VAcosα-θADθ˙AD=δAD=VDsinβ-θAD-VAsinα-θADRADα˙=uAVAβ˙=vDVD

where, VAD is the relative proximity rate for the relative distance RAD between the attacker A and the defender D; δAD is the line-of-sight angle rate for the line-of-sight θAD.

The relative motion equation between attacker A- target T is:5 R˙AT=VAT=VTcosγ-θAT-VAcosα-θATθ˙AT=δAT=VTsinγ-θAT-VAsinα-θATRATα˙=uAVAγ˙=wTVT

where, VAT is the relative proximity rate for the relative distance RAT between the attacker A and the target T; δAT is the line-of-sight angular rate for the line-of-sight angle θAT.

When all three parties, attacker A, defender D and target T, are no longer maneuvering, the corresponding minimum distance expression is22:6 RminADt=RAD2δADVAD2+RAD2δAD2

7 RminATt=RAT2δATVAT2+RAT2δAT2

where, RminAD and RminAT are the final minimum distances between the attacker A and the defender D, the target T, respectively.

Remark 1

From the Eqs. (6) and (7), it can be seen that the attacker A, who requires to minimize the final distance to the target T, has to choose the suitable control strategy in order to make the line-of-sight angular rate δAT converge to zero. Similarly, the goal of the defender D is to minimize the final distance to the attacker A. A suitable control strategy is to be chosen so that the line-of-sight angular rate δAD tends to zero.

In the attacker-defender-target dynamic confrontation game, the attacker pursues the target by choosing the optimal control strategy. At the same time, the attacker has to avoid the defender. The defender needs to choose the optimal control strategy to intercept the attacker. The target party has to choose the optimal control strategy to avoid the attacking party.

Nonlinear differential game control systems

Based on the above analysis of the tripartite confrontation game, the UAV state variables are selected as x1=δAD, x2=δAT, and the line-of-sight angular rates δAD and δAT are derived, respectively, then8 δ˙AD=-2VDcosβ-θAD-VAcosα-θADRADθ˙AD+cosβ-θADRADvD-cosα-θADRADuA

9 δ˙AT=-2VTcosγ-θAT-VAcosα-θATRATθ˙AT-cosα-θATRATuA+cosγ-θATRATwT

Substitute x1=δAD, x2=δAT and VAD, VAT in Eqs. (4) and (5) into Eqs. (8) and (9). Then, the state equations for the attacker A-defender D and the attacker A-target T can be expressed as follows, respectively.10 x˙1=-2VADRADx1+cosβ-θADRADvD-cosα-θADRADuA

11 x˙2=-2VATRATx2-cosα-θATRATuA+cosγ-θATRATwT

Defining f1x1=-2VADRADx1,k1x1=cosβ-θADRAD,g1x1=-cosα-θADRAD,f2x2=-2VATRATx2, g2x2=-cosα-θATRAT,h2x2=cosγ-θATRAT. Then, the Eqs. (10) and (11) can be described as nonlinear systems as shown below, respectively.12 x˙1=f1x1+k1x1vD+g1x1uA

13 x˙2=f2x2+g2x2uA+h2x2wT

From the Eqs. (12) and (13), RAD and RAT cannot be equal to 0. Therefore, in this paper, the minimum action distance Rmin1 and Rmin2 are defined. When the distance RAD<Rmin1 between the attacker A and the defender D, the defender D wins. When the distance RAT<Rmin2 between the attacker A and the target T, the attacker A wins, otherwise, the target T wins. At this point, the trilateral dynamic confrontation game ends.

When β-θAD=π2, α-θAD=π2, α-θAT=π2, and γ-θAT=π2, k1x1, g1x1, g2x2 and h2x2 are equal to 0, the two nonlinear systems are not controllable. Therefore, the feasible domains of the nonlinear systems (12) and (13) are derived as respectively.14 Ψ1=x1β-θAD≠π2,α-θAD≠π2,RAD≠0

15 Ψ2=x2α-θAT≠π2,γ-θAT≠π2,RAT≠0

Trilateral confrontation fuzzy differential game design

In this paper, the attack-defense-target tripartite dynamic confrontation game is divided into two parts of the confrontation game between the attacker A-defender D and the attacker A-target T.

In an attacker A- defender D confrontation game, the optimal control strategy is chosen by attacker A, and the interception by defender D is avoided. The defender D then needs to intercept the attacker A. In an attacker A-target T confrontation game, the optimal control strategy is chosen by the attacker A. Target T is pursued. Target T needs to avoid attacker A.

Fuzzy evaluation design

In tripartite dynamic confrontation game, the attacker A-defender D and attacker A-target T components of the game need to be considered by attacker A simultaneously. The control strategy for capturing the target T corresponding to attacker A and the control strategy for avoiding defender D are derived separately. The weighted combination of the above two strategies is constituted as the optimal control strategy for the attacker A. The fuzzy evaluation method23 is introduced by this section. The current state is evaluated. The corresponding weights of the two control strategies are obtained.

The control strategy of attacker A can be formulated as follows.16 uAt=b1uAet+b2uApt

where, uAet is the control strategy of the attacker A evading the defender D. uApt is the control strategy of the attacker A capturing the target T. b1, b2 are the corresponding weight coefficients satisfying b1+b2=1, and b1≥0, b2≥0.

When b1>b2, avoiding the defender D is given more importance by the attacker A. The distance to the defending party is desired to be increased. The defending party is kept away. When b1<b2, capturing the target T is emphasized by the attacker A. The distance to the target is expected to decrease. The target party is moved closer.

In this paper, the fuzzy evaluation method is used to determine the values of b1 and b2, b1 and b2 are categorized into 11 levels, as shown in Table 1.Table 1 Coefficient evaluation ratings.

level	1	2	3	4	5	6	7	8	9	10	11	
b1	1	0.9	0.8	0.7	0.6	0.5	0.4	0.3	0.2	0.1	0	
b2	0	0.1	0.2	0.3	0.4	0.5	0.6	0.7	0.8	0.9	1	

Remark 2

The factors affecting the evaluation rating can be categorized into two types. One is the urgency of the attacker A to avoid the defender D at the current moment, denoted by μ1. The other is the urgency of the attacker A to pursue the target T at the current moment, denoted by μ2. As the distance between the attacker and the defender gets smaller, the greater the effect of μ1 on the evaluation rating, and thus the value of b1. As the distance between the attacker and the target gets smaller, the greater the effect of μ2 on the rank, the greater the value of b2 will be.

It is assumed that the weighting coefficient of the influence factor μ1 is c1, and the weight coefficient of influence factor μ2 is c2, then the formulas for c1 and c2 are shown below:17 c1=1-c2c2=1-RATtRATt+RADt3

Thus, the weighting coefficients of the factors μ1 and μ2 can be expressed as C=c1,c2.

In order to calculate the membership values of factors μ1 and μ2 to the rating scale, the following nonlinear membership function is defined:18 μ1x=1-μ2xμ2x=kx-13

where, k=0.1, x is evaluation levels, x∈1,11, and x is a positive integer.

Let μ^ij=μij,i=1,2,j=1,2,⋯11, then the fuzzy evaluation matrix is O=μ^ij2×11.

Then, the formula for the integrated fuzzy evaluation vector Q is shown below:19 Q=C∘O=qj1×11,j=1,2,⋯11

where, “∘“ is the weighted average type fuzzy evaluation synthesis operator. It can fully utilize the information of the fuzzy evaluation matrix O and has high comprehensive performance; qj is a specific element in the vector Q, and its specific expression is shown follows:20 qj=min1,∑i=12ciμ^ij,j=1,2,⋯11

All the combined fuzzy evaluation results are subjected to a normalization operation.21 q~j=qj∑j=111qj,j=1,2,⋯11

Defining the normalized vectors as Q~=q~j1×11,j=1,2,⋯11. The corresponding components in Q~ are weighted and summed using the weighted average method to obtain the composite evaluation value as shown below:22 q=∑j=111q~jα⌢1·j∑j=111q~jα⌢1

where, α⌢1=10 is an exponential coefficient. The role played by q~j can be controlled. The larger the value of α⌢1, the more pronounced will be the effect of the larger term in q~j.

Remark 3

The specific values of b1 and b2 can be obtained from the fuzzy evaluation scale differences corresponding to the composite evaluation value q. If q is equal to an evaluation level, then b1 and b2 are the values corresponding to the level in the table. If q is between two evaluation levels, the values of b1 and b2 are obtained by prorating the weighting factors corresponding to the two evaluation levels.

Fuzzy differential game control design

Based on the fuzzy evaluation method in the previous section, the control strategy of the attacker can be designed. The nonlinear control systems (12) and (13) can be rewritten as:23 x˙1=f1x1+k1x1vD+g1x1b1uAet+b2uApt

24 x˙2=f2x2+g2x2b1uAet+b2uApt+h2x2wT

where, the Eqs. (23) and (24) are nonlinear systems of attacker A-defender D and attacker A-target T, respectively. vD and wT denote the respective control strategies of the defender D and the target T, respectively. uAet and uApt are the control strategies of the attacker A avoiding the defender D and capturing the target T, respectively.

Considering the nonlinear systems (23) and (24), the performance index functions of the two-part game are defined as:25 J1x1=∫0∞Q1x1+vDTtR1vDt-uATtR2uAetdt

26 J2x2=∫0∞Q2x2+uATtR3uApt-wTTtR4wTtdt

where, J1x1 is the cost function between the attacker A and the defender D; J2x2 is the cost function between the attacker A and the target T. Q1x1 and Q2x2 are semi-positive definite functions, Q1x1=x1TQ1x1, Q2x2=x2TQ2x2; Q1, Q2, R1, R2, R3 and R4 are symmetric positive definite matrices.

Further, the Hamiltonian functions of the two-part game are defined as:27 H1x1,vD,uAe,∇J1x1=Q1x1+vDTtR1vDt-uATtR2uAet+∇J1x1Tf1x1+k1x1vD+g1x1b1uAet+b2uApt

28 H2x2,uAp,wT,∇J2x2=Q2x2+uATtR3uApt-wTTtR4wTt+∇J2x2Tf2x2+g2x2b1uAet+b2uApt+h2x2wT

where, H1· is the Hamiltonian function between the attacker A and the defender D; H2· is the Hamiltonian function between the attacker A and the target T; ∇J1x1=∂J1x1∂x1 is the partial derivative of the performance function J1x1 with respect to x1; ∇J2x2=∂J2x2∂x2 is the partial derivative of the performance function J2x2 with respect to x2.

According to the minimax principle, the necessary conditions for the existence of Nash equilibrium solutions vD∗,uAe∗ and uAp∗,wT∗ for two partial differential games are:29 H1x1,vD∗,uAe,∇J1∗x1≤H1x1,vD∗,uAe∗,∇J1∗x1≤H1x1,vD,uAe∗,∇J1∗x1

30 H2x2,uAp∗,wT,∇J2∗x2≤H2x2,uAp∗,wT∗,∇J2∗x2≤H2x2,uAp,wT∗,∇J2∗x2

where, ∇J1∗x1 and ∇J2∗x2 are the optimal cost functions of the attacker A- defender D and attacker A-target T games, respectively. Their values satisfy the following equational relationship:31 J1∗x1=maxuAetminvDt∫0∞Q1x1+vDTtR1vDt-uATtR2uAetdt

32 J2∗x2=maxwTtminuApt∫0∞Q2x2+uATtR3uApt-wTTtR4wTtdt

By Hamilton–Jacobi–Isaacs (HJI) theory, the optimal cost functions J1∗x1 and J2∗x2 can be obtained by solving the following HJI Eqs. (33) and (34) for two partial games.33 maxuAetminvDtH1x1,vDt,uAet,∇J1∗x1=0

34 maxwTtminuAptH2x2,uApt,wTt,∇J2∗x2=0

According to the necessary conditions for optimal control:35 ∂H1∂vD=0∂H1∂uAe=0and∂H2∂uAp=0∂H2∂wT=0

The control strategy of the fuzzy differential game with two partial games can be obtained as:36 vD∗x1=-12R1-1k1Tx1∇J1∗x1uAe∗x1=b12R2-1g1Tx1∇J1∗x1

37 uAp∗x2=-b22R3-1g2Tx2∇J2∗x2wT∗x2=12R4-1h2Tx2∇J2∗x2

where, ∇J1∗x1=∂J1∗x1∂x1, ∇J2∗x1=∂J2∗x2∂x2.

Then, the HJI Eqs. (33) and (34) can be further formulated as:38 Q1x1+vD∗TR1vD∗-uAe∗TR2uAe∗+∇J1∗x1Tf1x1+k1x1vD∗+g1x1b1uAe∗+b2uAp∗=0

39 Q2x2+uAp∗TR3uAp∗-wT∗TR4wT∗+∇J2∗x2Tf2x2+g2x2b1uAe∗+b2uAp∗+h2x2wT∗=0

The specific HJI equations can be obtained by substituting the fuzzy differential countermeasure control strategies (36) and (37) into the Eqs. (38) and (39):40 Q1x1+∇J1∗x1Tf1x1-14∇J1∗x1Tk1x1R1-1k1Tx1∇J1∗x1+b124∇J1∗x1Tg1x1R2-1g1Tx1∇J1∗x1-b222∇J1∗x1Tg1x1R3-1g2Tx2∇J2∗x2=0

41 Q2x2+∇J2∗x2Tf2x2-b224∇J2∗x2Tg2x2R3-1g2Tx2∇J2∗x2+14∇J2∗x2Th2x2R4-1h2Tx2∇J2∗x2+b122∇J2∗x2Tg2x2R2-1g1Tx1∇J1∗x1=0

The optimal cost functions J1∗x1 and J2∗x2 for the two-part differential game can be solved through the two HJI equations in (40) and (41). However, the HJI equations are nonlinear partial differential form and it is very difficult to solve them directly. The HJI equations can be solved by the neural network approximation based adaptive dynamic programming (ADP)24–27.

Adaptive control design

To realize the fuzzy differential game control strategy proposed in the previous section, two evaluation neural networks are used in this section to approximate the optimal cost functions J1∗x1 and J2∗x2 of the two-part game, respectively.42 J1∗x1=Wc1∗Tφc1x1+ξc1

43 J2∗x2=Wc2∗Tφc2x2+ξc2

where, Wc1∗∈Rl1 and Wc2∗∈Rl2 are the ideal weight vectors of the two neural networks, respectively. φc1x1∈Rl1 and φc2x2∈Rl2 are the activation functions of the two evaluated neural networks, respectively. l1 and l2 are the number of neurons in the hidden layer of the two evaluated neural networks, respectively. ξc1 and ξc2 are the approximation errors of the two evaluated neural networks, respectively.

Substituting the Eqs. (42) and (43) into the Eqs. (36) and (37), respectively, the fuzzy differential game control strategy for the two-part game containing the ideal weights of the neural network is obtained:44 vD∗x1=-12R1-1k1Tx1∇φc1x1TWc1∗+∇ξc1uAe∗x1=b12R2-1g1Tx1∇φc1x1TWc1∗+∇ξc1

45 uAp∗x2=-b22R3-1g2Tx2∇φc2x2TWc2∗+∇ξc2wT∗x2=12R4-1h2Tx2∇φc2x2TWc2∗+∇ξc2

Since the ideal weights Wc1∗ and Wc2∗ of the neural network for two partial games are all unknown in practical control engineering. Setting W^c1 and W^c2 as estimates of the corresponding ideal weights, respectively. Then the estimation errors of the ideal weights for the two evaluation networks can be defined as respectively:46 W~c1=Wc1∗-W^c1

47 W~c2=Wc2∗-W^c2

Then the actual output of the two evaluation neural networks during the actual iteration is:48 J^1x1=W^c1Tφc1x1

49 J^2x2=W^c2Tφc2x2

Substituting the Eqs. (48) and (49) into the Eqs. (36) and (37), respectively, the fuzzy differential game control strategy for the two-part game in the estimated state can be obtained:50 v^Dx1=-12R1-1k1Tx1∇φc1x1TW^c1u^Aex1=b12R2-1g1Tx1∇φc1x1TW^c1

51 u^Apx2=-b22R3-1g2Tx2∇φc2x2TW^c2w^Tx2=12R4-1h2Tx2∇φc2x2TW^c2

Substituting J^1x1, J^2x2, v^Dx1, u^Aex1, u^Apx2 and w^Tx2 in the estimated state into the Eqs. (27) and (28) , then the approximation errors ec1 and ec2 are:52 ec1=Q1x1+v^DTx1R1v^Dx1-u^ATx1R2u^Aex1+∇J^1x1Tf1x1+k1x1v^Dx1+g1x1b1u^Aex1+b2u^Apx2

53 ec2=Q2x2+u^ATx2R3u^Apx2-w^TTx2R4w^Tx2+∇J^2x2Tf2x2+g2x2b1u^Aex1+b2u^Apx2+h2x2w^Tx2

In order to the two evaluation neural network weights W^c1 and W^c2 in the state converge to the ideal weights Wc1∗ and Wc2∗, respectively, two objective functions are designed as shown below, respectively28–30:54 Ec1=12ec1Tec1

55 Ec2=12ec2Tec2

Using the gradient descent method, the weight update laws for the two evaluation neural networks can be derived as follows:56 W^˙c1=-ηc1∂Ec1∂W^c1=-ηc1ψ11+ψ1Tψ12ec1

57 W^˙c2=-ηc2∂Ec2∂W^c2=-ηc2ψ21+ψ2Tψ22ec2

where, ηc1 and ηc2 are the learning rates of the two evaluated neural networks, respectively. The expressions for ψ1 and ψ2 are shown below:58 ψ1=∇φc1x1f1x1+k1x1v^Dx1+g1x1b1u^Aex1+b2u^Apx2

59 ψ2=∇φc2x2f2x2+g2x2b1u^Aex1+b2u^Apx2+h2x2w^Tx2

System stability analysis

Assumption 1

In the nonlinear systems (23) and (24), f1x1, k1x1, g1x1, f2x2, g2x2 and h2x2 satisfy the local Lipschitz continuity conditions, f10=0, f20=0; There exist positive constants f1M, k1M, g1M, f2M, g2M and h2M such that f1x1≤f1Mx1, k1x1≤k1M, g1x1≤g1M, f2x2≤f2Mx2, g2x2≤g2M and h2x2≤h2M hold.

Assumption 2

The two evaluated neural network ideal weights Wc1∗ and Wc2∗, the corresponding neural network excitation functions φc1x1 and φc2x2, and their corresponding gradients are bounded. That is, there exist positive constants Wc1m, Wc1M, Wc2m, Wc2M, φc1M, φc2M, dφ1M and dφ2M such that Wc1m≤Wc1∗≤Wc1M, Wc2m≤Wc2∗≤Wc2M, φc1x1≤φc1M, φc2x2≤φc2M, ∇φc1x1≤dφ1M and ∇φc2x2≤dφ2M hold.

Theorem 1

Under Assumptions 1 and Assumption 2, considering the nonlinear control systems (23) and (24) and the cost functions (25) and (26). Then, for the fuzzy differential game control strategies (50) and (51) designed for the two-part game and the evaluation network weight update laws (56) and (57) provided, the estimation errors W~c1 and W~c2 of the ideal weights of the two evaluation networks can be guaranteed to be ultimately consistent and bounded.

Proof

The Lyapunov function is chosen to be60 L1=12W~c1TW~c1

61 L2=12W~c2TW~c2

Lyapunov function Eqs. (60) and (61) are derived, respectively:62 L˙1=W~c1TW~˙c1=W~c1Tηc1ψ11+ψ1Tψ12ec1

63 L˙2=W~c2TW~˙c2=W~c2Tηc2ψ21+ψ2Tψ22ec2

Substituting the fuzzy differential game control strategies (50) and (51) for the two-part game into the Eqs. (52). One can obtain64 ec1=Q1x1+W^c1T∇φc1x1f1x1-14W^c1T∇φc1x1k1x1R1-1k1Tx1∇φc1x1TW^c1+b124W^c1T∇φc1x1g1x1R2-1g1Tx1∇φc1x1TW^c1-b222W^c1T∇φc1x1g1x1R3-1g2Tx2∇φc2x2TW^c2

Let A¯=∇φc1x1k1x1R1-1k1Tx1∇φc1x1T, B¯=∇φc1x1g1x1R2-1g1Tx1∇φc1x1T, C¯=∇φc1x1g1x1R3-1g2Tx2∇φc2x2T. The estimation errors (46) and (47) of the ideal weights are substituted into the Eq. (64).65 ec1=Q1x1+Wc1∗T∇φc1x1f1x1-14Wc1∗TA¯Wc1∗+b124Wc1∗TB¯Wc1∗-b222Wc1∗TC¯Wc2∗-W~c1T∇φc1x1f1x1+14Wc1∗TA¯W~c1+14W~c1TA¯Wc1∗-14W~c1TA¯W~c1-b124Wc1∗TB¯W~c1-b124W~c1TB¯Wc1∗+b124W~c1TB¯W~c1+b222Wc1∗TC¯W~c2+b222W~c1TC¯Wc2∗-b222W~c1TC¯W~c2

From the HJI Eq. (40):66 0=Q1x1+Wc1∗T∇φc1x1f1x1-14Wc1∗T∇φc1x1k1x1R1-1k1Tx1∇φc1x1TWc1∗+b124Wc1∗T∇φc1x1g1x1R2-1g1Tx1∇φc1x1TWc1∗-b222Wc1∗T∇φc1x1g1x1R3-1g2Tx2∇φc2x2TWc2∗+∇ξc1Tf1x1+k1x1vD∗x1+g1x1b1uAe∗x1+b2uAp∗x2+14∇ξc1Tk1x1R1-1k1Tx1∇ξc1-b124∇ξc1Tg1x1R2-1g1Tx1∇ξc1-b222Wc1∗T∇φc1x1g1x1R3-1g2Tx2∇ξc2

Define the residual error εJ1 due to the approximation error of the evaluation neural network as:67 εJ1=∇ξc1Tf1x1+k1x1vD∗x1+g1x1b1uAe∗x1+b2uAp∗x2+14∇ξc1Tk1x1R1-1k1Tx1∇ξc1-b124∇ξc1Tg1x1R2-1g1Tx1∇ξc1-b222Wc1∗T∇φc1x1g1x1R3-1g2Tx2∇ξc2

Then Eq. (66) can be further simplified as:68 Q1x1+Wc1∗T∇φc1x1f1x1-14Wc1∗TA¯Wc1∗+b124Wc1∗TB¯Wc1∗-b222Wc1∗TC¯Wc2∗=-εJ1

Substituting the Eq. (68) into the Eq. (65), ec1 can be described as:69 ec1=-W~c1T∇φc1x1f1x1+14Wc1∗TA¯W~c1+14W~c1TA¯Wc1∗-14W~c1TA¯W~c1-b124Wc1∗TB¯W~c1-b124W~c1TB¯Wc1∗+b124W~c1TB¯W~c1+b222Wc1∗TC¯W~c2+b222W~c1TC¯Wc2∗-b222W~c1TC¯W~c2-εJ1=-W~c1T∇φc1x1f1x1+Wc1∗TA¯4-b124B¯W~c1+W~c1TA¯4-b124B¯Wc1∗-W~c1TA¯4-b124B¯W~c1+b222Wc1∗TC¯W~c2+b222W~c1TC¯Wc2∗-b222W~c1TC¯W~c2-εJ1

Substituting the Eq. (69) into the Eq. (62), the expression for L˙1 is given by:70 L˙1=β¯c1-W~c12∇φc1x1f1x1+W~c12A¯2-b122B¯Wc1∗-W~c13A¯4-b124B¯+β¯c1b222Wc1∗TC¯W~c1W~c2+b222W~c12C¯Wc2∗-b222W~c12C¯W~c2-εJ1W~c1

where, β¯c1=ηc1ψ11+ψ1Tψ12.

Based on Assumption 1 and Assumption 2, there exist positive constants ƛβm, ƛβM, ƛ1m, ƛ1M, ƛ2m, ƛ2M, ƛ3m, ƛ3M, ƛ4m and ƛ4M, such that ƛβm≤β¯c1≤ƛβM, ƛ1m≤A¯2-b122B¯≤ƛ1M, ƛ2m≤C¯≤ƛ2M, ƛ3m≤∇φc1x1f1x1≤ƛ3M and ƛ4m≤εJ1≤ƛ4M hold.

Applying the inequality ab≤12a2+b2, one can obtain31,32:71 L˙1≤-12ƛβm2W~c14-18ƛ1m2W~c12-ƛβmƛ3mW~c12+ƛβMƛ1MWc1MW~c12+b224ƛβM2Wc1M2W~c12+b224ƛ2MW~c22+b222ƛβMƛ2MWc2MW~c12-b222ƛβmƛ2mW~c2W~c12+12ƛ4M2+12ƛβM2W~c12

where, let72 Θ5=-18ƛ1m2-ƛβmƛ3m+ƛβMƛ1MWc1M+b224ƛβM2Wc1M2+b222ƛβMƛ2MWc2M-b222ƛβmƛ2mW~c2+12ƛβM2

73 Θ6=b224ƛ2MW~c22+12ƛ4M2

Then, L˙1 can be further expressed as:74 L˙1≤-12ƛβm2W~c14+Θ5W~c12+Θ6

Therefore, when W~c1 satisfies Eq. (75),75 W~c1≥Θ5+Θ52+2ƛβm2Θ6ƛβm2

L˙1<0 holds.

Similarly, substituting the fuzzy differential game control strategies (50) and (51) for the two-part game into the Eq. (53), the approximation error ec2:76 ec2=Q2x2+W^c2T∇φc2x2f2x2-b224W^c2T∇φc2x2g2x2R3-1g2Tx2∇φc2x2TW^c2+14W^c2T∇φc2x2h2x2R4-1h2Tx2∇φc2x2TW^c2+b122W^c2T∇φc2x2g2x2R2-1g1Tx1∇φc1x1TW^c1

Let D¯=∇φc2x2g2x2R3-1g2Tx2∇φc2x2T, E¯=∇φc2x2h2x2R4-1h2Tx2∇φc2x2T, F¯=∇φc2x2g2x2R2-1g1Tx1∇φc1x1T; Substitute the estimation errors (46) and (47) of the ideal weights of the two evaluation networks into the Eq. (76).77 ec2=Q2x2+Wc2∗T-W~c2T∇φc2x2f2x2-b224Wc2∗T-W~c2TD¯Wc2∗-W~c2+14Wc2∗T-W~c2TE¯Wc2∗-W~c2+b122Wc2∗T-W~c2TF¯Wc1∗-W~c1=Q2x2+Wc2∗T∇φc2x2f2x2-b224Wc2∗TD¯Wc2∗+14Wc2∗TE¯Wc2∗+b122Wc2∗TF¯Wc1∗-W~c2T∇φc2x2f2x2+b224Wc2∗TD¯W~c2+b224W~c2TD¯Wc2∗-b224W~c2TD¯W~c2-14Wc2∗TE¯W~c2-14W~c2TE¯Wc2∗+14W~c2TE¯W~c2-b122Wc2∗TF¯W~c1-b122W~c2TF¯Wc1∗+b122W~c2TF¯W~c1

From the HJI Eq. (41):78 0=Q2x2+Wc2∗T∇φc2x2+∇ξc2Tf2x2-b224Wc2∗T∇φc2x2+∇ξc2Tg2x2R3-1g2Tx2∇φc2x2TWc2∗+∇ξc2+14Wc2∗T∇φc2x2+∇ξc2Th2x2R4-1h2Tx2∇φc2x2TWc2∗+∇ξc2+b122Wc2∗T∇φc2x2+∇ξc2Tg2x2R2-1g1Tx1∇φc1x2TWc1∗+∇ξc1=Q2x2+Wc2∗T∇φc2x2f2x2-b224Wc2∗T∇φc2x2g2x2R3-1g2Tx2∇φc2x2TWc2∗+14Wc2∗T∇φc2x2h2x2R4-1h2Tx2∇φc2x2TWc2∗+b122Wc2∗T∇φc2x2g2x2R2-1g1Tx1∇φc1x1TWc1∗+εJ2

where, εJ2 is the residual error due to the approximation error of the evaluation neural network:79 εJ2=∇ξc2Tf2x2+g2x2b1uAe∗x1+b2uAp∗x2+h2x2wT∗x2+b224∇ξc2Tg2x2R3-1g2Tx2∇ξc2+b122Wc2∗T∇φc2x2g2x2R2-1g1Tx1∇ξc1-14∇ξc2Th2x2R4-1h2Tx2∇ξc2

Then Eq. (78) can be simplified:80 Q2x2+Wc2∗T∇φc2x2f2x2-b224Wc2∗TD¯Wc2∗+14Wc2∗TE¯Wc2∗+b122Wc2∗TF¯Wc1∗=-εJ2

Substituting the Eq. (80) into the Eq. (77), ec2 can be described as:81 ec2=-W~c2T∇φc2x2f2x2+Wc2∗Tb224D¯-14E¯W~c2+W~c2Tb224D¯-14E¯Wc2∗-W~c2Tb224D¯-14E¯W~c2-b122Wc2∗TF¯W~c1-b122W~c2TF¯Wc1∗+b122W~c2TF¯W~c1-εJ2

Substituting the Eq. (81) into the Eq. (63), the expression for L˙2 can be described as follows:82 L˙2=β¯c2-W~c22∇φc2x2f2x2+W~c22b222D¯-12E¯Wc2∗-W~c13b224D¯-14E¯+β¯c2-b122Wc2∗TF¯W~c1W~c2-b122W~c22F¯Wc1∗+b122W~c22F¯W~c1-εJ2W~c2

where β¯c2=ηc2ψ21+ψ2Tψ22.

Based on Assumption 1 and Assumption 2, there exist positive constants ƛ5m, ƛ5M, ƛ6m, ƛ6M, ƛ7m, ƛ7M, ƛ8m, ƛ8M, ƛ9m and ƛ9M, such that ƛ5m≤β¯c2≤ƛ5M, ƛ6m≤b222D¯-12E¯≤ƛ6M, ƛ7m≤F¯≤ƛ7M, ƛ8m≤∇φc2x2f2x2≤ƛ8M and ƛ9m≤εJ2≤ƛ9M hold.

Then L˙2 can be expressed as:83 L˙2≤-12ƛ5m2W~c24-18ƛ6m2W~c22-ƛ5mƛ8mW~c22+ƛ5Mƛ6MWc2MW~c22+b124ƛ5M2Wc2M2W~c22+b124ƛ7MW~c12-b122ƛ5mƛ7mWc1mW~c22+b122ƛ5Mƛ2MW~c1W~c22+12ƛ9M2+12ƛ5M2W~c22

where, let84 Θ7=-18ƛ6m2-ƛ5mƛ8m+ƛ5Mƛ6MWc2M+b124ƛ5M2Wc2M2-b122ƛ5mƛ7mWc1m+b122ƛ5Mƛ2MW~c1+12ƛ5M2

85 Θ8=b124ƛ7MW~c12+12ƛ9M2

Then, L˙2 can be further expressed as:86 L˙2≤-12ƛ5m2W~c24+Θ7W~c22+Θ8

Therefore, when W~c2 satisfies Eq. (87).87 W~c2≥Θ7+Θ72+2ƛ5m2Θ8ƛ5m2

L˙2<0 holds.

In summary, according to the Lyapunov stability theory, the neural network weight estimation errors W~c1 and W~c2 can be guaranteed to be ultimately consistent and bounded.

The proof is completed.

From the control strategies (44) and (45) for the fuzzy differential game in the two-part game and the control strategies (50) and (51) for the two-part game in the estimation state, one can obtain33,3488 v^Dx1=vD∗x1-12R1-1k1Tx1∇φc1x1TW^c1+12R1-1k1Tx1∇φc1x1TWc1∗+12R1-1k1Tx1∇ξc1

89 u^Aex1=uAe∗x1+b12R2-1g1Tx1∇φc1x1TW^c1-b12R2-1g1Tx1∇φc1x1TWc1∗-b12R2-1g1Tx1∇ξc1

90 u^Apx2=uAp∗x2-b22R3-1g2Tx2∇φc2x2TW^c2+b22R3-1g2Tx2∇φc2x2TWc2∗+b22R3-1g2Tx2∇ξc2

91 w^Tx2=wT∗x2+12R4-1h2Tx2∇φc2x2TW^c2-12R4-1h2Tx2∇φc2x2TWc2∗-12R4-1h2Tx2∇ξc2

Let ς¯1x1=12R1-1k1Tx1, ς¯2x1=b12R2-1g1Tx1, ς¯3x2=b22R3-1g2Tx2,ς¯4x2=12R4-1h2Tx2.

Then92 v^Dx1=vD∗x1+ς¯1x1W~c1+12R1-1k1Tx1∇ξc1u^Aex1=uAe∗x1-ς¯2x1W~c1-b12R2-1g1Tx1∇ξc1

93 u^Apx2=uAp∗x2+ς¯3x2W~c2+b22R3-1g2Tx2∇ξc2w^Tx2=wT∗x2-ς¯4x2W~c2-12R4-1h2Tx2∇ξc2

Assumption 3

In the given compact set, ς¯1x1, ς¯2x1, ς¯3x2, ς¯4x2, J1∗x1 and J2∗x2 satisfy the local Lipschitz continuity condition. Then there exist positive constants p¯1, p¯2, p¯3, p¯4, p¯5 and p¯6 such that ς¯1x1≤p¯1, ς¯2x1≤p¯2, ς¯3x2≤p¯3, ς¯4x2≤p¯4, ∇J1∗x1≤p¯5, ∇J2∗x2≤p¯6 holds.

Theorem 3

Under the Assumption 1, Assumption 2, Assumption 3, considering nonlinear differential control systems (23) and (24). The optimal fuzzy differential game control strategies (50) and (51) are applied to the corresponding nonlinear systems, respectively. Then, when the estimation errors W~c1 and W~c2 are ultimately consistent and bounded, the nonlinear systems (23) and (24) can be asymptotically stable.

Proof

The Lyapunov function is chosen as L3=J1∗x1 and L4=J2∗x2, respectively.94 L˙3=J˙1∗x1=∇J1∗x1Tf1x1+k1x1v^Dx1+g1x1b1u^Aex1+b2u^Apx2=∇J1∗x1Tf1x1+k1x1vD∗x1+ς¯1x1W~c1+1122R1-1k1Tx1∇ξc1+∇J1∗x1Tg1x1b1uAe∗x1-ς¯2x1W~c1-b1b122R2-1g1Tx1∇ξc1+∇J1∗x1Tg1x1b2uAp∗x2+ς¯3x2W~c2+b2b222R3-1g2Tx2∇ξc2=∇J1∗x1Tf1x1+k1x1vD∗x1+g1x1b1uAe∗x1+b2uAp∗x2+∇J1∗x1Tk1x1ς¯1x1W~c1-∇J1∗x1Tg1x1b1ς¯2x1W~c1+∇J1∗x1Tg1x1b2ς¯3x2W~c2+εH1

where, εH1 is the residual error caused by the approximation error of the evaluation neural network.95 εH1=12∇J1∗x1Tk1x1R1-1k1Tx1∇ξc1-b122∇J1∗x1Tg1x1R2-1g1Tx1∇ξc1+b222∇J1∗x1Tg1x1R3-1g2Tx2∇ξc2

From the HJI Eq. (38):96 ∇J1∗x1Tf1x1+k1x1vD∗+g1x1b1uAe∗+b2uAp∗=-Q1x1-vD∗TR1vD∗+uAe∗TR2uAe∗

Then, the expression for L˙3 is further described as:97 L˙3≤-x1TQ1x1-vD∗TR1vD∗+uAe∗TR2uAe∗+12∇J1∗x1Tk1x1k1Tx1∇J1∗x1+12W~c1Tς¯1Tx1ς¯1x1W~c1+b122∇J1∗x1Tg1x1g1Tx1∇J1∗x1+12W~c1Tς¯2Tx1ς¯2x1W~c1+b222∇J1∗x1Tg1x1g1Tx1∇J1∗x1+12W~c2Tς¯3Tx2ς¯3x2W~c2+εH1

Based on Assumption 1, Assumption 2, Assumption 3 and Theorem 1, there exist positive constants W1M, W2M and p¯7 such that W~c1≤W1M, W~c2≤W2M and εH1≤p¯7 hold.

98 L˙3≤-λminQ1x12+14R1-1p¯52k1M2+b124R2-1p¯52k1M2+12p¯52k1M2+b12+b222p¯52g1M2+12W1M2p¯12+12W1M2p¯22+12W2M2p¯32+p¯7

where, λminQ1 is the smallest eigenvalue of the matrix Q1.

Let Ξ1=12+14R1-1p¯52k1M2+b124R2-1+b12+b222p¯52g1M2+12W1M2p¯12+p¯22+12W2M2p¯32+p¯7.

Then, the expression for L˙3 is:99 L˙3≤-λminQ1x12+Ξ1

Therefore, when x1 satisfies Eq. (100):100 x1>Ξ1λminQ1

L˙3≤0 holds. Similarly101 L˙4=J˙2∗x2=∇J2∗x2Tf2x2+g2x2b1u^Aex1+b2u^Apx2+h2x2w^Tx2=∇J2∗x2Tf2x2+g2x2b1uAe∗x1-ς¯2x1W~c1-b1b122R2-1g1Tx1∇ξc1+∇J2∗x2Tg2x2b2uAp∗x2+ς¯3x2W~c2+b2b222R3-1g2Tx2∇ξc2+∇J2∗x2Th2x2wT∗x2-ς¯4x2W~c2-1122R4-1h2Tx2∇ξc2=∇J2∗x2Tf2x2+g2x2b1uAe∗x1+b2uAp∗x2+h2x2wT∗x2-∇J2∗x2Tg2x2b1ς¯2x1W~c1+∇J2∗x2Tg2x2b2ς¯3x2W~c2-∇J2∗x2Th2x2ς¯4x2W~c2+εH2

where, εH2 is the residual error caused by the approximation error of the evaluation neural network.102 εH2=-b122∇J2∗x2Tg2x2R2-1g1Tx1∇ξc1+b222∇J2∗x2Tg2x2R3-1g2Tx2∇ξc2-12∇J2∗x2Th2x2R4-1h2Tx2∇ξc2

From the HJI Eq. (39):103 ∇J2∗x2Tf2x2+g2x2b1uAe∗+b2uAp∗+h2x2wT∗=-Q2x2-uAp∗TR3uAp∗+wT∗TR4wT∗

Then, the expression for L˙4 is:104 L˙4≤-x2TQ2x2-uAp∗TR3uAp∗+wT∗TR4wT∗+b122∇J2∗x2Tg2x2g2Tx2∇J2∗x2+12W~c1Tς¯2Tx1ς¯2x1W~c1+b222∇J2∗x2Tg2x2g2Tx2∇J2∗x2+12W~c2Tς¯3Tx2ς¯3x2W~c2+12∇J2∗x2Th2x2h2Tx2∇J2∗x2+12W~c2Tς¯4Tx2ς¯4x2W~c2+εH2

Let εH2≤p¯8, where p¯8 is a positive constant.105 L˙4≤-λminQ2x22+b224R3-1p¯62g2M2+14R4-1p¯62h2M2+b122p¯62g2M2+12W1M2p¯22+b222p¯62g2M2+12W2M2p¯32+12p¯62h2M2+12W2M2p¯42+p¯8

Let Ξ2=b224R3-1+b12+b222p¯62g2M2+14R4-1+12p¯62h2M2+12W2M2p¯32+p¯42+12W1M2p¯22+p¯8.

Then, the expression for L˙4 can be further described as:106 L˙4≤-λminQ2x22+Ξ2

Therefore, when x2 satisfies Eq. (107):107 x2>Ξ2λminQ2

L˙4≤0 holds.

By the Lyapunov stability, when the estimation errors W~c1 and W~c2 are ultimately consistent and bounded, the nonlinear systems (23) and (24) can be asymptotically stable. Proof is completed.

Simulation study

In this section, the simulation study for the tripartite UAV confrontation game is done to verify the validity of the proposed fuzzy differential game control strategy.

Assuming that the initial position of the attacker is xA,yA=0,0, the velocity is VA=500m/s, and the initial heading angle is α0=60o. The initial position of the defender is xD,yD=5000,0, the speed is VD=300m/s, and the initial heading angle is β0=80o. The initial position of the target is xT,yT=5500,0, the velocity is VT=200m/s, and the initial heading angle is γ0=80o.

In the fuzzy differential game control strategy, let the semipositive definite functions Q1x1=40x1Tx1, Q2x2=40x2Tx2. The control parameters R1=1, R2=10, R3=1, R4=10. The excitation functions of the evaluation neural networks are φc1x1=x1,x2,x12,x22 and φc1x2=x1,x2,x12,x22, respectively. The number of hidden layer neurons of the two evaluation neural networks are l1=4 and l2=4, respectively. The initial weights of the evaluation neural networks are W^c10=0,0,0,0T and W^c20=0,0,0,0T, respectively. The learning rates of the neural networks are ηc1=0.5 and ηc2=0.535. The simulation results are represented as follows.

The flight trajectories of the trilateral UAV confrontation game between the attacker, the defender and the target are shown in Fig. 2. The control inputs for the attacker, the defender and the target are shown in Figs. 3, 4 and 5 respectively.Fig. 2 Tripartite UAV confrontation game.

Fig. 3 Control inputs of the attacker.

Fig. 4 Control inputs of the defender.

Fig. 5 Control inputs of the target.

As can be seen from the Fig. 2, after 6.4 s in dynamic confrontation game, the attacker UAV is successfully intercepted by the defender UAV and the target is successfully escaped. The variation in the coefficient weights of the control strategies selected by the attacker during its participation in the two-part game is shown in Fig. 6. As can be seen from Fig. 6, at the beginning of the simulation, the attacker chooses to approach the target quickly because the attacker is far away from the defender. The coefficient b1 is smaller and the coefficient b2 is larger. As the target continues to be approached by the attacker, the attacker continues to be approached by the defender. The need for the attacker to evade the defender becomes more and more pressing. The coefficient b1 starts to get larger, while the coefficient b2 gets smaller.Fig. 6 Fuzzy evaluation coefficients.

The relative distances between the attacker and defender are shown in Fig. 7. The relative distances between the attacker and the target are shown in Fig. 8. As can be seen from Figs. 7 and 8, the distance between the attacker and the defender is small after 6.4 s. Then the attacker is successfully intercepted by the defender. The relative distance between the attacker and the target is large. The target is successfully escaped.Fig. 7 Distance between attacker and defender.

Fig. 8 Distance between attacker and target.

The relative approach rates between the attacker and defender are shown in Fig. 9. The relative rate of approach between the attacker and target is shown in Fig. 10. The relative rate change of the confrontational game between the attacker and the defender, the target, is captured by Figs. 9 and 10. The line-of-sight angular rate between the attacker and defender is shown in Fig. 11. The angular rate of line of sight between the attacker and target is shown in Fig. 12. As can be seen from Fig. 11 and Fig. 12, the line-of-sight angle rates between the attacker and the defender, the target all converge to 0. The design principle is satisfied.Fig. 9 Approach rates of attacker and defender.

Fig. 10 Approach rates of attacker and target.

Fig. 11 Angular rates of attackers and defenders.

Fig. 12 Angular rates of attacker and target.

The estimated weights of the evaluation neural network for the attacker and defender confrontation game are shown in Fig. 13. The estimated weights of the evaluation neural network for the attacker and target confrontation game are shown in Fig. 14. As can be seen from Figs. 13 and 14, the evaluation neural network weights can all converge to the desired weights by iterating from the initial weight 0.Fig. 13 Estimation of first network weights.

Fig. 14 Estimation of second network weights.

The estimation errors of the evaluation neural networks of the attacker and defender are shown in Fig. 15. The estimation errors of the evaluation neural networks of the attacker and target are shown in Fig. 16. As can be seen from Figs. 15 and 16, the estimation errors of the evaluation neural networks for the confrontation game converge to 0. Both parts of the system can reach the steady state to be verified.Fig. 15 First network estimation error.

Fig. 16 Second network estimation error.

Simulation comparison experiment

In order to further verify the effectiveness and superiority of the fuzzy differential game control strategy method proposed in this paper. In36, the proportional guidance method is used as a comparison method with a proportional guidance coefficient 10. The initial position, speed, and heading angle of the tripartite UAV are kept in the same way as the corresponding parameters of the fuzzy differential game control in the above paper.

The results of the tripartite UAV confrontation game using the fuzzy differential game control strategy and proportional guidance method are shown in Fig. 17. Where “FDG” is used as the acronym of fuzzy differential game method and “PG” is used as the acronym of proportional guidance method. As can be seen in Fig. 17, with the proportional guidance method, the attacker UAV is intercepted by the defender in 5.7 s, however, when the fuzzy differential game control method is used, the attacker UAV is intercepted by the defender only in 6.4 s. By comparing the interception time, the fuzzy differential game method is adopted and the better control strategy is obtained by the attacker UAV.Fig. 17 Comparison of tripartite UAV game.

The control inputs of the attacker, the defender, and the target using the fuzzy differential game control strategy and the proportional guidance method are shown in Figs. 18, 19 and 20, respectively. The tripartite UAV position, velocity and other states are synthesized by the fuzzy differential game method, and its control curve changes are more complex compared to the proportional guidance method. The comparison of the distances between the attacker and the defender under the above two methods is shown in Fig. 21. The proportional guidance method is employed and the distance between the attacker and the defender converges to 0 at 5.7 s, and the victory is gained prematurely by the defender. The strategy adopted by the attacker is not optimal. The comparison of the distances between the attacker and the target under the adoption of the above two methods is shown in Fig. 22. The fuzzy differential game method is employed and the attacker UAV is closer to the target UAV and the more optimal strategy is adopted.Fig. 18 Comparison of Attacker Inputs.

Fig. 19 Comparison of Defender Inputs.

Fig. 20 Comparison of target inputs.

Fig. 21 Distance between attacker and defender.

Fig. 22 Distance between attacker and target.

The performance comparison between the fuzzy differential game algorithm and the proportional guidance algorithm proposed in this paper is shown in Table 2. The comparison parameters mainly include the simulation time, the final distance between the attacker and the target party to each other, the final distance between the attacker and the defender to each other, and the estimation error of the evaluation neural network. As can be seen from Table 2, the proportional guidance method is adopted, the attacking UAV is intercepted by the defending party in 5.7 s, and the distance from the attacking party to the target party is farther. Meanwhile, the estimation error of its evaluation neural network is larger. However, the fuzzy differential game control method is employed and the attacking party UAV is intercepted by the defending party in 6.4 s. The distance between the attacking side and the target side is closer than the proportional guidance method, and the attacking side UAV obtains a better control strategy. Meanwhile, the fuzzy differential game control method is adopted to evaluate the neural network with a small estimation error, and the error accuracy is achieved to 7.873e-09.Table 2 Comparison of the two algorithms.

Algorithm comparison	Simulation time/s	RAT/m	RAD/m	network estimation error	
Proportional guidance	5.7	2955	0.8123	8.079e−05	
fuzzy differential game	6.4	2217	0.5148	7.873e−09	

Through the above analysis, compared with the proportional guidance method, the fuzzy differential game method proposed in this paper seeks a better control strategy. By using the fuzzy differential game method, the complex confrontation environment can be better coped with and more dynamic change factors can be considered in the decision-making process. Therefore, the fuzzy differential game method proposed in this paper provides a more optimal solution to the tripartite UAV confrontation game problem.

Conclusion

A game control strategy based on fuzzy differential game is proposed for the problem of attack, defense and target confrontation of tripartite UAVs. The confrontation model between the three-party UAVs was firstly established by this research, and the corresponding nonlinear differential control system for the three-party UAVs was designed. Second, combining the fuzzy evaluation method and differential game theory, a fuzzy differential game control method is innovatively proposed to solve the adversarial game problem of the tripartite UAVs.

In this method, the tripartite UAV system is divided into two subsystems: attacker-defender and attacker-target. The optimal control strategies of the UAVs in the two subsystems are derived separately through differential countermeasure game theory. As the attacker is involved in both subsystem parts of the game at the same time. A fuzzy evaluation method is used to determine the weight coefficients of the attacking party in the two parts of the game. In order to achieve this goal, two Hamilton–Jacobi–Isaacs (HJI) equations are constructed by defining the Hamiltonian functions of the two subsystems. Since the HJI equations are difficult to solve directly, an adaptive dynamic programming approach is used. Two evaluation neural networks are introduced to approximate the optimal value functions of the two subsystems separately. The convergence of the two evaluated neural network weights and the stability of the two nonlinear differential control systems are proved.

Finally, the effectiveness and superiority of the proposed fuzzy differential countermeasure game control strategy is verified by simulation comparison experiments with the traditional proportional guidance control method. This research provides an innovative control strategy for solving the complex three-party UAV confrontation problem. And it is provided a certain reference value for military fields such as attacking the target party and intercepting enemy UAVs. Future research directions include further optimization of the algorithm to improve real-time and robustness. Meanwhile, it can be further extended to the study of more complex multi-party UAV game confrontation.

Acknowledgements

This work is supported by National Natural Science Foundation of China under Grant 62103057.

Author contributions

All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by Xingjian Fu and Hang Yan. The first draft of the manuscript was written by Hang Yan and all authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.

Data availability

All data generated or analysed during this study are included in this published article.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Xingjian Fu and Hang Yan.
==== Refs
References

1. Liu C Sun S Tao C Sliding mode control of multi-agent system with application to UAV air combat Comput. Electr. Eng. 2021 96 107 121 10.1016/j.compeleceng.2021.107491
Liu, C., Sun, S. & Tao, C. Sliding mode control of multi-agent system with application to UAV air combat. Comput. Electr. Eng. 96, 107–121 (2021).10.1016/j.compeleceng.2021.107491
2. Ma Y Wang G Hu X Cooperative occupancy decision making of Multi-UAV in Beyond-Visual-Range air combat: A game theory approach IEEE Access. 2019 8 11624 11634 10.1109/ACCESS.2019.2933022
Ma, Y., Wang, G. & Hu, X. Cooperative occupancy decision making of Multi-UAV in Beyond-Visual-Range air combat: A game theory approach. IEEE Access. 8, 11624–11634 (2019).10.1109/ACCESS.2019.2933022
3. Zhang C Zhu Y Yang L An optimal guidance method for free-time orbital pursuit-evasion game J. Syst. Eng. Electron. 2022 33 6 1294 1308
Zhang, C., Zhu, Y. & Yang, L. An optimal guidance method for free-time orbital pursuit-evasion game. J. Syst. Eng. Electron. 33(6), 1294–1308 (2022).
4. Fei LU Cooperative differential games guidance laws for multiple attackers against an active defense target Chin. J. Aeronaut. 2022 35 5 374 389 10.1016/j.cja.2021.07.033
Fei, L. U. et al. Cooperative differential games guidance laws for multiple attackers against an active defense target. Chin. J. Aeronaut. 35(5), 374–389 (2022).10.1016/j.cja.2021.07.033
5. English JT Wilhelm JP Defender-aware attacking guidance policy for the target–attacker–defender differential game J. Aerosp. Inf. Syst. 2021 18 6 366 376
English, J. T. & Wilhelm, J. P. Defender-aware attacking guidance policy for the target–attacker–defender differential game. J. Aerosp. Inf. Syst. 18(6), 366–376 (2021).
6. Gomoyunov M Solution to a zero-sum differential game with fractional dynamics via approximations Dyn. Games Appl. 2020 10 2 417 443 10.1007/s13235-019-00320-4
Gomoyunov, M. Solution to a zero-sum differential game with fractional dynamics via approximations. Dyn. Games Appl. 10(2), 417–443 (2020).10.1007/s13235-019-00320-4
7. Shiri H Park J Bennis M Communication-efficient massive UAV online path control: Federated learning meets mean-field game theory IEEE Trans. Commun. 2020 68 11 6840 6857 10.1109/TCOMM.2020.3017281
Shiri, H., Park, J. & Bennis, M. Communication-efficient massive UAV online path control: Federated learning meets mean-field game theory. IEEE Trans. Commun. 68(11), 6840–6857 (2020).10.1109/TCOMM.2020.3017281
8. Zeng X Yang L Zhu Y Yang F Comparison of two optimal guidance methods for the long-distance orbital pursuit-evasion game IEEE Trans. Aerosp. Electron. Syst. 2020 57 1 521 539 10.1109/TAES.2020.3024423
Zeng, X., Yang, L., Zhu, Y. & Yang, F. Comparison of two optimal guidance methods for the long-distance orbital pursuit-evasion game. IEEE Trans. Aerosp. Electron. Syst. 57(1), 521–539 (2020).10.1109/TAES.2020.3024423
9. Xu J Deng Z Song Q Chi Q Wu T Huang Y Gao M Multi-UAV counter-game model based on uncertain information Appl. Math. Comput. 2020 366 674 684
Xu, J. et al. Multi-UAV counter-game model based on uncertain information. Appl. Math. Comput. 366, 674–684 (2020).
10. Liu MJ Wu XQ Wang HY Wang. UAV attack and defense game decision making based on quantum particle swarm optimization Fire Control Command Control 2022 47 73 78
Liu, M. J., Wu, X. Q., Wang, H. Y. & Wang.,. UAV attack and defense game decision making based on quantum particle swarm optimization. Fire Control Command Control 47, 73–78 (2022).
11. Kim J Oh H Yu B Kim S Optimal task assignment for UAV swarm operations in hostile environments Int. J. Aeronaut. Space Sci. 2021 22 2 456 467 10.1007/s42405-020-00317-z
Kim, J., Oh, H., Yu, B. & Kim, S. Optimal task assignment for UAV swarm operations in hostile environments. Int. J. Aeronaut. Space Sci. 22(2), 456–467 (2021).10.1007/s42405-020-00317-z
12. Mirzaei, M., Kosari, A., & Maghsoudi, H. Optimal path planning for two UAV in a pursuit-evasion game. In 2021 IEEE International Conference on Automation/XXIV Congress of the Chilean Association of Automatic Control (ICA-ACCA). 1–7 (2021).
13. Sun J Liu C Zhao X Backstepping-based zero-sum differential games for missile-target interception systems with input and output constraints IET Control Theory Appl. 2018 12 2 243 253 10.1049/iet-cta.2017.0501
Sun, J., Liu, C. & Zhao, X. Backstepping-based zero-sum differential games for missile-target interception systems with input and output constraints. IET Control Theory Appl. 12(2), 243–253 (2018).10.1049/iet-cta.2017.0501
14. Cheng T Zhou H Dong FX Multi-vehicle surprise strike integrated differential countermeasure guidance law design J. Beijing Univ. Aeronaut. Astronaut. 2022 48 898 909
Cheng, T., Zhou, H. & Dong, F. X. Multi-vehicle surprise strike integrated differential countermeasure guidance law design. J. Beijing Univ. Aeronaut. Astronaut. 48, 898–909 (2022).
15. Wang YZ Tang JS Guo J Three-dimensional guidance with adaptive differential countermeasures for hypersonic attack and defense games Acta Armamentarii. 2023 44 8 2342 2353
Wang, Y. Z., Tang, J. S. & Guo, J. Three-dimensional guidance with adaptive differential countermeasures for hypersonic attack and defense games. Acta Armamentarii. 44(8), 2342–2353 (2023).
16. Casbeer DW Garcia E Pachter M The target differential game with two defenders J. Intell. Rob. Syst. 2018 89 87 106 10.1007/s10846-017-0563-0
Casbeer, D. W., Garcia, E. & Pachter, M. The target differential game with two defenders. J. Intell. Rob. Syst. 89, 87–106 (2018).10.1007/s10846-017-0563-0
17. Tao CHAO Xin WANG Song WANG Ming YANG Linear-quadratic and norm-bounded differential game combined guidance strategy against active defense aircraft in three-player engagement Chin. J. Aeronaut. 2023 36 8 331 350 10.1016/j.cja.2023.04.012
Tao, C. H. A. O., Xin, W. A. N. G., Song, W. A. N. G. & Ming, Y. A. N. G. Linear-quadratic and norm-bounded differential game combined guidance strategy against active defense aircraft in three-player engagement. Chin. J. Aeronaut. 36(8), 331–350 (2023).10.1016/j.cja.2023.04.012
18. Garcia E Casbeer DW Pachter M Design and analysis of state-feedback optimal strategies for the differential game of active defense IEEE Trans. Autom. Control 2018 64 2 553 568
Garcia, E., Casbeer, D. W. & Pachter, M. Design and analysis of state-feedback optimal strategies for the differential game of active defense. IEEE Trans. Autom. Control 64(2), 553–568 (2018).
19. Liu K Zheng SX Lin MY Optimal strategy design for fugitive problem based on differential games J. Autom. 2021 47 1840 1854
Liu, K., Zheng, S. X. & Lin, M. Y. Optimal strategy design for fugitive problem based on differential games. J. Autom. 47, 1840–1854 (2021).
20. Sun LJ Research on adaptive dynamic planning and its application in missile interception and guidance 2019 Nanjing University of Aeronautics and Astronautics
Sun, L. J. Research on adaptive dynamic planning and its application in missile interception and guidance (Nanjing University of Aeronautics and Astronautics, 2019).
21. Zhao L Zhou FJ Liu Y A tripartite differential countermeasure method for “pursuit-flight-prevention” in three-dimensional space Syst. Eng. Electron. Technol. 2019 41 322 335
Zhao, L., Zhou, F. J. & Liu, Y. A tripartite differential countermeasure method for “pursuit-flight-prevention” in three-dimensional space. Syst. Eng. Electron. Technol. 41, 322–335 (2019).
22. Vamvoudakis KG Non-zero sum Nash Q-learning for unknown deterministic continuous-time linear systems Automatica 2015 61 274 281 10.1016/j.automatica.2015.08.017
Vamvoudakis, K. G. Non-zero sum Nash Q-learning for unknown deterministic continuous-time linear systems. Automatica 61, 274–281 (2015).10.1016/j.automatica.2015.08.017
23. Xiao G Zhang H Qu Q Jiang H General value iteration based single network approach for constrained optimal controller design of partially-unknown continuous-time nonlinear systems J. Frankl. Inst. 2018 355 5 2610 2630 10.1016/j.jfranklin.2018.02.001
Xiao, G., Zhang, H., Qu, Q. & Jiang, H. General value iteration based single network approach for constrained optimal controller design of partially-unknown continuous-time nonlinear systems. J. Frankl. Inst. 355(5), 2610–2630 (2018).10.1016/j.jfranklin.2018.02.001
24. Wu H Li M Gao Q Wei Z Zhang N Tao X Eavesdropping and anti-eavesdropping game in UAV wiretap system: A differential game approach IEEE Trans. Wireless Commun. 2022 21 11 9906 9920 10.1109/TWC.2022.3180395
Wu, H. et al. Eavesdropping and anti-eavesdropping game in UAV wiretap system: A differential game approach. IEEE Trans. Wireless Commun. 21(11), 9906–9920 (2022).10.1109/TWC.2022.3180395
25. Wang D He H Liu D Adaptive critic nonlinear robust control: A survey IEEE Trans. Cybern. 2017 47 10 3429 3451 10.1109/TCYB.2017.2712188 28682269
Wang, D., He, H. & Liu, D. Adaptive critic nonlinear robust control: A survey. IEEE Trans. Cybern. 47(10), 3429–3451 (2017).28682269 10.1109/TCYB.2017.2712188
26. Qin C Qiao X Wang J Zhang D Hou Y Hu S Barrier-critic adaptive robust control of nonzero-sum differential games for uncertain nonlinear systems with state constraints IEEE Trans. Syst. Man Cybern. Syst. 2024 54 1 50 63 10.1109/TSMC.2023.3302656
Qin, C. et al. Barrier-critic adaptive robust control of nonzero-sum differential games for uncertain nonlinear systems with state constraints. IEEE Trans. Syst. Man Cybern. Syst. 54(1), 50–63 (2024).10.1109/TSMC.2023.3302656
27. Qin C Wang J Zhu H Zhang J Hu S Zhang D Neural network-based safe optimal robust control for affine nonlinear systems with unmatched disturbances Neurocomputing 2022 506 228 239 10.1016/j.neucom.2022.07.072
Qin, C. et al. Neural network-based safe optimal robust control for affine nonlinear systems with unmatched disturbances. Neurocomputing 506, 228–239 (2022).10.1016/j.neucom.2022.07.072
28. Qin C Shang Z Zhang Z Zhang D Zhang J Parallel learning-based security robust tracking control for nonlinear systems with uncertainties: An event-triggered design Eng. Appl. Artif. Intell. 2024 133 108 117 10.1016/j.engappai.2024.108077
Qin, C., Shang, Z., Zhang, Z., Zhang, D. & Zhang, J. Parallel learning-based security robust tracking control for nonlinear systems with uncertainties: An event-triggered design. Eng. Appl. Artif. Intell. 133, 108–117 (2024).10.1016/j.engappai.2024.108077
29. Wei X Yang J Optimal strategies for multiple unmanned aerial vehicles in a pursuit/evasion differential game J. Guidance Control Dyn. 2018 41 8 1799 1806 10.2514/1.G003480
Wei, X. & Yang, J. Optimal strategies for multiple unmanned aerial vehicles in a pursuit/evasion differential game. J. Guidance Control Dyn. 41(8), 1799–1806 (2018).10.2514/1.G003480
30. Jiang H Research on nonlinear control theory and optimization method based on adaptive dynamic programming 2019 Northeastern University
Jiang, H. Research on nonlinear control theory and optimization method based on adaptive dynamic programming (Northeastern University, 2019).
31. Chen NY Research on finite-time adaptive dynamic planning guidance based on differential countermeasures 2019 Nanjing University of Aeronautics and Astronautics
Chen, N. Y. Research on finite-time adaptive dynamic planning guidance based on differential countermeasures (Nanjing University of Aeronautics and Astronautics, 2019).
32. Zhang Y Zhang P Wang X Song F Li C Hao J An open loop Stackelberg solution to optimal strategy for UAV pursuit-evasion game Aerosp. Sci. Technol. 2022 129 107 130 10.1016/j.ast.2022.107840
Zhang, Y. et al. An open loop Stackelberg solution to optimal strategy for UAV pursuit-evasion game. Aerosp. Sci. Technol. 129, 107–130 (2022).10.1016/j.ast.2022.107840
33. Sun J Liu C Distributed zero-sum differential game for multi-agent systems in strict-feedback form with input saturation and output constraint Neural Netw. 2018 106 8 19 10.1016/j.neunet.2018.06.007 30007124
Sun, J. & Liu, C. Distributed zero-sum differential game for multi-agent systems in strict-feedback form with input saturation and output constraint. Neural Netw. 106, 8–19 (2018).30007124 10.1016/j.neunet.2018.06.007
34. Salmon JL Willey LC Casbeer D Garcia E Moll AV Single pursuer and two cooperative evaders in the border defense differential game J. Aerosp. Inf. Syst. 2020 17 5 229 239
Salmon, J. L., Willey, L. C., Casbeer, D., Garcia, E. & Moll, A. V. Single pursuer and two cooperative evaders in the border defense differential game. J. Aerosp. Inf. Syst. 17(5), 229–239 (2020).
35. Yuan Y Deng Y Luo S Duan H Distributed game strategy for unmanned aerial vehicle formation with external disturbances and obstacles Front. Inf. Technol. Electron. Eng. 2022 23 7 1020 1031 10.1631/FITEE.2100559
Yuan, Y., Deng, Y., Luo, S. & Duan, H. Distributed game strategy for unmanned aerial vehicle formation with external disturbances and obstacles. Front. Inf. Technol. Electron. Eng. 23(7), 1020–1031 (2022).10.1631/FITEE.2100559
36. Chen LB Liu SC Yuan RF An active defense guidance method based on self-learning differential countermeasures Electro-Opt. Control 2023 30 8 14
Chen, L. B., Liu, S. C. & Yuan, R. F. An active defense guidance method based on self-learning differential countermeasures. Electro-Opt. Control 30, 8–14 (2023).
