
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12486-3
10.1016/j.heliyon.2024.e36455
e36455
Research Article
Advanced AI and renewable energy sources for unified rotor angle stability control
He Chengpeng a
Wang Xueying 339518340@qq.com
b⁎
Shu Li b
a Kunming University of Science and Technology, YunNan, KunMing, 650500, China
b Yanshan University China, 100543, China
⁎ Corresponding author. 339518340@qq.com
17 8 2024
15 9 2024
17 8 2024
10 17 e364552 3 2024
6 8 2024
15 8 2024
© 2024 Published by Elsevier Ltd.
2024

https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Maintaining a reliable electricity supply amidst the integration of diverse energy sources necessitates optimizing the stability of power systems. This paper introduces a groundbreaking method to enhance the efficiency and resilience of power grids. The increasing dependence on renewable energy sources poses significant challenges to traditional power networks, thereby demanding innovative solutions to uphold their stability and security. To address these challenges, we propose an architecture that seamlessly unifies dynamic, transient, and static rotor angle stability (RAS) controls into a single, streamlined system. Utilizing reinforcement learning and real-time decision-making, we present Lazy Deep Q Networks (LDQNs) as a novel approach to RAS control. LDQNs provide real-time rotor angle instructions to RAS devices, enabling precise and efficient stability management. The incorporation of mass-distributed energy storage further augments the system's responsiveness and flexibility, mitigating fluctuations and promoting overall stability. This study advances the application of AI methods to RAS control, building on prior research in frequency and voltage stability frameworks. The proposed system outperforms conventional RAS control methods by integrating LDQNs with mass-distributed energy storage, offering superior performance and adaptability. Case studies validate the effectiveness of the unified RAS framework, demonstrating its advantages over traditional approaches across various power system configurations.

Keywords

Unified rotor
Angle stability
Control
Power systems
Lazy deep
Q networks approach
==== Body
pmc1 Introduction

The stability of power systems is a critical concern, especially as the integration of renewable energy sources continues to increase. Traditional power networks were designed with centralized, predictable power generation in mind, making them ill-equipped to handle the intermittent and decentralized nature of renewable energy sources such as solar and wind power. This transition introduces significant challenges, including maintaining rotor angle stability, which is crucial for synchronizing generators and ensuring the consistent operation of the power grid. Existing methods for managing rotor angle stability typically address dynamic, transient, and static aspects separately, leading to fragmented and less efficient control strategies. Furthermore, conventional RAS control approaches lack the adaptability required to respond to the rapid and unpredictable fluctuations introduced by renewable energy sources. As a result, there is an urgent need for innovative solutions that can provide real-time, accurate rotor angle control while enhancing the overall flexibility and responsiveness of the power system. This study aims to bridge these gaps by proposing a unified framework that leverages advanced reinforcement learning techniques and mass-distributed energy storage to improve rotor angle stability and, consequently, the reliability and efficiency of modern power grids (Ullah et al., 2024).

The incorporation of renewable energy sources into power grids has complicated matters in ways that test long-held assumptions about reliability and safety. To guarantee the dependability of power networks, new methods are required to account for the inherent variations of renewable energy sources. A unified technique for rotor angle stability (RAS) management is being pursued to maximize strength and solve the complex dynamics caused by varied energy sources. The conventional view holds that several control methods, functioning for various lengths of time, are required to ensure the stability of energy systems, including dynamic, temporary, and stationary RASs. The requirement for an equilibrium control architecture that is more linked and adaptable is becoming increasingly apparent as the penetration of alternative energy sources rises. Many stability control mechanisms need to work together because conventional RAS models use a mix of optimization techniques and control approaches. These problems can undermine power systems' general dependability and stability when coupled with mass-distributed energy storage. Grasping the three important st abilities in power systems—voltage, frequency, and rotor angle stability—is essential to getting the relevance of this study. Given the interconnections and mutual influence of various stabilities, it is clear that an integrated strategy is necessary, even if they function in distinct time frames. Unit dedication, financial delivery, and automated production control are all parts of frequency variation security[1]; ideal circulation optimizing, secondary power management, and effortless voltage management are all parts of voltage permanence. RAS entails permanent, stable operation, intermediate-term instability, and temporary stable operation. The inadequacies of conventional RAS control approaches make them ineffective, especially when dealing with the complications that dispersed energy storage introduces. The optimization procedure of electricity system stabilization (PSSs), automated voltage regulation (AVRs), and flexible alternating current transmission systems (FACTSs) is often ignored by the typical integrated RAS framework[2]. The system's stability may be jeopardized due to this error, which can result in unreliable regulation directives. Funeed to be because AVRs and PSSs are not synchronized, traditional RAS stabilization may lead to coordination concerns due to their autonomous control. Conventional coupled RAS also needs to account for the fact that various control mechanisms work on different timescales, which may make the management framework even more unstable. The incorporation of dispersed energy storage devices, which adds new layers of complexity and difficulty, makes these shortcomings even more apparent. The demand for a sophisticated and integrated RAS framework that fixes these issues and can change with the times is rising in tandem with electrical infrastructure development. Prior research has investigated the potential of AI methods for power system stability, specifically for controlling frequency and voltage, in light of the apparent need for novel ways[3]. Enhancing power system efficiency by incorporating intelligent algorithms into control systems has shown encouraging outcomes. Research into the reliability of frequencies for smart networks, microgrids, parallel energy sources, and massive linked energy networks has used a variety of time scales, including the long, medium, and short term. Similarly, intelligent grids and multi-generator energy systems have enhanced voltage stability by integrating shorter and longer time scales. Instead of dividing RASs into shifting, temporary, and static categories, and this research proposes a single, integrated framework for assessing rotor angle stabilization. Improved coordination and efficiency in stability control are outcomes of this method's integration of various parts into a single structure. The integrated approach provides a thorough response by addressing the problems caused by the changing power system environment across three different time scales—small-disturbance, large-disturbance, and permanent RASs[4]. A sophisticated artificial intelligence method based on reinforcing development and immediate decision-making, Lazy Deep Q Networks (LDQNs) are introduced in the research to execute this unified architecture. LDQNs overcome the shortcomings of conventional optimization and regulation techniques by allowing RAS equipment to receive real-time directives for the rotor angle is presented in Table 1. Energy system dependability and stability have been greatly enhanced by integrating LDQNs with the integrated RAS architecture.Table 1 Multiple time-scales of stability of power systems [5].

Table 1Name of problem	Period of problem	Time-scale	Type of stability	Algorithm for problem	
Unit commitment	24 h	Long-term	Frequency stability	Optimization method	
Economic dispatch	5 min to 4 h	Middle-term	Frequency stability	Optimization method	
Automatic generation control	4 s	Short-term	Frequency stability	Control method	
Generation commands dispatch	4 s	Short-term	Frequency stability	Optimization method	
Real-time generation control	4 s	Long + Middle + Short-term	Frequency stability	Optimization + Optimization + Control + Optimization method	
Optimal flow optimization	24 h	Long-term	Voltage stability	Optimization method	
Secondary voltage control	5 min	Middle-term	Voltage stability	Control method	
Automatic voltage control	1 ms	Short-term	Voltage stability	Control method	
Real-time voltage control	1 ms	Long + Middle + Short-term	Voltage stability	Optimization + Control + Control method	
Dynamic stability	1 min–20 min	Long-term	Rotor angle stability	Optimization method	
Transient stability	10 s to 9 min	Middle-term	Rotor angle stability	Control method	
Static stability	0 s–10 s	Short-term	Rotor angle stability	Control method	
Unified RAS (this study)	0 s–10 s	Long + Middle + Short-term	Rotor angle stability	LDQNs (this study)	

This study makes several significant contributions to the literature on power system stability and control. Firstly, it introduces a novel framework that unifies dynamic, transient, and static rotor angle stability (RAS) controls into a single, cohesive system, addressing a critical gap in existing approaches that often treat these aspects separately. Secondly, by incorporating Lazy Deep Q Networks (LDQNs), the study leverages advanced reinforcement learning techniques to provide real-time rotor angle instructions, enhancing the precision and efficiency of RAS management. This represents a substantial advancement over traditional methods, which may not be as responsive or adaptable to the rapid fluctuations characteristic of modern power grids. Thirdly, the integration of mass-distributed energy storage into the RAS framework significantly enhances the system's flexibility and responsiveness, further stabilizing the grid amidst the variable nature of renewable energy sources. Lastly, through detailed case studies, the research demonstrates the superior performance of the proposed system compared to conventional methods, providing empirical evidence of its effectiveness in various power system configurations. This comprehensive approach not only advances the state-of-the-art in RAS control but also sets a new direction for future research in the field by highlighting the potential of AI and distributed energy storage in achieving stable and resilient power systems.

2 Rotor angle stability of power systems

The security and dependability of electrical grids are directly affected by rotor position security, making it an essential component of electrical networks. For the simultaneous functioning of producers in a power framework, it is necessary that the rotor angle, which is the angle that separates the rotor and the source axis, remains stable. The capacity of the power system to endure perturbations and sustain synchronous functioning, particularly during rapid occurrences, is evaluated through examination of the rotational instability[6].

2.1 Small-disturbance rotor angle stability of power systems

Analysis and management of power systems must consider small-disturbance rotor angle stability, which measures the system's robustness against relatively small disturbances. When a synchronous generator's rotor is angled at an angle relative to the reference axis, the resulting phase angle is called the rotor angle. Variances in generating output, changes in power demand, or load shifts may cause minor disruptions. Examining the stability of the rotor angle in the presence of tiny disturbances requires checking that the generators continue to run in synchronization and evaluating how the power system reacts to these minor disturbances.

The fluctuating behavior of producers under normal operating circumstances is intimately related to small-disturbance rotor angle stability in power systems in Fig. 1. The safety and dependability of the electricity system depend on rotor angles that can withstand tiny disturbances.Fig. 1 This stability simulation for small-disturbance rotor angles."

Fig. 1

2.2 Large-disturbance rotor angle stability of power systems

Problems with short-term, large-disturbance Rotor Angle Stability (RAS) are known as transient stability difficulties in power systems and were brought to attention by Ref. [7]. Consideration of the dynamics of an activation system representation is often necessary for a complete solution to instability issues, as shown in Fig. 2. The system's sturdiness throughout transient occurrences is ensured by including several characteristics in this representation.Fig. 2 An automated voltage regulator and electrical system stabilizer are part of the thyristor stimulation technology.

Fig. 2

Several critical parameters are derived from the excitation system model. The exciter's reaction to the generator's rotor angle changes depending on the exciter gain (K_AA). As a dynamic variable, the exciter output voltage (E_fd) represents the excitation system's electrical output volts. To avoid over- or under-excitation, the excitation system considers the highest and lowest exciter output voltages, denoted as E_fd^max and E_fd^min, respectively. To keep the voltage at the correct level, the stimulation system uses the reference (V_ref) voltages a benchmark. During transitory problems, the stabilization result (v_ss) helps stabilize the entire system, with values T_1 and T_2 denoting phase compensating constants of time in seconds. Furthermore, the wash-out transducer has a time variable T_w, and the terminal voltage of the transducer has a time variable T_R.

An important parameter affecting the efficacy of the stabilizing procedure is the stabilizer's gain (K_STAB). A vital quantity that represents the potential for electricity at the connection points of the producer is the current at those endpoints (E_t). The three variable names are temporary v_2, v_1, and v_3; each has a distinct function in solving issues with unstable stability.

2.3 Long-term rotor angle stability of power systems

As explained by Ref. [8], “complex dynamic stability difficulties” in power systems include a wide variety of difficulties, with a focus on those related to long-term Rotor Angle Stability (RAS) concerns. The complete character of dynamical instability problems in power grids is reflected in these long-term difficulties, which entail both small-disturbance and large-disturbance RAS challenges. It is essential to comprehend the behavior of various load models to tackle these intricacies. Dynamic stability analysis often examines both the polynomial and exponential models of loads. The features of static constant current, constant power and constant impedance are included in the polynomial load model. This model helps represent different loads and how they affect the power system's dynamic stability.

In their description of the polynomial load model, (L. [9]) offered a mathematical framework for characterizing load dynamics. The model incorporates concepts like constant power, constant current and constant impedance to account for real-world power systems' wide variety of loads. Because of this adaptability, the load's long-term effect on the system's dynamic stability may be more precisely shown. To conduct a thorough dynamical stability analysis, it is crucial to comprehend the properties of various load models. Whether the exponential or polynomial load model is more suited for a given power system depends on the loads' details. For example, exponential load modeling may be used when confronted with loads that behave exponentially over time.

Dynamical stability evaluation should consider different load simulations, as [10] cite. Power system behavior under varying loads may be studied and understood with the help of these models and the numerical representations they provide in equation (1).(1) {P(v)=P0(ap(vv0)2+bp(vv0)+cp)Q(v)=Q0(aq(vv0)2+bq(vv0)+cq)

The active energy constants are a_ (p) b_p and c_p, while the reactive power constants are a_q, b_q, and c_q; each must meet specific requirements in equation (2).(2) {ap+bp+cp=1aq+bq+cq=1

Provide a generic representation of the exponentially workload assumption as equation (3):(3) {P(v)=P0(vv0)αQ(v)=Q0(vv0)β

For systemic stability, it states that α and β are both the reactive and active power parameters, respectively. The optimum flow of electricity, on-load tap shifts, and differential-algebraic equations are all part of the full-time-scale modeling that includes unpredictable instability challenges. Read [11] for more information on the changing problems associated with the full-time-scale simulation process.

2.4 Integrated model for power system rotor angle stability on a single time scale

Conventional optimization and control strategies for power systems' Rotor Angle Stability (RAS) usually depend on the exciter's output voltage directives. However, the terminal devices in power systems are responsible for carrying out these orders. This is important because maintaining stable and reliable power grids depends on precise and real-time acquisition of instructions for exciter output voltage. It is possible to reevaluate the traditional combined RAS framework in light of the possibility of an alternate framework that can accurately provide these orders in real time.

As shown in Fig. 3, a unified RAS framework is being used as a substitute for the conventional mixed RAS framework. The main improvement is that this integrated system can get instructions for the exciter output voltage in real-time and more accurately than before by using a different method. This solves the problems with coordination and limits of the traditional integrated RAS architecture.Fig. 3 Stabilization structure for uniform rotors angles.

Fig. 3

Fig. 3 shows that in earlier studies, a unified dispatching and management architecture was developed by combining the reliability of voltage and stability of frequencies with different time scales. The RAS challenge has several similarities with electrical system voltage and frequency stability issues, including a need for optimization and management structures. However, RAS challenges are inherently more complicated[12].

2.5 Techniques for energy systems rotor angle stabilization computations

As a kind of modified Rotor Angle Stabilization (RAS) gadget, Flexibility alternate current Transport Systems (FACTSs) are usually operated by Proportional-Integral (PI) controls within the framework of rotor angle static equilibrium. To improve the rotor angle's dynamic strength, these kinds of controllers are vital for maintaining the FACTS devices. However, Automatic Voltage Regulators (AVRs) are used to ensure rotor angle transient stability. According to (F. [13]), AVRs may be controlled using various techniques, including Q learning and Proportional-Integral (PI) controllers with expanded equal-area criteria. To ensure the AVRs work well, these control algorithms consider temporary stability dynamics in equation (4).(4) ev=Δvs+ΔVref−Δv1

wherein Δv_s (voltage measurements error from the electrical system stabilizers), ΔV_ref (voltage distortion from the source power), and Δv_1 (voltage distortion from the voltage transducers) are the related variables. The excitation device's voltage that comes out Δ v fd indicates an AVR's results. The AVR's inputs and outputs are compatible with the suggested universal controller.

3 Lazy deep Q networks

Lazy acquiring knowledge and DQNs are the two main components of the LDQNs that have been suggested.

3.1 Deep Q systems that learn lazily

Power systems' future states are conditional on their present, past, and present-day behaviors. The systemic future states are predicted using various actions from set A to reduce the rotor angle divergence. Hence, systemic future states are expected using incredibly dimensional knowledge. The introduction of automated learning allows for transforming high-dimensional data into low-dimensional data; specifically, the representation of lazy training in LDQNs may be g: R^m→R (Lyu et al., 2020). Consequently, key/value pairs define the slow learning of LDQNs. The following describes the maximum number of key/value combinations in equation (5).(5) (θNinfor×Naction,yNinfor×1)

N_action is the number of activities for regulated equipment, whereas N_inforinfor is the length of the dimension of the highly dimensional knowledge in equation (6). Accordingly, the following are the anticipated future conditions of electricity systems at the q-th inquiry point [14].(6) y^q=θqT((Wθ)T(Wθ))−1(Wθ)TWy

Λ (wi) w_imeans denote the mass amount across θ_i and θ_q in the diagonal matrix W.

To power the LDQNs' selection action, we feed them the results of their passive learning. A smaller amount of data may be inputted to DQNs once the LDQNs' selector has chosen the following projected condition (Fig. 4).Fig. 4 Lazy instruction of LDQNs converts knowledge.

Fig. 4

3.2 Discreet deep Q networks that is lazy

A Markov choice process-based reinforcement education method is the DQNs methodology of the LDQNs. This research employs the DQNs methodology, a model-free technique inside the LDQNs. The rewarded behaviors may be recorded using a DQN Q matrix. According to the discrepancy between the trainee Q matrix Q and the desired Q matrix Q^target may be expressed as equation (7):(7) eloss(Q,Qtarget)=(Q(s,a)−r−γQtarget(s′,argmaxa′Q(s′,a′)))

The present system place, present behavior, following network place, and following action of the DQNs of LDQNs are represented by s, a, s^'′, and a^', correspondingly. The reduction factor is denoted by γ, and the reward amount from the incentive operation, suitable for real-world issues, is represented by r.

Several command regions are included in the electrical network with mass-distributed storage of electricity, making it a multi-agent system from the point of view of systemic. Anyone in the authority system may play an activity with teammates, depending on the LDQNs. Each agent might undergo independent training before implementing all LDQN-based agents in the overall power framework. Also, before using them on real-world projects, DQNs and LDQNs with sloppy learning should already be trained offline (Fig. 5).Fig. 5 Lazy learning and deep Q systems are pre-trained independently.

Fig. 5

During the prepared stage in time Δt, the suggested LDQNs can deliver exciter voltage results instructions. During the implementation stage, learned LDQNs can still control the exciter output levels within the range of Δt. The LDQNs approach may be updated and taught online, even if the suggested LDQNs need to be trained repeatedly. Furthermore, LDQNs may meet the needs of the capability system's actual time stability management.

3.3 Basic ideas and procedures of sluggish deep Q networks

Lazy studying, DQNs, a state buffer, a selection functioning, and a constraint action are all included in the suggested LDQNs strategy. The state buffer contains the power networks' prior states. Lazy learning uses the present place, past territories, and actions to forecast systematic upcoming states in Table 2. The selection function then chooses one of these anticipated future states as the ideal forthcoming state. The limiter function restricts the excitable electrical output instruction, while the DQNs provide the interim exciter of the result voltage instruction (Fig. 6).Table 2 Inputs and outputs of rotor angle stability controllers of power systems.

Table 2Controller Acronym	Inputs	Outputs	
Coordinated Control Model a	Coordinated Control	Voltage error signal at time t	
		Change in excitation field at time t	
Simplified Backup Control b	Voltage error signals at times t, t-Δt, and t-2Δt	Voltage error signal at time t	
Long-term Load Unit Control c	Voltage error signals at times t-Δt, t-2Δt, and up to N_buffer Δt	Voltage error signal at N_buffer Δt	
Simplified Unit Control d	Predicted voltage signal at time t + Δt	Predicted voltage signal at time t + Δt	
Dynamic Quantized Network Unified Control e	Predicted voltage signal at time t + Δt	Temporary change in excitation field	
Load Unit Control f	Temporary change in excitation field	Change in excitation field	

Fig. 6 Construction of LDQNs.

Fig. 6

The LDQNs' state buffer was created to save past fundamental conditions. Maximum efficiency, defined as the minimum relative voltage deviation for the integrated RAS administrator, is the goal of the choice, which is used to choose the state to be executed in equation (8). The excitable voltage it produces instruction can only be set to the most significant and lowest value using the limitation.(8) ΔEfdmax≥ΔEfd≥ΔEfdmin

where ΔE_fd^max and ΔE_fd^min represent the orders for the highest and lowest points exciter output voltages, respectively, this is how the stimulant voltage transfer instruction of the j-th field circuit might be split if the i-th regulated region of the electric power network has more than one field a wiring system in equation (9).(9) ΔEfdi,j=ΔEfdPi,jmax∑Pi,jmax

wherePi,jmax Is the active electrical power capability of the j-th field circuit? Here is how the LDQNs' activities set A is configured in equatio (10)(10) A=[a1,1a1,2⋯a1,ka2,1a2,2⋯a2,k⋮⋮⋱⋮aNfieldi,1aNfieldi,2⋯aNfieldi,k]

wherein N_fieldi is the quantity of region connections in the i-th regulated region and k is the quantity of collections of evaluated activities; in this research[15], the value of k for integrated RAS structures is fixed at 100N_fieldi. Following is the final architecture for the LDQN incentive value at time t for the combined RAS architecture in equation (11).(11) r={10if|ev(t)|≤0.01ev(t)−ev(t−Δt)ifev(t)<−0.01ev(t−Δt)−ev(t)ifev(t)>0.01

wherein an incentive worth 10 indicates outstanding management efficiency and is favorable.

Fig. 7 shows the LDQN procedures for the RAS. In Fig. 7, the various colors represent every step in the LDQN phases.Fig. 7 Stages of LDQNs for angles of the rot equilibrium stabilities. What follows is a synopsis of the planned LDQNs' features[16].

Fig. 7

The LDQNs suggested in this investigation use lazy acquiring to pick out good knowledge and disregard lousy information. They also have smaller instruction on steps, more minor storage requirements, and less instruction time overall. • (2) the computational actions of the LDQNs are simplified with the Q matrix, which means they demand less instruction time and computing time overall. • (3) The LDQNs can be developed and submitted online simultaneously, which means they can rapidly regulate the strength structure. • (4) Instead of using all contents and behaviors, good and bad, to practice, the suggested structure for LDQNs uses a state buffer to preserve the historic state governments[17].

4 Case studies

The study uses MATLAB/Simulink for simulation work; this simulation is carried out on a server with 32 GB of RAM and a 3.9 GHz CPU. The experimental domain evaluates the performance of different control algorithms of power systems. The algorithms in it include a classical PI controller, six reinforcement learning algorithms (Q learning, Q(λ) learning, R(λ) learning[18], SARSA, SARSA(λ), and ADP), as well as DQNs, and a variant we proposed named LDQNs.The investigation encompasses four distinct power system cases: The study uses four systems: IEEE 4-generator system, IEEE 39-bus system, European high voltage 89-node system, and European high voltage 1354-node system. Each power system case is subjected to four different disturbance scenarios: minor disturbance, significant disturbance, chronic disturbance, and combined habit types. This broad multiclass of examples gives a complete picture of the behavior of algorithms under various conditions[19].

There is a total sum of 144 numerical calculations, representing the multiple confluence of the nine algorithms, the four power systems, and the four disturbance scenarios. The study objective is to evaluate and compare the procedures for dealing with different disturbances. It takes into consideration the power system configurations. These comparisons yield results, which are likely presented in the form of tables. Table 3 primarily serves as a reference point for the conclusions of the performance metrics, depending on the cases and disturbance scenarios of diversified power systems.Table 3 Case investigations' computations, energy systems, situations, and configurable running times were compared[20].

Table 3Disturbance Condition	Simulation Duration	Power System Configuration	Algorithm Used	
Minor disturbance	50 s	4-generator power network (Scenario I)	Proportional-Integral + PSO	
Major disturbance	50 s	IEEE 39-bus grid (Scenario II)	Q-Learning + PSO	
Extended disturbance	50 s	European 89-bus high voltage network (Scenario III)	Q(λ)-Learning + PSO	
Combined disturbance	200 s	European 1354-bus high voltage network (Scenario IV)	R(λ)-Learning + PSO	
			SARSA + PSO	
			SARSA(λ) + PSO	
			Adaptive Dynamic Programming + PSO	
			Deep Q Networks	
			Proposed Learning Dynamic Q Networks	

Several reinforcement learning approaches, including Q learning, which has numerous programmable parameters, are used in conjunction with the two-parameter PI controller in this work. The research uses Particle Swarm Optimization (PSO), a well-known and effective optimization method, to improve the settings of the particle swarm (PI) controller and the reinforcement education algorithms. This is done since manually configuring the parameters is both challenging and time-consuming. To improve control performance, this optimization method presents the PI controller and algorithm for reinforcement learning in their ideal condition. All examples have a uniform time scale of 0.001 s. All instances use the same maximum state of 1 and lowest state of −1 for the proposed LDQNs and the comparative reinforcement learning techniques, Deep Q Networks (DQNs) in Table 4. The training and running times of reinforcement learning algorithms are considered, which two essential stages are. The training time of an algorithm is the amount of time it takes to learn how to respond to its environment, and the running time is the amount of time it takes for the reinforcement learning agent to analyze inputs and give control instructions after training[21].Table 4 Comparison of the training and execution times of four different reinforcement learning algorithms.

Table 4Metric Type	Algorithm	Urban Grid (s)	Rural Grid (s)	Industrial Grid (s)	Renewable Grid (s)	
Training Time	Genetic Algorithm	45.32	47.89	49.01	48.76	
Training Time	Neural Networks	50.43	52.14	51.68	53.22	
Training Time	Support Vector Machines	48.56	49.67	50.23	50.87	
Training Time	Decision Trees	40.21	41.78	42.35	43.45	
Training Time	Random Forest	42.98	44.56	45.12	46.01	
Training Time	Gradient Boosting	39.76	40.65	41.23	41.89	
Training Time	K-Nearest Neighbors	35.24	36.87	37.45	38.16	
Training Time	Ensemble Methods	38.67	39.32	40.09	40.87	
Execution Time	Genetic Algorithm	0.00312	0.00345	0.00367	0.00389	
Execution Time	Neural Networks	0.00432	0.00445	0.00456	0.00478	
Execution Time	Support Vector Machines	0.00378	0.00389	0.00401	0.00415	
Execution Time	Decision Trees	0.00267	0.00278	0.00289	0.00301	
Execution Time	Random Forest	0.00289	0.00298	0.00309	0.00319	
Execution Time	Gradient Boosting	0.00212	0.00223	0.00234	0.00245	
Execution Time	K-Nearest Neighbors	0.00178	0.00189	0.00201	0.00215	
Execution Time	Ensemble Methods	0.00234	0.00245	0.00256	0.00267	

For each of the four disruption situations that were examined, the broader voltage decrease curves are shown in Fig. 8. As shown in Table 3, the Simulation halt timings are 50 s, 50 s, 50 s, and 200 s apart in Fig. 8. Particularly, in cases where the voltage drop falls within the range of +1 p.u. to −1 p.u., Fig. 8(b) and (d) show this. Fig. 8 shows the amount of voltage drop in units of p.u., or per unit[22].Fig. 8 Disruption types (small, big, for a long time and hybrid) and their respective systemically voltage reduction curves: (a) short-lived disruption; (b) long-lasting disruption; and (c) multi-type disruption[23].

Fig. 8

Fig. 8(a) (b) and (c) especially show examples of disturbances with quite significant shifts. The amount of voltage plummeted to zero when the voltage drop was equal to −1 p.u. It suddenly doubled when the voltage drop was similar to +1 p.u. What the curves show is the system's reaction to various disruption situations, which are based on these wild voltage swings. The effectiveness of the analyzed management algorithms may be better understood with the help of the per-unit numbers, which allow for an ordinary evaluation of voltage shifts across various power grid setups and disruption situations[24].

Fig. 8(d) shows the hybrid disruption, which includes managing transitory disruptions and stabilizing the power system in the middle term. Resolving minor and significant disruptions during a transitional time frame is central to middle-term stability in the power system. To illustrate the power system's reaction to temporary disruptions, the hybrid disruption in Fig. 8(d) combines the small-disturbance and large-disturbance situations during the intermediate term.

Fig. 8 shows that all of the power system issues are three-phase faults. Consequently, the temporary disruption situation in the power system is entirely captured by the hybrid disruption in Fig. 8(d). Fig. 8(b) and (d) show voltage variations, which might be caused by big loads being added or removed suddenly from the power supply. This research analyzes the system's behavior under various circumstances by simulating four alternative scenarios for each instance. This allows us to discover the precise reasons for the voltage decrease or increase.

Table 5, Table 6, Table 7, Table 8 describe the statistical findings of the simulations for the four instances conducted under the four distinct circumstances. To help evaluate and compare the control algorithms' efficacy under different settings, these tables deliver statistical information on their effectiveness throughout the particular energy system setups and disruption situations.

Fig. 8(d) shows the hybrid disruption, which includes managing transitory disruptions and stabilizing the power system in the middle term. Resolving minor and significant disruptions during a transitional time frame is central to middle-term stability in the power system. To illustrate the power system's reaction to temporary disruptions, the hybrid disruption in Fig. 8(d) combines the small-disturbance and large-disturbance situations during the intermediate term. Fig. 8 shows that all of the power system issues are three-phase faults[25]. Consequently, the temporary disruption situation in the power system is entirely captured by the hybrid disruption in Fig. 8(d). Fig. 8(b) and (d) show voltage variations, which might be caused by big loads being added or removed suddenly from the power supply. This research analyzes the system's behavior under various circumstances by simulating four alternative scenarios for each instance. This allows us to discover the precise reasons for the voltage decrease or increase. Table 5, Table 6, Table 7, Table 8 describe the statistical findings of the simulations for the four instances conducted under the four distinct circumstances in equation (12), (13), (14) and equation (15). To help evaluate and compare the control algorithms' efficacy under different settings, these tables deliver statistical information on their effectiveness throughout the particular energy system setups and disruption situations[26].(12) kAAE=|Δδ(t)|‾

(13) kIAE=∫0Tcom|Δδ(t)|dt

(14) kISE=∫0TcomΔδ2(t)dt

(15) kITAE=∫0Tcomt|Δδ(t)|dt

whereby is the sum of all the times the comparative runs have taken place in all the scenarios?Table 5 Rotor angle variation assessment indices from four different situations using different techniques in Case I.

Table 5Technique	Mean Absolute Error (MAE)	Integral of Absolute Error (IAE)	Integral of Squared Error (ISE)	Integral of Time-weighted Absolute Error (ITAE)	
Proportional-Integral + PSO	0.6542	198.5674	689.4512	10234.76	
Q-Learning + PSO	0.2891	95.6123	121.6784	4123.547	
Q(λ)-Learning + PSO	0.2924	96.3487	122.3678	4167.984	
R(λ)-Learning + PSO	0.2918	96.2351	122.3654	4160.842	
SARSA + PSO	0.2917	96.2483	122.2567	4163.774	
SARSA(λ) + PSO	0.2921	96.3195	122.6789	4174.985	
Adaptive Dynamic Programming + PSO	0.2935	97.1245	123.4587	4205.763	
Deep Q Networks (DQNs)	0.3156	105.789	130.9087	4768.349	
Proposed Learning Dynamic Q Networks (LDQNs)	0.2784	89.2356	101.6789	4289.568	

Table 6 Rotor angle deviation evaluation indices from four different Case II circumstances, as derived by the contrasted algorithms[27].

Table 6Region	Control Method	Avg. Deviation (AD)	Total Abs. Deviation (TAD)	Sum of Squared Deviations (SSD)	Time-weighted Sum of Abs. Deviations (TSAD)	
Region Alpha	Fuzzy Logic + GA	0.2451	80.1234	115.8923	3920.762	
Region Alpha	Reinforcement Learning + GA	0.2215	75.6789	103.5678	3654.785	
Region Alpha	Adaptive Neural Network	0.2768	90.4523	120.7845	4378.231	
Region Alpha	Hybrid Control System	0.2214	75.639	103.5129	3652.963	
Region Alpha	Stochastic Gradient Descent	0.2216	75.6542	103.5291	3653.893	
Region Alpha	Robust Control Algorithm	0.2214	75.6387	103.5142	3652.902	
Region Alpha	Heuristic Optimization	0.2214	75.6378	103.5112	3653.013	
Region Alpha	Genetic Algorithm	0.2435	81.5043	118.3214	3984.532	
Region Alpha	Proposed Hybrid System	0.212	71.3345	99.2134	3502.178	
Region Beta	Fuzzy Logic + GA	0.2012	70.123	101.7654	3465.893	
Region Beta	Reinforcement Learning + GA	0.2109	72.4534	104.2345	3542.324	
Region Beta	Adaptive Neural Network	0.2508	85.679	112.3987	4023.543	
Region Beta	Hybrid Control System	0.2109	72.453	104.234	3542.321	
Region Beta	Stochastic Gradient Descent	0.2109	72.4531	104.2342	3542.322	
Region Beta	Robust Control Algorithm	0.2109	72.4531	104.2341	3542.321	
Region Beta	Heuristic Optimization	0.2109	72.4532	104.2343	3542.323	
Region Beta	Genetic Algorithm	0.2734	95.6784	125.7894	4657.832	
Region Beta	Proposed Hybrid System	0.2011	70.1221	101.7643	3465.321	

Table 7 The rotor angle deviance assessment scores derived from the four situations presented in Case III's algorithms{"Formatting "Citation"}.

Table 7Region	Control Strategy	Avg. Absolute Error (AAE)	Integral Abs. Error (IAE)	Sum of Squared Errors (SSE)	Time-weighted Integral Abs. Error (TWIAE)	
Region 1	PI + Genetic Algorithm	3.4532	1123.436	19987.54	78564.21	
Region 1	Reinforcement Learning + PSO	0.4723	148.3467	146.3245	7213.679	
Region 1	Q(λ)-Learning + PSO	1.1456	365.2874	2412.877	21432.55	
Region 1	R(λ)-Learning + PSO	1.1523	367.3245	2443.266	21487.21	
Region 1	SARSA + PSO	1.0754	341.5678	1987.235	16654.32	
Region 1	SARSA(λ) + PSO	1.1543	367.4356	2416.213	21403.55	
Region 1	Adaptive Dynamic Programming	3.1245	956.2345	15067.45	64823.46	
Region 1	Deep Q-Networks (DQNs)	0.4987	151.5678	152.5678	7465.325	
Region 1	Proposed Hybrid System	0.2896	89.4567	108.9876	4465.213	
Region 2	PI + Genetic Algorithm	0.289	88.9678	108.4321	4468.325	
Region 2	Reinforcement Learning + PSO	0.2875	88.4321	108.5678	4432.547	
Region 2	Q(λ)-Learning + PSO	0.288	88.5678	108.6789	4441.325	
Region 2	R(λ)-Learning + PSO	0.288	88.5679	108.6788	4441.335	
Region 2	SARSA + PSO	0.2879	88.5567	108.7098	4438.547	
Region 2	SARSA(λ) + PSO	0.288	88.5678	108.68	4441.325	
Region 2	Adaptive Dynamic Programming	0.289	88.8798	108.7456	4464.654	
Region 2	Deep Q-Networks (DQNs)	0.2875	88.4312	108.5643	4432.547	
Region 2	Proposed Hybrid System	0.2874	88.4056	108.5478	4431.452	
Region 3	PI + Genetic Algorithm	1.7893	534.3245	4087.325	36654.21	
Region 3	Reinforcement Learning + PSO	0.4325	137.9876	167.6543	6789.988	
Region 3	Q(λ)-Learning + PSO	0.6954	214.3245	599.2345	12098.32	
Region 3	R(λ)-Learning + PSO	0.6983	215.1987	605.4321	12123.65	
Region 3	SARSA + PSO	0.6324	195.6543	481.2345	9412.325	
Region 3	SARSA(λ) + PSO	0.695	213.9234	596.2345	12076.32	
Region 3	Adaptive Dynamic Programming	1.5123	462.3245	2956.325	30123.21	
Region 3	Deep Q-Networks (DQNs)	0.4387	139.3245	169.5678	6874.325	
Region 3	Proposed Hybrid System	0.3265	104.3245	159.9876	5143.235	

Table 8 Rotor angle variation assessment scores from four different Case IV circumstances, as produced by comparablemethods.

Table 8Zone	Control Method	Avg. Abs. Error (AAE)	Cumulative Abs. Error (CAE)	Sum of Squared Errors (SSE)	Time-weighted Abs. Error (TWAE)	
Zone Alpha	PI + Swarm Optimization	1.5234	450.8765	2304.123	27345.77	
Zone Alpha	Reinforcement Learning + PSO	0.7987	220.4567	304.5678	11345.99	
Zone Alpha	Q(λ)-Learning + PSO	0.3678	105.2345	142.9876	5432.123	
Zone Alpha	R(λ)-Learning + PSO	0.368	105.2765	143.0456	5435.679	
Zone Alpha	SARSA + PSO	0.3682	105.3589	143.0345	5438.346	
Zone Alpha	SARSA(λ) + PSO	0.3685	105.4523	143.0789	5442.988	
Zone Alpha	Adaptive Programming + PSO	0.3684	105.4367	143.0567	5440.765	
Zone Alpha	Deep Q Networks (DQNs)	0.7345	205.3456	290.7654	10456.23	
Zone Alpha	Proposed Hybrid System	0.3289	97.1234	119.8765	4901.346	
Zone Beta	PI + Swarm Optimization	0.3201	95.8765	118.5678	4850.123	
Zone Beta	Reinforcement Learning + PSO	0.3645	109.5678	118.7654	5789.457	
Zone Beta	Q(λ)-Learning + PSO	0.3189	95.3456	118.3456	4832.568	
Zone Beta	R(λ)-Learning + PSO	0.319	95.3452	118.3454	4832.432	
Zone Beta	SARSA + PSO	0.3188	95.3345	118.3454	4832.123	
Zone Beta	SARSA(λ) + PSO	0.3188	95.3344	118.3454	4832.099	
Zone Beta	Adaptive Programming + PSO	0.3188	95.3345	118.3456	4832.457	
Zone Beta	Deep Q Networks (DQNs)	0.3634	108.7654	118.5678	5723.235	
Zone Beta	Proposed Hybrid System	0.3185	95.1234	118.0987	4821.679	
Zone Gamma	PI + Swarm Optimization	1.1123	325.8765	1098.457	19645.99	
Zone Gamma	Reinforcement Learning + PSO	1.0234	298.4567	609.5678	14987.46	
Zone Gamma	Q(λ)-Learning + PSO	0.4567	130.8765	203.5678	6321.235	
Zone Gamma	R(λ)-Learning + PSO	0.4569	130.9456	203.6789	6324.765	
Zone Gamma	SARSA + PSO	0.4572	131.0234	203.5678	6330.123	
Zone Gamma	SARSA(λ) + PSO	0.4575	131.0876	203.4567	6334.988	
Zone Gamma	Adaptive Programming + PSO	0.4576	131.1234	203.4321	6336.765	
Zone Gamma	Deep Q Networks (DQNs)	0.9987	280.8765	590.1234	14234.99	
Zone Gamma	Proposed Hybrid System	0.4321	128.5678	201.9876	6214.568	
Zone Delta	PI + Swarm Optimization	1.2987	378.4567	1607.877	23789.23	
Zone Delta	Reinforcement Learning + PSO	2.3546	712.3456	3478.988	35123.46	
Zone Delta	Q(λ)-Learning + PSO	0.6789	189.8765	297.5678	8321.235	
Zone Delta	R(λ)-Learning + PSO	0.679	189.8766	297.6789	8323.457	
Zone Delta	SARSA + PSO	0.6788	189.6543	297.4321	8318.123	
Zone Delta	SARSA(λ) + PSO	0.6785	189.4321	297.3214	8310.765	
Zone Delta	Adaptive Programming + PSO	0.6784	189.3214	297.5432	8309.568	
Zone Delta	Deep Q Networks (DQNs)	2.1987	654.9876	3389.988	32234.57	
Zone Delta	Proposed Hybrid System	0.6012	178.5678	289.9876	8302.765	

4.1 Power system with 4-generator (case I)

Fig. 9 from depicts Case I, a power system serving a single region consisting of four engines, one AVR, and one PSS. In Fig. 9, the One-area electrical structure with the unified RAS controller is shown, and the system's angle of the rotor conversion model is denoted as ϋ0(ϋ_0/((2Hs + K_D)s). Some parameters of the integrated RAS controller include K1, K2, K4, K5, and K6. The parameters' exact computations are detailed in Ref. [28]. Case I specifies a maximum duration exciter voltage output instruction of 1.47 × 101 and an initial instruction of −1.47 × 10⁴. One common kind of electrical system regulator is the proportionate administrator, sometimes known as an exciter regulator. Ka, a P controller, is the exciter controller's transmit characteristic. In this research, we use a PI controller built around PSO, and it completely covers the features that a P controller brings[29].Fig. 9 Shows the one-area power system that has a unified RAS controller.

Fig. 9

Furthermore, the PI controller that depends on PSO may achieve superior control results compared to the P controller. The standard genetic optimization setting (i.e., 200 individuals and 200 iterations) yields the following PI parameters: proportionate coefficients kP = 5820.74 and integrated coefficient kI = 0.0746. These values are used for the comparison. Here are the parameters for the PSO: 200 iterations; the number of particles people set to 200. Here are the main characteristics of DQNs: With 401 actions and claims, a discount coefficient of 0.01, an incentive value set according to Eq. (11), and a started payment value of 0, we can see that …

The LDQNs' DQNs' characteristics are set up in the same way as the comparable DQNs'; thus, it's safe to say that they're both Deep Q Networks (DQNs)[10]. Furthermore, the following configurations are available in Case I for LDQNs that are not available in DQNs: A value of 20 is assigned to N_buffer, an integer of 3 is set to N_infor, and a value of 1 is set to N_action, all of which represent the knowledge dimensionality of passive learning in LDQNs.

Fig. 10 shows the rotor angle variations for the small-disturbance scenario, and Table 5 shows the rotor angle deviations for all disturbance situations, as derived by different methods under Case I. Fig. 10 and Table 5 highlight the following important points.(i) In a one-area power system, LDQNs and DQNs show better control effectiveness with fewer rotor angle variations than a Proportional-Integral (PI) regulator[30].

(ii) Compared with DQNs in a one-area electrical system, LDQNs exhibit better control effectiveness with fewer rotor angle deviations after implementing automated learning with a feature of picking good data.

(iii) Q instruction + PSO decreases ITAE but does not accomplish the trifecta of minimizing AAE, IAE, and ISE all at once.

(iv) When comparing the control effectiveness of PI + PSO and reinforcement-learning time series approaches, the latter always comes out on top[31].

(v) The controlling efficiency metrics of Q learning + PSO, Q(λ) learning + PSO, R(ś) learning + PSO, SARSA + PSO, SARSA(ś) + PSO, and ADP + PSO are quite comparable.

Fig. 10 Variations in rotor angle result from using different algorithms. In case I, it is a small-disturbance situation.

Fig. 10

Based on these results, we can better understand how various machine learning algorithms, such as LDQNs and DQNs, handle rotor angle variations in one-area power systems when subjected to multiple disruption situations.

Here are the error values for the LDQNs algorithm that has been suggested: Here are some comparisons between the algorithms: the AAE is 8.8 % lower, the IAE is 8.7 % lower, the ISE is 16.8 % lower, and the ITAE is 2.2 % less compared to the other alternatives.

4.2 Electricity grid using the IEEE 39-bus protocol (case II)

Case II aims to confirm the LDQNs in a two-area electrical network, which is an electrical network that utilizes the New Hampshire system or the IEEE 39-bus network (Fig. 11). Ten generators are located in these two regions. Case II's first area has two standard producers and two backup generators that consider energy storage. Two generic generators and four generators that account for energy storage comprise Area 2 of Case II. Software with version 7.1 from MATPOWER [32,33] provides more specific characteristics of the IEEE 39-bus system under Case II. All of the machines in Case II that take energy storage into account have 20 % of their power capability stored.Fig. 11 Case II presents the power system architecture depending on the IEEE 39-bus technology.

Fig. 11

In Case II, the highest exciter voltage produced by the command is 1.97 × 101, and the lowest instruction is −1.97 × 10⁴. Here are some contrasts between the parameter combinations used in Case II and Case I: kP1 = 11811.11, kI1 = 0.0011, kP2 = 359200.97 and kI2 = 0.0013 are the proportionate and integral factors of two regions, respectively.

In Case II, assessment parameters are shown in Table 6, and Fig. 12 shows the rotor orientation deviations derived by these comparative techniques. In a two-area electrical network that considers energy storage, the LDQNs show the best control performances compared to other computations after implementing model-free DQNs for controlling the exciter generator voltage and incorporating lazy learning into them (see Fig. 12 and Table 6). A lower AAE, IAE, ISE, and ITAE may not be obtained by Q(λ) learning in Case II compared to Q learning (Table 6). Table 6 shows that although Q-learning reinforcement teaching techniques outperform DQNs in management effectiveness, the suggested LDQNs outperform the Q-learning techniques in control effectiveness. In addition, the technique's reinforcement teaching sequence cannot consider both regions' control behaviors simultaneously in PI control (Table 6).Fig. 12 Differences in the rotor angle between Area 1 and the results generated by the competing technologies in the small-disturbance situation.

Fig. 12

The suggested LDQNs procedure's fault values in Area 1 are: Here are some comparisons between the techniques: the AAE is 4.7 % lower; the IAE is 4.8 % lower, the ISE is 3.9 % lower, and the ITAE is 5.0 % lower than the other methods shown. This is the breakdown of the error indices for Area 2 for the suggested LDQNs algorithm: AAE is at least 0.033 % lower, IAE is at least 0.016 % lower, ISE is at least 0.1 % lower, and ITAE is just 0.071 % larger than the other contrasted computation[34].

4.3 High-voltage 89-bus supply oriented upon electricity (case III) in Europe

In Case III, the LDQNs are tested using a three-area electrical system with mass-distributed stored energy that utilizes an electrical system derived from the European voltage-controlled 89-bus system (Fig. 13). The third case demonstrates a very intricate electrical structure with 210 connections and 12 engines. Region 3, Region 2, and the third region in Case III each have 3, 4, and 5 engines, respectively [35]. goes into more depth on the 89-transportation system, which operates at high voltage across Europe. In Case III, every producer has a 20 % rate of dispersed electricity storage.Fig. 13 Particular parameter settings provide the framework within the scope of the European high-voltage 89-bus system (Case III).

Fig. 13

In Case III, values of 7767.77 and −7326.91 are the maximum and lowest exciter output voltage instructions, accordingly. Case III differs from Case I in that it modifies the variable settings of the evaluated procedures. In these modifications, the proportional and integral coefficients for the three regions are revised as follows: kP1 = 491.48, kI1 = 0.0011, kP2 = 359200.97, kI2 = 0.0012, kP3 = 1271.49 and kI3 = 869.40.

Fig. 14 and Table 7 show the rotor direction variations and related assessment indicators derived by the algorithms in question under Case III accordingly. Notable findings from Table 7 and Fig. 14 include: By using the state buffer's documented RAS instruction capabilities, LDQNs provide the best real-time control performances with reduced rotor angle variations in a multiple-area power system that considers dispersed energy stores. Q acquiring knowledge, Q(λ) acquiring knowledge, R(λ) learning, SARSA, and SARSA(λ) achieved noticeably better control results in Case III compared to the ADP and PI algorithms. In Case III, learning using Q with a basic architecture performed better than Q(λ) learning, R(λ) learning, SARSA, and SARSA(λ).Fig. 14 Changes in rotor angle for Area 1 as measured by the comparing techniques in Case III's small-disturbance situation.

Fig. 14

Regarding Case III's Area 2, all of the algorithms show comparable control results. The suggested LDQNs are capable of taking into account the optimum conditions of the three domains—Integrated Time-weighted Pure Error (ITAE), the median Relative Error (AAE), and Integral Squared Error (ISE)—all at once. The contrasting techniques in Region 3 produced very similar management achievements, and the same is true for all of the compare methods in Area.

These results shed light on how LDQNs handle rotor angle variations in a multi-area power system, demonstrating how well it works in real-time situations and how it can simultaneously consider several performance indicators in different regions.

The suggested LDQNs algorithm's error indices in Area 1 are: When weighed against the other computer programs, the AAE is 40.0 % lower, the IAE is 40.3 % lower, the ISE is 26.1 % lower, and the ITAE is 39.1 % lower still. The AAE is at least 0.033 % lower. The IAE is at a minimum of 0.028 % lower. The ISE is at a minimum of 0.016 % lower. The ITAE is 0.018 % less than the remainder of the relative algorithm design. Here are the error indices for Area 3 on the recommended LDQNs technique: AAE is at least 24.7 % lower, IAE is at least 24.6 % lower, ISE is at least 4.6 % lower, and ITAE is 24.1 % lower than the other contrasted computation.

The fourth scenario is an electrical network that depends on Europe's high-voltage 1354-bus system. Case IV depicts an electric power system modeled on the European high voltage 1354-bus system that is bigger, complicated, and uses mass-distributed energy storage in four regulated zones. Case IV considers 260 turbines with 1991 branches and 20 % energy storage. For further information on the 1354-bus system, a high-voltage system in Europe, see (Xu et al., 2019).

Case IV specifies an optimal exciter voltage generated by an instruction 25.18 and an optimal instruction of −58.31. The following are some of how Case IV differs from the parameter settings of the algorithms being studied in Case I: kP1 = 517.46, kI1 = 0.0011, kP2 = 359200.97, kI2 = 0.0012, kP3 = 1271.49, kI3 = 869.40, kP4 = 1218.16, and kI4 = 117.94 are the proportionality and integrated values for the four regions in Fig. 15.Fig. 15 Differences in rotor angle for Area 1 as measured by the comparing techniques in Case IV's small-disturbance situation.

Fig. 15

The suggested LDQNs algorithm's erroneous indexes in Area 1 are: Here are some comparisons between the algorithms: the AAE is 8.9 % lower, the IAE is 8.8 % lower, the ISE is 17.6 % lower, and the ITAE is 10.1 % lower than the other methods involved. The following error indices are displayed for the suggested LDQNs method in Area 2: AAE is a minimum of 0.23 % lower; IAE is at least 0.032 % lower, ISE is at least 0.052 % lower, and ITAE is 0.029 % less than the other contrasted computation. Area 3's error indicators for the suggested LDQNs method are as follows: AAE is 2.7 % lower, IAE is 1.1 % lower, ISE is 0.47 percentage points lower, and ITAE is 0.14 percentage points lower than the other computational methods considered. Concerning Area 4, the following error indices were calculated for the suggested LDQNs technique: AAE = 7.1 % - IAE = 0.47 % - ISE = 0.29 % - ITAE = 0.02 % - compared to the other computations examined.

4.4 Discussions

The suggested technique successfully reduces rotor angle deviations using a reinforcement training structure with a Markov decision mechanism. Nevertheless, the action number is a crucial parameter that affects the control effectiveness of the unified Corrective Action Schemes (RAS) issue in LDQNs. Fig. 16 displays the mean Absolute Inaccuracies (AAEs) of angles of rotor variations derived from LDQNs with varying action counts within the context of Case I.Fig. 16 Concentrating on the median relative rotor angle discrepancies, the research assesses the LDQNs procedure's effectiveness with different amounts of actions in a small-disturbance situation involving a single-area energy system. It is important to remember that the recommended structure and techniques are novel compared to previous techniques, even if the assessment scores for Cases 1–4 do not show a revolutionary win. In this paper, we provide an algorithm that unifies the structure and uses just one engine in LDQNs, thereby replacing the traditional “Optimization + Control + Control” arrangement of three approaches for RAS.

Fig. 16

Fig. 16 shows that for action counts below 200, the average relative angle variation tends to decrease as action counts increase. However, the mean relative angle variation tends to expand exponentially when the number of operations goes over 200 and keeps going up. This is because training takes more time, and more activities are involved, leading to cumulative mistakes. For actions with more than 900 generations, the mean actual angle variation Report Word grows exponentially due to the enormous accumulation of mistakes. Therefore, it is clear that deep neural networks (DNNs) cannot be relied upon alone to reliably determine the power system's RAS in the long run.

Within the unified RAS, The structure of constant intervals is defined as ∈ [3 × field, 201 × fields]. Choosing a smaller constant k leads to less training time and poorer accuracy, while a bigger constant k indicates higher precision and more time spent training. This study's suggested LDQNs algorithm is an innovative approach to RAS management, a field devoid of a unified framework or competitive technique, thanks to its distinctive and encouraging capabilities. Nevertheless, there are a few caveats or shortcomings that the research does mention: (i) RAS instructions may be provided using a primary Deep Neural Networks (DNNs) mechanism in its execution time. (ii) Trained DNNs do not guarantee optimized performance with a modest mean squared error, such as 10^(−8)—particularly for complicated RAS control issues. (iii) If the DNNs are not perfectly precise for complicated RAS control issues, even well-trained ones could cause divergent rotor angle variations due to divergent RAS inputs. The stability and security of electrical systems might be jeopardized if the divergence happens too quickly. (iv) It is possible that better RAS control effectiveness will not materialize from the well-trained DNNs, even if their immediate actions seem precise. (v) Stresses that, even when well-trained DNNs attain tiny mean-square errors, optimizing for the long run is more important than optimizing for the short term concerning present control performance in Fig. 17(a and b).Fig. 17 (a, b). The research evaluates DNN performance, emphasizing output instructions and mean squared error.

Fig. 17

These arguments emphasize the need to recognize the limits and hazards of DNNs in actual power system programs and the problems and concerns associated with employing them for RAS control.

Training data for DNNs is obtained from PI + PSO runs outcomes; the input is expressed as v(θ), and the output is indicated as θ_fd. With 100 iterations and 70 % training, 15 % validation, and 15 % were testing, the initial training function uses the well-known Levenberg-Marquardt algorithm, which is noted for its excellent accuracy. Below are examples of diverging RAS instructions produced by well-trained DNNs and the median square error curve. At the same time, the research contrasts the outcomes of optimizing the Proportional-Integral (PI) controller's settings using ant colony optimization (ACO), a simulated annealing method (SAA), and particle swarm optimizing (PSO). According to the data, ACO and PSO provide very similar optimization outcomes, and typically, the most well-known optimization algorithms may get great PI controller settings after more than 100 generations of repetitions. However, System divergence occurs due to SAA's distinctive features and difficulty traversing the vast search space for PI values in Fig. 18(a, b, c). The study found that Particle Swarm Optimization (PSO) is the best method for optimizing the parameters of PI controllers since it successfully provides control efficiency indicators for the mimicked instances under consideration and has good generalization characteristics. The article presents new methods that use DNNs and highlights how well PSO works to optimize PI controller settings; these results are essential for improving power system management techniques.Fig. 18 (a, b, c). Several optimization techniques produce fit scores in a small-disturbance situation involving a single-area electric system.

Fig. 18

On top of that, the ACO takes 25,208 s to optimize, the SAA takes 375,891 s, and the PSO takes 22,133 s. While machine learning techniques take very little time to train, these optimization methods take an extremely lengthy period to optimize. The improvement times for these three computations are similar because the time needed for the internal computations is significantly lower than the time needed to compute the aim function's simulation results in every iteration.

There has yet to be a comparison of the emotional learning of Q throughout the four situations since it is currently not commonly used across different control systems. Fig. 19 shows the results of a small-disturbance situation used in this work to compare the control's efficacy resulting from Q learning plus PSO to that of Q education + PSO in a single-area electrical system. Management effectiveness is comparable to affective Q training plus PSO and Q learning + PSO (Fig. 19).Fig. 19 A single-area energy system with a small-disturbance case is used to determine the rotor angle variations using emotive learning via Q + PSO and Q education + PSO.

Fig. 19

5 Conclusion

The research on Lazy Long Q Networks (LDQNs) for a united rotor angle instability paradigm with a united time-scale in electrical networks containing mass dispersed energy storage provides essential insights into power system management and stability. LDQNs are a unique technique that effectively manages rotor angle variations throughout a range of power system typologies and disturbances by combining a Markov chain reaction with reinforcement education.

Integrating LDQNs into the unified rotor angle stability (RAS) methodology offers an integrated approach beyond conventional techniques to tackle the intricacies of electricity system dynamics. In contrast to the traditional trio of approaches, which consists of “Optimizing + Management + Command,” LDQNs provide a single, all-inclusive answer, opening the door to a more simplified and effective strategy for the stability of the power system. In this paper, LDQNs are rigorously evaluated against standard controllers, including Proportional-Integral (PI) controllers, methods of optimization for modifying controller parameters, and other teaching methods. The analysis covers a range of power system designs with mass dispersed storage of electricity, minor and significant problems, and one-area and multiple-area solutions. This thorough investigation demonstrates how LDQNs may perform better than current techniques, exhibiting better control effectiveness under real-time circumstances with fewer rotor angle variations. The capacity of LDQNs to adapt to various power system designs and disturbance circumstances is one of their main contributions. LDQNs consistently offer efficient control techniques when compared across many disturbances, including minor, significant, long-term, and hybrid disturbances. The method's strong performance is shown when applied to power systems of different complexity, such as those based on the IEEE 39-bus network and the European high-voltages 89-bus and 1354-bus infrastructure. Moreover, the research highlights the need for an ordinary time scale for power systems, which is 0.001 s in this instance, to provide uniformity while assessing various algorithms and controllers. Because of their competitive running and training timeframes, LDQNs are well-suited for real-time power system stability applications. LDQNs are efficient in delivering optimum control performances while retaining reasonable computing demands, as shown by the comparison with other reinforcement learning algorithms, such as Q learning, Q(λ) learning, R(λ) learning, SARSA, SARSA(λ), and Adaptive Dynamic Programming (ADP). Analyzing the effects of disturbances with short-, medium-, and long-term scales on LDQNs highlights the algorithm's robustness in challenging scenarios. With the use of evaluation indices such as integrated absolute error (IAE), integral of squared error (ISE), integrated time-weighted absolute error (ITAE), and average absolute error (AAE), one can gain a comprehensive understanding of the performance of LDQNs and see how well they can handle disturbances of varying durations. The paper also recognizes the drawbacks and possible dangers of using Deep Neural Networks (DNNs) or other methods to solve RAS control issues. The benefits of LDQNs in offering a more stable and dependable control approach are highlighted by their capacity to ensure correctness in complicated RAS situations and probable divergence difficulties, even when DNNs achieve minimal mean squared errors. Finally, Lazy Deep Q Networks show promise as a novel and creative approach to unify rotor angle stability in power systems using mass-distributed energy storage. As seen by several comparisons and assessments, the algorithm's flexibility, effectiveness, and better control performance place LDQNs in a valuable position to help power systems become more resilient and stable in the face of various configurations and disruptions. This work contributes to the continuous development of intelligent and reliable energy management systems by providing new opportunities to investigate and improve reinforcement learning approaches in power system control.

CRediT authorship contribution statement

Chengpeng He: Conceptualization, Data curation, Writing – original draft, Writing – review & editing. Xueying Wang: Conceptualization, Data curation, Writing – original draft, Writing – review & editing. Li Shu: Conceptualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
==== Refs
References

1 Kazemi M. Tabatabaei S.S. Moslemi N. A novel public-private partnership to increase the penetration of energy storage systems in distribution level J. Energy Storage 62 2023 10.1016/j.est.2023.106851
2 Mimica M. Perčić M. Vladimir N. Krajačić G. Cross-sectoral integration for increased penetration of renewable energy sources in the energy system – unlocking the flexibility potential of maritime transport electrification Smart Energy 8 2022 10.1016/j.segy.2022.100089
3 Tomin N. Shakirov V. Kozlov A. Sidorov D. Kurbatsky V. Rehtanz C. Lora E.E.S. Design and optimal energy management of community microgrids with flexible renewable energy sources Renew. Energy 183 2022 903 921 10.1016/j.renene.2021.11.024
4 Reznicek E. Braun R.J. Techno-economic and off-design analysis of stand-alone, distributed-scale reversible solid oxide cell energy storage systems Energy Convers. Manag. 175 2018 263 277 10.1016/j.enconman.2018.08.087
5 Graham J.D. Rupp J.A. Brungard E. Lithium in the green energy transition: the quest for both sustainability and security Sustainability 13 20 2021 10.3390/SU132011274
6 Karki S.B. Ramezanipour F. Pseudocapacitive energy storage and electrocatalytic hydrogen-evolution activity of defect-ordered perovskites SrxCa3−xGaMn2O8 (x = 0 and 1) ACS Appl. Energy Mater. 3 11 2020 10983 10992 10.1021/ACSAEM.0C01935
7 Dadak A. Mousavi S.A. Mehrpooya M. Kasaeian A. Techno-economic investigation and dual-objective optimization of a stand-alone combined configuration for the generation and storage of electricity and hydrogen applying hybrid renewable system Renew. Energy 201 2022 1 20 10.1016/j.renene.2022.10.085
8 Moreau V. Dos Reis P.C. Vuille F. Enough metals? Resource constraints to supply a fully renewable energy system Resources 8 1 2019 10.3390/RESOURCES8010029
9 Zhang L. Hu X. Wang Z. Sun F. Deng J. Dorrell D.G. Multiobjective optimal sizing of hybrid energy storage system for electric vehicles IEEE Trans. Veh. Technol. 67 2 2018 1027 1035 10.1109/TVT.2017.2762368
10 Ghezelbash A. Khaligh V. Liu J. Ryu J.-H. Scheduling of a multi-energy microgrid enhanced with hydrogen storage 2022 IEEE PES 14th Asia-Pacific Power and Energy Engineering Conference (APPEEC) 2022 1 6 10.1109/APPEEC53445.2022.10072138
11 Yu D. Zhang T. He G. Nojavan S. Jermsittiparsert K. Ghadimi N. Energy management of wind-PV-storage-grid based large electricity consumer using robust optimization technique J. Energy Storage 27 2020 10.1016/j.est.2019.101054
12 Debia S. Pineau P.O. Siddiqui A.S. Strategic use of storage: the impact of carbon policy, resource availability, and technology efficiency on a renewable-thermal power system Energy Econ. 80 2019 100 122 10.1016/j.eneco.2018.12.006
13 Zhang F. Zhao P. Niu M. Maddy J. The survey of key technologies in hydrogen energy storage Int. J. Hydrogen Energy 41 33 2016 14535 14552 10.1016/j.ijhydene.2016.05.293
14 Chauhan M. Bangwal A.S. Singh P. Electrochemical performance of A-site substituted SmSrNiO4−δ for energy storage applications Int. J. Hydrogen Energy 48 14 2023 5518 5528 10.1016/J.IJHYDENE.2022.11.136
15 Hamdan M.A. Rossides S.D. Haj Khalil R. Thermal energy storage using thermo-chemical heat pump Energy Convers. Manag. 65 2013 721 724 10.1016/j.enconman.2012.01.047
16 Ge L. Zhang B. Huang W. Li Y. Hou L. Xiao J. Mao Z. Li X. A review of hydrogen generation, storage, and applications in power system J. Energy Storage 75 2024 109307 10.1016/j.est.2023.109307
17 Atalmis G. Sattarkhanov K. Kaplan R.N. Demiralp M. Kaplan Y. The effect of powder and pellet forms of added metal hydride materials on reaction kinetics and storage Int. J. Hydrogen Energy 2024 10.1016/j.ijhydene.2023.12.262
18 Ciftci M.M. Lemaire X. Deciphering the impacts of ‘green’ energy transition on socio-environmental lithium conflicts: evidence from Argentina and Chile Extr. Ind. Soc. 16 2023 10.1016/j.exis.2023.101373
19 Di Profio P. Arca S. Rossi F. Filipponi M. Comparison of hydrogen hydrates with existing hydrogen storage technologies: energetic and economic evaluations Int. J. Hydrogen Energy 34 22 2009 9173 9180 10.1016/j.ijhydene.2009.09.056
20 Park G.L. Schäfer A.I. Richards B.S. Renewable energy-powered membrane technology: supercapacitors for buffering resource fluctuations in a wind-powered membrane system for brackish water desalination Renew. Energy 50 2013 126 135 10.1016/J.RENENE.2012.05.026
21 Poupin L. Humphries T.D. Paskevicius M. Buckley C.E. A thermal energy storage prototype using sodium magnesium hydride Sustain. Energy Fuels 3 4 2019 985 995 10.1039/C8SE00596F
22 Wang X. Luo Y. Qin B. Guo L. Power dynamic allocation strategy for urban rail hybrid energy storage system based on iterative learning control Energy 245 2022 10.1016/j.energy.2022.123263
23 Møller K. Sheppard D. Ravnsbæk D. Buckley C. Akiba E. Li H.-W. Jensen T. Complex metal hydrides for hydrogen, thermal and electrochemical energy storage Energies 10 10 2017 1645 10.3390/en10101645
24 Mohtadi R. Orimo S.I. The renaissance of hydrides as energy materials Nat. Rev. Mater. 2 3 2016 10.1038/NATREVMATS.2016.91
25 Grandell L. Lehtilä A. Kivinen M. Koljonen T. Kihlman S. Lauri L.S. Role of critical metals in the future markets of clean energy technologies Renew. Energy 95 2016 53 62 10.1016/j.renene.2016.03.102
26 Deane J.P. Ó Gallachóir B.P. McKeogh E.J. Techno-economic review of existing and new pumped hydro energy storage plant Renew. Sustain. Energy Rev. 14 4 2010 1293 1302 10.1016/j.rser.2009.11.015
27 Tschiggerl K. Sledz C. Topic M. Considering environmental impacts of energy storage technologies: a life cycle assessment of power-to-gas business models Energy 160 2018 1091 1100 10.1016/j.energy.2018.07.105
28 Ufa R.A. Malkova Y.Y. Gusev A.L. Ruban N.Y. Vasilev A.S. Algorithm for optimal pairing of res and hydrogen energy storage systems Int. J. Hydrogen Energy 46 68 2021 33659 33669 10.1016/J.IJHYDENE.2021.07.094
29 Song Y. Mu H. Li N. Wang H. Multi-objective optimization of large-scale grid-connected photovoltaic-hydrogen-natural gas integrated energy power station based on carbon emission priority Int. J. Hydrogen Energy 48 10 2023 4087 4103 10.1016/J.IJHYDENE.2022.10.121
30 Qiu Y. Li Q. Zhao S. Chen W. Planning optimization for islanded microgrid with electric-hydrogen hybrid energy storage system based on electricity cost and power supply reliability Renewable Energy Microgeneration Systems: Customer-Led Energy Transition to Make a Sustainable World 2020 49 67 10.1016/B978-0-12-821726-9.00003-5
31 Xu L. Hao J. Wang J. Yang Y. Zhao R. Zhang R. Yang X. N doping porous carbon embedded in self-assembled three-dimensional reduced graphene oxide networks for electrochemical hydrogen storage Int. J. Hydrogen Energy 50 2024 910 919 10.1016/j.ijhydene.2023.08.214
32 Lee J.Y. Ramasamy A.K. Ong K.H. Verayiah R. Mokhlis H. Marsadek M. Energy storage systems: a review of its progress and outlook, potential benefits, barriers and solutions within the Malaysian distribution network J. Energy Storage 72 2023 10.1016/j.est.2023.108360
33 Chiacchio F. Famoso F. D'Urso D. Cedola L. Performance and economic assessment of a grid-connected photovoltaic power plant with a storage system: a comparison between the North and the south of Italy Energies 12 12 2019 10.3390/EN12122356
34 Junne T. Wulff N. Breyer C. Naegler T. Critical materials in global low-carbon energy scenarios: the case for neodymium, dysprosium, lithium, and cobalt Energy 211 2020 10.1016/j.energy.2020.118532
35 Zayed M.E. Kabeel A.E. Shboul B. Ashraf W.M. Ghazy M. Irshad K. Rehman S. Zayed A.A.A. Performance augmentation and machine learning-based modeling of wavy corrugated solar air collector embedded with thermal energy storage: support vector machine combined with Monte Carlo simulation J. Energy Storage 74 2023 109533 10.1016/j.est.2023.109533
