
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)13185-4
10.1016/j.heliyon.2024.e37154
e37154
Research Article
Improving remote sensing scene classification using dung Beetle optimization with enhanced deep learning approach
Alamgeer Mohammad a
Al Mazroa Alanoud b
S. Alotaibi Saud c
Alanazi Meshari H. Meshari.alanazi@nbu.edu.sa
d⁎
Alonazi Mohammed e
S. Salama Ahmed f
a Department of Information Systems, Applied College at Mahayil, King Khalid University, Saudi Arabia
b Department of Information Systems, College of Computer and Information Sciences, Princess Nourah Bint Abdulrahman University (PNU), P.O. Box 84428, Riyadh, 11671, Saudi Arabia
c Department of Computer Science and Artificial Intelligence, College of Computing, Umm Al-Qura University, Saudi Arabia
d Department of Computer Science, College of Sciences, Northern Border University, Arar, Saudi Arabia
e Department of Information Systems, College of Computer Engineering and Sciences, Prince Sattam Bin Abdulaziz University, Al-Kharj, 16273, Saudi Arabia
f Department of Electrical Engineering, Faculty of Engineering & Technology, Future University in Egypt, New Cairo, 11845, Egypt
⁎ Corresponding author. Meshari.alanazi@nbu.edu.sa
30 8 2024
30 9 2024
30 8 2024
10 18 e371545 3 2024
14 8 2024
28 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Remote sensing (RS) scene classification has received significant consideration because of its extensive use by the RS community. Scene classification in satellite images has widespread uses in remote surveillance, environmental observation, remote scene analysis, urban planning, and earth observations. Because of the immense benefits of the land scene classification task, various approaches have been presented recently for automatically classifying land scenes from remote sensing images (RSIs). Several approaches dependent upon convolutional neural networks (CNNs) are presented for classifying brutal RS scenes; however, they could only partially capture the context from RSIs due to the problematic texture, cluttered context, tiny size of objects, and considerable differences in object scale. This article designs a Remote Sensing Scene Classification using Dung Beetle Optimization with Enhanced Deep Learning (RSSC-DBOEDL) approach. The purpose of the RSSC-DBOEDL technique is to categorize different varieties of scenes that exist in the RSI. In the presented RSSC-DBOEDL technique, the enhanced MobileNet model is primarily deployed as a feature extractor. The DBO method could be implemented in this study for hyperparameter tuning of the enhanced MobileNet model. The RSSC-DBOEDL technique uses a multi-head attention-based long short-term memory (MHA-LSTM) technique to classify the scenes in the RSI. The simulation evaluation of the RSSC-DBOEDL approach has been examined under the benchmark RSI datasets. The simulation results of the RSSC-DBOEDL approach exhibited a more excellent accuracy outcome of 98.75 % and 95.07 % under UC Merced and EuroSAT datasets with other existing methods regarding distinct measures.

Keywords

Remote sensing images
Scene classification
Deep learning
Dung beetle optimization
Transfer learning
==== Body
pmc1 Introduction

Remote sensing image (RSI) is a valuable information source for earth observation that can support us in evaluating and monitoring intricate structures on the surface of Earth [1]. According to the developments of earth observation technology, the quantity of RSIs has significantly increased. It provides specific emergency to the quest for how to attain maximized utilization of constantly expanding RSIs for intelligent earth observation [2]. Therefore, it is highly significant for understanding complex and large RSIs. While a main and challenging issue for efficiently analyzing RSI, scene classification of RSIs can be an active research field. RSI scene classification is to properly label specified RSIs with predetermined semantic classes [3]. Scene classification implies the distinctiveness of the various semantic features of RSIs, which means that the spatial distributions and various features are considered in the RSIs [4]. By comparison with image processing dependent upon the pixels and scene levels, object-level image classification reflects the various spatial distribution methods for the diverse objects from the higher spatial resolution (HSR) RSIs. Several applications of scene classification, particularly for land-use detection and urban development [5]. However, scene classification is quite a challenging task because of the intricate structural and spatial patterns of RSIs.

Developments in computer vision (CV) comprise determining exact features that can be effective in-memory applications and computational time [6]. Further, these methods are desired to have an excellent capacity for generalization while still achieving higher effective performance rates. RSI characterization is an important research domain. The feature-based approach has another stage from data mining techniques, which is to denote high levels of performance on RSI analysis [7]. Image classification has a wide-ranging application in the field of CV difficulty. Satellite images could be categorized, and the features in the images, namely landscape, building, desert, and buildings, are considered with time-based variations. In the subcategory of Artificial Intelligence (AI), machine learning (ML) has attained substantial achievements and recently launched in remote detection [8]. Supporting the application of deep convolutional features, the approaches reliant on deep learning (DL) have achieved higher effectiveness in image classification, and the accuracy is continuously being advanced with the development of novel techniques. Currently, DL-based algorithms play a crucial part in the extraction of higher-level features and developed as a leading model in CV and pattern recognition [9]. In the DL method, convolutional neural networks (CNNs) are defined as standard data-driven algorithms and robust tools that could be employed for identifying complex structures and extracting useful data RSIs besides the hierarchical convolutional features of hyperspectral image (HSI) [10].

This study designs a Remote Sensing Scene Classification using Dung Beetle Optimization with Enhanced Deep Learning (RSSC-DBOEDL) approach. The purpose of the RSSC-DBOEDL technique is to categorize different varieties of scenes that exist in the RSI. In the presented RSSC-DBOEDL technique, the enhanced MobileNet model is primarily utilized as a feature extractor. The DBO method will be implemented in this study for hyperparameter tuning of the enhanced MobileNet model. To classify the scenes in the RSI, the RSSC-DBOEDL technique uses a multi-head attention-based long short-term memory (MHA-LSTM) system. The experimental evaluation of the RSSC-DBOEDL approach could be examined under the benchmark RSI datasets.● The study contributes by utilizing an improved MobileNet method as a feature extractor. This improvement potentially enhances the quality and discriminative power of the extracted features, which is significant for scene categorization in RSI.

● The study also employs the DBO method for hyperparameter tuning of the improved MobileNet approach. This contribution ensures that the method is fine-tuned efficiently, optimizing its accomplishment for the specific task of scene categorization within RSI.

● The study also presents a novel classification model namely MHA-LSTM. This model implements both attention mechanisms and LSTM networks, potentially comprehending long-range reliabilities in the RSI scenes, thus improving the accuracy of the classification.

● The study contributes to the field by computing the proposed RSSC-DBOEDL technique on benchmark RSI datasets. This analysis provides valuable insights into the efficiency and generalizability of the proposed technique, establishing its accomplishment relative to present techniques in the literature.

2 Related works

Khan and Basalamah [11] developed a multi-branch DL technique that effectively incorporates global contextual features with multiscale features. This technique contains 2 phases. An initial phase mines global contextual data, and the next phase employs an FCN to extract multiscale local features. Ma et al. [12] introduced a method for scene classification network framework search dependent upon multiobjective neural development (SceneNet). In SceneNet, the network model searching and coding have been attained by utilizing an evolutionary method. Furthermore, the multiobjective optimizer algorithm is used, and the modest neural structure has been acquired in a Pareto solution set. In Ref. [13], a lightweight RSSC approach was developed like RSCNet. Primarily, the lightweight ShuffleNet-v2 network was implemented for extracting features. Secondarily, the weights of the backbone could be modified by employing transfer learning (TL). Finally, a normalization approach was also applied.

Chen et al. [14] presented an effectual RSSC architecture named as BiShuffleNeXt. At first, the sandglass bottleneck was employed for extraction. Second, the spatial path contains three sandglass bottlenecks, each with a step or two. In Ref. [15], a multi-grouped-CNN (MGCNN) for self-adapting, which can be the ability to support the effectiveness of CNN has been introduced, and the technique of combining many convolution layers proficient of being implemented hence, as a plug-in framework has been designed. In the meantime, a hyperparameter C in MGCNN was presented to extract features. Wang et al. [16] developed a triplet-metric-guided multiscale attention (TMGMA) was suggested. Primarily, the multiscale attention module (MAM) is conducted by multiscale feature maps. Secondarily, to take task-based salient features the triplet metric (TM) was employed. In Ref. [17], a deep nearest neighbor-NN dependent up attention mechanism (DN4AM) was developed to determine the few-shot scene classification method of RSIs. Scene classification-based attention maps have been utilized in this technique.

Xu et al. [18] designed a hierarchical features fusion of the CNN (HFFCNN) model. Firstly, an adaptive spatial-wise attention-related multiscale non-linear bag-of-visual-words (ASA-MNBoVW) framework was developed. Subsequently, a weighted image pyramid framework was selected. Lastly, a linear classifier was implemented. Dong, Lin, and Xie [19] present a novel technique. Firstly, the method builds a distortion magnitude space using diverse features and uses distortion adjustments on the support set samples utilizing the Optimal Search for Distortion Magnitude (ODS) approach. Furthermore, a Dual-Path Classification (DC) decision strategy is employed. In Ref. [20], a Data Customization-based Multiobjective Optimization Pruning (DCMOP) framework is proposed. The multiobjective evolutionary algorithms (MOEAs) and a Data Customization-based Proxy Mechanism (DCPM) method are also utilized. Ganashree et al. [21] propose an enhanced artificial bee colony optimization algorithm combined with a Convolutional Neural Network (IABC-CNN) model. Also, the Multiclass-Support Vector Machine (MSVM) method is utilized for classification.

3 The proposed model

In this paper, an automated scene classification called the RSSC-DBOEDL method on RSI is developed. The major aim of the RSSC-DBOEDL methodology is to categorize different varieties of scenes that exist in the RSI. In the introduced RSSC-DBOEDL model, 3 stages of functions will be included, such as enhanced MobileNet model-based feature extractor, DBO-based hyperparameter tuning, and MHA-LSTM-based classification. Fig. 1 shows the workflow of the RSSC-DBOEDL approach.Fig. 1 Workflow of RSSC-DBOEDL algorithm.

Fig. 1

3.1 Feature extraction: enhanced MobileNet model

The enhanced MobileNet architecture is used to develop the feature vectors. CNNs are generally employed in the field of DL for distinct CV tasks. CNNs are a class of DNNs that are planned especially for processing grid-like data, namely videos and images. It is simulated by the human visual system, which is recognized for its capability to recognize designs and features from visual data. CNNs contain several layers comprising convolution, pooling, and fully connected (FC) layers. The convolution layer is a major issue of CNNs and includes the application of convolution filters (kernels) for extracting local features in the input image. The pooling layer becomes smaller, and the spatial sizes of mapping features and support make the network further computationally effective. FC layer has been employed to generate forecasts dependent upon the extraction features. Examples of important CNN structures are GoogLeNet (Inception), AlexNet, ResNet, VGGNet, and EfficientNet. These designs vary in terms of complexity and depth. MobileNet is a certain structure planned for mobile and embedded devices with restricted computational resources. It can be established to enable real-time object classification and detection on smartphones and other lower-power devices.

MobileNet utilizes depthwise separable convolutions that are computationally effectual related to typical convolutional. Depthwise separable convolutions contain a depthwise convolutional (that filters channels separately) and then a pointwise convolutional (that integrates data across channels). This structure decreases the count of parameters and computations but preserves the optimum solution.

MobileNet is an effective DNN model designed mainly for embedded visual applications and mobile devices, and it is dependent upon depthwise convolutional, involving 28 convolutional layers [22]. MobileNet was adopted as a base model in this study, and improvements were made to it. Depthwise convolution drastically decreases the number of parameters and the computational difficulty compared to traditional convolution layers, thus accomplishing a shorter training time and more compact model architecture, which makes it well suited for utilization on resource-limited devices. Because of the limitations of MobileNet, such as restricted robustness to variations in input data and relatively low accuracy, it is necessary to establish an alternative solution while maintaining efficiency to overcome these limitations.

Several changes are made to its underlying structure to improve the robustness of MobileNet. Firstly, custom data pre-processing is applied based on Lab Space and CLAHE into the input images to enhance the detail performance and image contrast, so making the model recognize and learn features easier. The unique 28-layer MobileNet framework was adapted to a further efficient and compact 15-layer design, but all the layers used a ReLU activation function, depthwise separable convolution, and Batch Normalization. These modifications greatly enhance the computation speed, alleviate the overfitting risk, and considerably lower the model complexity. The classical CNN model has some shortcomings in generalization ability, resulting in overfitting when it comes to insufficient data. The study introduced a dropout layer to evade over-fitting by considering the relatively smaller number of keratosis images. In the training model, dropout can eliminate specific neurons in the hidden state with the given probability, which causes modifications from the network architecture. The whole dropout procedure corresponds to averaging the NN. To a certain degree, it reduces the mutual adaptation between neurons, thereby accomplishing regularization effects. It was found that dropout fixed at 0.25 was well-performed after several tests. The last output layer will be a softmax activation function and a full connection layer with two output units.

3.2 Hyperparameter tuning: DBO algorithm

The DBO method is employed in optimum choice of the hyperparameters compared to the enhanced MobileNet model. The DBO algorithm tunes the hyperparameters of the MobileNet model such as learning rate, number of epochs, and batch size. The DBO method is a metaheuristic algorithm inspired by the rolling, foraging, dancing, reproducing, and stealing behaviours of dung beetles (DBs) [23]. The DBO technique stimulates the activity of DBs to traverse the searching range and find a better solution.

3.2.1 Rolling ball dung Beetle

DB used to roll in 2 dissimilar manners: once there were difficulties in its path, and if there weren't. While the DB is moving forward and doesn't hit an obstacle, the DB should apply a celestial cue (sunrise direction) to keep the dung ball rolling from the straight line. The position alters with light intensity of DBs, and the position can be upgraded by using Eq. (1):(1) xit+1=xit+α×k×xit−1+b×|xit−xw|

Without obstacle mode. Where t specifies the existing iteration count, and xti shows the location of ith DBs in the population at tth permutation. k∈(0,0.2] indicates a constant of deflection coefficients, b denotes a constant that goes to zero and one, and α indicates the natural co-efficient within [−1, 1], where −1 means variation in the original direction and 1 implies no deviation. xw refers to the worse location in the existing populationt and |xit−xw| simulates the changes.

With obstacle mode. The DB should dance to relocate towards the dissimilar path once it obtains a block and is incapable of proceeding. The tangent function generates a novel rolling direction between [0andπ] by simulating the dancing behaviour of DBs. The dung ball is able to roll once the DB finds a new path. Consequently, the succeeding location determines the dancing behaviour of DBs:(2) xit+1=xit+tan(θ)|xit−xit−1|

If θ=0,π2 or π,tan(θ) is 0, hence the DB location remains the same.

3.2.2 Breeding dung Beetle

The dung balls remain rolled toward the place hidden by DBs and suitable for spawning. DBs should choose the best position to lay their eggs to provide a safe environment:(3) {Lb*=max{x*×(1−R),Lb}Ub*=min{x*×(1+R),Ub}

Where R=1−tTmax and Tmax refer to the maximal iteration counter. Lb and Ub define the low and up boundaries of the optimizer problems, and x* indicates the optimum location of the existing population.

If the female DB finds an appropriate region for egg positioning, then it can place its eggs in this region. During each iteration of DBO, the female DB lays a single brood ball. Furthermore, the boundary range of the spawning area changes dynamically, and the R‐value is the most important factor that defines these variabilities. Therefore, throughout the iterative process, the brood ball's location is dynamic as follows:(4) Bit+1=x*+b1×(Bit−Lb*)+b2×(Bit−Ub*)

Here Bit denotes the location information of ith brood balls at the tth iterations, b1 and b2 indicate the two independent arbitrary vectors of dimension 1 x D, and D shows the dimensionality of the optimizer problems.

3.2.3 Foraging dung Beetle

3.2.3.1 Lows

(5) xit+1=xl+S×g×(|xit−x*|+|xit−xl|)

In Eq. (5), xlt shows the location information of the ith stealing DBs at tth iterations, g denotes the random vector of size 1xD which follows the uniform distribution, and S represents a constant.Algorithm 1 Pseudocode of DBO algorithmInput: The maximal iteration Tmax, the population count N.	
Output: Optimum location Xb and its fitness value fb.	
Initialize the population i←1,2,3,…N and determine the pertinent parameter,	
while t≤Tmax do	
 if i=1,2,…,N do	
 if i== Ball-Rolling Dung Beetles then	
 δ= rand (1) if δ<0.9 then	
 Update Ball-Rolling Dung Beetles.	
 else	
 Update Ball-Rolling Dung Beetles.	
 end	
 end	
 The R-Value is evaluated using R=1−tTmax;	
 if i== Breeding Dung Beetles then	
 Update Breeding Dung Beetles.	
 end	
 if i== Foraging Dung Beetles then	
 Update Foraging Dung Beetles.	
 end	
 if i== Stealing Dung Beetles then	
 Update Stealing Dung Beetles.	
 end	
 end	
end	
return Xb and its fitness value fb.	

The DBO algorithm improves an FF to realize the best classifier solution. This expresses a positive integer to define the better solution of candidate efficiencies. The decreases in the classifier errors will be supposed to be FF, as signified in Eq. (6).fitness(xi)=ClassifierErrorRate(xi)

(6) =No.ofmisclassifiedinstancesTotalno.ofinstances*100

3.3 Scene classification: MHA-LSTM model

At this stage, the classification of the scenes occurs utilizing the MHA-LSTM approach. LSTM is a kind of RNN model that is designed to model and capture data sequences while handling gradient disappearing problems [24]. Particularly, LSTM is effective at modelling and locating long‐term dependency in data sequences. LSTM consists of multiple interconnected cells, each having its memory cells and its own set of gates. Fig. 2 illustrates the framework of MHA-LSTM. The important parts of LSTM cells are:Fig. 2 Architecture of MHA-LSTM

Fig. 2

Forget Gate (ft): Controls that several data from the prior cell layer (Ct−1) need to be kept or discarded. This is obtained by the prior cell layer and the existing input (xt) and generates an output of the forget gate.(7) ft=σ(Wf⋅[ht−1,xt]+bf)

Input Gate (it): Decide whether several data need to be included in the cell layer. It will take the prior cell state and the existing input and generate an output of the input gate.(8) it=σ(Wi⋅[ht−1,xt]+bi)

Candidate Cell State (C˜t): It is the cell state of the new candidate, calculated by the existing input and tanh activation function.(9) C˜t=tanh(Wc⋅[ht−1,xt]+bc)

Cell State Update (Ct): It is upgraded by merging the data taken in the newest candidate cell state (it⋅C˜t) and the prior cell state (ft⋅Ct−1).(10) Ct=ft⋅Ct−1+it⋅C˜t

Output Gate (ot): Decides which portion of the cell state needs to be outputted as the last prediction. This achieves the upgrade cell state and the present input and generates a resultant output gate.(11) ot=σ(Wo⋅[ht−1,xt]+bo)

Hidden State (ht): The hidden state represents the output of the LSTM cell, employed as a prediction and entered into the next time step. This can be evaluated through the output gate to the cell status.(12) ht=0t⋅tanh(Ct)

The MHA-LSTM model excels at capturing long‐term dependency in the data sequence. MHA expands the concept of the self‐attention module by using numerous attention heads simultaneously. The attention head considers dissimilar shares of the input series and allows the model to capture different kinds of data and dependencies in parallel. The Query (Q), Key (K), and Value (V) Projections are the basic elements of MHA. All the attention heads calculate attention scores among the keys (K) and the query (Q) of the input sequences and use the score to weigh the values (V). The attention score is calculated by the scaled dot product:(13) Attention(Q,K,V)=softmax(QKTdk)⋅V

Now, dk denotes the dimension of the key vector.

Concatenation and Linear Transformation: The model applied a linear conversion and concatenated them to attain the last MHA output after calculating the attention output for all the attention heads:(14) MultiHead(Q,K,V)=Concat(head1,head2,…,headhh)Wo

Let Concat concatenate the output from each head, and WO denotes the learned linear conversion.

4 Results and discussion

The simulation analysis of the RSSC-DBOEDL system can be tested utilizing UC Merced and EuroSAT datasets. The UCM database [25] holds 2100 samples with 21 classes. Next, the EuroSAT dataset comprises 1000 instances with ten classes as described in Table 1. Fig. 3 demonstrates the sample images.Table 1 Details on the UCM dataset.

Table 1UC-Merced Dataset	
Class	Labels	No. of Samples	
Agricultural	C-1	100	
Airplane	C-2	100	
Baseball Diamond	C-3	100	
Beach	C-4	100	
Buildings	C-5	100	
Chaparral	C-6	100	
Dense Residential	C-7	100	
Forest	C-8	100	
Freeway	C-9	100	
Golf Course	C-10	100	
Harbor	C-11	100	
Intersection	C-12	100	
Medium Residential	C-13	100	
Mobile Home Park	C-14	100	
Overpass	C-15	100	
Parking Lot	C-16	100	
River	C-17	100	
Runway	C-18	100	
Sparse Residential	C-19	100	
Storage Tanks	C-20	100	
Tennis Court	C-21	100	
Total Number of Samples	2100	

Fig. 3 Sample images.

Fig. 3

Fig. 4 illustrates the classifier analysis of the RSSC-DBOEDL system on the UCM database. Fig. 4a and b showcases the confusion matrices accomplished by the RSSC-DBOEDL technique at 70:30TRAPH/TESPH. These accomplished findings signified that the RSSC-DBOEDL technique has accurately identified and categorized each of the 21 classes. Furthermore, Fig. 4c examines the PR outcome of the RSSC-DBOEDL technique. The simulation result defined that the RSSC-DBOEDL technique attained improved PR performance in 21 classes. However, Fig. 4d exhibits the ROC result of the RSSC-DBOEDL method. The simulation values of the RSSC-DBOEDL system resulted in efficient experimental data with higher values of ROC at 21 class labels.Fig. 4 UCM dataset (a–b) Confusion matrices, (c) PR_curve, and (d) ROC.

Fig. 4

An extensive classification experimental analysis of the RSSC-DBOEDL model with the UCM database is in Table 2 and Fig. 5. These acquired findings denote that the RSSC-DBOEDL method properly recognizes 21 classes. According to 70%TRAPH, the RSSC-DBOEDL method attains average accuy, precn, recal, and Fscore of 98.74 %, 86.96 %, 86.78 %, and 86.79 %, respectively. At the same time, based on a 30%TESPH, the RSSC-DBOEDL system attains average accuy, precn, recal, and Fscore of 98.75 %, 87.28 %, 86.86 %, and 86.77 % correspondingly.Table 2 Classifier outcome of RSSC-DBOEDL technique on UCM database.

Table 2Class Labels	Accuy	Precn	Recal	Fscore	
TRAPH (70 %)	
Agricultural (C-1)	98.71	87.84	86.67	87.25	
Airplane (C-2)	99.12	92.06	87.88	89.92	
Baseball Diamond (C-3)	98.91	89.71	87.14	88.41	
Beach (C-4)	98.71	85.14	88.73	86.90	
Buildings (C-5)	98.64	86.36	83.82	85.07	
Chaparral (C-6)	98.98	89.23	87.88	88.55	
Dense Residential (C-7)	98.78	88.89	83.58	86.15	
Forest (C-8)	98.78	83.75	93.06	88.16	
Freeway (C-9)	98.37	75.71	88.33	81.54	
Golf Course (C-10)	98.98	90.41	89.19	89.80	
Harbor (C-11)	98.71	85.07	86.36	85.71	
Intersection (C-12)	98.71	90.16	80.88	85.27	
Medium Residential (C-13)	98.71	84.42	90.28	87.25	
Mobile Home Park (C-14)	98.37	85.29	80.56	82.86	
Overpass (C-15)	98.84	90.41	86.84	88.59	
Parking Lot (C-16)	98.78	88.24	85.71	86.96	
River (C-17)	98.16	78.31	87.84	82.80	
Runway (C-18)	98.98	91.18	87.32	89.21	
Sparse Residential (C-19)	98.57	83.82	85.07	84.44	
Storage Tanks (C-20)	99.05	91.18	88.57	89.86	
Tennis Court (C-21)	98.78	89.04	86.67	87.84	
Average	98.74	86.96	86.78	86.79	
TESPH (30 %)	
Agricultural (C-1)	98.89	82.14	92.00	86.79	
Airplane (C-2)	98.89	93.55	85.29	89.23	
Baseball Diamond (C-3)	98.73	89.29	83.33	86.21	
Beach (C-4)	99.37	96.30	89.66	92.86	
Buildings (C-5)	99.05	96.43	84.38	90.00	
Chaparral (C-6)	98.89	96.55	82.35	88.89	
Dense Residential (C-7)	98.57	83.33	90.91	86.96	
Forest (C-8)	99.21	89.66	92.86	91.23	
Freeway (C-9)	98.73	90.00	90.00	90.00	
Golf Course (C-10)	98.57	84.00	80.77	82.35	
Harbor (C-11)	98.89	96.55	82.35	88.89	
Intersection (C-12)	97.30	69.23	84.38	76.06	
Medium Residential (C-13)	98.89	92.00	82.14	86.79	
Mobile Home Park (C-14)	98.73	88.46	82.14	85.19	
Overpass (C-15)	98.89	81.48	91.67	86.27	
Parking Lot (C-16)	98.57	81.82	90.00	85.71	
River (C-17)	98.73	84.62	84.62	84.62	
Runway (C-18)	98.41	75.68	96.55	84.85	
Sparse Residential (C-19)	99.21	91.18	93.94	92.54	
Storage Tanks (C-20)	98.57	92.00	76.67	83.64	
Tennis Court (C-21)	98.57	78.57	88.00	83.02	
Average	98.75	87.28	86.86	86.77	

Fig. 5 Average of RSSC-DBOEDL model at UCM database.

Fig. 5

To calculate the effectiveness of the RSSC-DBOEDL technique under the UCM database, TRA and TES accuy curves can be described as presented in Fig. 6. The TRA and TES accuy curves shows the RSSC-DBOEDL algorithm through several epochs. It provides important details regarding the generalization capabilities and learning processes of the RSSC-DBOEDL method. At upgrading in epoch count, it will be perceived that the TRA and TES accuy curves acquire enhanced. It was noticed that the RSSC-DBOEDL system gains maximum TES accuy, which allows it to recognize the patterns in the data of TRA and TES.Fig. 6 Accuy curve of RSSC-DBOEDL technique on UCM dataset.

Fig. 6

Fig. 7 displays the overall TRA and TES loss values of the RSSC-DBOEDL algorithm on the UCM dataset at epochs. The TRA loss gets decreased through epochs. Primarily, it can be diminished as the model changes the weight to decrease the predictive error under the data of TRA and TES. The loss curves exhibit the level at which the model fits the TRA data. Gradually decreased the TRA and TES loss, and the RSSC-DBOEDL methodology excellently learns the patterns demonstrated in the data of TRA and TES. It will be noticed that the RSSC-DBOEDL technique modifications the parameters to reduce the dissimilarity between the actual and predicted TRA labels.Fig. 7 Loss curve of RSSC-DBOEDL technique on UCM dataset.

Fig. 7

The EuroSAT dataset [26] holds 1000 samples with ten classes. Subsequently, the EuroSAT database contains 100 samples with 10 class labels as mentioned in Table 3.Table 3 Details of the EuroSAT database.

Table 3EuroSAT Dataset	
Class	Labels	No. of Instances	
Annual Crop	C-1	100	
Forest	C-2	100	
Herbaceous	C-3	100	
Highway	C-4	100	
Industrial	C-5	100	
Pasture	C-6	100	
Permanent Crop	C-7	100	
Residential	C-8	100	
River	C-9	100	
Sea Lake	C-10	100	
Total No. of Instances	1000	

Fig. 8 displays the classifier outcome of the RSSC-DBOEDL system on the EuroSAT database. Fig. 8a and b denotes the confusion matrices accomplished by the RSSC-DBOEDL methodology at 70:30TRAPH/TESPH. The accomplished outcome represented that the RSSC-DBOEDL system has appropriately identified and categorized ten classes. Also, Fig. 8c illustrates the PR study of the RSSC-DBOEDL method. The outcome value implied that the RSSC-DBOEDL system has gained maximized PR value on ten classes. Then, Fig. 8d defines the ROC curve of the RSSC-DBOEDL technique. These outcomes display that the RSSC-DBOEDL technique provides excellent ROC outcomes at 10 class labels.Fig. 8 EuroSAT dataset (a–b) Confusion matrices, (c) PR_curve, and (d) ROC.

Fig. 8

Table 4 and Fig. 9 represent the overall classification experimental study of the RSSC-DBOEDL algorithm on the EuroSAT database. The results stated that the RSSC-DBOEDL methodology properly recognizes ten classes. With 70%TRAPH, the RSSC-DBOEDL technique attains average accuy, precn, recal, and Fscore of 93.80 %, 69.69 %, 69.21 %, and 68.88 %, respectively. At the same time, with a 30%TESPH, the RSSC-DBOEDL model reaches average accuy, precn, recal, and Fscore of 95.07 %, 75.14 %, 75.27 %, and 74.90 % correspondingly.Table 4 Classifier outcome of RSSC-DBOEDL technique on EuroSAT dataset.

Table 4Class Labels	Accuy	Precn	Recal	Fscore	
TRAPH (70 %)	
Annual Crop (C-1)	93.00	79.17	49.35	60.80	
Forest (C-2)	92.14	60.26	66.20	63.09	
Herbaceous (C-3)	94.43	74.29	71.23	72.73	
Highway (C-4)	93.14	66.13	60.29	63.08	
Industrial (C-5)	94.00	70.27	72.22	71.23	
Pasture (C-6)	92.71	61.54	69.57	65.31	
Permanent Crop (C-7)	94.14	65.85	80.60	72.48	
Residential (C-8)	93.86	63.24	70.49	66.67	
River (C-9)	96.00	85.25	73.24	78.79	
Sea Lake (C-10)	94.57	70.89	78.87	74.67	
Average	93.80	69.69	69.21	68.88	
TESPH (30 %)	
Annual Crop (C-1)	96.33	73.08	82.61	77.55	
Forest (C-2)	94.67	69.70	79.31	74.19	
Herbaceous (C-3)	94.67	76.19	59.26	66.67	
Highway (C-4)	93.00	72.00	56.25	63.16	
Industrial (C-5)	95.33	73.33	78.57	75.86	
Pasture (C-6)	94.00	72.41	67.74	70.00	
Permanent Crop (C-7)	95.00	76.47	78.79	77.61	
Residential (C-8)	95.33	80.49	84.62	82.50	
River (C-9)	96.67	82.76	82.76	82.76	
Sea Lake (C-10)	95.67	75.00	82.76	78.69	
Average	95.07	75.14	75.27	74.90	

Fig. 9 Average of RSSC-DBOEDL technique on EuroSAT dataset.

Fig. 9

To determine the performance of the RSSC-DBOEDL system on the EuroSAT dataset, TRA and TES accuy curves have been described as presented in Fig. 10. The TRA and TES accuy curves denote the effectiveness of the RSSC-DBOEDL method over several epochs. This figure provides significant details for the generalization capabilities and learning process of the RSSC-DBOEDL technique. With an increasing epoch, it will be perceived that the gains in the TRA and TES accuy curves have increased. It could be observed that the RSSC-DBOEDL method gets enhanced TES accuy and is proficient in identifying the patterns at the data TRA and TES.Fig. 10 Accuy curve of RSSC-DBOEDL technique on EuroSAT dataset.

Fig. 10

Fig. 11 shows the wide-ranging TRA and TES loss values of the RSSC-DBOEDL algorithm with the EuroSAT dataset over epochs. The TRA loss shows a decrease over epochs. Mostly, the loss values acquired diminished as the model adjusted the weight for the reduction of the estimated error under the data TRA and TES. The loss curves exhibit the range where the model will be fitted with the TRA data. The TRA and TES loss is gradually decreased and defined that the RSSC-DBOEDL methodology efficaciously learns the patterns demonstrated in the TRA and TES data. It could also be demonstrated that the RSSC-DBOEDL technique alters the parameters to lessen the distinction between the original and predictive TRA label.Fig. 11 Loss curve of RSSC-DBOEDL technique on EuroSAT dataset.

Fig. 11

Table 5 signifies a comparative result of the RSSC-DBOEDL method with current models on two datasets [11]. Fig. 12 illustrates a comparison study of the RSSC-DBOEDL system on the UCM database. The experimental values specified that the SMS-EMOA technique gains poor performance, whereas the EfficientNet model reaches slightly improved results. Along with that, the MobileNet, ResNet-50, ResNet-101, ShuffleNet, GoogleNet, DenseNet, and VGG-16 techniques obtain moderate performance. Nevertheless, the RSSC-DBOEDL technique has outperformed the other existing methods with a higher accuy of 98.75 %.Table 5 Accuy outcome of RSSC-DBOEDL approach with recent methods on two datasets.

Table 5Accuracy (%)	
Method	UC Merced Dataset	EuroSAT Dataset	
EfficientNet	88.00	85.23	
MobileNet	93.00	87.52	
ResNet-50	95.00	90.34	
ResNet-101	95.00	90.51	
ShuffleNet	91.00	88.68	
SMS-EMOA	78.45	73.02	
GoogleNet	94.00	88.51	
DenseNet, VGG-16	96.00	91.68	
RSSC-DBOEDL	98.75	95.07	

Fig. 12 Accuy outcome of RSSC-DBOEDL approach on UCM dataset.

Fig. 12

Fig. 13 demonstrates a comparison analysis of the RSSC-DBOEDL model on the EuroSAT datasets. The observational data specified that the SMS-EMOA algorithm has reduced performance, whereas the EfficientNet technique obtains moderately better outcomes. Besides that, the MobileNet, ResNet-50, ResNet-101, ShuffleNet, GoogleNet, DenseNet, and VGG-16 methods gain modest performance. However, the RSSC-DBOEDL approach has exceeded the other present systems with a higher accuy of 95.07 %.Fig. 13 Accuy outcome of RSSC-DBOEDL approach on EuroSAT dataset.

Fig. 13

In Fig. 14, the comparative cumulative distribution function (CDF) results of the RSSC-DBOEDL model with existing models are depicted. The CDF offers a detailed depiction of the probability distribution of a random variable. By inspecting the CDF, the probability that the variable takes on a value less than or equal to any particular value can be interpreted. This aid to visualize and understand the behavior of the data. The figure shown that the RSSC-DBOEDL model reaches better performance over other models with enhanced CDF values. These results show a better outcome of the RSSC-DBOEDL methodology in the scene classification process.Fig. 14 CDF Results of proposed RSSC-DBOEDL model with existing models.

Fig. 14

5 Conclusion

In this article, an innovative automatic scene classification method named RSSC-DBOEDL method on RSI is developed. The main objective of the RSSC-DBOEDL system is to categorize different varieties of scenes that exist in the RSI. In the presented RSSC-DBOEDL method, three functions, namely enhanced MobileNet model-based feature extractor, DBO-based hyperparameter tuning, and MHA-LSTM-based classification, are employed. The DBO algorithm could be implemented to optimize hyperparameter tuning of the enhanced MobileNet architecture. To classify the scenes in the RSI, the RSSC-DBOEDL technique used the MHA-LSTM model. The experimental evaluation of the RSSC-DBOEDL approach can be examined under the benchmark RSI database. The experimentation outcomes highlight the excellent performance of the RSSC-DBOEDL system compared to other methods with respect to distinct measures.

Data availability statement

Data sharing is not applicable to this article as no datasets were generated during the current study.

CRediT authorship contribution statement

Mohammad Alamgeer: Writing – original draft, Methodology, Conceptualization. Alanoud Al Mazroa: Formal analysis, Data curation. Saud S. Alotaibi: Methodology, Investigation. Meshari H. Alanazi: Writing – review & editing, Resources, Project administration, Funding acquisition. Mohammed Alonazi: Visualization, Validation, Supervision, Software. Ahmed S. Salama: Writing – review & editing, Visualization, Validation, Supervision, Software.

Declaration of competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Acknowledgments

The authors extend their appreciation to the Deanship of Research and Graduate Studies at King Khalid University for funding this work through Large Research Project under grant number RGP2/158/45 . Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2024R719 ), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors extend their appreciation to the Deanship of Scientific Research at Northern Border University, Arar, KSA for funding this research work through the project number “NBU-FFR-2024-1180-07 . This study is partially funded by the Future University in Egypt (FUE).
==== Refs
References

1 Wang C. Tang X. Li L. Tian B. Zhou Y. Shi J. IDN: inner-class dense neighbours for semi-supervised learning-based remote sensing scene classification Remote Sensing Letters 14 1 2023 80 90
2 Zhao Y. Liang J. Huang S. Huang P. Hierarchical deep features progressive aggregation for remote sensing images scene classification IEEE J. Sel. Top. Appl. Earth Obs. Rem. Sens. 2024
3 Akila G. Gayathri R. Weighted multi-deep feature extraction for hybrid deep convolutional LSTM-based remote sensing image scene classification model Geocarto Int. 37 27 2022 18217 18253
4 Wang J. Li W. Zhang M. Tao R. Chanussot J. Remote sensing scene classification via multi-stage self-guided separation network IEEE Trans. Geosci. Rem. Sens. 2023
5 Jin J. Zhou W. Ye L. Lei J. Yu L. Qian X. Luo T. DASFNet: dense-Attention–Similarity-Fusion Network for scene classification of dual-modal remote-sensing images Int. J. Appl. Earth Obs. Geoinf. 115 2022 103087
6 Zhang N. Wang G. Wang J. Chen H. Liu W. Chen L. All adder neural networks for on-board remote sensing scene classification IEEE Trans. Geosci. Rem. Sens. 2023
7 Chen W. Li X. Wang L. Mine remote sensing scene classification using deep learning Remote Sensing Intelligent Interpretation for Mine Geological Environment: from Land Use and Land Cover Perspective 2022 Springer Nature Singapore Singapore 165 176
8 Zhao Y. Liang J. Huang S. Huang P. Hierarchical deep features progressive aggregation for remote sensing images scene classification IEEE J. Sel. Top. Appl. Earth Obs. Rem. Sens. 2024
9 Zheng F. Lin S. Zhou W. Huang H. A lightweight dual-branch swin transformer for remote sensing scene classification Rem. Sens. 15 11 2023 2865
10 Xu Q. Ouyang C. Jiang T. Yuan X. Fan X. Cheng D. MFFENet and ADANet: a robust deep transfer learning method and its application in high precision and fast cross-scene recognition of earthquake-induced landslides Landslides 19 7 2022 1617 1647
11 Khan S.D. Basalamah S. Multi-branch deep learning framework for land scene classification in satellite imagery Rem. Sens. 15 13 2023 3408
12 Ma A. Wan Y. Zhong Y. Wang J. Zhang L. SceneNet: remote sensing scene classification deep learning network using multiobjective neural evolution architecture search ISPRS J. Photogrammetry Remote Sens. 172 2021 171 188
13 Chen Z. Yang J. Feng Z. Chen L. RSCNet: an efficient remote sensing scene classification model based on lightweight convolution neural networks Electronics 11 22 2022 3727
14 Chen Z. Yang J. Feng Z. Chen L. Li L. BiShuffleNeXt: a lightweight bi-path network for remote sensing scene classification Measurement 209 2023 112537
15 Wu X. Zhang Z. Zhang W. Yi Y. Zhang C. Xu Q. A convolutional neural network based on grouping structure for scene classification Rem. Sens. 13 13 2021 2457
16 Wang H. Gao K. Min L. Mao Y. Zhang X. Wang J. Hu Z. Liu Y. Triplet-metric-guided multiscale attention for remote sensing image scene classification with a convolutional neural network Rem. Sens. 14 12 2022 2794
17 Chen Y. Li Y. Mao H. Chai X. Jiao L. A novel deep nearest neighbor neural network for few-shot remote sensing image scene classification Rem. Sens. 15 3 2023 666
18 Xu K. Deng P. Huang H. Mining hierarchical information of CNNs for scene classification of VHR remote sensing images IEEE Transactions on Big Data 9 2 2022 542 554
19 Dong Z. Lin B. Xie F. Optimizing few-shot remote sensing scene classification based on an improved data augmentation approach Rem. Sens. 16 3 2024 525
20 Hu Z. Gong M. Lu Y. Li J. Zhao Y. Zhang M. Data customization-based multiobjective optimization pruning framework for remote sensing scene classification IEEE Trans. Geosci. Rem. Sens. 2023
21 Ganashree K.C.G. Hemavathy R. Anala M.R. Land scene classification from remote sensing images using improved artificial bee colony optimization algorithm Int. J. Electr. Comput. Eng. 14 1 2024 347 357
22 Li S. Li C. Liu Q. Pei Y. Wang L. Shen Z. An actinic keratosis auxiliary diagnosis method based on an enhanced MobileNet model Bioengineering 10 6 2023 732 37370662
23 Zhao K. Guo D. Sun M. Zhao C. Shuai H. Short-term traffic flow prediction based on VMD and IDBO-LSTM IEEE Access 2023
24 Quansah P.K. Short-term load forecasting using A particle-swarm optimized multi-head attention-augmented CNN-LSTM network arXiv preprint arXiv:2309.03694 2023
25 http://weegee.vision.ucmerced.edu/datasets/landuse.html.
26 https://www.kaggle.com/datasets/apollo2506/eurosat-dataset.
