
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39294200
71726
10.1038/s41598-024-71726-3
Article
Effective clinical decision support implementation using a multi filter and wrapper optimisation model for Internet of Things based healthcare data
Robert Vincent Asir Chandra Shinoo racshinoo2022@gmail.com

Sengan Sudhakar sudhasengan@gmail.com

https://ror.org/02h9pt147 0000 0004 0422 9275 Department of Computer Science and Engineering, PSN College of Engineering and Technology, Tirunelveli, Tamil Nadu 627451 India
18 9 2024
18 9 2024
2024
14 218204 6 2024
30 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Feature Selection (FS) is essential in the Internet of Things (IoT)-based Clinical Decision Support Systems (CDSS) to improve the accuracy and efficiency of the system. With the increasing number of sensors and devices used in healthcare, the volume of data generated is vast and complex. Relevant FS from this data is crucial in reducing computational overhead, improving the system’s interpretability, and enhancing the Decision-Making System (DMS) quality. FS also aids in addressing the problems of data redundancy and noise, which can negatively impact the system’s performance. FS is critical to developing practical and dependable CDSS in IoT-based healthcare sectors. This research proposes a two-phase FS model. Phase-I employs an ensemble of five Filter Methods (FM), followed by a Pearson Correlation Method (PCM). Phase-II uses the Binary Optimized Genetic Grey Wolf Optimization Algorithm (BOGGWOA) as a Wrapper Method (WM). This recommended model integrates the most valuable features of each filter. Then, it uses the Pearson Correlation Coefficient (PCC) to get rid of features that aren’t needed, a Support Vector Machine (SVM) to guess how accurate their classification will be, and BOGGWOA as the Wrapper Method (WM) to pick the most essential features with the best CA.

Keywords

Internet of Things
Clinical decision support systems
Filter methods
Accuracy
Deep learning
Feature selection
Subject terms

Health care
Energy science and technology
issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

The Internet of Things (IoT) is crucial in healthcare, enabling remote patient monitoring and predicting health issues1. Healthcare applications facilitate information exchange among professionals, patients, and organizations, improving cost efficiency and service quality. IoT supports big data analysis in healthcare systems using wearable devices and sensors2. For instance, smartwatches can detect healthcare anomalies like irregular heartbeats in older individuals, triggering alarms when detected3. Medical applications exploit Bluetooth, Radio Frequency Identification (RFID), and Body Area Networks (BAN). Clothing sensors measure vital signs like Electrocardiography (ECG), electromyography (EMG), electroencephalography (EEG), and blood pressure. Smartwatches gather position, movement, body temperature, and oxygen saturation data. This information is transferred to cloud storage, accessible to doctors, hospitals, and relatives. The Internet of Medical Things (IoMT) and Healthcare Management Systems (HMS) benefit from the popularity of wearable devices4. Constructing an efficient Clinical Decision Support System (CDSS) is challenging due to heterogeneous data from multiple sensors.

Despite these advancements, the construction of an effective CDSS faces significant challenges due to the heterogeneous and high-dimensional nature of data collected from multiple sensors. Such data often contains redundant and irrelevant attributes that not only fail to contribute but can also degrade the performance of CDSS. Pre-processing stages, including noise removal and resolution of data inconsistencies, are essential yet remain computationally intensive. The necessity for dimensionality reduction to enhance system efficiency and manage vast volumes of data is more pressing than ever.

Further dimensionality reduction is crucial in developing an effective CDSS due to challenges like missing data, outliers, and inconsistent attributes that may change over time. High-dimensional clinical datasets often contain redundant and irrelevant attributes that do not contribute to the CDSS design and can even degrade performance5. Dealing with medical data’s enormous volume and heterogeneity requires data pre-processing before applying data mining techniques6. This pre-processing stage involves noise removal, resolving inconsistencies, and addressing compatibility issues to improve the computation speed. Dimensionality reduction is commonly employed to enhance efficiency, reduce computational costs, and extract relevant information from medical data7. This study commonly employs dimensionality reduction to enhance efficiency, reduce computational costs, and extract relevant information from medical data7.

Two standard methods for reducing the dimensionality of medical data are record sampling and Feature Selection (FS)8. Record sampling randomly selects a subset of records for data mining, whereas FS determines relevant and non-redundant features. It is standard procedure to choose FS when the sum of features is significantly higher than the quantity of data. Additionally, Feature Extraction (FE) can be employed when combining existing features to produce meaningful new features, typically in Image Processing (IM) and Natural Language Processing (NLP) domains. By obtaining new features and reducing the previously available ones, FE attempts to decrease the aggregate volume of features. FS methods ensure that the dataset size is minimized while retaining sufficient data for the learning task. Many researchers have invented several FS methods to manage clinical datasets, with ongoing research aiming to enhance clinical training using CDSS9. FS methods play a crucial role in handling high-dimensional datasets. It is classified into distinct types, including Wrapper Method (WM), Filter Method (FM), Hybrid Method (HM), and Embedded Method (EM)10. WM, such as Support Vector Regression-Recursive Feature Elimination (SVM-RFE), utilizes prediction accuracy for FS and has shown superior performance compared to other approaches11.

On the other hand, FM exhibits good generalization ability and is well-suited for large-scale datasets due to its low computational complexity12. To maximize the benefits of both FM and WM, HM combines their strengths13. EM incorporates FS into the training process of learning algorithms, such as Artificial Neural Network (ANN) and Decision Tree (DT), to achieve optimal classification performance with minimal features14,15.

While FM, including Information Gain (IG), Gain Ratio (GR), Chi-square (CS), and Relief-F (RF), are recommended for their straightforward ranking procedures in high-dimensional data16, they may encounter challenges when dealing with an excessively large or small number of FS. They diminish Classification Accuracy (CA), such as SVM, by assessing features individually without dealing with their associations or correlations17. An ensemble FS technique using multiple filters has been proposed to focus on this issue. The ensemble method employs Correlation-based Feature Selection (CFS) to generate a preliminary feature subset18. Subsequently, IG, CS, RF, and Symmetrical Uncertainty (SU) filter rankings are employed to evaluate the best FS. The ensemble FS method removes irrelevant and redundant features while emphasizing the most important ones, improving CA through dimensionality reduction19. Moreover, partitioning and evaluating ranked features based on SU and correlation enable the identification of significant features, further enhancing the scalability and efficiency of performance11. This combination of multiple filter algorithms improves CA and addresses the challenges posed by redundancy, highly correlated features, and weak independent features that are strong when considered in groups20.

Current methods for dimensionality reduction, such as record sampling and FS, are standard, yet they encounter specific challenges. Traditional FS methods, including IG, GR, CS, and RF, often struggle with large or small feature sets and typically do not account for feature interdependencies, which can lead to diminished CA. Although advancements like the ensemble FS technique using CFS have improved CA by removing irrelevant and redundant features, these methods still fall short in terms of scalability and performance consistency across diverse datasets.

Different optimization algorithms have been developed for FS in ML. One such algorithm is the Weighted Model (WM), which considers the impact of FS on the algorithm and offers better classification performance despite being computationally expensive. Genetic Algorithm (GA) is another optimization technique that optimizes filter models and improves CA21. A Spatial Cooperative Search Method (SCSM) employing the Firefly Algorithm (FA) and rough set theory for the subset of FS has been invented by22. Another approach is Particle Swarm Optimization for Feature Selection (PSO-FS), which effectively handles pre-processing tasks using rough sets23,24. 25 integrated rough set theory with the Fish School Search Algorithm (FSA) to propose an efficient FS method. Additionally, heuristic algorithms with accelerators have been shown to improve search techniques and FE26. Fitness measures like k-NN classification error can also enhance the performance of classifiers27.

GA-based FS methods have demonstrated improved CA on diverse datasets28. 29 combined GA with SVM to achieve better CA. 30 applied PSO-FS with the Local Search (LS) approach and multiple FM and WM, resulting in higher accuracy. Besides GA and PSO, other optimization techniques are being proposed for FS. 31 introduced Chicken Swarm Optimization (CSO), which outperforms GA and PSO on standard datasets. 32 proposed a method that combines Deep Neural Network (DNN) with Elephant Search Optimization (ESO) for analyzing microarray data. 33 presented an FS approach using the Bat Algorithm (BA) and Optimum-Path Forest. Grey Wolf Optimizer (GWO), inspired by the predation behavior of grey wolves, has also gained attention as a Swarm Intelligence Technique (SIT)5.

GeFes is a brand-new wrapper-based FS that was proposed by25. The mutation and crossover operators’ performance in GeFes was improved by adopting a new operator in GA. The GeFes employs an integrated nested cross-validation procedure and a KNN classifier. Out of 135 FS, they were 99.02% accurate. The Hybrid Improved Dragonfly Algorithm (HIDA) is a suggested FS method by25, which integrates the Minimum Redundancy Maximum Relevance (mRMR) method’s merits and the Improved Dragonfly Algorithm (IDA). They first eliminated irrelevant features using mRMR before introducing local quantum optimums, dynamic swarming factors, and global optimums to increase IDA’s exploitation potential. For the arrhythmia dataset, they chose 169.4 features on average and had an CA of 74.77%. A II-stage FS approach was introduced by34. For this purpose, the GA was used to select the optimal feature set in the first stage, which involved selecting 72 features. For classification and evaluation purposes, they finally constructed a bootstrap aggregating (Bagging) based SVM Ensemble. With 92 features, they achieved 88.72% CA. SMOTE oversampling was used by35 for balancing the dataset and then came to the K-part Lasso to remove any remaining redundant features. Recursive feature elimination and RF classifiers were combined in the last phase to create the FS method known as RF-RFE. With a total of 89 features, this method produced an accuracy of 98.68%.

Building upon the theory of optimization algorithms in FS, this study presents an II-phase multi-filter and wrapper combined FS model. The model aims to enhance the predictive performance of Machine Learning (ML) models by leveraging the strengths of different FS techniques. In the first phase, the framework utilizes a combination of five Filter Methods (FM): Information Gain (IG), Random Forest Feature (RFF), Mutual Information (MI), Correlation-based Feature (CF), and Symmetrical Uncertainty (SU). These filters collaborate to select the best feature subset, and the accuracy of each feature is computed using the SVM algorithm. The framework incorporates a Weighted Model (WM) in the second phase to optimize the FS process further. Specifically, the Binary Optimized Genetic Grey Wolf Optimization Algorithm (BOGGWOA), a meta-heuristic approach, is employed as the WM to obtain the ideal feature subset. The WM operates after the initial five filter methods in the first phase. This two-phase approach ensures that only the most relevant and non-redundant FS is used, resulting in the highest CA for the predictive model.

The primary contributions to this work are:Introduction of a Hybrid Feature Selection Model This study proposes an II-phase FS approach that integrates multiple filter methods with a meta-heuristic wrapper—Binary Optimized Genetic Grey Wolf Optimization Algorithm (BOGGWOA). This research built this approach for large-scale, heterogeneous healthcare datasets typical of IoT environments.

Advanced Optimization Techniques The use of BOGGWOA, a novel optimization algorithm in the context of FS, optimizes the selection process to identify the most predictive features with minimal redundancy, significantly enhancing the overall system performance.

Comprehensive Validation The effectiveness of the proposed model is demonstrated through extensive testing across multiple healthcare datasets, showing superior performance in terms of CA and computational efficiency compared to existing methods.

The numerous sections of the research paper have been organized in the following manner: The resources and approaches that have been employed in this study will be discussed in “Materials and methods” section, the proposed approach has been provided in “Proposed method” section, the collection of data that was used and the specifics of implementation are shown in “Dataset and implementation” section, the results and analysis are stated in “Experiment analysis” section, and the conclusion and future work is disclosed in “Conclusion and future work” section.

Materials and methods

Correlation-based feature selection (CFS)

Due to the requirement that it is an FM, the CFS approach is not based on the selected classification approach; instead, it primarily deals with assessing feature subsets on an assumption of core features of the data. It explores feature subsets with a minimal feature-feature correlation to avoid redundant and high feature-class correlation and enhance predictability.

The algorithm calculates the merit of a feature-rich subset using Eq. (1):1 Merits=krcf¯k+k(k-1)rff¯.

rff¯: Average feature-feature correlation, rcf¯: Average feature-class correlation, k: Number of features of that subset.

Since the features in this study are continuous, the Feature Correlation (FC) is achieved using the standard Pearson Correlation Coefficient (PCC).

A best-first search methodology intended by36 is employed to search for the subset with the highest merit. An infinite subset of features is searched for, and each feature’s value is determined. In this initial evaluation phase, the feature-class correlation (FCC) is considered, and the merit Eq. (2) is simplified:2 krcf¯k+k(k-1)rff¯=1rcf¯1+1(1-1)rff¯=rcf¯1=rcf¯

The evaluation in this phase focuses solely on the Feature-Class Correlation (FCC). The algorithm expands the currently empty subset by adding the feature with the highest FCC. It subsequently selects the best subset feature with the feature included after reassessing every other feature33. This process continues, excluding the features already added until the algorithm cannot find any improvement in the FE. This work set a limit to avoid redundant backtracking. This technique generates the best feature subset within the specified limit.

Gain ratio (GR)

GR is an evaluation metric used in DT algorithms for FS. It extends Information Gain (IR) to handle features with many distinct values. The GR is calculated using the Eq. (3):3 GR(A)=(Info(D)-InfoA(D))SplitInfoA(D)

where Info(D) is the information content of the target variable, InfoA(D) is the information content of the target variable given feature A and SplitInfoA(D) is the intrinsic information of feature A. Info(D) and InfoA(D) are calculated using the same equation and SplitInfoA(D) is based on the probabilities of feature A. A higher GR indicates a more informative feature for predicting the target variable. Thus, features with higher GR values are preferred for DT models37.

Information gain (IG)

IG is a feature evaluation metric used in DT algorithms. An Eq. (4) feature indicates the objective variable’s entropy or unpredictability loss:4 Gain(A)=Info(D)-InfoA(D)

Gain (A) thus informs us of the gain resulting from branching on A. It shows the predicted decrease in the information needed to be brought on by understanding the value of feature A. Higher IG values indicate that the feature provides more helpful information for predicting the target variable.

Symmetrical uncertainty (SU)

Filter-based FS techniques were frequently employed, including MI, Pearson correlation, chi-squared test, IG, GR, and relief. The Fast Correlation-Based Filter (FCBF) approach was proposed by38 to eliminate redundant and irrelevant features. The redundancy was defined as Eqs. (5 and 6) in the SU measurement:5 IG(X∣Y)=E(X)-E(X∣Y)

6 SU(X,Y)=2×IG(X∣Y)E(X)+E(Y)

where IG(X|Y) is the IG of X after identifying Y, and E(X) and E(Y) are the information of features X and Y, respectively. To determine the correlation between features, C-correlation and F-correlation are defined based on SU.C-correlation the relationship SUi,c between any feature Fi and class C.

F-correlation represents the SU between any two features, Fi and Fj(i≠j) is indicated by the symbol SUi,j.

Mutual information (MI)

MI analyses two parameters relying on ML feature analysis. In FS, MI measures feature data related to the objective variable. The technique involves entropy, which evaluates the unpredictability of a parameter. Eq. (7) derives the feature A-target parameter MI:7 MIXi;C=∑c∈C∑xi∈XiPxi,yLogPxi,yPxiP(y)

where Xi is the collection of features at the location, and i and C are the set of classes. High MI values signify a strong relationship between a feature and a class. MI = 0 when there is independence. MI is defined as Eq. (8), the degree of uncertainty in a feature that is reduced when its class is known:8 MIXi;C=HXi-HXi∣C,

where HXi denotes the entropy inside Xi and HXi∣C denotes the entropy of Xi following knowledge of C. As a result, if knowing C gives much knowledge, HXi∣C will be low, and MI will be high. In FS and sequence alignments, MI is frequently used. Typically, the top-k features with the highest MI are selected.

Chi-square (χ^2)

Chi-square (χ^2) is a statistical test and feature evaluation metric used to measure the independence between concrete variables. In the context of FS, the CS test assesses the relationship between a feature and the target variable by comparing the observed frequencies of their shared occurrences with the expected frequencies under the theory of self-rule.

The Chi-square statistic for feature A and target variable Y is calculated using the following Eq. (9):9 χ2r,ci=NPr,ciPr¯,c¯i-Pr,c¯iPr¯,ci2P(r)P(r¯)PciPc¯i

where ‘N’ stands for the complete dataset, and ‘r’ stands for the feature’s presence (r¯ its absence), and ′ci′ stands for the class. P(r,ci) represents the probability that feature ‘r’ will appear in class ‘ci’, while P(r¯,ci) represents the probability that feature ‘r’ will not appear in class ‘ci′. The probabilities that a feature will occur or not in a class that is not labeled ci are Pr,c¯i and Pr¯,c¯i, respectively. P(r) represents the probability that a feature appears in the dataset, whereas P(r¯) represents the probability that the feature does not appear. The odds that a dataset is or is not labelled with class ci are Pci and Pc¯i.

Proposed method

II-Phases complete the proposed method: (i) Ensemble multi-filters and (ii) Feature optimization. There are III-stages in Phase I. Five filters are used in stage I of the EFS process, employing a balance of the IG, MI, SU, CF, and RF algorithms. The filter-FS are then sorted using the SVM and CA. In Stage-II, the PCC balances the top-ranked features; in stage III, the RF classifier determines the top-j features from the blended features. The output from Phase-I is accepted to Phase-II, where the BOGGWOA is employed for a more optimized feature set. The proposed method’s workflow is depicted in Fig. 1.Fig. 1 Proposed FS architecture.

Phase I: ensemble multi filters feature selection and ranking (EMFFSR)

The EMFFSR technique is developed in Phase 1 over III stages. First, Phase I discusses using many filters to rank IG, MI, SU, CF, and RF and performing feature ranking to determine the initial top-ranked features from the dataset based on each IG, Correlation, Feature Score, GR value, and Feature Weight (FW), respectively39. The assembly of ranking outputs using PCC to produce ensemble features with the highest correlation is described in Stage-II’s second part. The top-j FS procedure from the ensemble output is described in Phase III.

Stage 1: Multi (N) filters algorithm ranking

Using their scores in the Entropy Value (EV), Feature Score, FW, and GR value, respectively, the ranking scores of IG, CF, SU, MI, and RF are applied to each feature in the first stage to determine the relevance level of each feature. Each ranking set’s top features were combined. Depending on the size of each dataset, ‘m’ had a distinct value. The accuracy attained by each feature was reported after applying an SVM classification algorithm to each feature from the union set for a preliminary assessment. According to the descending value of the CS, the features were sorted using the CS for each feature. For the following phase, the highest accuracy-based top-k features were then selected. Running the SVM classifier in the sorted union set yielded a value for ‘k’ that ranged from 1 to the sum of features in the sorted union set.

After a particular value of ‘k,’ the selection of the final feature subset included as many ‘k’ features as possible because accuracy was not improving (or was the same or had decreased) for those values. Only pertinent features were chosen in this stage after many features were eliminated40.

Stage 2: ensembling the top-k features from filter models

This step accepts the Phase-I k-features as input. The correlation between the features is assessed at this phase. All features’ pairwise correlations were determined using the Pearson Correlation Coefficient (PCC). A measurement of a linear relationship’s strength between two features or variables is what PCC is. It is typically indicated by the letter ‘r’, where r = 1 denotes the perfect positive correlation between the features, r = − 1 implies the ideal negative correlation, and r = 0 means the absence of any connection between the attributes. Let us assume that two variables ‘a’ and ‘b’ have ‘I’ instances each, in which case the PCC between (a, b) can be computed as Eq. (10)10 r=∑ai-αbi-βΣai-α2Σbi-β2

Here, (a, b) are the mean values for any two randomly selected instances, respectively; ‘r’ is the PCC value between {(a, b), (a, b)} are the mean values for the features (a, b). All features and classes are processed together in a correlation matrix. Only the list of features that are not correlated is passed to the subsequent stage if two features are highly correlated.

Stage 3: using random forest to find the top-j features

The top-j features are then selected by applying the RF classifier to the feature set. Then, Phase-II receives these ‘j’ numbers of features. The less informative features are diminished as a result of this approach. Ensemble Learning Methods (ELM) produce many classifiers and combine their output. CT bagging41 and boosting42 are well-known techniques. The method of assigning scores that were incorrectly predicted by prior researchers greater significance is referred to as “boosting”.

The anticipated result is a weighted vote. Bagging sequential trees originate from bootstrap testing of the data set and are distinct from past trees. A simple, higher vote is used to make a prediction. The RF proposed by43 increases the unpredictability of bagging. The Classification Trees (CT) or Regression Trees (RT) are built differently by RF, and each tree is constructed using a different bootstrap sample of the data. The optimal partitioning throughout every factor splits each node in conventional trees. Each RF node has been divided, utilizing the best predictor from a randomly selected group. This illogical methodology outperforms several other classifiers, such as discriminant analysis, SVM, and NN, and is resistant to overfitting. Two parameters—the number of parameters in each node’s random subgroup and the forest’s trees—drive the easy-to-use approach.

For the S training set, the F-feature (where F is constant) and the M-trees (where M is constant) are the inputs for this approach. Select an M-tree randomly from the dataset, and then create DT using that tree. B times are made for DT. At every single node, make a minimal subset of each feature F and extract the best features. The D feature was selected as the highest score as a result. Mathematically, the RF algorithm is stated as Eq. (11).11 RF=ArgMaxj∈{1,2,⋯N}∑m=1mDTm,j

class ‘j’ stands for the data classes, and ‘m’ stands for the number of DTs, starting with the first example and going up to the ith example. The most significant value of the function, or the majority vote, is referred to by the term ArgMax.

Phase II: wrapper model-based feature optimization

To select the best feature set from the top-j input feature subset, a unique Hybrid Genetic Grey Wolf Optimization Algorithm (HGGWOA) applying Binary Optimization (BO) is proposed in this phase. The following section goes through the proposed model:

Grey wolf optimizer (GWO)

GWO imitates the actions of wolves as they search for prey. In the wild, wolves typically live in packs of 5 to 12. Alpha, beta, delta, and omega wolves comprise one pack of four distinct 8 breeds40.

Each pack’s alpha wolf decides what to do. Beta wolves help alphas in decision-making processes. Delta wolves accept alpha and beta. Omegas get other wolves.

Alpha (Gα→) is the best mathematical solution, whereas beta (Gβ→) and delta (Gδ→) are the second and the third best solutions. Omega (Gω→) is used to indicate alternative options. Eqs. (12–15) suggest alpha, beta, and delta wolves direct other wolves as they catch their prey, as depicted in Fig. 1.12 G→(t+1)&=G→p(t)-A→·D→

13 D→&=C→·G→p(t)-G→(t)

where the current iteration is denoted by ‘t’, the coefficient paths are A→, C→, the location of the target is G→p, and the location of the wolf is denoted by G→. The calculation of the A→, C→ Vectors are:14 A→=2a→·r1→-a→

15 C→=2r2→

where the elements of a→ are linearly reducing from 2 to 0 during iterations, and the paths r1→ , r2→ are random values ∈ [0, 1]. The parameter controls the balance of the exploration and exploitation processes a→, which is updated. According to the following Eq. (39) 45, the a→ values are determined as follows:16 a→=2-t·2Mt

where the optimizer’s maximum number of iterations is Mt.

As depicted in Fig. 2, the three best solutions, Gα→,Gβ→, andGδ→, direct other individuals (Gω→) to adjust their locations toward the presumed location of the prey.Fig. 2 GWO algorithm’s position update.

Equations (17) and (18) illustrate the procedure for updating position.17 Dα→=C1→·Gα→-G→Dβ→=C2→·Gβ→-G→Dδ→=C3→·Gδ→-G→

18 G1→=Gα→-A1→·Dα→G2→=Gβ→-A2→·Dβ→G3→=Gδ→-A3→·Dδ→

where Eq. (19) is used to calculate A1→,A2→,A3→ , and C1→,C2→,C3→ are calculated. As an average of the three solutions of G1→,G2→andG3→, the updated locations for the population, G→(t+1):19 G→(t+1)=G1→+G2→+G3→3

The Hybrid GGWO

The proposed the Hybrid GGWO46 algorithm. The fundamental concept of the GGWO is to address premature convergence due to inadequate solution exploitation in the regular GWO by incorporating the crossover and mutation operations from the GA into the algorithm. It will improve the method of searching for identifying the wolves’ successive positions and optimal solutions during subsequent iterations. It will improve the searching technique for determining the wolves’ following positions and optimal solutions during subsequent iterations47.

In subsequent iterations, it will aid in streamlining the search process to identify the wolves’ future stopping points and the best solutions. Additionally, the exploration process is quickened by this hybridization, and the optimum solution is found quickly.

The grey wolves’ pack is made by search agents (wolves) initially generated randomly as part of the GGWO operational sequence. In the vector Vim,→ each wolf in the pack is assigned the following position: Eq. (20)20 Vi(x)→=vi1(x)..vi2(x)..vicurnti(x)Ti=1,2⋯..N

‘x’ ranges from 1 to a predetermined maximum of iterations, xmax, where the sum of search agents (wolves) is N, the ith wolf in the pack’s location is vii(x), the index of the current iteration is curnti.

Initial Population The representation of each solution happens in the early population stage by several genes. GWij, a first step in producing random solutions48. According to Eq. (21), based on the upper and lower boundaries, GWij is applied.21 GWij=GWLOwj+ı×GWUpj-GWLOwj

where between 0 and 1, W is a random number, i = (1,2… PSize), j=(1,2,⋯gn), is the population size (depending on the number of wolves).

Fitness Function In the optimization process, the most essential phase is choosing the FF because it serves as a benchmark for assessing the solutions of the population. Each gene’s Fitness Value (FV) is determined as following Eq. (22).22 FitGwi=∑i=1gn1yi-y^i2

Different weights are assigned to FV. The FF is then assessed for each solution, as displayed below, Eq. (23)23 FitGwij=∑i=1PSize∑j=1gn1yi-y^ij2

The search space to which the FV is assigned is shared by the search agents (wolves). After returning to the original solutions and selecting new ones, the search agents’ updated FV is FitNGwij are assessed. These updated FV are then contrasted with the earlier ones. Fresh values are of greater significance than previous ones; consequently, search agents substitute them with new values. In any case, traditional approaches endure49. Alpha or leader wolves have the most FF, next to beta and delta wolves. The lowest-ranked wolves, the omega wolves, and all other solutions follow the three distinct classifications. The search method is divided into two key stages for each iteration: exploration and exploitation.

Exploration Stage Every sample performs GA crossover and mutation to improve the search range while avoiding initial convergence and disregarding the traditional solution’s GWO technique50. The position of the prey and the alpha wolves are determined by the usual GWO algorithm by varying the value of a random variable, a→, from [0 to 2].

Crossover The parent solution Ps and neighbor solution Ns are crossed over. Depending on the value of a predetermined parameter called the Crossover Rate (CR), the Ns and Ps solutions are merged to produce one or more children. Based on a→. and r→1, GGWO determines D’s value using Ns to find the solution closest to the parent.

Mutation Mutation methods are employed in the following stages of prey (solution) search:

Algorithm for Mutation ( ).

where the newly generated solution is Newij, the parent solution is PSij, λijis a random number in the range [− 1,1], RSij is the selected random solution.

(d) Exploitation Stage The optimum prey location is selected by the top three wolves (α,β, and δ) during the exploitation stage. The following Eq. (24) is used to select the top three solutions:24 viS(x)=v1S,(x)⋯v2S,(x)⋯viS,(x),S∈{α,β,δ}

The following Eqs. (25–27) describes how the three-vector solutions are selected:25 Fitαvα(x)=Maxi=1...nFitαiviα(x)∣vα(x)∈Vi(x)

26 Fitβvβ(x)=Maxi=1..nFitβiviβ(x)∣vβ(x)∈Vi(x)vα(x)

27 Fitvδ(x)=Maxi=1..nFitδiviδ(x)∣vδ(x)∈Vi(x)\vα(x),vβ(x)

where the requisite requirement must be applied:Fitαvα(x)>Fitβvβ(x)>Fitδvδ(x)

They choose and refresh the three variables and the prey’s points using an identical crossover method. Finally, the best three wolves follow their best positions to update the position of other wolves (omega). The top three wolves then force the other wolves (omega wolves) to adjust their positions following their positions.

The proposed binary optimized genetic grey wolf optimization algorithm (BOGGWOA)

In general, GWO has succeeded in addressing issues with continuous optimization. By definition, FS is a binary problem. As a result, without adjustments, the algorithm described in this section cannot be used to tackle such issues. To solve the FS problem, a binary version of the GGWOA needs to be created. In the original GGWOA51, agents can continuously move around the search space since their location vectors have a continuous domain.

The update process of wolves, according to52, is a function of three vector positions, namely G1, G2, which advances each wolf to the top three solutions. The position updating can be changed to the following Eq. (28) so that the agents can operate in a binary space:28 Gdt+1=1ifsigmoidG1+G2+G33≥rand0otherwise

Sigmoid (a) is denoted as following Eq. (29). Gdt+1 is the binary updated location at iteration ‘t’ in dimension ‘d’, ‘r’ is a random number selected from a uniform distribution ∈ [1,0], and ‘r’ is a random number.29 sigmoid(a)=11+e-10(x-0.5)

The following Eq. (30) is updated and used to determine G1 , G2 , G3.30 G1d=1ifGαd+bstepαd≥10OtherwiseG2d=1ifGβd+bstepβd≥10OtherwiseG3d=1ifGδd+bstepδd≥10Otherwise

where Gα,β,δd denotes the alpha, beta, and delta wolves’ position vectors in ‘d’ dimensions and bstepα,β,δd denotes a binary step in ‘d’ dimensions, the Eq. (31), which is as follows:31 bstepα,β,δd=1ifcstepα,β,δd≥rand0otherwise

where ‘a rand’ is a random number generated by a uniform distribution ∈ [1,0], dimension is indicated by ‘d’, and for ‘d’, cstepα,β,δd is a constant value. The following Eq. (32) is used to determine this component:32 cstepα,β,δd=11+e-10A1dDα,β,δ-0.5d

In BGGWO, the exploration and exploitation are directed by an FF that is Eq. (33) represented as follows53 based on the top three solutions’ positions updated in46:33 Dα→=C1→·FitαvαG-viSGDβ→=C2→·FitβvβG-viSGDδ→=C3→·FitδvδG-viSG

A 1-D vector serves as an illustration of the study’s solution. The number of features determines this vector’s length. In this binary vector, the meaning of 0 and 1 correspond to the following:0 → No Feature Selected

1 → Feature Selected

The FS problem is, by definition, bi-objective. One objective is finding the fewest possible features, and increasing the CA is another. The following Eq. (34) is employed as an FF (the classifier is KMN) to consider:34 Fitness=αρR(D)+β|S||T|

where ρR (D) is the KNN classifier’s error rate and α = [0,1] and β = (1 − α) are limits modified from54. Additionally, |S| represents the features’ selected subset, and |T| represents all of the dataset’s features.

The obtaining of mathematical equations in this part should be emphasized. This work suggested using the GGWOA55 in combination with the concepts to tackle binary problems. The following algorithm presents the intricate steps in the BOGGWOA.

Algorithm for BOGGWOA.

Dataset and implementation

Table 1 lists the number of test cases and features for each of the five data sets used in our study, which include (i) Arrhythmia, (ii) Heart Disease, (iii) SPECT-F, (iv) Statlog, and (v) HCC. Using Python 3.6, the method and tests were applied in Jupyter Notebook. Python 3.6 is used to program Jupyter Notebook, where the process and experiments are implemented. The personal computer used for the trials has a 2.4 GHz Intel Core i7 processor, 16 GB of RAM, and OSX 10.11.6. The majority of FS, ML, and data processing methods are developed with the use of pre-existing libraries.Table 1 Description of benchmark data sets.

S. No	Dataset	Instance	Attribute	Class	Attribute type	
1	HCC	165	49	2	Integer, Real	
2	SPECT F	267	44	2	Integer	
3	Statlog	270	13	2	Real, Categorical	
4	Heart-C	303	75	2	Categorical, Integer, Real	
5	Arrhythmia	452	279	16	Categorical, Integer, Real	

Parameter tuning

Phase-I

The kernel function, its parameter ‘g’, and the penalty parameter ′cp′ throughout the ranking process has essential effects on the results, and the SVM classifier participates in Stage-I in order to determine the top ‘k’ features from the initial filter selected feature list. The most popular kernel function, the Radial Basis Function (RBF), was chosen since it has benefits like less optimum parameters and good classification efficiency. Its manifestation is Eq. (35).35 Kxi,yi=Exp-xi-yi22g2

A cross-validation method was used to identify the ideal pairing of g and cp. Within the range of cp∈2-10,210 and g∈2-10,210, the search was conducted with a step length of 0.5. The SVM parameters used in this model are displayed in Table 2 below.Table 2 SVM classifier’s parameters.

Parameter	Value	
SVM Model	C-SVM	
Kernel Function	RBF	
Parameter of Kernel Function,g	32	
Penalty Parameter,cp	1.414	
SVM, support vector machine; RBF, radial basis function.

The PCC (r), which has several names and is the most often used correlation coefficient in Stage II, is as follows:Pearson’s ‘r’

Bivariate correlation

Pearson Product-Moment Correlation Coefficient (PPMCC)

The correlation coefficient

Summary statistics, including the PCC sum datasets. The degree and direction of a linear relationship between the two numerical variables are marked. Even though how a relationship’s strength (effect size) is explained varies across disciplines, the following general guidelines are specified in Table 3:Table 3 PCC’s parameters.

PCC (r) Value	Strength	Direction	
Greater Than .5	Strong	Positive	
Between .3 and .5	Moderate	Positive	
Between 0 and .3	Weak	Positive	
0	None	None	
Between 0 and -.3	Weak	Negative	
Between -.3 and -.5	Moderate	Negative	
Less Than -.5	Strong	Negative	

The PCC can be used to evaluate statistical hypotheses because it is an inferential statistic. This work used a r-value between 0.3 and 0.5 in our model.

The RF algorithm is used in Stage III to extract the top-j features from the PCC merged feature list. A RF scatters them across various randomly selected subsets of the training set as an ensemble of DT. After all the predictors have been trained, the ensemble can predict a new instance by averaging all the single-tree predictions. Random Forest (RF) also introduces extra randomness in the tree-growing process by selecting the best feature from a random subset of features. This differs from choosing the feature that primarily reduces the overall error when split.

Very few presumptions regarding the training data are presented to DT. Therefore, if left unrestrained, they will closely fit—and overfit—the training data to their structure, overfitting them and failing to make reliable predictions about new data. Regularization can be accomplished by specifying several hyperparameters that impose constraints on the topology of the trees used to build the RF in order to prevent overfitting. Table 4 lists the hyperparameters that were assumed into consideration and the cross-validation search values that were experimented with.Table 4 Consideration of hyperparameters for RF algorithm.

Hyperparameter	Meaning	Sampled values	
No. of estimators	No. trees in the forest	10–25	
Max depth	Max depth of the tree	1–5	
Max of leaf nodes	Ma of leaf nodes in the DT	200–5000	
Max of features	There are some features to consider when looking for the best split	1–4	
Min of tests to split	The minimum of tests required to split an internal node	1–20	
Min of tests for a leaf	The minimum of tests required to be at a leaf node	1–5	

Phase-II

The proposed BOGGWOA-based WM optimizes the filter results. The standard parameters, such as the maximum of iterations (Maxiter), the dimension of the functions (D) and the population sizes (N) were set to identical values for all of the datasets. This research examines the proposed method’s optimization performance on problems with D = 100, 500, and 1000 for all test problems. The population size is 50, and the maximum of iterations is 1000. Table 5 displays the BOGGWOA algorithm’s ideal parameter settings. To ensure a fair comparison between the various datasets, 30 independent runs are made for each algorithm experiment on a benchmark function.Table 5 BOGGWOA’s parameter settings.

Parameters	Definition	Value	
N	Population Size	50	
Pc	Cross Probability	0.8	
Pm	Mutation Probability	0.01	
Maxiter	Maximum Number of Iterations	1000	
k	Nonlinear Adjustment Coefficient	0.5	

Experiment analysis

Table 6 lists all the features used and acquired throughout Phases 1 and 2. Each filter method (IG, CF, MI, RFF, and SU) was included in the union for five experimental datasets, which took the form of a union of features in Phase 1. As previously noted, this study sorted the non-increasing mean accuracy based on features after applying the SVM classification method. Thus, the accuracy would be the same when additional features are fed in the feature subset for all the datasets if their corresponding accuracies varied sine wave-like up to a specific number of feature subsets. As a result, from the first stage, there were a certain number of features for all five datasets where the accuracy was not growing, and these features were then passed to the next stage.Table 6 FS at the proposed model’s every phase.

Dataset	No. of features	FS by the FM union	SVM FS (k)	No. of non-correlated features	RF FS (j)	BOGGWOA FS	
HCC	165	IG	63	182	45	19	9	
CF	58	
MI	43	
RFF	68	
SU	72	
SPECT F	267	IG	77	232	58	24	17	
CF	68	
MI	56	
RFF	89	
SU	96	
Statlog	270	IG	88	232	58	24	13	
CF	57	
MI	61	
RFF	93	
SU	88	
Heart-C	303	IG	93	247	62	26	14	
CF	76	
MI	73	
RFF	68	
SU	101	
Arrhythmia	452	IG	95	287	72	30	21	
CF	102	
MI	91	
RFF	78	
SU	112	

The second stage involved calculating the correlation value and eliminating significantly connected features. For the HCC dataset, 137 out of the 182 features from stage 1 that were deemed to be strongly associated were discarded in this phase. The remaining 45 features were subjected to a stage using the RF method, and it was found that the accuracy was variable for the first 19 features before declining. These 19 traits were selected and passed on to the second phase. Similarly, the SPECT F, Statlog, Heart-C, and Arrhythmia datasets yielded the first 24, 24, 26, and 30 top features, respectively. The final set of features, 9, 17, 13, 14, and 21 for the HCC, SPECT F, Statlog, Heart-C, and Arrhythmia datasets, were obtained in the second phase by applying the BOGGWOA to the ‘top-j’ features obtained in the second phase.

Analysis in terms of classification

The proposed FS model is compared against various FS models for the selected five datasets, including:ReliefF-MA For classification,56 proposed combining the FM (ReliefF) and a WM (Memetic Algorithm (MA). The method’s objective is to remove the unimportant features and choose the most crucial feature subsets. The author first applied an MA for FS after calculating and updating the scores of each feature for each data set using the ReliefF algorithm. A classifier is used to assess CA using the k-Nearest Neighbor (k-NN) method with Leave-One-Out Cross-Validation (LOOCV).

B-BPSO The Binary BPSO (B-BPSO) method57 proposes a uniform mix of local decision-makers and global researchers to determine the best feature subset. In order to prevent gene loss, it upgrades particle leaders with reinforced memory and generates object CAs utilizing the 1-NN classifier.

IGWO-KELM IGWO-FS was developed to identify the optimal medical data feature subset58. The proposed approach employs GA to build a range of starting positions and GWO to refresh the population’s discrete searching space locations to identify the most suitable feature subset for a more accurate KELM classification objective.

The proposed model is combined with an SVM classifier, BPNN, and RNN, projecting three variations to compare the abovementioned models and determine the most excellent CA the recommended FS technique can achieve. Table 7 displays the parameter setting for the BPNN and RNN classifiers.Table 7 BPNN and RNN based on hyper-parameters.

Parameter	BPNN setting	RNN setting	
No. of nodes in the hidden layer	5	3	
Activation function	Sigmoid	Sigmoid	
Network training function	trainlm	traingdx	
Epoch	100	100	

The Accuracy, Sensitivity, Specificity, Precision, and F1-score of each of the models, as mentioned above, are assessed using Eqs. (36–41), which are used to generate the True Positive (TP), True Negative (TN), False Positive (FP), and False Negative (FN) values.36 Accuracy=TP+TNTP+TN+FP+FN.

In the equation provided above, TP indicates the total amount of positive instances noticed as being positive by the classifier, TN symbolizes negative in value cases, FP reflects the number of positive instances, and FN signifies negative instances.37 Accuracy=(TP+TN)(TP+FP+TN+FN)

38 Sensitivity=TPTP+FN,

39 Specificity=TNTN+FP

40 Precision=TPTP+FP

41 F-score=2TP2TP+(FP+FN)

Figure 3 and Table 8 display the outcomes attained by all models for the specified performance metrics.Fig. 3 Accuracy, sensitivity, specificity, precision, and F1-score performance of the comparison models (a) HCC, (b) SPECT F, (c) Statlog, (d) Heart-C, and (e) Arrhythmia Dataset.

Table 8 Comparison of the model's performance for different datasets.

Models	No. of FS	Accuracy	Sensitivity	Specificity	Precision	F1-score	
HCC	
ReliefF-MA	18	0.9679	0.9532	0.9858	0.9879	0.9702	
B-BPSO	14	0.9776	0.9708	0.9858	0.9881	0.9794	
IGWO-KELM	11	0.9679	0.9649	0.9716	0.9763	0.9706	
Proposed-SVM	9	0.9776	0.9709	0.9857	0.9882	0.9795	
Proposed-BPNN	9	0.9744	0.9593	0.9929	0.9940	0.9763	
Proposed-RNN	9	0.9872	0.9882	0.9859	0.9882	0.9882	
SPEC-F	
ReliefF-MA	27	0.8148	0.7872	0.8361	0.7872	0.7872	
B-BPSO	24	0.8241	0.7660	0.8689	0.8182	0.7912	
IGWO-KELM	19	0.8426	0.8085	0.8689	0.8261	0.8172	
Proposed-SVM	17	0.8333	0.8085	0.8525	0.8085	0.8085	
Proposed-BPNN	17	0.8624	0.8367	0.8833	0.8542	0.8454	
Proposed-RNN	17	0.8807	0.8750	0.8852	0.8571	0.8660	
Statlog	
ReliefF-MA	24	0.7283	0.5000	0.9038	0.8000	0.6154	
B-BPSO	18	0.8043	0.7750	0.8269	0.7750	0.7750	
IGWO-KELM	11	0.8152	0.6750	0.9231	0.8710	0.7606	
Proposed-SVM	13	0.8478	0.7895	0.8889	0.8333	0.8108	
Proposed-BPNN	13	0.8587	0.8108	0.8909	0.8333	0.8219	
Proposed-RNN	13	0.8696	0.8378	0.8909	0.8378	0.8378	
Heart-C	
ReliefF-MA	34	0.8424	0.7722	0.9170	0.9125	0.8365	
B-BPSO	28	0.9010	0.8846	0.9184	0.9200	0.9020	
IGWO-KELM	16	0.8614	0.7692	0.9592	0.9524	0.8511	
Proposed-SVM	14	0.8713	0.7963	0.9574	0.9556	0.8687	
Proposed-BPNN	14	0.8922	0.8333	0.9583	0.9574	0.8911	
Proposed-RNN	14	0.9118	0.8679	0.9592	0.9583	0.9109	
Arrhythmia	
ReliefF-MA	37	0.8361	0.7778	0.8824	0.8400	0.8077	
B-BPSO	28	0.8115	0.7963	0.8235	0.7818	0.7890	
IGWO-KELM	24	0.7705	0.7593	0.7794	0.7321	0.7455	
Proposed-SVM	21	0.8525	0.8246	0.8769	0.8545	0.8393	
Proposed-BPNN	21	0.8852	0.8596	0.9077	0.8909	0.8750	
Proposed-RNN	21	0.9098	0.8947	0.9231	0.9107	0.9027	

For the HCC dataset, the proposed FS method with RNN classifier achieved 98.7% accuracy on an average with 9 features, while the proposed model with SVM classifier came in second with a score of 97.7%. The proposed BPNN-classified FS model came in third with a score of 97.4% accuracy but a higher precision score of 99.4% than the SVM and RNN variants, which both scored 98.8%. Notably, although the B-BPSO model is proposed to employ more features than this method, its 97.7% agreement rate with the Proposed-SVM model is comparable. For the HCC dataset, the ReliefF-MA and IGWO-KELM models achieved 96.7% accuracy. The proposed model has made 83.3%, 86.2%, and 88% accuracy with SVM, BPNN, and RNN models for the SPECT-F dataset. Regarding all metrics, the proposed model, based on an RNN classifier, performs better than all the other models. With an accuracy of 84.2%, superior precision, and an F1-score of 82.6% and 81.7%, the IGWO-KELM outperformed the proposed SVM model. However, the model had 19 selected features as opposed to the proposed FS model’s 17 features for the SPECT-F dataset. Lower accuracy (81.48%) and more features, respectively, were achieved using ReliefF-MA. The Proposed-RNN model has an accuracy of 87% for the Statlog dataset, followed by the Proposed models based on BPNN and SVM, with accuracy rates of 85.8% and 84.8%, respectively. It is noteworthy that, among all the models, the IGWO-KELM had selected the limited features. It is described by the dataset’s definite nature, demonstrated by its higher Specificity of 92.3% compared to other models.

Only the proposed model with the RNN classifier outperformed the B-BPSO in terms of accuracy for the Heart-C dataset with a score of 90%. Compared to the proposed model, which has a better FS of 14, the B-BPSO model has the second-highest FS with a count of 28. The worst model is the ReliefF-MA, which has 34 features and an accuracy rate of 84.2%. Again, the proposed model performs better in accuracy, with an RNN classifier accuracy of 91.1%, compared to the proposed BPNN and presented SVM accuracy of 89.2% and 87.1%, respectively. The Arrhythmia dataset has a similar FS pattern, with ReliefF-MA performing the lowest with 37 features and 83.6% accuracy. The IGWO-KELM model, which has 24 features, is closely followed by the proposed model, which has just 21 features. However, compared to all the models, the IGWO-KELM has the lowest accuracy of 77%. It is clear from the above study that the proposed model, which includes all three classifier combinations, performs better than the models that were compared. A more focused RNN-based classifier for medical diagnosis is designed using the proposed FS model as pre-processing since the RNN classifier outperformed the other two variants in every metric.

Statistical significance test

To demonstrate whether the proposed method’s results were statistically important in comparison with other innovative algorithms presented in this work, the statistical significance test was conducted. Making quantitative decisions about any process is possible with the help of statistical testing59. A conjecture or hypothesis about the process was to be evaluated to see if there was enough evidence to “reject” it60. The null hypothesis is the hypothesis in question. In our situation, a claim of null hypothesis says that the distribution of the findings of the two sets was similar. This work used a one-sample t-test with two distinct ranges of significance to assess whether the null hypothesis was true, and the results are shown in Table 9. About CA, p-values, and t-values were computed. According to Table 10, this proposed strategy is statistically significant at a significance of 0.10 level for all five datasets.Table 9 Summary of the statistical significance test results.

Dataset	t-Value	p-Value	Significance level (0.10)	
HCC	− 2.56739	0.017671	Significant	
SPECT F	− 1.57008	0.080728	Significant	
Statlog	− 1.69134	0.068714	Significant	
Heart C	− 2.70692	0.018632	Significant	
Arrhythmia	− 1.92809	0.061859	Significant	

Table 10 Comparison of AUC Scores across different models for five datasets.

Dataset	Proposed model	ReliefF-MA	B-BPSO	IGWO-KELM	
HCC	0.962	0.948	0.955	0.941	
SPECT F	0.889	0.870	0.882	0.865	
Statlog	0.930	0.912	0.921	0.900	
Heart C	0.945	0.929	0.938	0.920	
Arrhythmia	0.978	0.960	0.970	0.955	

Area under curve analysis

The AUC scores detailed in Table10 demonstrate the better performance of the proposed model across all datasets when compared with existing models such as ReliefF-MA, B-BPSO, and IGWO-KELM. For instance, in the HCC dataset, the proposed model achieved an AUC of 0.962, outperforming the ReliefF-MA (0.948), B-BPSO (0.955), and IGWO-KELM (0.941). This trend is also consistent across other datasets, with the proposed model showing an extreme performance in the Arrhythmia dataset, achieving an AUC of 0.978, significantly higher than the next best model, B-BPSO, which scored 0.970. The minor lead is observed in the SPECT F dataset, where the proposed model’s AUC of 0.889 still surpasses the closest competitor, B-BPSO, at 0.882.

The time complexity comparison across different models, as shown in Table 11, highlights the computational efficiency of our proposed model compared to existing methods such as ReliefF-MA, B-BPSO, and IGWO-KELM. Across all datasets—HCC, SPECT F, Statlog, Heart C, and Arrhythmia—this model consistently exhibits a time complexity of O(n log n), indicating a more efficient processing capability in handling large datasets. In contrast, the ReliefF-MA and B-BPSO models show a higher complexity of O(n2), which can lead to significantly slower performance as the dataset size increases. The IGWO-KELM model matches our proposed model in efficiency but does not consistently outperform it in other metrics like AUC and CA. This efficient computational profile underscores our model’s suitability for real-time clinical decision support systems, where speed and accuracy are paramount.Table 11 Comparison of time complexity across different models for five datasets.

Dataset	Proposed model	ReliefF-MA	B-BPSO	IGWO-KELM	
HCC	O(n log n)	O(n2)	O(n2)	O(n log n)	
SPECT F	O(n log n)	O(n2)	O(n2)	O(n log n)	
Statlog	O(n log n)	O(n2)	O(n2)	O(n log n)	
Heart C	O(n log n)	O(n2)	O(n2)	O(n log n)	
Arrhythmia	O(n log n)	O(n2)	O(n2)	O(n log n)	

Conclusion and future work

In this work, a novel II-Phase Feature Selection (FS) framework combining multiple Filter Methods (FM) and a Wrapper Method (WM) is proposed. Phase I uses an ensemble of five FMs, including Mutual Information (MI), ReliefF, Information Gain (IG), Symmetrical Uncertainty (SU), and Correlation Filter (CF). Pearson Correlation (PC) is used in the second step of Phase I. In Phase II, a meta-heuristic called Binary Optimized Genetic Grey Wolf Optimization (BOGGWO) is employed as the WM to obtain the optimal feature subset. The union of the best-chosen FS by the 5-filters individually is used to leverage their common strength, and each feature’s accuracy is computed using the popular Machine Learning (ML) algorithm, SVM. This step ensures that if a specific FM mistakenly eliminates an essential feature, other FMs can include it in the feature set. Next, highly correlated features are removed to ensure that redundant attributes are neglected, and finally, BOGGWOA ensures that only the prime attributes are selected to achieve the highest accuracy. The model was compared with different models, and the result has proven the model’s efficiency in feature reduction through different metrics. Despite its promising results, the model has limitations, including computational intensity with large datasets, which may extend computation time. Additionally, its performance can vary across different datasets, potentially affecting its generalizability to various Internet of Things (IoT) healthcare applications.

Future research will focus on enhancing the computational efficiency of the model and validating its effectiveness across a broader range of IoT healthcare datasets. Further exploration into integrating additional learning models with the BOGGWOA could also enhance the model’s efficacy and applicability in real-world scenarios.

Author contributions

Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing—Original Draft, Writing—Review & Editing, Project administration: “R. Asir Chandra Shinoo, S. Sudhakar”.

Data availability

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Kumar S Internet of Things, is a revolutionary approach for future technology enhancement: A review J. Big Data 2019 6 1 21 10.1186/s40537-019-0268-2
Kumar, S. et al. Internet of Things, is a revolutionary approach for future technology enhancement: A review. J. Big Data 6, 1–21 (2019).
2. Dias D Paulo Silva Cunha J Wearable health devices: Vital sign monitoring, systems, and technologies Sensors 2018 18 8 2414 10.3390/s18082414 30044415
Dias, D. & Paulo Silva Cunha, J. Wearable health devices: Vital sign monitoring, systems, and technologies. Sensors 18(8), 2414 (2018).30044415
3. Joyia G Internet of medical+ings (IOMT): Applications, benefits and future challenges in healthcare domain J. Commun. 2017 12 240 247
Joyia, G. et al. Internet of medical+ings (IOMT): Applications, benefits and future challenges in healthcare domain. J. Commun. 12, 240–247 (2017).
4. Paulraj D An automated exploring and learning model for data prediction using balanced CA-SVM J. Ambient Intell. Humaniz. Comput. 2020 12 1 12
Paulraj, D. An automated exploring and learning model for data prediction using balanced CA-SVM. J. Ambient Intell. Humaniz. Comput. 12, 1–12 (2020).
5. Hall M Correlation-Based Feature Selection for Machine Learning 1999 Waikato University
Hall, M. Correlation-Based Feature Selection for Machine Learning (Waikato University, 1999).
6. Cios KJ Moore GW Uniqueness of medical data mining Artif. Intell. Med. 2002 26 1–2 1 24 10.1016/S0933-3657(02)00049-0 12234714
Cios, K. J. & Moore, G. W. Uniqueness of medical data mining. Artif. Intell. Med. 26(1–2), 1–24 (2002).12234714
7. Esfandiari N Knowledge discovery in medicine: Current issue and future trend Expert Syst. Appl. 2014 41 9 4434 4463 10.1016/j.eswa.2014.01.011
Esfandiari, N. et al. Knowledge discovery in medicine: Current issue and future trend. Expert Syst. Appl. 41(9), 4434–4463 (2014).
8. Kawamoto K Improving clinical practice using clinical decision support systems: A systematic review of trials to identify features critical to success Biomed. J. 2005 330 7494 765 772
Kawamoto, K. et al. Improving clinical practice using clinical decision support systems: A systematic review of trials to identify features critical to success. Biomed. J. 330(7494), 765–772 (2005).
9. Vergara JR A review of feature selection methods based on mutual information Neural Comput. Appl. 2014 24 1 175 186 10.1007/s00521-013-1368-0
Vergara, J. R. et al. A review of feature selection methods based on mutual information. Neural Comput. Appl. 24(1), 175–186 (2014).
10. Kohavi R John GH Wrappers for feature subset selection Artif. Intell. 1997 97 1 273 324 10.1016/S0004-3702(97)00043-X
Kohavi, R. & John, G. H. Wrappers for feature subset selection. Artif. Intell. 97(1), 273–324 (1997).
11. Bolón-Canedo V A review of feature selection methods on synthetic data Knowl. Inf. Syst. 2013 34 3 483 519 10.1007/s10115-012-0487-8
Bolón-Canedo, V. et al. A review of feature selection methods on synthetic data. Knowl. Inf. Syst. 34(3), 483–519 (2013).
12. Peng H Long F Ding C Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy IEEE Trans. Pattern Anal. Mach. Intell. 2005 27 1226 1238 10.1109/TPAMI.2005.159 16119262
Peng, H., Long, F. & Ding, C. Feature selection based on mutual information: Criteria of max-dependency, max-relevance, and min-redundancy. IEEE Trans. Pattern Anal. Mach. Intell. 27, 1226–1238 (2005).16119262
13. Dreiseitl S Ohno-Machado L Logistic regression, and artificial neural network classification models: A methodology review J. Biomed. Inform. 2002 35 5 352 359 10.1016/S1532-0464(03)00034-0 12968784
Dreiseitl, S. & Ohno-Machado, L. Logistic regression, and artificial neural network classification models: A methodology review. J. Biomed. Inform. 35(5), 352–359 (2002).12968784
14. Quinlan JR Induction of Decision Trees Machine Learning 1986 Kluwer Academic Publishers 81 106
Quinlan, J. R. Induction of Decision Trees. In Machine Learning Vol. 1 81–106 (Kluwer Academic Publishers, 1986).
15. Bommert A Benchmark for filter methods for feature selection in high-dimensional classification data Comput. Stat. Data Anal. 2020 143 106839 10.1016/j.csda.2019.106839
Bommert, A. et al. Benchmark for filter methods for feature selection in high-dimensional classification data. Comput. Stat. Data Anal. 143, 106839 (2020).
16. Ali H Imbalance class problems in data mining: A review Indones. J. Electr. Eng. Comput. Sci. 2019 14 3 1560 1571
Ali, H. et al. Imbalance class problems in data mining: A review. Indones. J. Electr. Eng. Comput. Sci. 14(3), 1560–1571 (2019).
17. Dongare SA A feature selection approach for enhancing the cardiotocography classification performance Int. J. Eng. Tech. 2018 4 2 222 226
Dongare, S. A. et al. A feature selection approach for enhancing the cardiotocography classification performance. Int. J. Eng. Tech. 4(2), 222–226 (2018).
18. Alirezanejad M Heuristic filter feature selection methods for medical datasets Genomics 2019 112 2 1173 1181 10.1016/j.ygeno.2019.07.002 31276753
Alirezanejad, M. et al. Heuristic filter feature selection methods for medical datasets. Genomics 112(2), 1173–1181 (2019).31276753
19. Singh B A feature subset selection technique for high dimensional data using symmetric uncertainty J. Data Anal. Inf. Process 2014 2 95 105
Singh, B. et al. A feature subset selection technique for high dimensional data using symmetric uncertainty. J. Data Anal. Inf. Process 2, 95–105 (2014).
20. Ren YG Rough set attribute reduction algorithm based on GA Comput. Eng. Sci. 2006 47 134 136
Ren, Y. G. et al. Rough set attribute reduction algorithm based on GA. Comput. Eng. Sci. 47, 134–136 (2006).
21. Long NC Attribute reduction based on rough sets and the discrete firefly algorithm Recent Advances in Information and Communication Technology 2014 Springer 13 22
Long, N. C. et al. Attribute reduction based on rough sets and the discrete firefly algorithm. In Recent Advances in Information and Communication Technology 13–22 (Springer, 2014).
22. Inbarani HH Supervised hybrid feature selection based on PSO and rough sets for medical diagnosis Comput. Methods Programs Biomed. 2014 113 175 185 10.1016/j.cmpb.2013.10.007 24210167
Inbarani, H. H. et al. Supervised hybrid feature selection based on PSO and rough sets for medical diagnosis. Comput. Methods Programs Biomed. 113, 175–185 (2014).24210167
23. Bae C Feature selection with intelligent dynamic swarm and rough set Expert Syst. Appl. 2010 37 7026 7032 10.1016/j.eswa.2010.03.016
Bae, C. et al. Feature selection with intelligent dynamic swarm and rough set. Expert Syst. Appl. 37, 7026–7032 (2010).
24. Chen Y Finding rough set reducts with fish swarm algorithm Knowl. Based Syst. 2015 81 22 29 10.1016/j.knosys.2015.02.002
Chen, Y. et al. Finding rough set reducts with fish swarm algorithm. Knowl. Based Syst. 81, 22–29 (2015).
25. Yu L Liu H Efficient feature selection via analysis of relevance and redundancy J. Mach. Learn. Res. 2004 5 1205 1224
Yu, L. & Liu, H. Efficient feature selection via analysis of relevance and redundancy. J. Mach. Learn. Res. 5, 1205–1224 (2004).
26. Babatunde O A genetic algorithm-based feature selection Int. J. Electr. Commun. Comput. Eng. 2014 5 4 899 905
Babatunde, O. et al. A genetic algorithm-based feature selection. Int. J. Electr. Commun. Comput. Eng. 5(4), 899–905 (2014).
27. Gunavathi C Performance analysis of genetic algorithm with kNN and SVM for feature selection in tumor classification Int. Sch. Sci. Res. Innov. 2014 8 8 1490 1497
Gunavathi, C. et al. Performance analysis of genetic algorithm with kNN and SVM for feature selection in tumor classification. Int. Sch. Sci. Res. Innov. 8(8), 1490–1497 (2014).
28. Tan KC A hybrid evolutionary algorithm for attribute selection in data mining Expert Syst. Appl. 2009 36 4 8616 8630 10.1016/j.eswa.2008.10.013
Tan, K. C. et al. A hybrid evolutionary algorithm for attribute selection in data mining. Expert Syst. Appl. 36(4), 8616–8630 (2009).
29. Moradi P hybrid particle swarm optimization for feature subset selection by integrating a novel local search strategy Appl. Soft Comput. 2016 43 117 130 10.1016/j.asoc.2016.01.044
Moradi, P. et al. hybrid particle swarm optimization for feature subset selection by integrating a novel local search strategy. Appl. Soft Comput. 43, 117–130 (2016).
30. Hafez, A. I. et al. An innovative approach for feature selection based on chicken swarm optimization. In 2015, 7th International Conference of Soft Computing and Pattern Recognition (SoCPaR) 19–24.
31. Panda M Elephant search optimization combined with the deep neural network for microarray data analysis J. King Saud Univ. Comput. Inf. Sci. 2017 32 940
Panda, M. Elephant search optimization combined with the deep neural network for microarray data analysis. J. King Saud Univ. Comput. Inf. Sci. 32, 940 (2017).
32. Douglas R A wrapper approach for feature selection based on bat algorithm and optimum-path forest Expert Syst. Appl. 2014 41 5 2250 2258 10.1016/j.eswa.2013.09.023
Douglas, R. et al. A wrapper approach for feature selection based on bat algorithm and optimum-path forest. Expert Syst. Appl. 41(5), 2250–2258 (2014).
33. Mirjalili, S.; et al., A. Grey wolf optimizer. Adv. Eng. Softw. 2014, 69, 46–61
34. Breiman L Bagging predictors Mach. Learn. 1996 24 2 123 140 10.1007/BF00058655
Breiman, L. Bagging predictors. Mach. Learn. 24(2), 123–140 (1996).
35. Shapire R Boosting the margin: A new explanation for the effectiveness of voting methods Ann. Stat. 1998 26 5 1651 1686
Shapire, R. et al. Boosting the margin: A new explanation for the effectiveness of voting methods. Ann. Stat. 26(5), 1651–1686 (1998).
36. Breiman L Random forests Mach. Learn. 2001 45 1 5 32 10.1023/A:1010933404324
Breiman, L. Random forests. Mach. Learn. 45(1), 5–32 (2001).
37. Tawhid MA A Hybrid grey wolf optimizer and genetic algorithm for minimizing potential energy function Memetic Comp. 2017 9 347 359 10.1007/s12293-017-0234-5
Tawhid, M. A. et al. A Hybrid grey wolf optimizer and genetic algorithm for minimizing potential energy function. Memetic Comp. 9, 347–359 (2017).
38. Emary E Binary grey wolf optimization approaches for feature selection Neurocomputing 2016 172 371 381 10.1016/j.neucom.2015.06.083
Emary, E. et al. Binary grey wolf optimization approaches for feature selection. Neurocomputing 172, 371–381 (2016).
39. ZorarpacI E A hybrid approach of differential evolution and artificial bee colony for feature selection Expert Syst. Appl. 2016 62 91 103 10.1016/j.eswa.2016.06.004
ZorarpacI, E. et al. A hybrid approach of differential evolution and artificial bee colony for feature selection. Expert Syst. Appl. 62, 91–103 (2016).
40. Elghamrawy SM A hybrid Genetic-Grey Wolf Optimization algorithm for optimizing Takagi–Sugeno–Kang fuzzy systems Neural Comput. Appl. 2022 34 17051 17069 10.1007/s00521-022-07356-5
Elghamrawy, S. M. et al. A hybrid Genetic-Grey Wolf Optimization algorithm for optimizing Takagi–Sugeno–Kang fuzzy systems. Neural Comput. Appl. 34, 17051–17069 (2022).
41. Altman NS An introduction to kernel and nearest-neighbor nonparametric regression Am. Stat. 1992 46 3 175 185 10.1080/00031305.1992.10475879
Altman, N. S. An introduction to kernel and nearest-neighbor nonparametric regression. Am. Stat. 46(3), 175–185 (1992).
42. Yang, C. S. et al. Feature selection using memetic algorithms. In Proceedings of the Third International Conference on Convergence and Hybrid Information Technology, Busan 416–423 (2008).
43. Zhang Y Feature selection algorithm based on bare-bones particle swarm optimization Neurocomputing 2015 148 150 157 10.1016/j.neucom.2012.09.049
Zhang, Y. et al. Feature selection algorithm based on bare-bones particle swarm optimization. Neurocomputing 148, 150–157 (2015).
44. Qiang Li An enhanced grey wolf optimization based feature selection wrapped kernel extreme learning machine for medical diagnosis Comput. Math. Methods Med. 2017 2017 9512741 1 15
Qiang, Li. et al. An enhanced grey wolf optimization based feature selection wrapped kernel extreme learning machine for medical diagnosis. Comput. Math. Methods Med. 2017(9512741), 1–15 (2017).
45. Sahebi G Movahedi P Ebrahimi M Pahikkala T Plosila J Tenhunen H GeFeS: A generalized wrapper feature selection approach for optimizing classification performance Comput. Biol. Med. 2020 125 103974 10.1016/j.compbiomed.2020.103974 32890978
Sahebi, G. et al. GeFeS: A generalized wrapper feature selection approach for optimizing classification performance. Comput. Biol. Med. 125, 103974 (2020).32890978
46. Cui X Li Y Fan J Wang T Zheng Y A hybrid improved dragonfly algorithm for feature selection IEEE Access 2020 8 155619 155629 10.1109/ACCESS.2020.3012838
Cui, X., Li, Y., Fan, J., Wang, T. & Zheng, Y. A hybrid improved dragonfly algorithm for feature selection. IEEE Access 8, 155619–155629 (2020).
47. Kadam V Jadhav S Yadav S Bagging-based ensemble of Support Vector Machines with improved elitist GA-SVM features selection for cardiac arrhythmia classification Int. J. Hybrid Intell. Syst. 2020 16 25 33
Kadam, V., Jadhav, S. & Yadav, S. Bagging-based ensemble of Support Vector Machines with improved elitist GA-SVM features selection for cardiac arrhythmia classification. Int. J. Hybrid Intell. Syst. 16, 25–33 (2020).
48. Wang T Chen P Bao T Li J Yu X Arrhythmia classification algorithm based on SMOTE and Feature Selection IJPE 2021 17 263 10.23940/ijpe.21.03.p2.263275
Wang, T., Chen, P., Bao, T., Li, J. & Yu, X. Arrhythmia classification algorithm based on SMOTE and Feature Selection. IJPE 17, 263 (2021).
49. Luo J Ahmad SF Alyaemeni A Ou Y Irshad M Alyafi-Alzahri R Alsanie G Unnisa ST Role of perceived ease of use, usefulness, and financial strength on the adoption of health information systems: The moderating role of hospital size Human. Soc. Sci. Commun. 2024 11 1 516 10.1057/s41599-024-02976-9
Luo, J. et al. Role of perceived ease of use, usefulness, and financial strength on the adoption of health information systems: The moderating role of hospital size. Human. Soc. Sci. Commun. 11(1), 516. 10.1057/s41599-024-02976-9 (2024).
50. Cao P Pan J Understanding Factors influencing geographic variation in healthcare expenditures: A small areas analysis study INQUIRY J. Health Care Organ. Provis. Financ. 2024 1 2 10.1177/00469580231224823
Cao, P. & Pan, J. Understanding Factors influencing geographic variation in healthcare expenditures: A small areas analysis study. INQUIRY J. Health Care Organ. Provis. Financ. 1, 2. 10.1177/00469580231224823 (2024).
51. Xue Q Xu DR Cheng TC Pan J Yip W The relationship between hospital ownership, in-hospital mortality, and medical expenses: an analysis of three common conditions in China Arch. Public Health 2023 81 1 19 10.1186/s13690-023-01029-y 36765426
Xue, Q., Xu, D. R., Cheng, T. C., Pan, J. & Yip, W. The relationship between hospital ownership, in-hospital mortality, and medical expenses: an analysis of three common conditions in China. Arch. Public Health 81(1), 19. 10.1186/s13690-023-01029-y (2023).36765426
52. Zhu C An adaptive agent decision model based on deep reinforcement learning and autonomous learning J. Logist. Inform. Serv. Sci. 2023 10 3 107 118 10.33168/JLISS.2023.0309
Zhu, C. An adaptive agent decision model based on deep reinforcement learning and autonomous learning. J. Logist. Inform. Serv. Sci. 10(3), 107–118. 10.33168/JLISS.2023.0309 (2023).
53. Zhang M Wei E Berry R Huang J Age-dependent differential privacy IEEE Trans. Inf. Theory 2024 70 2 1300 1319 10.1109/TIT.2023.3340147
Zhang, M., Wei, E., Berry, R. & Huang, J. Age-dependent differential privacy. IEEE Trans. Inf. Theory 70(2), 1300–1319. 10.1109/TIT.2023.3340147 (2024).
54. Hu F Qiu L Zhou H Medical device product innovation choices in Asia: An empirical analysis based on product space Front. Public Health 2022 10 871575 10.3389/fpubh.2022.871575 35493362
Hu, F., Qiu, L. & Zhou, H. Medical device product innovation choices in Asia: An empirical analysis based on product space. Front. Public Health 10, 871575. 10.3389/fpubh.2022.871575 (2022).35493362
55. Sahoo KK Ghosh R Mallik S Wrapper-based deep feature optimization for activity recognition in the wearable sensor networks of healthcare systems Sci. Rep. 2023 13 965 10.1038/s41598-022-27192-w 36653370
Sahoo, K. K. et al. Wrapper-based deep feature optimization for activity recognition in the wearable sensor networks of healthcare systems. Sci. Rep. 13, 965. 10.1038/s41598-022-27192-w (2023).36653370
56. Sahebi G Movahedi P Ebrahimi M Pahikkala T Plosila J Tenhunen H GeFeS: A generalized wrapper feature selection approach for optimizing classification performance Comput. Biol. Med. 2020 125 103974 10.1016/j.compbiomed.2020.103974 32890978
Sahebi, G. et al. GeFeS: A generalized wrapper feature selection approach for optimizing classification performance. Comput. Biol. Med. 125, 103974. 10.1016/j.compbiomed.2020.103974 (2020).32890978
57. Ahmed SF Alam MSB Afrin S Rafa SJ Rafa N Gandomi AH Insights into Internet of medical things (IoMT): Data fusion, security issues and potential solutions Inf. Fusion 2024 102 102060 10.1016/j.inffus.2023.102060
Ahmed, S. F. et al. Insights into Internet of medical things (IoMT): Data fusion, security issues and potential solutions. Inf. Fusion 102, 102060. 10.1016/j.inffus.2023.102060 (2024).
58. Mandal AK Nadim M Saha H Sultana T Hossain MD Huh E-N Feature subset selection for high-dimensional, low sampling size data classification using ensemble feature selection with a wrapper-based search IEEE Access 2024 12 62341 62357 10.1109/ACCESS.2024.3390684
Mandal, A. K. et al. Feature subset selection for high-dimensional, low sampling size data classification using ensemble feature selection with a wrapper-based search. IEEE Access 12, 62341–62357. 10.1109/ACCESS.2024.3390684 (2024).
59. Moran M Gordon G Deep curious feature selection: A recurrent, intrinsic-reward reinforcement learning approach to feature selection IEEE Trans. Artif. Intell. 2024 5 3 1174 1184 10.1109/TAI.2023.3282564
Moran, M. & Gordon, G. Deep curious feature selection: A recurrent, intrinsic-reward reinforcement learning approach to feature selection. IEEE Trans. Artif. Intell. 5(3), 1174–1184. 10.1109/TAI.2023.3282564 (2024).
60. Nie F Ma Z Wang J Li X Fast sparse discriminative k-means for unsupervised feature selection IEEE Trans. Neural Netw. Learn. Syst. 2024 35 7 9943 9957 10.1109/TNNLS.2023.3238103 37022041
Nie, F., Ma, Z., Wang, J. & Li, X. Fast sparse discriminative k-means for unsupervised feature selection. IEEE Trans. Neural Netw. Learn. Syst. 35(7), 9943–9957. 10.1109/TNNLS.2023.3238103 (2024).37022041
