
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)13172-6
10.1016/j.heliyon.2024.e37141
e37141
Research Article
Application of deep ensemble learning for palm disease detection in smart agriculture
Savaş Serkan serkansavas@kku.edu.tr

Department of Computer Engineering, Kırıkkale University, 71450, Kırıkkale, Turkey
29 8 2024
15 9 2024
29 8 2024
10 17 e3714110 8 2023
27 8 2024
28 8 2024
© 2024 The Author
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Agriculture has notably become one of the fields experiencing intensive digital transformation. Leveraging state-of-the-art techniques in this domain has provided numerous advantages for agricultural activities. Deep learning (DL) algorithms have proven beneficial in addressing various agricultural challenges. This study presents a comprehensive investigation into applying DL models for palm disease detection and classification in the context of smart agriculture. The research aims to address the limitations observed in previous studies and improve the robustness and generalizability of the results. To achieve this, a two-stage optimization methodology is employed. First, transfer learning and fine-tuning techniques are applied using various pre-trained deep neural network models. The experiments show promising results, with all models achieving high accuracy rates during training and validation. Furthermore, their performance on unseen test data is also assessed to ensure practical applicability. The top-performing models are MobileNetV2 (92.48 %), ResNet (92.42 %), ResNetRS50 (92.30 %), and DenseNet121 (92.01 %). Second, a deep ensemble learning approach is applied to enhance the models' generalization capability further. The best-performing models with different criteria are combined using the ensemble technique, resulting in remarkable improvements in disease detection tasks. DELM1 emerges as the most successful ensemble model, achieving an ROC AUC Score of 99 %. This study demonstrates the effectiveness of deep ensemble learning models in palm disease detection and classification for smart agriculture applications. The findings contribute to advancing disease detection systems and emphasize the potential of ensemble learning. The study provides valuable insights for future research, guiding the application of DL techniques to address critical agricultural challenges and improve crop health monitoring systems. Another contribution is combining various plant diseases and insect pest classes using diverse datasets. A comprehensive classification system is achieved by considering different disease classes and stages within the white scale category, improving the model's robustness.

Keywords

Smart agriculture
Smart farming
Palm disease
Deep ensemble learning
Transfer learning
==== Body
pmc1 Introduction

The digital transformation, brought about by information and communication technologies (ICT), marks the latest evolution in working life and social life, alongside the advancements in science and technology. The impact of this transformation is evident in various sectors, including agriculture, production, industry, education, health, service sector, defense, shelter, nutrition, and social life. The continuous development of ICTs has led to their increasing adoption in various fields. As a result, the volume of digital data generated in these areas has been growing exponentially over time [1].

In this new digital transformation era, the widespread use of ICT has led to massive amounts of data in all fields, thanks to the "input-output" relationship inherent in digital technologies. This data holds immense value for institutions when processed according to their intended purpose. To unlock this treasure trove, machine learning (ML) and deep learning (DL) applications under the scope of artificial intelligence (AI) play a vital role. Accessing and managing information from raw data has become possible by applying methods and techniques in AI. Consequently, making informed decisions based on data and information has become increasingly advantageous, particularly for institutions. Integrating automation technologies with digital transformation allows machines to perform physically demanding tasks and steers individuals towards skilled jobs, optimizing human resources. This shift towards qualified employment is just one of the many advantages of digital transformation, including energy savings, time savings, and resource conservation [2]. The widespread adoption of computer systems across all industries, especially since the 2000s, made manual calculations obsolete. This led to the increase in popularity of AI, ML, and DL, as they were proven effective in producing meaningful outcomes.

Agriculture has notably become one of the fields experiencing intensive digital transformation in recent times. Leveraging state-of-the-art techniques in this domain has provided numerous advantages for agricultural activities. The digital evolution of the agricultural sector, while rooted in a long history of data utilization, is primarily driven by innovation in data collection, processing, and management. Advances in sensor technology have enabled real-time monitoring of specific metrics, facilitating more efficient automation of robotic processes in farming. Concurrently, the increasing accessibility and cost-effectiveness of computing power have fostered the development of novel decision-support tools, enhancing agricultural management practices. Of particular significance is the improved quality of farm-level information and the sophisticated technologies employed for its collection, storage, processing, management, and distribution. This technological progress has markedly transformed the agricultural landscape, integrating data-driven approaches into various aspects of farming operations [3]. Particularly, DL algorithms have proven beneficial in addressing various agricultural challenges. A study by Ref. [4] examined 38 different DL research efforts, revealing that 63 % focused on disease detection, 11 % on weed detection, and 26 % on yield prediction. The research highlights that the advantages of deep learning are encouraging and can pave the way for smarter, more sustainable agriculture and safer food production. Similarly [5], investigated 32 different DL studies and found that DL outperformed other technologies in performance. Moreover, considering advancements in computer hardware, the researchers expect that DL will receive increased attention and find broader applications in future research. Additionally, they mentioned that most of their investigations were based on convolutional neural network (CNN) approaches. CNN is extensively utilized in agricultural research due to its robust image-processing capabilities. One of the primary applications of DL in agriculture is plant and crop classification, which offers valuable insights for yield prediction, pest control, disaster monitoring, robotic harvesting, and more. Traditionally, manual plant disease detection was time-consuming. However, with the advancements in AI, image-processing techniques can now be employed for plant disease detection. Plant disease detection models typically utilize leaf pattern recognition and image classification to identify and diagnose diseases accurately [6].

These activities, also known as smart farming or smart agriculture, have triggered the beginning of a new era in the agricultural sector. With the rise of digital transformation, farm fields are generating significant amounts of data. Moreover, there is currently great potential for data generation in various other domains.

Smart farming plays a crucial role in addressing various challenges in agriculture, including productivity, environmental impact, food security, and sustainability. With a growing global population, it is essential to significantly increase food production while ensuring its availability and nutritional quality worldwide. Smart farming also aims to protect natural ecosystems by adopting sustainable farming practices. To overcome these challenges, it is crucial to understand the complex and unpredictable agricultural ecosystems better. This can be achieved by continuously monitoring, measuring, and analyzing various physical aspects and phenomena. The analysis of extensive agricultural data and the utilization of new ICT play a vital role in managing crops and farms at both small and large scales. These technologies enhance decision-making processes by providing context, situation, and location awareness. Remote sensing, utilizing satellites, airplanes, and uncrewed aerial vehicles (UAVs) such as drones, facilitates large-scale observation of agricultural environments. This method offers several advantages in agriculture, as it is a well-established, non-destructive approach for gathering information about land features and enables systematic data collection over extensive geographical areas [7]. This study implemented a smart farm application using a new Mendeley dataset and the Kaggle dataset, leveraging the potential offered by these datasets.

Palm trees are vital in tropical regions worldwide, serving many purposes beyond their visual appeal. Their significance spans various aspects, including food and beverage production, construction materials, clothing, and medicinal applications. Palm trees have been integral to human lives for centuries, symbolizing abundance, prosperity, and the essence of life in many cultures and regions around the globe. Palm trees possess a diverse range of applications. Their leaves are utilized for thatching roofs and crafting baskets, while the sturdy trunk serves construction and furniture-making purposes. Certain palm tree fruits, such as coconuts and dates, offer edible and nutritious options. Palm oil, derived from the fruit, finds application in cooking, cosmetics, and biofuels. Furthermore, people frequently plant palm trees for their aesthetic value in landscaping and ornamentation. Additionally, many use smaller palm tree varieties for decorative purposes. Palm plants offer numerous advantages to their surroundings. People highly value them for their capacity to enhance the visual attractiveness of both indoor and outdoor spaces, creating a tropical ambiance. Beyond their aesthetic appeal, palm trees improve air quality by absorbing carbon dioxide and releasing oxygen through photosynthesis. Moreover, these trees offer shade, effectively cooling the surrounding environment and reducing energy consumption. The cultivation of palm trees results in the production of a diverse array of products. One of the most prominent is palm oil, extracted from the fruit. Palm oil is extensively applied in the food industry, cosmetics, soaps, and biodiesel fuel. Other notable products derived from palm trees include palm sugar, obtained from the sap, palm wine, palm fiber utilized in the creation of mats and ropes, and palm leaves used for thatching roofs or crafting decorative items. Five common uses of palm products include cooking oil, cosmetics, biofuels, food additives, and soap & detergent manufacturing. Five products derived from palm trees are coconuts, dates, palm oil, palm sugar, and palm leaves. Palm leaves are used for thatching roofs and making baskets, mats, and other decorative items [8].

Due to all these benefits, preventing diseases and factors that could harm palm trees during their growth is essential. In agriculture, digital transformation enables DL algorithms to carry out such tasks effectively. Based on this, the study's research questions (RQs) were determined as follows.RQ1 How can DL algorithms be utilized to detect diseases in palm trees at an early stage?

RQ2 How does the two-stage analysis method, specifically transfer learning and ensemble learning, demonstrate its effectiveness in detecting palm tree diseases?

RQ3 What are the implications and potential benefits of the comprehensive analysis method for future research in smart farming?

The objectives of this study carried out in line with these RQs are as follows.• To demonstrate the effectiveness of DL algorithms in smart agriculture applications.

• To implement a smart farming application using DL algorithms for early detection of palm tree diseases.

• To develop a two-stage analysis method using optimization techniques to enhance productivity and efficiency in smart farming.

• To save resources by reducing the need for manual inspection.

• To establish a foundation for future research in smart farming by providing a comprehensive analysis method and contributing to the development of larger-scale agricultural management systems and predictive models.

The contributions of this work are summarized as follows.• An example study of smart farming.

• Increasing productivity through early detection of palm tree diseases is significant in the production industry: By implementing smart farming techniques that utilize AI and image recognition algorithms, early detection of diseases in palm trees can be achieved. Monitoring the health of palm trees through visual analysis of leaf images can help detect diseases at their early stages, allowing for timely intervention and increasing overall productivity.

• Resource savings by reducing the number of experts required for manual control: Implementing smart farming technologies minimizes the dependency on manual control and human intervention. By automating tasks through AI algorithms, the need for many experts for manual monitoring and decision-making decreases, leading to resource and cost savings.

• Two-stage analysis method using optimization techniques: In the context of smart farming, a two-stage analysis method can be applied to optimize resource utilization and enhance productivity. The first stage involves transfer learning and fine-tuning. In the second stage, an ensemble learning technique was implemented.

• A comprehensive analysis method that lays the groundwork for future research: The proposed two-stage analysis method can serve as a thorough foundation for future research in smart farming. Researchers can further explore integrating AI and optimization techniques to enhance agricultural practices' efficiency and sustainability. Additionally, the collected data from such studies can contribute to developing larger-scale agricultural management systems and predictive models.

The remainder of the study is structured as follows: In the second section, a literature review summarizes relevant research in the field. The third section elaborates on the materials and methodology used in the study, explaining the data collection process and the techniques employed. The fourth section presents the experimental results, showcasing the findings obtained through the smart farming application and the two-stage analysis method. The fifth section comprises a discussion of the results, where the implications and significance of the findings are thoroughly examined. In the sixth and final section, the study is concluded, summarizing the key insights, highlighting the contributions to the field of smart farming, and suggesting potential avenues for future research.

2 Related works

Recently, there has been a growing focus on the automation of agricultural processes by utilizing AI techniques and robotic systems. ML has played a significant role in enhancing various agricultural tasks. The capacity for automatic feature extraction, inherent in DL methods, especially CNNs, has led to substantial advancements and achieved accuracy levels comparable to human performance in diverse agricultural applications. These applications include plant disease detection and classification, discrimination between weeds and crops, fruit counting, land cover classification, and crop/plant recognition [9]. DL, encompassing algorithms such as CNN, recurrent neural networks, and generative adversarial networks, has recently gained significant attention and application across various industries, including agriculture [10].

Among these studies, the utilization of image-processing algorithms for disease classification on different plant leaves has gained prominence. In the study conducted by Ref. [11], researchers employed two classifiers: CNN to differentiate between leaf spot and blight spot diseases and support vector machine (SVM) for red palm weevil (RPW) pest detection. The CNN algorithm's most basic function is the convolution process where features are extracted from the image using filters. The formula of the convolution process performed is shown in Equation (1) [12].(1) S(i,j)=(K*I)(i,j)=∑m∑nI(i+m,j+n)K(m,n)

where I represents an input feature map or image, K represents a convolutional filter (kernel), and S(i,j) denotes the output of the convolution operation at position (i,j) in the output feature map. In CNNs, convolution involves sliding the filter K over the input I, multiplying corresponding elements, and summing them up to produce a single value at each position (i,j) of the output feature map S. Each convolutional filter K is learned during the training process to extract specific features from the input data I. These features can represent various patterns or structures, depending on the task (e.g., edges, textures, or higher-level features in deeper layers). In multi-channel inputs (e.g., RGB images or multi-spectral data), the convolution operation is applied independently to each channel of I, using corresponding channels of K.

In CNN structure, the convolution layer is usually followed by the pooling layer. Here, the features extracted in the previous layer are selected and reduced according to the criteria and filter size. Although maximum pooling is generally used in the pooling layer, average pooling is also used in different studies. Maximum pooling formula is given Equation (2) and average pooling formula is given in Equation (3) [13].(2) MaxPooling(X)i,j,k=maxm,nXi·sx+m,j·sy+n,k

(3) AvgPooling(X)i,j,k=1fx·fy∑m,nXi·sx+m,j·sy+n,k

where X is the input, (i,j) are the indices of the output, k is the channel index, sx and sy are the stride values in the horizontal and vertical directions, respectively, and the pooling window is defined by the filter size fx and fx centered at the output index [13].

SVM aims to find a hyperplane that provides the widest margin between two classes. Mathematically, SVMs find the optimal hyperplane using the margin balance max(0,1−yi(wT.xi−b)) and generate solutions with Lagrange multipliers. To ensure that each data point is on the correct side of the margin, the SVM is obtained as in Equation (4) [14].(4) yi(wT.xi−b)≥1,i=1,2,…,n

where w is the normal vector of the hyperplane, x is the input value, y is the output, b is the amount of deviation of the hyperplane from the origin along the normal vector w.

Researchers used the Kaggle dataset in the study. The CNN and SVM algorithms achieved an accuracy ratio of 97.9 % and 92.8 %, respectively. A thermal images dataset for SVM was created, consisting of two classes: healthy palms and palms infected with RPW.

In the study by Ref. [15], principal component analysis (PCA) and artificial neural networks (ANNs) were sequentially used for palm leaf disease detection. PCA is part of the family of dimension reduction techniques, and it is valid, mainly if the data is multivariate and big and the variables are highly correlated. The aim is to identify the reduced size of the uncorrelated variable cluster that represents the original data with the minimum loss of information [16]. PCA has steps such as data standardization, computing the covariance matrix (cov(X,Y)=1n∑i=1n(x−x‾)(y−y‾) where X and Y are variables, x and y are members of X and Y, x‾ and y‾ are mean of X and Y, and n is the number of members) [17], eigenvalue decomposition (Av=λv where A is the matrix, v is a non-zero vector (the eigenvector), and λ is a scalar (the eigenvalue)), selecting principal components (vi (the eigenvectors)), and transforming the data. ANNs are computational models inspired by the structure and function of the human brain. They comprise a network of interconnected nodes, known as artificial neurons or nodes (y=f(∑(wi·xi)+b) where y is the output of the neuron, w is the weight, x is the input, and b is the bias), which process input data and produce an output based on learned patterns and connections within the network. The study employed a local dataset of 300 images and achieved sensitivity, specificity, and accuracy rates of 99.3 %, 100 %, and 99.67 %, respectively.

Another palm disease classification study was conducted by Ref. [18]. Researchers proposed a CNN model to classify four common diseases: Bacterial leaf blight, Brown spots, Leaf smut, white scale, and healthy leaves. Their model achieved an accuracy rate of 99.10 %, while VGG-16 and MobileNet achieved 99.35 % and 99.56 %, respectively. VGG16 consists of 16 layers (13 convolutional layers and 3 fully connected layers) and has approximately 138 million parameters. The convolutional layers use 3x3 filters with a stride of 1 and padding of 1, followed by max pooling layers with 2x2 windows and stride 2. The network maintains a consistent pattern of conv-conv-pool throughout, doubling the number of filters after each pooling layer (starting with 64 and ending with 512). The final layers consist of two fully connected layers with 4096 units each, followed by a Softmax output layer with 1000 units (for ImageNet classification) [19]. Unlike VGG16, MobileNet uses depthwise separable convolutions. A MobileNet architecture consists of a series of these depthwise separable convolutions, each followed by batch norm and ReLU activation. The network starts with a full convolution, followed by 13 depthwise separable convolution blocks. It ends with a global average pooling layer, a fully connected layer, and a Softmax classifier. MobileNet introduces two hyperparameters: width multiplier (α) and resolution multiplier (ρ), allowing further model shrinking [20]. The Date Palm and Leaf Disease 3 datasets were combined and utilized in the study.

In a study by Ref. [21], the researcher modified the ResNet model for Palm tree disease detection and classification. Two DL approaches, modified Residual Network (MResNet) and transfer learning of Inception ResNet, were implemented for Palm leaf disease classification. In the MResNet model, researchers replaced the 1x1 filters with 3x3 convolutional filters in all blocks for different stacks. Moreover, raising the size of 3x3 filters to 5x5 in the blocks will assume large local areas from images to extract many valuable features. Inception ResNet model integrates the Inception modules, which use multiple filter sizes and pooling operations in parallel, with residual connections from ResNet. The model typically has around 55.8 million parameters. It consists of a stem network followed by multiple Inception-ResNet blocks and reduction blocks. The Inception-ResNet blocks use a combination of 1x1, 3x3, and 5x5 convolutions, along with residual connections to facilitate better gradient flow. The reduction blocks use stridden convolutions and max pooling to reduce spatial dimensions. The network culminates in a global average pooling layer, followed by a dropout layer and a Softmax classifier [22]. The dataset used in the study was available on the Kaggle platform and consisted of images of white scales, brown spots, and healthy leaves. Among the implemented DL models, MResNet achieved higher validation accuracy and 100 % accuracy at different epochs.

There are different types of palm trees. Table 1 summarizes the DL studies conducted with palm trees in the literature to be suitable for this study's scope.Table 1 Previous studies for palm disease classification.

Table 1Study	Author	Year	Method/Model	Dataset	Performance Results (Acc.)	
[11]	H. Alaa, K. Waleed, M. Samir, M. Tarek, H. Sobeah, and M.A. Salam	2020	CNN and SVM	Kaggle dataset and a thermal images dataset	97.9 % and 92.8 %	
[15]	H. Hamdani, A. Septiarini, A. Sunyoto, S. Suyanto, and F. Utaminingrum	2021	PCA and ANN	A local dataset consisting of 300 images	99.67 %	
[18]	M. Abu-zanona, S. Elaiwat, S. Younis, N. Innab, and M.M. Kamruzzaman	2022	CNN	Date Palm and Leaf Disease 3 datasets	99.10 %	
[21]	M. Ahmed and A. Ahmed	2023	ResNet and MResNet	Kaggle dataset	100 %	
[23]	S.M.N. Nobel, M.A. Imran, N.Z. Bina, M.M. Kabir, M. Safran, S. Alfarhood, and M.F. Mridha	2024	ECA-Net with ResNet50 and DenseNet201	Image dataset of infected date palm leaves by Dubas insects	98.67 %	

The limitations of the studies obtained in the literature are discussed in the discussion section and evaluated in comparison with this study (see Section 5). Indeed, several studies have applied similar methodologies to different plant diseases in the domain of smart agriculture applications. Some of these studies include banana leaf disease classification [24], grape disease detection and identification [25], identification of disease and evaluation of bacteriosis in peach leaf [26], plant disease classification [[27], [28], [29], [30], [31]], and tomato leaf disease classification [32]. These studies demonstrate the versatility and applicability of similar methodologies in detecting and classifying diseases in various plant species, highlighting the potential of smart agriculture technologies to address disease-related challenges across different crops.

The common aspect of these studies mentioned in the literature is the application of DL algorithms with both novel and pre-trained models. Numerous ML and DL models have been developed in the domain of leaf disease identification. However, challenges persist in this area. Many research studies have employed pre-trained models like GoogLeNet, AlexNet, VGGNet, and ResNet, along with training data from ImageNet, yielding higher accuracy than other models. While the commonly used Kaggle dataset is suitable for training CNN models, some researchers have worked with fewer original images, resorting to artificial expansion techniques such as data augmentation. These techniques involve rotation, translation, cropping, grayscale conversion, and more to enhance the model's performance. The main challenge lies in the limited availability of training data, which can result in models suffering from under-fitting and an inability to predict various leaf diseases accurately. Augmentation and annotation techniques can help address this issue by compensating for the small dataset size and hardware constraints [33].

3 Materials and Methods

This section describes the materials used in the study and the applied methodologies. The study's methodology is structured in two stages. In the first stage, transfer learning and fine-tuning processes are performed. In this stage, pre-trained models are utilized as a starting point, and then the models are fine-tuned on the specific palm leaf disease dataset. The fine-tuning process enables the model to adapt and learn features relevant to the target task of palm leaf disease classification.

In the second stage, ensemble models are constructed based on the results obtained from the first stage. Ensemble learning is used to combine multiple individual models, leveraging their diverse strengths to enhance overall prediction performance. Ensemble models can be created using different techniques, such as bagging or boosting.

Fig. 1 presents the overall methodology of the study in a block diagram. It showcases the two-stage process of transfer learning and fine-tuning in the first stage, followed by the testing of ensemble models formed in the second stage. This two-stage approach utilizes transfer and ensemble learning for robust and accurate classification of palm leaf diseases.Fig. 1 Block diagram of the study.

Fig. 1

Study data was subjected to various processes to overcome the limitations found in previous studies. Study data was obtained from two different sources: Mendeley and Kaggle.

3.1 Datasets

To expand the scope and generalizability, two different datasets are combined in this study. The utilization of both datasets in the study has addressed the limitations observed in previous research. This is because the dataset contains both disease and pest-related images. Furthermore, it includes data related to different stages of the disease. The number of images between classes in the datasets varies and a class imbalance problem arises. Data augmentation techniques have been employed during the pre-processing stage to overcome this imbalance and enhance the data. This approach has allowed the study to mitigate the variability in class sizes and create a more balanced dataset, leading to more robust and accurate results in the classification process. Fig. 2 presents the sample images of the classes used in the study.Fig. 2 Sample images from datasets. (a) Brown spots, (b) Bug, (c) Mixed (Dubas), (d) Honey, (e) White Scale Phase-1, (f) White Scale Phase-2, (g) White Scale Phase-3, (h) Healthy. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)

Fig. 2

Fig. 2(a) shows a sample image of the Brown spots class, while Fig. 2(b) shows the image of the Bug class. Fig. 2(c) shows an image example of the Mixed (Dubas) class, while Fig. 2(d) shows an image example of the Honey class. Fig. 2(e), (f), and 2(g) show sample images of Phase-1, Phase-2 and Phase-3 types of White Scale class respectively. Finally, the sample image of the Healthy class is shown in Fig. 2(h).

3.1.1 Mendeley image dataset

The infected date palm leaves by Dubas insects dataset [34] is provided by Abdullah Mazin and Haider Almayaly University of Karbala. They used drone cameras and captured a dataset of 3000 palm leaf images. Then, they categorized these images into four groups based on the health status of the leaves and the presence of insects: healthy, infected with bugs only, infected with honeydew only, and infected with both bugs and honeydew. The images of leaves infected with insects depict a range of insect life cycle stages, from the third generation of nymphs to the adult stage in the fifth nymph stage. The non-bug categories have 800 images, while the bug category has 600 images. This dataset holds significant value in evaluating the severity of infestation, estimating insect populations, and determining the extent of damage. A survey of agricultural lands infested with Dubas insects was conducted to obtain images of infected palm leaves. The survey covered a period of the insect life cycle, including spring and autumn generations. With the help of an agricultural guide, infected palms were identified during land scanning. A drone captured images of the leaves from 1 to 2 m. Small bugs with eggs sometimes appeared on the leaves in autumn, while apparent bugs appeared in spring. Honeydew was observed on the leaves at the end of spring and summer. Researchers screened the images for noise, shadow, and dust. All imaging was performed with the consent and knowledge of the orchard owners. This dataset holds valuable information for effectively monitoring and managing infestations in agricultural lands [34].

3.1.2 Kaggle date palm dataset

The Kaggle dataset involves images of healthy leaves, brown spot disease, and white-scale infection [35]. The dataset used in the study consists of 470 images of the brown spots class, 1203 images of the healthy class, and 958 images of the white scale class. Additionally, within the white scale class, there are three subdivisions: Phase 1, Phase 2, and Phase 3. The white scale class contains 321 images of Phase 1, 350 images of Phase 2, and 287 images of Phase 3. To improve the scope of the study, each phase of the white scale class is evaluated individually.

3.2 Data augmentation

Data augmentation is a technique used to artificially increase the size of a dataset by creating new data points from existing ones. This is done by applying various transformations to the existing data points, such as rotation, flipping, scaling, translation, color space transformations, kernel filters, image blending, random deletion, cropping, and fading. Data augmentation aims to create a more diverse and representative dataset that can be used to train machine learning models. By applying various transformations to the data, we can help to ensure that the models do not overfit the training data and that they can generalize to new data [36,37].

Rotation augmentation was applied to augment the data. Images belonging to each class were rotated at different angles to balance the images among the classes in the dataset.

Different augmentation techniques can be applied to different images. For example, contrast change or rotation can be applied to a medical image. Similarly, the most appropriate augmentation technique for the dataset in this study is rotation. Leaves can grow at very different angles. During the imaging process, the leaves of a tree can be recorded at very different angles relative to each other. For this reason, the images in this study were augmented with the rotation technique. The images in the dataset were first rotated by 45-degree angles. To ensure equality between the images in the classes, the rotation process continued by increasing 45-degree angles until the images of all classes were 2160. Fig. 3 presents examples of images obtained after data augmentation.Fig. 3 Sample images from augmented dataset. (a) Brown spots, (b) Bug, (c) Mixed (Dubas), (d) Honey, (e) White Scale Phase-1, (f) White Scale Phase-2, (g) White Scale Phase-3, (h) Healthy. (For interpretation of the references to color in this figure legend, the reader is referred to the Web version of this article.)

Fig. 3

After data augmentation process, Fig. 3(a) shows a sample rotated image of the Brown spots class, while Fig. 3(b) shows the rotated sample image of the Bug class. Fig. 3(c) shows a rotated image example of the Mixed (Dubas) class, while Fig. 3(d) shows a rotated image example of the Honey class. Fig. 3(e), (f), and 3(g) show sample rotated images of Phase-1, Phase-2 and Phase-3 types of White Scale class respectively. Finally, the sample rotated image of the Healthy class is shown in Fig. 3(h).

3.3 Data pre-processing

After data augmentation, the data pre-processing stage involved splitting the dataset into three groups: train, validate, and test data. 10 % of the augmented images were set aside for the test process, while the remaining 90 % were divided into two groups: 80 % for the train set and 20 % for the validation set. All images in the study were resized to a resolution of 224x224 pixels. The images were kept in the red-green-blue (RGB) color space, and no color space transformation was applied. Considering all this information, details about the study dataset are presented in Table 2, providing an overview of the data split and pre-processing steps.Table 2 Information about the dataset.

Table 2Class	Original Data	Augmented Data	Train Data	Validation Data	Test Data	
Brown Spots	470	2160	1556	388	216	
Bug	600	2160	1556	388	216	
Mixed (Dubas)	800	2160	1556	388	216	
Honey	800	2160	1556	388	216	
White Scale P1	321	2160	1556	388	216	
White Scale P2	350	2160	1556	388	216	
White Scale P3	287	2160	1556	388	216	
Healthy	2003	2160	1556	388	216	
Total	5631	17280	12448	3104	1728	

3.4 Transfer learning and fine tuning

DL models have gained popularity as a powerful ML approach to tackle many AI problems. Their success can be attributed to their ability to adapt and generalize to various tasks.

3.4.1 Transfer learning

One essential technique that leverages this adaptability is transfer learning. In transfer learning, a pre-trained model that has already successfully solved a previous problem is reused and applied to a different situation. By utilizing the knowledge and weights gained from solving the initial problem, the model can be fine-tuned or retrained to address the new task. This approach offers several advantages [38,39].• It can save time and resources, as it does not require training a new model from scratch.

• It can improve the model's performance on the new problem, as the model will already have some knowledge of the relevant features.

• It can solve problems with limited data, as the model can be transferred from a problem with a large dataset.

In this study, transfer learning was implemented using the TensorFlow Keras Applications module [40]. Within this module, various pre-trained model architectures are available, including ConvNeXt, DenseNet, EfficientNet, EfficientNet_V2, Inception_ResNet_V2, Inception_V3, MobileNet, MobileNet_V2, MobileNet_V3, NASNet, RegNet, ResNet, ResNet_RS, ResNet_V2, VGG16, VGG19, and Xception. For each module, there are different pre-trained models with varying numbers of layers (functions/models). In the study, the base model of each module was transferred and used. By doing so, the successful models served as a base for future analyses, where more complex models with additional layers from the same successful module could be explored. The models used in the study are listed in Table 3, providing an overview of the different pre-trained architectures employed in the transfer learning process.Table 3 Modules and models used in the study.

Table 3N	Module	Model	N	Module	Model	
1	xception	Xception	10	inception_v3	InceptionV3	
2	vgg16	VGG16	11	inception_resnet_v2	InceptionResNetV2	
3	resnet_rs	ResNetRS50	12	densenet	DenseNet121	
4	resnet	ResNet50	13	mobilenet_v2	MobileNetV2	
5	resnet_v2	ResNet50V2	14	efficientnet	EfficientNetB0	
6	regnet	RegNetX002	15	efficientnet_v2	EfficientNetV2B0	
7	regnet	RegNetY002	16	convnext	ConvNeXtTiny	
8	nasnet	NASNetMobile	17	mobilenet_v3	MobileNetV3Small	
9	mobilenet	MobileNet				

The models presented in Table 3 can be briefly explained as follows. The Xception model extends the Inception architecture using deep discrete convolutions developed for DL. VGG16 is a multilayer DL model named after its 16-layer depth, each equipped with 3x3 convolution filters. ResNet50 uses residual connections to alleviate difficulties in training depending on the depth. It has 50 layers and is widely used in classification tasks. ResNet50V2 is an updated version that provides better training and accuracy (in some tasks) than the original ResNet50. It includes batching normalization and changes in the ordering of activation functions. ResNetRS50 is a variation of the ResNet architecture and facilitates the training of deeper networks using Residual Structure. RegNetX002 is part of the RegNet family and aims to improve performance by optimizing the width and depth of the network structure. It uses extended convolution filters. RegNetY002, like RegNetX002, has extended convolution filters and an optimized structure. The difference is that it uses more complex block structures and SE (Squeeze-and-Excitation) modules. NASNetMobile is a model that Google has optimized using Neural Architecture Search (NAS). MobileNet is a lightweight and efficient CNN architecture designed for mobile and embedded devices. It uses deep discrete convolutions to reduce the number of parameters and computational cost. InceptionV3 is an expanding DL model that uses multiple convolution filters and variable filter sizes. InceptionResNetV2 combines the Inception architecture with residual connections. DenseNet121 is based on the principle that each layer receives input from all previous layers. These dense connections increase parameter efficiency and improve gradient propagation. MobileNetV2 is an update of MobileNet and uses inverted residual structures. EfficientNetB0 optimizes performance by scaling model sizes in width, depth, and resolution. EfficientNetV2B0 is an enhanced version of EfficientNet that uses a more optimized structure and techniques for faster and more efficient training. ConvNeXtTiny is part of the ConvNeXt family and aims to improve performance using modern convolutional network techniques. MobileNetV3Small is a model optimized for mobile devices. Its narrow block structures and SE modules provide lower computational costs and high efficiency [19,20,[41], [42], [43], [44], [45], [46], [47], [48], [49], [50], [51]].

Deep Neural Networks (DNNs) are artificial neural network architecture characterized by multiple hidden layers. These layers progressively extract features from the input data, allowing DNNs to model complex relationships between features and outputs. Each layer performs a non-linear transformation on the weighted sum of its inputs from the previous layer, progressively extracting higher-order features. Each layer's choice of activation function plays a crucial role in determining the network's ability to model non-linear relationships. The specific activation function chosen depends on the problem type, data structure, and network architecture. CNNs are a specialized type of DNN architecture particularly well-suited for image classification tasks. Convolutional layers apply learnable filters (kernels) of small dimensions (e.g.,3x3, 5x5) across the input image, capturing local features. Pooling layers then down sample the output of these convolutions, typically using techniques like max pooling, which selects the maximum value within a defined window. Dropout is a regularization technique commonly employed in DNNs to prevent overfitting. During training, a random subset of neurons is temporarily dropped with a predefined probability.

3.4.2 Fine tuning

The performance of DNN models is heavily influenced by hyperparameters such as the number of hidden layers, neuron count per layer, learning rate, and activation function choice. The models used in this study were initially successful in the ImageNet competition and have been adapted to the specific conditions of that competition. As a result, the classification layers of these models were fine-tuned to be tailored to the current research problem. Hyperparameter optimization techniques involve systematically evaluating different configurations and selecting the combination that yields the best performance on a validation dataset. This iterative process is crucial for achieving optimal model performance. Fine-tuning parameters were determined through an optimization process using the test-retest method in this study. The fine-tuning process for the classification layers in the study is as follows: To optimize the loss value, dense layers with 1000, 512, 256, and 128 neurons were sequentially incorporated into the classification stage. Additionally, dropout layers with a rate of 12 were inserted between these dense layers. The activation function used in the dense layers was rectified linear unit (ReLU). The number of neurons in the output layer of the models was determined as 8, which is the number of classes in the dataset, and the Softmax activation function was used in this layer. The study used Adam as the optimizer, and the learning rate was determined to be 1x10−5. For the multi-classification problem, the loss value was chosen as categorical_crossentropy. Epoch numbers were chosen as 30, and the batch size values were determined as 20.

By fine-tuning the classification layers, the models were better adapted to the specific problem of palm disease classification, leading to improved performance and accuracy in the study. The ReLU activation function is a pivotal landmark in the DL revolution and is widely used in neural networks. Its simplicity and effectiveness have contributed significantly to the success of DL models. ReLU works by thresholding values at 0, returning 0 when the input (x) is negative, and outputting the input value (x) when it is non-negative (x ≥ 0). The ReLU activation function has an output range from 0 to positive infinity. This characteristic makes the function computationally efficient, facilitates the backpropagation process during training, and allows the DNN to learn more robust features and perform well in various tasks. The calculation of ReLU is shown in Equation (5).(5) f(x)=max(0,x)

Softmax function squashes a vector in the range (0, 1), and all the resulting elements add up to 1. It is applied to the output scores s. As elements represent a class, they can be interpreted as class probabilities. The Softmax function cannot be used independently for each si since it depends on all s elements. For a given class si, the Softmax function can be computed as Equation (6) [52]:(6) f(s)i=esi∑iCesj

where esj are the scores inferred by the net for each class in C.

Categorical Cross-Entropy loss is also called Softmax Loss because it is a Softmax activation plus a Cross-Entropy loss. Using this loss, a CNN was trained to output a probability over the C classes for each image. The calculation of categorical cross-entropy is shown in Equation (7) [52].s→Softmax→CrossEntropyLoss

(7) f(s)i=esi∑iCesj→CE=−∑iCtilog(f(s)i)

In multi-class classification, the labels are typically represented using one-hot encoding. In this encoding scheme, each target label (t) is a vector of zeros, except for the positive class (Cp), which has a value of 1. Consequently, only the positive class (Cp) contributes to the loss calculation. To express this mathematically (Equation (8)):(8) CE=−log(esp∑iCesj)

where sp is the model score for the positive class.

When Softmax loss is used in a multi-label problem, it contains an element for each positive class. The positive and negative classes can be derived after some calculations, shown as Equation (9), Equation (10), and Equation (11).(9) CE=1M∑pM−log(esp∑iCesj)

(10) ∂∂spi(1M∑pM−log(esp∑iCesj))=1M((espi∑iCesj−1)+(M−1)espi∑iCesj)

(11) ∂∂sn(1M∑pM−log(esn∑iCesj))=esn∑iCesj

where sn is the score of any negative class in C that is different from Cp, M are the positive classes of a sample, each sp in M is the model score for each positive class, and spi is the score of any positive class [52].

3.5 Deep ensemble learning

Ensemble learning uses multiple algorithms to obtain better predictive performance than any single one of its constituent algorithms could. With the growing popularity of DL technologies, researchers have started to ensemble these technologies for various purposes [53]. The deep ensemble learning technique was applied by creating different numbers of combinations according to varying criteria among the deep transfer learning models used in the first stage.

The DeepStack module was used to implement the Dirichlet Ensemble technique. DeepStack is a Python module designed to construct DL ensembles. It was initially built on the Keras framework and is distributed under the MIT license. The Dirichlet Ensemble technique involves weighting the ensemble members to optimize a specific metric or score based on a validation dataset. Ensemble weights are optimized using a randomized search method based on the Dirichlet distribution. During the fitting process of the Dirichlet ensemble model, the module calculates the ensemble weights by optimizing the Area Under the Curve (AUC) Binary Classification Metric. This optimization is achieved through randomized search utilizing the properties of the Dirichlet distribution [54,55].

3.6 Evaluation metrics

Especially in medical research, the proper selection and use of evaluation metrics play a vital role in assessing the performance of a study. Therefore, not only accuracy metrics but also precision, recall (sensitivity), and f1-score of the models were evaluated in the study. Additionally, Receiver Operating Characteristic (ROC) and AUC values were obtained, and the performance of the study was assessed based on these metrics within the ensemble learning technique. True positive (TP), true negative (TN), false positive (FP), and false negative (FN) values of each class were used for evaluation.

The accuracy rate represents the ratio of correct predictions to all examples, and a high value indicates that the model predicted samples in each class accurately overall. Precision is used to measure the accuracy of positive predictions and shows the ability of a classifier to avoid false positive errors. A high precision value indicates that the model is precise and selective in its positive predictions, minimizing the false positive errors. Recall is used to measure the ability of a model to identify positive instances correctly. It quantifies the proportion of correctly predicted positive instances out of all actual positive instances in the dataset. Recall represents the model's ability to avoid false negatives and capture all positive instances. A high recall value indicates that the model is sensitive and effective in identifying positive instances. It is also known as true positive rate (TPR). The F1 score combines precision and recall into a single value and provides a balanced measure of a model's performance by considering the model's ability to avoid false positives and capture all positive instances. The F1 score is calculated using the harmonic mean of precision and recall. The F1 score provides a more reliable evaluation of the classifier's performance by considering false positives and negatives. A higher F1 score indicates a better trade-off between precision and recall, suggesting that the model has achieved a good balance in correctly predicting positive instances while minimizing false positives and false negatives [55]. The calculations of these metrics are given in Equation (12), Equation (13), Equation (14), and Equation (15), respectively.(12) Accuracy=TP+TNTP+TN+FP+FN

(13) Precision=TPTP+FP

(14) Recall(Sensitivity)=TPTP+FN

(15) F1score=2xPrecisionxRecallPrecision+Recall

The ROC curve is a graphical tool used to evaluate the performance of a model by illustrating the trade-off between sensitivity (true positive rate) and specificity (true negative rate) at various threshold values. It helps visualize the model's classification performance changes with different threshold settings. The AUC is a single scalar metric derived from the ROC curve. It quantifies the overall performance of the model in a binary classification task. The AUC represents the probability that a randomly chosen positive instance will be ranked higher by the model than a randomly chosen negative instance. In other words, it measures the classifier's ability to distinguish between positive and negative instances. A higher AUC value indicates better discrimination power and overall performance of the classifier. A perfect classifier has an AUC value of 1, indicating that it can ideally separate positive and negative instances. On the other hand, a random classifier will have an AUC value of 0.5, as it performs no better than random chance. The ROC curve and AUC provide valuable insights into the model's performance across different threshold settings, allowing us to understand its ability to discriminate between positive and negative instances regardless of the specific threshold chosen for classification. This makes them valuable tools for evaluating and comparing different classifiers in binary classification tasks.

3.7 Experimental Setup

This study utilized a MacBook M2 Pro with 16 GB RAM, a 16-core GPU, and a 10-core CPU with a 512 SSD for all experimental procedures. The research environment was established using the Anaconda platform and JupyterLab. Python version 3.8.16 served as the primary programming language, complemented by TensorFlow library version 2.12.0 for DL functionalities. Data visualization and DL tasks were facilitated by Matplotlib (version 3.7.1) and Sklearn (version 1.2.2) libraries, respectively. Additionally, the study incorporated ensemble learning techniques through the deepstack module (version 0.0.9). All experiments were conducted within this computational framework to ensure consistency and reproducibility of results.

4 Results

This section presents the analysis results. As mentioned in the Materials and Methods (see Section 3), a base model was selected for each of the numerous pre-trained DNNs available under the Tensorflow Keras Applications modules. This selection is intended to make inferences about the modules' performance and provide a basis for future studies.

During the training process, feature extraction is performed between the architectural layers of the CNN algorithm, where the models learn the features of each class in the dataset used in the study. This feature extraction process occurs in the representation layers of a representative model across different channels (RGB: 0-1-2), as illustrated in Fig. 4.Fig. 4 Feature extraction process.

Fig. 4

As shown in Fig. 4, the input image progresses through the RGB channels, and different filters are applied to learn the features of the input image by the algorithm. Furthermore, convolutional operations occur naturally within the CNN layers during feature extraction. As a result, the input image undergoes convolution across the layers, leading to smaller image dimensions between the layers based on the learning logic of the representation. These convolutional operations are performed using multiple filters and techniques between the layers. A representative model is selected from the used models, and the progression of convolution and image transformation across different layers is presented in Fig. 5.Fig. 5 Changes in the input image after feature extraction and convolution in different CNN layers (a) Input image, (b) Layer 5, (c) Layer 25, (d) Layer 75, (e) Layer 125, and (f) Layer 170.

Fig. 5

As shown in Fig. 5, the input image with a resolution of 224x224 pixels is given to the model (Fig. 5(a)), it is reduced to 110x110 at the 5th layer (Fig. 5(b)) and at the 25th layer, it is reduced to 55x55 pixels (Fig. 5(c)). While processing between layers, the resolution of the image continued to decrease and while it was 26x26 at layer 75 as shown in Fig. 5(d), the image was reduced to 13x13 resolution at layer 125 (Fig. 5(e)). In the 170th layer, the image resolution was obtained as 6x6 pixels as shown in Fig. 5(f). Due to the varying layer structures in different models, the resolution ratios differ across the layers for each model used in this study.

4.1 Experimental results of stage 1

The models listed in Table 2 were first trained and validated with training and validation data in the first stage of the study, which is the transfer learning and fine-tuning stage. Subsequently, the obtained weights were used to evaluate the models on the test data, and the resulting scores were recorded. These results are presented in Table 4.Table 4 Transfer learning and fine-tuning results.

Table 4N	Model	Train Acc.	Train Loss	Val. Acc.	Val. Loss	Test Acc.	
1	Xception	99.74	0.0067	77.48	1.83	90.39	
2	VGG16	97.04	0.1013	76.68	1.90	89.53	
3	ResNetRS50	99.79	0.0053	84.34	1.27	92.30	
4	ResNet50	99.78	0.0056	79.96	1.96	92.42	
5	ResNet50V2	99.78	0.0057	77.58	2.05	90.11	
6	RegNetX002	98.82	0.038	78.54	1.13	90.05	
7	RegNetY002	99.79	0.006	75.42	1.52	88.66	
8	NASNetMobile	99.55	0.015	75.45	1.49	89.00	
9	InceptionV3	99.74	0.0067	77.09	1.46	87.85	
10	InceptionResNetV2	99.59	0.014	75.74	1.44	88.95	
11	DenseNet121	99.77	0.0073	80.79	1.49	92.01	
12	MobileNet	99.77	0.0055	80.51	1.49	91.72	
13	MobileNetV2	99.79	0.0048	80.48	1.68	92.48	
14	MobileNetV3Small	99.78	0.0055	75.03	1.61	89.76	
15	EfficientNetB0	99.77	0.0056	79.83	1.49	91.90	
16	EfficientNetV2B0	99.78	0.0055	78.74	1.55	90.68	
17	ConvNeXtTiny	99.81	0.0059	81.54	1.37	90.91	

Table 4 shows that all models, except VGG16, have a training accuracy of approximately 99 %. The train loss values for these models are also very low. However, the validation accuracy of the models ranges from ∼75 % to ∼85 %. Upon evaluating the models on unseen data, the test accuracy values are found to be in the range of ∼87 %–∼92 %. Fig. 6 presents the ranking of the models based on their test accuracy.Fig. 6 Test accuracies of models.

Fig. 6

As shown in Fig. 6, the MobileNetV2 model achieved the highest test accuracy with a value of 92.48 %. The second-highest accuracy was obtained by the ResNet50 model, which reached 92.42 %. Following closely, the ResNetRS50 model achieved the third-highest test accuracy of 92.30 %. Another successful model within the DenseNet module, DenseNet121, achieved an accuracy of 92.01 %. The training and validation loss graphs obtained during the training of these models are presented in Fig. 7, respectively.Fig. 7 Training and validation loss values.

Fig. 7

As shown in Fig. 7, all models have achieved training loss values close to zero, while validation loss values range between 1 and 2. Although the difference between these two data groups is as expected, future studies can be conducted further to reduce the validation loss values of the models.

In the initial phase of the study, confusion matrix results for the test process conducted with the saved weights during training have been obtained. The precision, recall, and F1 score values for each model are presented in Table 5.Table 5 Confusion matrices of models.

Table 5Class	Brown Spots	Bug	Mixed (Dubas)	Healthy	Honey	White Scale P1	White Scale P2	White Scale P3	
Model	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	Pre.	Rec.	F1 S.	
Xception	1.00	0.96	0.98	0.89	0.78	0.83	0.76	0.78	0.77	0.94	1.00	0.97	0.77	0.81	0.79	0.95	0.94	0.94	0.95	0.98	0.96	1.00	1.00	1.00	
VGG16	1.00	0.95	0.97	0.91	0.80	0.85	0.73	0.79	0.76	0.90	1.00	0.95	0.79	0.73	0.75	0.94	0.92	0.93	0.92	0.98	0.95	1.00	1.00	1.00	
ResNetRS50	1.00	0.95	0.97	0.91	0.88	0.89	0.85	0.76	0.80	0.93	1.00	0.96	0.78	0.86	0.82	0.98	0.95	0.96	0.96	0.98	0.97	0.99	1.00	1.00	
ResNet50	0.99	0.96	0.97	0.91	0.87	0.89	0.82	0.84	0.83	0.97	1.00	0.98	0.82	0.82	0.82	0.96	0.93	0.94	0.93	0.98	0.96	1.00	1.00	1.00	
ResNet50V2	0.99	0.94	0.97	0.88	0.83	0.86	0.82	0.75	0.78	0.91	1.00	0.95	0.77	0.82	0.80	0.93	0.92	0.93	0.94	0.97	0.95	0.99	1.00	0.99	
RegNetX002	1.00	0.95	0.97	0.93	0.86	0.89	0.75	0.75	0.75	0.93	1.00	0.96	0.75	0.76	0.76	0.93	0.91	0.92	0.92	0.98	0.95	1.00	1.00	1.00	
RegNetY002	0.99	0.93	0.96	0.89	0.76	0.82	0.71	0.80	0.75	0.94	1.00	0.97	0.72	0.70	0.71	0.92	0.93	0.93	0.94	0.97	0.95	1.00	1.00	1.00	
NASNetMobile	0.99	0.92	0.95	0.92	0.84	0.88	0.73	0.78	0.76	0.96	0.99	0.97	0.74	0.73	0.74	0.89	0.93	0.91	0.93	0.94	0.94	0.97	0.99	0.98	
InceptionV3	0.99	0.95	0.97	0.89	0.80	0.84	0.69	0.76	0.73	0.94	0.97	0.95	0.74	0.74	0.74	0.92	0.89	0.90	0.91	0.93	0.92	0.96	0.99	0.98	
InceptionResNetV2	1.00	0.88	0.93	0.87	0.83	0.85	0.81	0.73	0.77	0.93	0.99	0.96	0.75	0.83	0.79	0.87	0.89	0.88	0.90	0.98	0.94	1.00	0.99	0.99	
DenseNet121	1.00	0.97	0.98	0.88	0.90	0.89	0.82	0.81	0.82	0.96	1.00	0.98	0.81	0.77	0.79	0.97	0.92	0.94	0.93	0.99	0.96	1.00	1.00	1.00	
MobileNet	1.00	0.96	0.98	0.90	0.83	0.86	0.80	0.84	0.82	0.93	1.00	0.96	0.83	0.80	0.81	0.96	0.93	0.94	0.93	0.99	0.96	1.00	1.00	1.00	
MobileNetV2	1.00	1.00	1.00	0.94	0.82	0.88	0.81	0.85	0.83	0.94	1.00	0.97	0.80	0.81	0.80	0.99	0.93	0.96	0.94	1.00	0.97	1.00	1.00	1.00	
MobileNetV3Small	0.99	0.94	0.96	0.90	0.84	0.87	0.75	0.78	0.77	0.87	1.00	0.93	0.79	0.73	0.76	0.97	0.91	0.94	0.93	0.98	0.95	1.00	1.00	1.00	
EfficientNetB0	1.00	0.94	0.97	0.90	0.89	0.90	0.86	0.78	0.82	0.93	1.00	0.96	0.82	0.87	0.84	0.96	0.89	0.93	0.91	0.99	0.94	0.99	1.00	1.00	
EfficientNetV2B0	1.00	0.93	0.96	0.90	0.83	0.86	0.78	0.80	0.79	0.92	1.00	0.96	0.82	0.80	0.81	0.93	0.90	0.92	0.91	0.99	0.95	0.99	1.00	1.00	
ConvNeXtTiny	0.99	0.94	0.97	0.90	0.85	0.88	0.79	0.83	0.81	0.91	0.98	0.94	0.84	0.79	0.81	0.95	0.89	0.92	0.91	0.98	0.95	0.99	1.00	1.00	
Support	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	216	

When Table 5 is analyzed, it is seen that ten models reached 100 % and seven models reached 99 % Precision for the Brown Spots class. The top 5 most successful models for the Recall value of this class are MobileNetV2 (100 %), DenseNet121 (97 %), Xception (96 %), ResNet50 (96 %), and MobileNet (96 %). In the F1 score value, MobileNetV2 (100 %), Xception (98 %), DenseNet121 (98 %), and MobileNet (98 %) models were the most successful models. Among the Precision values of the Bug class, the most successful models were MobileNetV2 with 94 %, RegNetX002 with 93 %, and NASNetMobile with 92 %. When the Recall results of this class are analyzed, it is seen that there are performance differences between the models. The most successful model was DenseNet121 with 90 %, followed by EfficientNetB0 and ResNetRS50 with 89 % and 88 %, respectively. The values of the Mixed (Dubas) class were lower than the other classes. This class's most successful models in Precision values were EfficientNetB0 (86 %) and ResNetRs50 (85 %). Three models achieved 82 %, and these are ResNet50, ResNet50V2, and DenseNet121 models. Recall values for this class are similarly slightly lower than those for the other classes. The most successful model in this metric was MobileNetV2 with 85 %, while the other two that achieved 84 % were MobileNet and ResNet50. The models that performed well in the F1 score of this class were MobileNetV2 and ResNet50, respectively. Both of them achieved 83 %. The Healthy class's Precision, Recall, and F1 scores are generally high. The most successful models in Precision are ResNet50 with 97 % and NasNetMobile and DenseNet121 with 96 %. In the Recall value, all models except four achieved 100 % success, and these four models achieved high rates between 97% and 99 %. The two models with an F1 score of 98 % are ResNet50 and DenseNet121. The most successful models of the Precision value of the Honey class are ConvNeXtTiny with 84 % and MobileNet with 83 %, while the three models with 82 % are EfficientNetB0, EfficientNetV2B0, and ResNet50. The most successful models of this class in Recall are EfficientNetB0 (87 %) and ResNetRS50 (86 %). In the F1 score, the EfficientNetB0 model is at the forefront with 84 %. This model is followed by ResNet50 and ResNetRS50 models with 82 %. In White Scale classes, the models obtained the most successful results in the P3, P2, and P1 classes, respectively. In White Scale P3 class, almost all models achieved 100 % and 99 % for all metrics. Only NasNetMobile and InceptionV3 models obtained results between 96 % and 99 %, which are also considered very high. In the White Scale P2 class, the most successful models in the Precision value were ResNetRS50 with 96 % and Xception with 95 %. All models achieved 90 % and above. MobileNetV2 achieved 100 % in the Recall value, while many other models achieved high rates, such as 99 % and 98 %. In the F1 score value, ResNetRS50 and MobileNetV2 were the two models that achieved 97 %. Finally, in the White Scale P1 class, the model with the highest Precision value was MobileNetV2, with 99 %, while the second most successful model was ResNetRS50, with 98 %. The other two models, with 97 %, were MobileNetV3Small and DenseNet121. The most successful models of this class with Recall values were ResNetRS50 (95 %) and Xception (94 %), respectively. The models with successful F1 scores are ResNetRS50 and MobileNetV2, which have 96 % success. As can be seen from these explanations, the models whose confusion matrix results are successful for all classes are among certain module families. These are the ResNet family, MobileNet family, and DenseNet family. Other models rarely achieved success in some values for some classes.

Table 5 shows that the initial findings indicate that the "White Scale Phase-3″ disease type is the most accurately classified by all models. On the other hand, the "Honey" class appears to be the most challenging class for all models in general.

In addition to these findings, the confusion matrix results obtained in the study were primarily evaluated based on the Recall (Sensitivity) metric, also known as the true positive rate, focusing on disease detection for each class. Subsequently, F1 score values, the harmonic mean of precision and recall, were evaluated. Based on these evaluations, the recall values for each class were considered first (in cases of equal ratios among models, the F1 score was considered, and in the case of further equality, the Test accuracy). The most successful models of this evaluation for each disease class are as follows.• For Brown spots detection: MobileNetV2

• For Bug detection: DenseNet121

• For Mixed (Dubas) disease detection: MobileNetV2

• For Healthy palm class detection: ResNet50

• For Honey detection: EfficientNetB0

• For White Scale Phase-1 detection: ResNetRS50

• For White Scale Phase-2 detection: MobileNetV2

• For White Scale Phase-3 detection: MobileNetV2

Based on the F1 score values and considering the subsequent criteria of Sensitivity (Recall) and Test accuracy in case of equality, the most successful models for each disease class are as follows.• For Brown spots detection: MobileNetV2

• For Bug detection: EfficientNetB0

• For Mixed (Dubas) disease detection: ResNet50

• For Healthy palm class detection: ResNet50

• For Honey detection: EfficientNetB0

• For White Scale Phase-1 detection: ResNetRS50

• For White Scale Phase-2 detection: MobileNetV2

• For White Scale Phase-3 detection: MobileNetV2

These models have shown superior performance in identifying specific disease classes, which can be valuable for effective disease management in smart agriculture. A high true positive rate (Recall (Sensitivity)) indicates the classifier's effectiveness in correctly identifying positive instances with a low rate of falsely classifying them as negative. This sensitivity allows the model to capture many positive instances, minimizing false negatives. In scenarios where the cost of false negatives is significant, such as disease diagnoses, recall takes precedence over precision, which measures the ability to avoid false positives. The F1 score is a dependable metric for evaluating the classifier's performance, as it considers both false positives and false negatives. A higher F1 score indicates that the model has struck a better balance between precision and recall, successfully predicting positive instances while minimizing false positives and negatives. These findings demonstrate the effectiveness of different models for specific disease classes, which can be valuable for future research and practical applications in smart agriculture.

The MobileNetV2 model exhibited superiority over the other models as its performance outshined them in many classes. It achieved better results than the other models in half of all classes, and its performance in the remaining classes was also exceptional. The confusion matrix graphics provide visual representations of the models' classification performance, highlighting MobileNetV2's overall excellence compared to the other models. Fig. 8 presents the comparison of five models based on their test accuracy and confusion matrix results.Fig. 8 Confusion matrices of most successful models (a) MobileNetV2, (b) ResNet50, (c) ResNetRS50, (d) DenseNet121, (e) EfficientNetB0.

Fig. 8

The confusion matrix plots shown in Fig. 8 belong to the five models that achieved the best test accuracy and classification performance. Fig. 8(a) shows the results of the MobileNetV2 model, while Fig. 8(b) shows the results of the ResNet50 model. Fig. 8(c) shows the confusion matrix results of the ResNetRS50 model. The confusion matrix results of DenseNet121 and EfficientNetB0 models are shown in Fig. 8(d) and (e), respectively. These graphs show the correct and incorrect model classifications for each disease class. Confusion matrix is an essential tool for analyzing classification performance and understanding which classes the model works best on. Examining these graphs can be used to understand which disease classes each model is more successful in and why it outperforms the others. In the confusion matrix results of all models, it is seen that the "Honey" class estimations have lower performance than the estimations belonging to other classes.

4.2 Experimental results for ensemble learning

In the study, different combinations of the tested models were created for the ensemble learning application, and their performances were evaluated based on the data. The Dirichlet Ensemble technique was employed using the Deepstack module to investigate whether running the models with different weights improved the disease detection performance. For the selection of deep ensemble learning models, the following criteria were used to achieve this goal.• Top 3 models based on test accuracy,

• Top 5 models based on recall evaluation,

• Top 4 models based on F1 score evaluation.

In the study's first stage, specific models (MobileNetV2, ResNet50, ResNetRS50) produced successful results in all metrics. For this reason, to provide diversity and use different evaluation metrics, the ensemble model criteria determined above were established. For this purpose, MobileNetV2, ResNet50, and ResNetRS50 models, the first three models with the highest success considering the accuracy metric, were used, and the Deep Ensemble Learning Model (DELM1) was created. The following criterion was the recall value for focusing on disease detection for each class. MobileNetV2, ResNet50, ResNetRS50, DenseNet121, and EfficientNetB0 were used as the five models with the highest recall value, and DELM2 was created. Finally, MobileNetV2, ResNet50, ResNetRS50, and EfficientNetB0 models were used as the four successful models in the F1 score metric, and DELM3 was created.

After identifying the models based on these criteria, their respective weights and the saved weights from the first stage were included as members of the algorithm for ensemble learning. ROC AUC Score was employed to evaluate the performance of these created ensemble learning models. The information and results of the ensemble learning models formed according to these criteria are presented in Table 6.Table 6 Deep ensemble learning results.

Table 6Criterion	Models	Ensemble Model Name	Result	
Test Accuracy	MobileNetV2
ResNet50
ResNetRS50	DELM1	model1 - Weight: 0.0006 - roc_auc_score: 0.6864
model2 - Weight: 0.9503 - roc_auc_score: 0.9914
model3 - Weight: 0.0492 - roc_auc_score: 0.9512
DirichletEnsemble roc_auc_score: 0.9900	
Recall Evaluation	MobileNetV2
ResNet50
ResNetRS50
DenseNet121
EfficientNetB0	DELM2	model1 - Weight: 0.0045 - roc_auc_score: 0.6864
model2 - Weight: 0.8648 - roc_auc_score: 0.9914
model3 - Weight: 0.1025 - roc_auc_score: 0.9512
model4 - Weight: 0.0061 - roc_auc_score: 0.5022
model5 - Weight: 0.0220 - roc_auc_score: 0.8580
DirichletEnsemble roc_auc_score: 0.9877	
F1-Score Evaluation	MobileNetV2
ResNet50
ResNetRS50
EfficientNetB0	DELM3	model1 - Weight: 0.0041 - roc_auc_score: 0.6864
model2 - Weight: 0.8566 - roc_auc_score: 0.9914
model3 - Weight: 0.1375 - roc_auc_score: 0.9512
model4 - Weight: 0.0018 - roc_auc_score: 0.8649
DirichletEnsemble roc_auc_score: 0.9886	

As shown in Table 6, the ensemble models have demonstrated remarkable performance, with the DELM1 model achieving the highest ROC AUC Score of 99 %. This model comprises the three top-performing individual models, MobileNet, ResNet50, and ResNetRS50, which were selected based on high-test accuracy. Similarly, DELM2 and DELM3 models have also shown promising results, closely following the performance of DELM1. All ensemble models showed very high performance when the overall performance was evaluated. The ROC AUC scores ranged between 0.9877 and 0.9900, indicating that the models succeeded in the classification task. The best-performing model is ResNet50. It has the highest weight and individual ROC AUC score (0.9914) in all ensemble models. This shows that ResNet50 is the most effective model for this task. When the contributions of the other models are evaluated, ResNetRS50 has the second highest contribution after ResNet50 (weights between 0.0492 and 0.1375). MobileNetV2, DenseNet121, and EfficientNetB0 contributed relatively less. In terms of the effectiveness of the ensemble strategy, the table shows that the Dirichlet Ensemble method achieves high performance by combining the strengths of different models. Even the weakest individual models (e.g., 0.5022 ROC AUC with DenseNet121) were included in the ensemble, improving the overall performance. For the diversity of evaluation criteria, separate ensemble models were created for test accuracy, recall, and F1 score, thus providing optimized models for different metrics. The similar high performance of the models for all criteria showed that the model was generally balanced. The consistently high performance of ResNet50 and ResNetRS50 in the results shows that these architectures are particularly suitable for palm tree disease detection. The low weights of MobileNetV2, DenseNet121, and EfficientNetB0 indicated that these models could be improved with further fine-tuning or different hyperparameters. The deep ensemble learning models have effectively enhanced the performance of the individual deep learning models, as evidenced by the table. This highlights the significance of deep ensemble learning models in smart agriculture research, as they can potentially improve the accuracy and reliability of disease detection and classification tasks.

5 Discussion

Smart agriculture activities gained momentum, especially during the last two decades, and intensive work continued throughout the last decade. DL models, which started to be applied in every field after AlexNet's ImageNet success in 2012, have also been applied to different problems in agriculture. One of these problems is disease detection in plants and trees. DL has played a crucial role in this progress, especially in data classification and image recognition tasks.

Integrating DL with such data acquisition technologies enhances the accuracy and effectiveness of agricultural instruments. CNNs have become popular in agricultural research due to their powerful image-processing capabilities. Plant and crop classification are among the most prevalent applications of DL in agriculture. These applications have proven to be highly beneficial for yield prediction, pest control, disaster monitoring, and robotic harvesting. One area where CNNs have significantly impacted is plant disease detection. Traditionally, manual plant disease detection was time-consuming and labor-intensive. However, with recent advances in AI and image processing, CNNs have revolutionized this process. Now, plant disease detection can be efficiently performed using AI-powered image analysis techniques, saving time and resources while enabling early detection and intervention to prevent the spread of diseases in crops and plants [6]. Some studies on disease classification for palm trees have also been carried out in the literature.

In the study conducted by Ref. [10], the CNN and SVM algorithms achieved an accuracy ratio of 97.9 % and 92.8 %, respectively. One study limitation is that researchers obtained the reported test results during the training and validation phases. They did not apply cross-testing and did not reserve unseen data for evaluation. Similarly, another limitation is the absence of using the exact data for CNN as was done for SVM. These limitations may affect the generalizability and robustness of the models in real-world scenarios. The study by Ref. [14] achieved sensitivity, specificity, and accuracy rates of 99.3 %, 100 %, and 99.67 %, respectively. The main limitation of the study is the small size of the dataset. Neural networks can produce high-accuracy results on a limited number of images, potentially leading to overfitting. Therefore, the obtained results might be misleading and not generalize well to more extensive and more diverse datasets. Further research with a larger and more varied dataset would be necessary to validate the findings and ensure the model's robustness. Another palm disease classification study was conducted by Ref. [17]. Study model achieved an accuracy rate of 99.10 %, while VGG-16 and MobileNet achieved 99.35 % and 99.56 %, respectively. The high accuracy rates observed in this study indicate that the results are based on training and validation data. The models may memorize data. The study did not utilize the cross-test method, meaning the models were trained on the same data, leading to potential overfitting and inflated accuracy results. In a study by Ref. [20], MResNet achieved higher validation accuracy and 100 % accuracy at different epochs. Despite gaining 100 % accuracy at various epoch rates, the mentioned study does have some limitations. One of these limitations is the issue of overfitting. Although the study employed data augmentation, the presence of a validation set makes it vulnerable to overfitting. Additionally, the fact that the stated accuracy is based on validation results rather than cross-testing poses another limitation.

When examining literature studies, it is evident that they generally achieve high accuracy rates. However, these rates can be misleading. The most significant limitation of the studies is that the presented accuracy rates are obtained during training and validation. A key characteristic of DL studies is their ability to be generalized to different problems. Therefore, the accuracy rates of the models should ideally be presented on a separate test dataset that was not used during training. Otherwise, models that achieve high accuracy may perform less well when faced with different data and may fail to detect diseases accurately. Another limitation of the studies is that each study focuses on detecting only a limited number of diseases on a specific dataset. In contrast, this study overcomes this limitation by combining different datasets, enabling the classification of a larger number of diseases.

The specific advantages of this study over other studies in the literature can be listed as follows. More robust results due to the two-stage strategy and results justified by many different model runs. Consistency in the results is achieved by obtaining high ratios in addition to obtaining these ratios using different ensemble techniques. Indication of generalizability of the models by combining multiple datasets and proving the generalizability of the model with the results obtained on unseen test data.

Despite the progress, there is still a vast untapped potential in agriculture for DL applications. Many agricultural tasks are yet to leverage the benefits of these new techniques, indicating numerous opportunities for further exploration and improvement in DL in agriculture. Both the successful results and the results with room for improvement of this study provide indications for future work in the field of smart farming, both promising and suggesting new areas of research. This is presented as an answer to RQ3. This study's two-stage optimization approach overcomes previous research's limitations, leading to a more robust and generalizable outcome, which addresses to the RQ2 of this study. By combining transfer learning and fine-tuning with deep ensemble learning, this approach improves the performance of the models. It enhances their ability to be applied to various datasets and disease classifications, which addresses to the RQ1 of this study. The use of multiple pre-trained models and ensemble learning techniques enables a more comprehensive analysis and better detection of diseases, ensuring higher accuracy and reliability in the results. As a result, this two-stage optimization method contributes to achieving a more robust and generalized solution in smart agriculture applications, particularly in palm disease detection and classification. This also addresses the RQ1 and the RQ2 together.

6 Conclusion

This study conducted a comprehensive investigation of the application of DL models for palm disease detection and classification in the context of smart agriculture. This research aimed to address the limitations observed in previous studies and improve the robustness and generalizability of the results. Therefore, a two-stage optimization methodology was applied in the study.

The first stage of our methodology involved using transfer learning and fine-tuning with various pre-trained DNN models. In the study, the base models of the modules under the TensorFlow Keras Application library were selected. Data augmentation equalized the number of images for each class, mitigating the class imbalance issue. The experiments yielded promising results, with all models achieving high accuracy rates during training and validation. However, it is essential to consider that these accuracy rates were based on training and validation data, and their performance on unseen test data might differ. Therefore, the results of the models on the unseen test data were also observed. The top-performing models based on test accuracy in the first stage were MobileNetV2 (92.48 %), ResNet (92.42 %), ResNetRS50 (92.30 %), and DenseNet121 (92.01 %). Regarding Recall (Sensitivity) evaluation, MobileNetV2 exhibited successful performance in Brown Spots, Mixed, White Scale P2, and P3 categories. At the same time, ResNet50 excelled in the Healthy class, ResNetRS50 in White Scale P1, DenseNet121 in Bug, and EfficientNetB0 in Honey categories. Furthermore, based on the F1 score assessment, MobileNetV2 demonstrated notable performance in Brown Spots, White Scale P2, and P3 classes, while ResNet50 excelled in Mixed and Healthy, ResNetRS50 in White Scale P1, and EfficientNetB0 in Bug and Honey disease categories.

To improve the models' ability to make accurate predictions across different datasets, a deep ensemble learning approach was employed. The best-performing models were combined using the Dirichlet ensemble technique, and their performance was evaluated based on the ROC AUC Score. The deep ensemble learning models, particularly DELM1, exhibited outstanding results, achieving an ROC AUC Score of 99 %. These models demonstrated improvement over individual DL models, indicating their potential for enhancing disease detection tasks in smart agriculture.

One noteworthy contribution of this study is using diverse datasets, combining various plant diseases and insect pest classes. A more comprehensive classification system was achieved by considering different classes and stages of diseases within the white-scale class. Including more disease classes allowed the creation of a more robust and generalizable model.

In conclusion, this two-stage approach involving DL and ensemble learning techniques has shown promising results in palm disease detection and classification. The combination of multiple DL models through ensemble learning proved effective in enhancing the overall performance. Nonetheless, it is crucial to consider the need to test these models on unseen data and extend the study to address a more extensive range of plant diseases and pests. These findings contribute to advancing smart agriculture applications, emphasizing the importance of deep ensemble learning models in this domain. Future research can build upon these findings to optimize disease detection systems for crops and agricultural settings.

6.1 Limitations

One of this study's limitations is the sheer number of pre-trained modules and their sub-models, which continue to increase over time. Testing these models requires significant time and computational resources, which may impose constraints on the train-validation-test process. Due to the rapidly evolving landscape of DL architectures, new models are continuously being introduced, making it challenging to evaluate and compare all possible combinations effectively.

Considering this limitation, a representative set of pre-trained models and sub-models based on specific criteria was selected, as mentioned in the methods section. Although this approach allowed to assess a diverse range of models, there may still be other models that could have shown different performances. It is crucial to recognize that the selection of models, though carefully considered, may only encompass some of the available options, potentially affecting the overall conclusions.

To overcome this limitation, future studies could investigate strategies for more efficient model selection. This could involve automated hyperparameter optimization techniques or advanced feature selection methods. Additionally, researchers may focus on the most promising pre-trained models based on preliminary evaluations, narrowing the model pool for more in-depth testing.

Another limitation in the study is the limitation of obtaining the dataset. There is a limited number of palm disease datasets worldwide. This makes the generalizability of the studies difficult. To support the developing artificial intelligence and image processing studies in the field of agriculture, more datasets are needed in this field. In future studies, the creation of new datasets by researchers will pave the way for further research.

Despite these limitations, this study's findings provide valuable insights into the potential of ensemble learning in the context of smart agriculture and disease detection. While the sheer number of pre-trained models poses challenges, this work demonstrates the effectiveness of the selected models and provides a foundation for further research in this domain. By building upon this study's findings, future studies can continue to explore the application of DL and ensemble learning techniques for addressing critical agricultural challenges and improving crop health monitoring systems.

Data availability statement

The Mendeley dataset used in this study is online available at: https://doi.org/10.17632/2NH364P2BC.2.

The Kaggle dataset used in this study is online available at: https://www.kaggle.com/datasets/hadjerhamaidi/date-palm-data.

Funding

Not applicable.

Ethical statements

Only data coming from publicly available datasets were used and ethics approval not applicable.

CRediT authorship contribution statement

Serkan Savaş: Writing – review & editing, Writing – original draft, Visualization, Validation, Resources, Methodology, Investigation, Formal analysis, Conceptualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
==== Refs
References

1 Savaş S. Topaloǧlu N. Ciylan B. Analysis of mobile communication signals with frequency analysis method Gazi University Journal of Science 25 2012
2 Savaş S. Digital transformation from data mining to big data and its effects on productivity Calp M.H. Bütüner R. Current Studies in Digital Transformation and Productivity 2022 ISRES Publishing Konya 55 67
3 Butuner R. Calp M.H. Butuner N. Butuner M. Robotic systems and artificial intelligence applications in agriculture Calp M.H. Butuner R. Current Studies in Technology, Innovation and Entrepreneurship, ISRES Publishing 2023 International Society for Research in Education and Science (ISRES), Konya 145 158
4 Bharman P. Ahmad Saad S. Khan S. Jahan I. Ray M. Biswas M. Deep learning in agriculture: a review Asian Journal of Research in Computer Science 13 2022 28 47 10.9734/ajrcos/2022/v13i230311
5 Ren C. Kim D.K. Jeong D. A survey of deep learning in agriculture: techniques and their applications Journal of Information Processing Systems 16 2020 1015 1033 10.3745/JIPS.04.0187
6 Magomadov V.S. Deep learning and its role in smart agriculture J Phys Conf Ser 1399 2019 044109 10.1088/1742-6596/1399/4/044109
7 Kamilaris A. Prenafeta-Boldú F.X. Deep learning in agriculture: a survey Comput. Electron. Agric. 147 2018 70 90 10.1016/J.COMPAG.2018.02.016
8 EcoCation, Palm Tree Uses Palm tree uses https://ecocation.org/palm-tree-uses/ 2023
9 Saleem M.H. Potgieter J. Arif K.M. Automation in agriculture by machine and deep learning techniques: a review of recent developments Precis. Agric. 22 2021 2053 2091 10.1007/S11119-021-09806-X
10 Zhu N. Liu X. Liu Z. Hu K. Wang Y. Tan J. Huang M. Zhu Q. Ji X. Jiang Y. Guo Y. Deep learning for smart agriculture: concepts, tools, applications, and opportunities Int. J. Agric. Biol. Eng. 11 2018 32 44 10.25165/IJABE.V11I4.4475
11 Alaa H. Waleed K. Samir M. Tarek M. Sobeah H. Salam M.A. An intelligent approach for detecting palm trees diseases using image processing and machine learning Int. J. Adv. Comput. Sci. Appl. 11 2020 434 441 10.14569/IJACSA.2020.0110757
12 Goodfellow I. Bengio Y. Courville A. Deep Learning 2016 MIT press
13 Nanos G. Neural networks: pooling layers, baeldung on computer science https://www.baeldung.com/cs/neural-networks-pooling-layers 2024
14 Karakış R. Destek vektör makinesi Savaş S. Buyrukoğlu S. Teori Ve Uygulamada Makine Öğrenmesi, Nobel Akademik Yayıncılık Eğitim Danışmanlık TİC 2022 LTD. ŞTİ. Ankara 93 118
15 Hamdani H. Septiarini A. Sunyoto A. Suyanto S. Utaminingrum F. Detection of oil palm leaf disease based on color histogram and supervised classifier Optik 245 2021 167753 10.1016/J.IJLEO.2021.167753
16 Buyrukoglu G. Temel Bilesenler Analizi Savas S. Buyrukoglu S. Teori ve Uygulamada Makine Ogrenmesi 2022 Nobel Akademik Yayincilik Egitim Danismanlik Tic. Ltd. Sti. Ankara 251 262
17 Getcalc, covariance {cov(X, Y)} calculator, covariance {cov(X, Y)} calculator https://getcalc.com/statistics-covariance-calculator.htm 2024
18 Abu-zanona M. Elaiwat S. Younis S. Innab N. Kamruzzaman M.M. Classification of palm trees diseases using convolution neural network Int. J. Adv. Comput. Sci. Appl. 13 2022 943 949 10.14569/IJACSA.2022.01306111
19 Simonyan K. Zisserman A. Very deep convolutional networks for large-scale image recognition 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings 2014 10.48550/arxiv.1409.1556
20 Howard A.G. Zhu M. Chen B. Kalenichenko D. Wang W. Weyand T. Andreetto M. Adam H. MobileNets: efficient convolutional neural networks for mobile vision applications ArXiv 2017 10.48550/arxiv.1704.04861
21 Ahmed M. Ahmed A. Palm tree disease detection and classification using residual network and transfer learning of inception ResNet PLoS One 18 2023 e0282250 10.1371/JOURNAL.PONE.0282250
22 Szegedy C. Ioffe S. Vanhoucke V. Alemi A.A. Inception-v4, inception-ResNet and the impact of residual connections on learning 31st AAAI Conference on Artificial Intelligence, AAAI 2017 2016 4278 4284 10.48550/arxiv.1602.07261
23 Nobel S.M.N. Imran M.A. Bina N.Z. Kabir M.M. Safran M. Alfarhood S. Mridha M.F. Palm leaf health management: a hybrid approach for automated disease detection and therapy enhancement IEEE Access 12 2024 9097 9111 10.1109/ACCESS.2024.3351912
24 Amara J. Bouaziz B. Algergawy A. A Deep Learning-Based Approach for Banana Leaf Diseases Classification 2017 Gesellschaft für Informatik e.V. https://dl.gi.de/items/13766147-8092-4f0a-b4e1-8a11a9046bdf
25 Math R.K.M. Dharwadkar N.V. Early detection and identification of grape diseases using convolutional neural networks J. Plant Dis. Prot. 129 2022 521 532 10.1007/S41348-022-00589-5/FIGURES/7
26 Yadav S. Sengar N. Singh A. Singh A. Dutta M.K. Identification of disease using deep learning and evaluation of bacteriosis in peach leaf Ecol Inform 61 2021 101247 10.1016/J.ECOINF.2021.101247
27 Barbedo J.G.A. Impact of dataset size and variety on the effectiveness of deep learning and transfer learning for plant disease classification Comput. Electron. Agric. 153 2018 46 53 10.1016/J.COMPAG.2018.08.013
28 Saleem M.H. Potgieter J. Arif K.M. Plant disease detection and classification by deep learning Plants 8 2019 468 10.3390/PLANTS8110468 Page 468 8 (2019 31683734
29 Ahmed I. Yadav P.K. Plant disease detection using machine learning approaches Expert Syst 40 2023 e13136 10.1111/EXSY.13136
30 Atila Ü. Uçar M. Akyol K. Uçar E. Plant leaf disease classification using EfficientNet deep learning model Ecol Inform 61 2021 101182 10.1016/J.ECOINF.2020.101182
31 Joshi R.C. Kaushik M. Dutta M.K. Srivastava A. Choudhary N. VirLeafNet: automatic analysis and viral disease diagnosis using deep-learning in Vigna mungo plant Ecol Inform 61 2021 101197 10.1016/J.ECOINF.2020.101197
32 Tan L. Lu J. Jiang H. Tomato leaf diseases classification based on leaf images: a comparison between classical machine learning and deep learning methods AgriEngineering 3 2021 542 558 10.3390/AGRIENGINEERING3030035 3 (2021) 542–558
33 Sarkar C. Gupta D. Gupta U. Hazarika B.B. Leaf disease detection using machine learning and deep learning: review and challenges Appl. Soft Comput. 145 2023 110534 10.1016/J.ASOC.2023.110534
34 Mazin A. Almayaly H. Image dataset of infected date palm leaves by dubas insects Mendeley 2 2023 10.17632/2NH364P2BC.2
35 Hajar Date Palm Data 2018 Kaggle https://www.kaggle.com/datasets/hadjerhamaidi/date-palm-data
36 Shorten C. Khoshgoftaar T.M. A survey on image data augmentation for deep learning J Big Data 6 2019 1 48 10.1186/S40537-019-0197-0/FIGURES/33
37 Alhudhaif A. Almaslukh B. Aseeri A.O. Guler O. Polat K. A novel nonlinear automated multi-class skin lesion detection system using soft-attention based convolutional neural networks Chaos, Solit. Fractals 170 2023 113409 10.1016/J.CHAOS.2023.113409
38 Güler O. Polat K. Classification performance of deep transfer learning methods for pneumonia detection from chest X-ray images Journal of Artificial Intelligence and Systems 4 2022 107 126 10.33969/AIS.2022040107
39 Cook D. Feuz K.D. Krishnan N.C. Transfer learning for activity recognition: a survey Knowl. Inf. Syst. 36 2013 537 556 10.1007/S10115-013-0665-3/TABLES/6 24039326
40 TensorFlow Module: tf.keras.applications | TensorFlow v2.12.0, Module: Tf.Keras.Applications | TensorFlow v2.12.0 2023 https://www.tensorflow.org/api_docs/python/tf/keras/applications
41 Chollet F. Xception: deep learning with depthwise separable convolutions Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 2017-January 2016 1800 1807 10.48550/arxiv.1610.02357
42 He K. Zhang X. Ren S. Sun J. Deep residual learning for image recognition Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016
43 Radosavovic I. Kosaraju R.P. Girshick R. He K. Dollár P. Designing network design spaces Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2020 10425 10433 10.1109/CVPR42600.2020.01044
44 Zoph B. Vasudevan V. Shlens J. Le Q.V. Learning transferable architectures for scalable image recognition Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2017 8697 8710 10.1109/CVPR.2018.00907
45 Tan M. Le Q.V. EfficientNet: rethinking model scaling for convolutional neural networks 36th International Conference on Machine Learning, ICML 2019 2019-June 2019 10691 10700 10.48550/arxiv.1905.11946
46 Szegedy C. Ioffe S. Vanhoucke V. Alemi A.A. Inception-v4, inception-ResNet and the impact of residual connections on learning 31st AAAI Conference on Artificial Intelligence, AAAI 2017 2016 4278 4284 10.1609/aaai.v31i1.11231
47 Huang G. Liu Z. Van Der Maaten L. Weinberger K.Q. Densely connected convolutional networks Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017 2017-January 2016 2261 2269 10.48550/arxiv.1608.06993
48 Sandler M. Howard A. Zhu M. Zhmoginov A. Chen L.C. MobileNetV2: inverted residuals and linear bottlenecks Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2018 4510 4520 10.1109/CVPR.2018.00474
49 Tan M. Le Q.V. EfficientNetV2: smaller models and faster training Proc Mach Learn Res 139 2021 10096 10106 https://arxiv.org/abs/2104.00298v3
50 Liu Z. Mao H. Wu C.Y. Feichtenhofer C. Darrell T. Xie S. A ConvNet for the 2020s Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition 2022-June 2022 11966 11976 10.1109/CVPR52688.2022.01167
51 Howard A. Sandler M. Chen B. Wang W. Chen L.C. Tan M. Chu G. Vasudevan V. Zhu Y. Pang R. Le Q. Adam H. Searching for MobileNetV3 Proceedings of the IEEE International Conference on Computer Vision 2019 1314 1324 10.1109/ICCV.2019.00140 2019-October
52 Gómez R. Understanding categorical cross-entropy loss, binary cross-entropy loss, Softmax loss, logistic loss, focal loss and all those confusing names Github 2018 https://gombru.github.io/2018/05/23/cross_entropy_loss/
53 An N. Ding H. Yang J. Au R. Ang T.F.A. Deep ensemble learning for Alzheimer's disease classification J Biomed Inform 105 2020 103411 10.1016/J.JBI.2020.103411
54 Borges J. DeepStack: ensembles for deep learning https://github.com/jcborges/DeepStack 2019
55 Savaş S. Enhancing disease classification with deep learning: a two-stage optimization approach for monkeypox and similar skin lesion diseases Journal of Imaging Informatics in Medicine 37 2024 778 800 10.1007/s10278-023-00941-7 38343247
