
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

39256414
70950
10.1038/s41598-024-70950-1
Article
A multi-view multi-label fast model for Auricularia cornea phenotype identification and classification
Xu Yinghang 15
Qu Shizheng qushizheng1211@163.com

1
Liu Huan 3
Zhang Lina 1
Liu Yunfei 1
Wang Lu 1
Li Zhuoshi zhuoshil@jlau.edu.cn

1245
1 https://ror.org/05dmhhd41 grid.464353.3 0000 0000 9888 756X College of Information Technology, Jilin Agricultural University, Changchun, 130118 China
2 https://ror.org/05dmhhd41 grid.464353.3 0000 0000 9888 756X College of Plant Protection, Jilin Agricultural University, Changchun, 130118 China
3 https://ror.org/018gks972 grid.443318.9 School of Data Science and Artificial Intelligence, Jilin Engineering Normal University, Changchun, China
4 https://ror.org/05dmhhd41 grid.464353.3 0000 0000 9888 756X Jilin Province Key Laboratory of Fungal Phenomics, Jilin Agricultural University, Changchun, China
5 https://ror.org/05dmhhd41 grid.464353.3 0000 0000 9888 756X International Cooperation Research Center of China for New Germplasm Breeding of Edible Mushrooms, Jilin Agricultural University, Changchun, China
10 9 2024
10 9 2024
2024
14 2113625 5 2024
22 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
The identification and classification of various phenotypic features of Auricularia cornea fruit bodies are crucial for quality grading and breeding efforts. The phenotypic features of Auricularia cornea fruit bodies encompass size, number, shape, color, pigmentation, and damage. These phenotypic features are distributed across various views of the fruit bodies, making the task of achieving both rapid and accurate identification and classification challenging. This paper proposes a novel multi-view multi-label fast network that integrates two different views of the Auricularia cornea fruiting body, enabling rapid and precise identification and classification of six phenotypic features simultaneously. Initially, a multi-view feature extraction model based on partial convolution was constructed. This model incorporates channel attention mechanisms to achieve rapid phenotypic feature extraction of the Auricularia cornea fruiting body. Subsequently, an efficient multi-task classifier was designed, based on class-specific residual attention, to ensure accurate classification of phenotypic features. Finally, task weights were dynamically adjusted based on heteroscedastic uncertainty, reducing the training complexity of the multi-task classification. The proposed network achieved a classification accuracy of 94.66% and an inference speed of 11.9 ms on an image dataset of dried Auricularia cornea fruiting bodies with three views and six labels. The results demonstrate that the proposed network can efficiently and accurately identify and classify all phenotypic features of Auricularia cornea.

Keywords

Multi-task learning
Multi-view learning
Phenotype identification and classification
Partial convolution
Auricularia cornea
Subject terms

Computer science
Biological techniques
Computational biology and bioinformatics
Mathematics and computing
the Department of Education of Jilin ProvinceJJKH20210331KJ; JJKH20210331KJ; Xu Yinghang Li Zhuoshi Jilin Province Science and Technology Development Plan20240305025YY Xu Yinghang Scientific Research Project of the national key research and development program of China2023YFD1301802 Li Zhuoshi the national key research and development program of China2023YFD1201600 Li Zhuoshi issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

The industrialization of edible mushroom production has resulted in high-yield, high-quality outputs1, making it the primary source of mushrooms in today’s market. Auricularia cornea, also known as wood ear mushroom, ranks as the third largest cultivated mushroom, comprising approximately 17% of global mushroom production. Widely cultivated in Asia, it is esteemed for its rich nutritional content and medicinal properties2. In the fields of biology and genetic breeding, phenotype refers to the complete set of characteristics that constitute an organism, including appearance, fundamental dimensions, morphology, and color. These characteristics result from the interaction between genotype and environmental factors3. Different qualities of Auricularia cornea exhibit variations in nutritional value, taste, and price. By identifying and classifying phenotypic traits related to quality, precise and detailed grading of edible mushrooms can be achieved, thereby enhancing their added value and generating substantial economic benefits4. Therefore, accurate and rapid identification and classification of the phenotypic characteristics has become an inevitable trend in the industrial cultivation and development of edible mushrooms.

The task of recognizing and classifying the phenotypes of Auricularia cornea fruiting bodies presents several challenges. First, the fruiting bodies curl during the drying process, resulting in irregular shapes that complicate phenotype extraction. Second, information such as damage, discoloration, and fruiting body count can easily be overlooked in a single view. Current methods for quality classification based on the phenotypes of Auricularia cornea fruiting bodies are often crude, relying on manual sorting or simple mechanical screening. Manual classification is time-consuming, labor-intensive, costly, and slow, with significant subjective bias. Mechanical screening typically uses different sizes of mesh grids, considering only the size of the Auricularia cornea, and can cause additional damage to the fruiting bodies, leading to further losses.

Computer vision and deep learning offer advantages such as speed, accuracy, and objectivity in the recognition and classification of phenotypic features of edible fungi. Many researchers have made attempts to identify phenotypic features and grade the quality of edible fungi using these methods. Wang et al. developed an image processing algorithm based on machine vision to measure the pileus diameter of fresh white button mushrooms, which were subsequently classified into three grades based on their size5. Liu et al. utilized a channel pruning algorithm and distillation method to reduce the size of the YOLOX model by more than half, enabling rapid detection of the surface texture of shiitake mushrooms, which were further categorized into Flower shiitake mushrooms and Smooth cap shiitake mushrooms6. Wu et al. integrated the YOLOv5 model for single-stage object detection with the PSPNet model for semantic segmentation to achieve real-time, non-contact grading of Antler Mushrooms into three grades based on their size7. Zhu et al. integrated Yolov5 with OpenCV to identify and detect the position and size of shiitake mushrooms. They used a resistance strain gauge sensor to obtain quality parameters, grading the mushrooms into three levels based on their shape and size8. Research on the identification and classification of phenotypes of edible fungi fruiting bodies has mainly focused on more regularly shaped species such as shiitake mushrooms and button mushrooms, with less attention given to the irregularly shaped Auricularia cornea fruiting bodies. In Xu et al., the CoTNet model served as the backbone network, incorporating depth-wise separable convolution along with coordinate attention (CA) and normalization-based attention module (NAM). However, the grading was only coarse, dividing the black wood ear mushroom into three levels without identifying or analyzing specific phenotypic features, thus lacking in credibility9.

In most of the previously mentioned methods for phenotype identification and classification of edible fungi, the focus has been placed on recognizing single phenotypes from a single perspective. However, Auricularia cornea exhibits multiple phenotypic traits. When classifying the quality of Auricularia cornea fruiting bodies, traits such as the number, color, and size of the fruiting bodies are involved. These phenotypic features are difficult to capture comprehensively from a single view. Therefore, to accurately and completely capture the phenotypic features of Auricularia cornea, images were collected from three different perspectives. Data annotation was performed on multiple phenotypic traits, including shape, size, color, number of ear pieces, and degree of damage, to support this study.

With the advancement of information technology, approaches based on multi-view and multi-task methodologies have made significant strides. In the study by Shi et al., overall features were delineated by capturing apple images from top, front, and back views, coupled with phenotypic information about size, resulting in advancements in the effectiveness of multi-view apple grading10. Chen et al. introduced a novel rice grading model that fused image features from three different views of rice grains, effectively determining the grading level of rice11. In the study by Chen et al., an enhanced YOLOv7-based multi-task deep convolutional neural network detection model was introduced, capable of simultaneously identifying tomato fruit clusters, evaluating fruit ripeness, and assessing the ripeness of the cluster12. Wang et al. introduced a dual-stream hierarchical bilinear pooling model based on different bilinear pooling approaches, enabling separate identification and classification of crop types and diseases13. These studies have explored the applications and advantages of multi-view and multi-task methods in crop phenotype identification and classification. By integrating image features from multiple perspectives, more comprehensive information is provided. Alternatively, considering various phenotypic traits of crops improves the accuracy and reliability of classification.

To improve the effectiveness and accuracy of phenotypic identification and classification of Auricularia auricula fruiting bodies, this study develops a lightweight, fast neural network model with multi-view image data input and multi-task learning output. First, this study utilizes partial convolution (PConv) as the foundational operator, combined with channel attention mechanisms, to construct a lightweight feature extraction network for multi-view image feature extraction. This approach enables real-time, precise extraction of phenotypic traits of Auricularia auricula fruiting bodies. Then, an efficient multi-task classifier, based on class-specific residual attention (CSRA), is developed14. Specifically, multiple parallel CSRA modules highlight the spatial distribution of different phenotypic traits from the overall characteristics of the fruiting bodies. This ensures that each phenotypic classification task focuses on its relevant features. Accurate classification of different phenotypes is then achieved through multiple fully connected layers. Finally, based on the principle of homoscedastic uncertainty, the weights among different tasks are automatically learned, enhancing training efficiency, reducing training difficulty, and improving the model’s accuracy.

The main contributions of this study can be summarized as follows: We have curated a multi-view, multi-label dataset for dried Auricularia cornea products. Using the FScan2000 edible mushroom phenotyping device, we collected multi-view data of dried Auricularia cornea products and meticulously recorded six phenotypic as labels, including size, number, shape, color, presence of pigmentation, and presence of damage. This dataset offers a more comprehensive set of classification criteria and richer information for phenotypic identification and classification of Auricularia auricula fruiting bodies.

A multi-view, multi-task fast model for the identification and classification of Auricularia auricula fruiting body phenotypes was developed. This model addresses the issues of incomplete information from single perspectives and low classification accuracy.

Optimization methods for multi-task learning were explored. A multi-task classifier, the class-specific attention classifier, was constructed. Incorporating homoscedastic uncertainty, the model automatically learns the weights between different tasks, enhancing classification accuracy.

We have validated the accuracy and efficiency of the model by comparing single-view and multi-view approaches, single-task and multi-task approaches, as well as with other lightweight models.

Materials and methods

Data acquisition

Data acquisition equipment

The images of dried Auricularia cornea fruiting body used in this study were captured using the FScan2000 edible mushroom phenotyping device. The device, depicted in Fig. 1, comprises two main components: the edible mushroom phenotyping box and the image processing software. Inside the box, three cameras of identical configuration are positioned at the top, left side, and bottom. The dimensions of the phenotyping box are 570 mm × 430 mm × 280 mm. The cameras have imaging pixels of 16 million (4608 × 3456), with a 1/2.3-inch CMOS sensor. The focal length of the camera lens is 8 mm. The internal light source of the box comprises 360∘ surround light and LED white light. The maximum shooting size of the phenotyping box is 400 mm × 300 mm, with a minimum accuracy of 0.12 mm.Fig. 1 Edible mushroom phenotype collection device: FScan2000.

Data collection methods

The experimental materials used in this study were obtained from Haotian Village, Najin Town, Taonan City, Jilin Province. These materials comprise a white variant strain of Auricularia cornea, a new edible mushroom variety bred by the team led by academician Yu Li from Jilin Agricultural University15. This strain exhibits higher protein and reducing sugar content compared to common Auricularia cornea varieties. As depicted in Fig. 2, in this study, the images collected from the side where the ear stalk of Auricularia auricula fruiting bodies is located are considered the top view, while the opposite side constitutes the bottom view. The fruiting bodies are arranged on the platform as shown in Fig. 2a, and the images collected from the right side are considered the side view of the Auricularia auricula fruiting bodies.Fig. 2 Three views of Auricularia cornea. (a) Top view, (b) side view, (c) bottom view.

After drying, Auricularia auricula tends to curl, resulting in significant differences in its features when observed from various angles. When placed on a flat surface, the observable angles become random. As shown in Fig. 3, each perspective provides partial phenotypic characteristics. Information such as the quantity, pigmentation, and damage of the fruiting bodies is distributed across different views, making it difficult to accurately and comprehensively analyze the phenotypes from a single perspective.Fig. 3 Graded features distribution in different views of Auricularia cornea images.

In this study, Auricularia auricula samples were placed on a transparent tray in the center of the FScan2000. The distance and angle of the camera were fixed, and images were captured from the top, bottom, and side views to ensure the completeness of phenotypic information. To further ensure the accuracy and reliability of the data, multi-label annotations were applied to the images of each sample. According to the dried Auricularia auricula standard DB22/T 2605-2016, six phenotypic characteristics of the fruiting bodies were recorded in detail: size, quantity, shape, color, presence of pigmentation, and any damage. The fruiting bodies were classified into four grades: Grade 1, Grade 2, Grade 3, and Substandard. The classification standards are provided in Table 1.

In this study, Auricularia cornea specimens were positioned on the transparent tray in the center of the FScan2000 device to ensure fixed shooting distances and angles. Top view images, bottom view images, and side view images of the Auricularia cornea were captured to ensure comprehensive coverage of its drying product grading features. To further ensure the accuracy and reliability of the Auricularia cornea data, the image data of each specimen underwent multi-label annotation. In accordance with the Auricularia cornea dried product standard DB22/T 2605-2016, this study meticulously recorded six indicators for Auricularia cornea, including size, number, shape, color, presence of pigmentation, and whether the fruiting bodies were damaged. These indicators serve as the basis for Auricularia cornea grading, categorizing it into four grades: first grade, second grade, third grade, and off-grade. Table 1 outlines the criteria for Auricularia cornea grading. Table 1 Phenotypic classification of Auricularia cornea.

Classification criteria	Grade	
First grade	Second grade	Third grade	Off-grade	
Size	2~3(cm)	3~4(cm)	4~5(cm)	>5(cm)	
Number	Single	Single	Single and multiple	Multiple	
Shape	Stretching or naturally curling	Natural curling	Natural curling	Curl into a ball	
Color	White to light yellow	Rice white to light yellow	Light yellow to beige	-	
Pigmentation	Absence	Absence	Presence	Presence	
Damage	Absence	Absence	Presence	Presence	

This study utilized standard ruler measurements to record the size information of Auricularia cornea fruiting bodies, categorizing them into five classes (grades: 1, 2, 3, 4, 5), namely: maximum diameter of 20 millimeters or less (20 mm-), 20 millimeters to 30 millimeters (20 mm+), 30 millimeters to 40 millimeters (30 mm+), 40 millimeters to 50 millimeters (40 mm+), and 50 millimeters or above (50 mm+). The quantity of Auricularia cornea fruiting bodies was classified into two categories (grades: 1, 2), representing either single or multiple. The shape of Auricularia cornea fruiting bodies was categorized into three classes based on the degree of curling after drying: fully expanded, naturally curled, and curled into a ball. The degree of color of the Auricularia cornea fruiting body was classified into three grades (grades: 1, 2, 3). Pigmentation status was divided into two categories (grades: 1, 2), indicating absence or presence of spots. Damage conditions were classified into two categories (grades: 1, 2), representing absence or presence of damage. The distribution of various label data for Auricularia cornea samples is presented in Table 2. Table 2 Distribution of data for each label.

Grade	Size	Number	Shape	Color	Pigmentation	Damage	
1	232	592	245	431	587	635	
2	287	99	146	229	104	56	
3	52	–	300	31	–	–	
4	64	–	–	–	–	–	
5	56	–	–	–	–	–	

The dataset comprises 691 sets of triview Auricularia cornea data, with each set including top view, bottom view, and side view images of the Auricularia cornea.

Data preprocessing

The model employed in this study utilizes images of size 224×224 as input data. Directly using the collected images as input would result in significant compression of Auricularia cornea images. As depicted in Fig. 4, this study crops out the Auricularia cornea from the original images. For all Auricularia cornea images, this study crops out the Auricularia cornea centered at its position. Specifically, the process involves initially annotating the positions of Auricularia cornea in the collected images using LabelMe software. Subsequently, a YOLOv8 model is fine-tuned based on a small portion of annotated data to enable detection and segmentation of Auricularia cornea.Fig. 4 Image cropping method.

The orientation of Auricularia cornea on a flat surface exhibits randomness, resulting in varying characteristics when capturing the same Auricularia cornea from different angles. As illustrated in Fig. 5, this study enhances data diversity by augmenting the cropped images of Auricularia cornea through angle flipping. By augmenting the dataset to four times its original size, a total of 2764 sets of triview Auricularia cornea data were obtained.Fig. 5 Data augmentation methods.

Dataset partitioning

From the dataset of 2764 sets of triview Auricularia cornea, this study obtained 2764 sets of Auricularia cornea from different angles, including top view, side view, bottom view, top and side view, bottom and side view, and top and bottom view. These datasets from different views were divided into training and testing sets in an 8:2 ratio.

In practical production applications, when capturing randomly placed Auricularia cornea from above, the images obtained contain both top view and bottom view images. Through experimentation, it has been determined that both combinations of top view images with side view images and bottom view images with side view images encompass the complete phenotypic features of Auricularia cornea. Can meet the recognition and classification tasks of Auricularia cornea phenotypes.Therefore, this study selected either top view or bottom view images, combined with side view images, to split a set of triview data into two sets of dual-view data. As a result, 5528 sets of dual-view Auricularia cornea data were obtained, and the dataset was divided into training and validation sets in an 8:2 ratio. Specifically, 4423 sets of dual-view Auricularia cornea data were allocated to the training set, while 1105 sets were allocated to the testing set.

Model structure and improvement

A multi-view, multi-task rapid grading model was developed in this study. The top-layer parameters of the backbone network were shared, whereas the bottom-layer parameters of the backbone network and the multi-task classifier parameters remained independent. Initially, a lightweight and fast CNN was designed as a feature extraction module for the various views. To differentiate the feature information required for different classification tasks, a multi-task classifier based on class-specific attention was constructed. Additionally, homoscedastic uncertainty was employed to measure the loss weights between tasks, and distinct loss function weights were designed for each classification task based on the data distribution of each phenotype. Figure 6 illustrates the overall structure of the model. The dual-view data of Auricularia cornea were input into the Lightweight Fast CNN to extract the comprehensive phenotypic features. Subsequently, the Class-specific Attention Classifier calculated and extracted the features needed for each specific phenotypic classification task, achieving precise identification and classification of Auricularia cornea phenotypes.Fig. 6 Multi-view and multi-task fast model structure.

The lightweight and fast CNN employed for multi-view feature extraction

The PConv method efficiently extracts spatial features by simultaneously reducing redundant computation and memory access16. Figure 7 depicts the operational principle of PConv. It involves applying conventional convolution solely to a part of input channels for spatial feature extraction, while maintaining the remaining channels unchanged. Thus, when Cp represents 1/4 of the total channel number, the FLOPs (Floating Point Operations) of PConv amount to only 1/16 of those of regular convolutional kernels. Each partial convolution divides the input feature map into two parts, selecting the initial Cp channels for conventional convolution operations. This approach to channel partitioning remains constant. However, different channels contribute variably to the identify classification task, potentially causing convergence difficulties and impacting the final grading performance of the model. This study introduces a channel attention mechanism to model the relationships among different channels in the input data. This mechanism aids the model in dynamically adjusting the weights of each channel to amplify significant features and mitigate irrelevant ones, thus enhancing the feature extraction ability of the model17.Fig. 7 Partial convolution structure diagram structure.

Figure 8 illustrates the operational principle of the channel attention module. Initially, features are extracted from the input data through operations such as convolutional layers, resulting in multiple feature maps, with each feature map corresponding to a channel. Subsequently, operations such as global average pooling, fully connected layers, and sigmoid functions are applied to calculate the importance weights of each channel. These weights signify the contribution of each channel to the final prediction. These channel weights are then applied to the corresponding feature maps, merging the features of different channels through weighted fusion to generate the final feature representation. Finally, the fused features are input into subsequent layers of the network.Fig. 8 Channel attention module structure.

Building upon SPConv, this study developed a lightweight and fast CNN for multi-view feature extraction. Table 3 displays the network architecture of the multi-view feature extraction network. The multi-view feature extraction network in this study is partitioned into four stages, with each stage comprising down-sampling layers and feature extraction layers. The down-sampling layer consists of 2D convolutional layers and batch normalization operations. The feature extraction layer employs a stacked Block approach, where each Block consists of PConv, a 1×1 convolution for dimensionality expansion, and another 1×1 convolution for dimensionality reduction. Furthermore, batch normalization and GRELU activation operations are applied following the first 1×1 convolution layer. Finally, skip connections are utilized to facilitate residual learning. In Stage 1 and Stage 2, the feature extraction layers consist of two parallel sets of Blocks for dual-channel processing, with each set dedicated to extracting features from two different views. In Stage 3 and Stage 4, the feature extraction layers are single-channel, responsible for integrating features from both views. Additionally, a channel attention module is introduced after the down-sampling layer in Stage 3. Table 3 Backbone network architecture table.

Layer	Layer	Group	Channel	Stride	Output size	
Input	–	2	6	–	224×224	
Stage1 down sampling	[4×4,40]×1	2	80	4	56×56	
Stage1 feature extraction	3×3,101×1,801×1,40×1	2	80	1	56×56	
Stage2 down sampling	[4×4,80]×1	2	160	2	28×28	
Stage2 feature extraction	3×3,201×1,1601×1,80×1	2	160	1	28×28	
Stage3 down sampling	[4×4,320]×1	1	320	2	14×14	
Stage3 channel attention	AvgPooling, 320Linear, 20Linear, 320×1	1	320	–	14×14	
Stage3 feature extraction	3×3,801×1,6401×1,320×2	1	320	1	14×14	
Stage4 down sampling	[4×4,320]×1	1	640	2	7×7	
Stage4 feature extraction	3×3,1601×1,12801×1,640×1	1	640	1	7×7	

During the drying process of Auricularia cornea, significant deformations occur, the distribution of phenotypic information in different perspectives when observed from various angles. In this study, images from two divergent perspectives are simultaneously inputted into the corresponding branches of a multi-view feature extraction network. After the feature extraction process, feature maps of size 7×7 with 640 dimensions are produced. By sharing the top-level parameters of the multi-view feature extraction network, the model tends to extract relevant to comprehensive information of multiple phenotypic features of Auricularia cornea. This approach reduces the risk of overfitting and ensures comprehensive training across all tasks. Moreover, parameter sharing significantly reduces the number of network parameters compared to the total parameters of N individual task-specific networks. This reduction implies higher efficiency of Multi-Task Learning models in real-time multi-task prediction scenarios18.

Class-specific attention classifier for multi-task learning

There are strong correlations among the phenotypes of Auricularia cornea fruiting bodies, such as size, number, and shape. The proposed model achieves Multi-Task Learning of these phenotypes by sharing partial parameters. However, a challenge arises: how to enable different classifiers to identify features relevant to their respective classification tasks more effectively and efficiently. To address this challenge, this research introduces a class-specific residual attention mechanism to capture distinct feature regions attended to by different tasks. A class-specific attention classifier for multi-task classification was thus constructed.

As shown in Fig. 9, the class-specific attention multi-task classifier developed in this study operates in two stages.

In Stage One, eight parallel CSRA modules are utilized to extract class-specific attention for various classification tasks. The Residual Attention module in Fig. 9 depicts the structure of CSRA. Each CSRA consists of parallel branches for average pooling and spatial pooling. The average pooling branch calculates the average feature over the entire input, resulting in class-agnostic average pool features, as shown in Eq. (1). The spatial pooling branch generates spatial pool features, yielding class-specific spatial pool features for each category, as illustrated in Eq. (3). By weighting and adding these two features, the importance of class-specific features in the global feature is enhanced, resulting in class-specific feature attention scores (Residual Attention), as depicted in Eq. (4).1 g=149∑k=149xk

2 sji=expTxjTmiΣk=149expTxkTmi

3 ai=∑k=149skixk

4 fi=g+λai

In Stage Two, the spatial features attended to by the spatial pooling branches of different CSRA modules differ. The class-specific attention features obtained from the eight CSRA modules are aggregated, and six parallel fully connected layers serve as classifiers for each task. Finally, classification predictions are made based on the feature vectors extracted through class-specific attention processing, enabling different task classifiers to identify the features they are focused on.Fig. 9 Class-specific attention classifier structure.

Uncertainty weight-based optimization contains multiple objective loss functions

On one hand, some phenotypic classification tasks exhibit an imbalance between positive and negative samples. When the number of samples in a certain category is significantly large, the value of the loss function is influenced by the category with a larger sample size, leading the model to bias towards the larger category during classification. In this study, different weight values are assigned to each phenotypic classification task individually to address this issue. The weights of smaller categories are increased, thereby regulating the loss value based on the sample quantity for different categories. This weighted approach based on categories helps alleviate the training challenges caused by sample imbalance.

On the other hand, the multi-task learning of Auricularia cornea phenotypes involves optimizing the model across six tasks, a common challenge in multi-task learning. A straightforward method for combining multi-task losses is to linearly weigh and sum the loss for each task. However, this approach poses several issues, given that the model’s performance is highly sensitive to the choice of weights. An effective and convenient method involves utilizing same-variance uncertainty learning for relative weight learning. Same-variance uncertainty is a form of uncertainty that is independent of the incidental uncertainties in the input data19. It is not an output of the model but a value that remains constant for all input data and varies across different tasks. Thus, it can be described as task-specific uncertainty.

The optimal weight for each task depends on the scale of measurement and, ultimately, on the magnitude of task noise. As shown in Eq. (5), where W represents the loss values for the six sub-tasks, σ1 and σ2 denote the magnitudes of noise.5 L(W,σ1,σ2)=12σ12L1(W)+12σ22L2(W)+logσ1σ2

Experimental setup

Experimental platform

All model training and testing were conducted on the same computer and server with the following specifications: Linux version 5.15.0-60-generic, server with NVIDIA A800 80G, 24 vCPU Intel(R) Xeon(R) Platinum 8255C CPU @ 2.50 GHz, computer graphics card driver MX150, programming platform Anaconda 3.5, CUDA 10.2, development environment PyTorch, and programming with Python 3.8.10.

Evaluation indices

To evaluate the performance of our multi-view and multi-task fast model, this paper applies evaluation metrics including precision, recall, multi-label F1 score(macro-F1), exact match ratio, and average inference time. Considering the multi-tasking nature, the average of the six grading tasks of Auricularia cornea is taken as the macro-F1, and the prediction of all grading tasks correctly is counted effectively as the Exact Match Ratio. The calculation of each evaluation metric is as follows:6 P=TPTP+FP

7 R=TPTP+FN

8 Accuracy=TP+TNTP+FN+FP+TN

9 F1-score=2×Precision×RecallPrecision+Recall

10 macro-F1=F1-score1+F1-score2+⋯+F1-scoren3

Results

Comparison between single-view and multi-view

To demonstrate the superiority of multi-view approaches in identifying and classifying Auricularia cornea phenotypes, this study established identical classification models for each classification task related to the size, number, shape, color, pigmentation, and damage of Auricularia cornea fruiting bodies. Building on this foundation, the study trained six models separately using images from different perspectives as training sets and compared the accuracy of classification tasks across these perspectives. The perspectives compared include the top view, side view, bottom view, top and side view, bottom and side view, and top and bottom view. ï»¿ Table 4 presents the classification performance of the models for Auricularia cornea fruiting bodies across six tasks under both single-view and multi-view conditions. It demonstrates that the Exact Match Ratio achieved by the models under multi-view conditions consistently exceeds that achieved under single-view conditions. Table 4 The accuracy rates for each single-task grading under different views.

Different view	Exact match ratio (%)	Accuracy of each indicator	
Size (%)	Number (%)	Shape (%)	Colour (%)	Pigmentation (%)	Damage (%)	
Top view	20.80	53.89	84.81	73.42	75.77	75.59	81.56	
Side view	13.56	64.38	84.09	56.24	66.73	80.29	72.33	
Bottom view	20.98	63.65	83.36	68.17	67.81	73.06	90.78	
Top and side view	28.93	99.28	81.56	79.75	72.88	69.44	81.19	
Bottom and side view	30.02	99.46	84.81	76.13	70.52	74.86	79.75	
Top and Bottom view	21.52	67.99	72.69	72.69	85.35	74.50	84.27	

In the three single-view comparisons, the model achieved higher Exact Match Ratios in the top and bottom views, with percentages of 20.8% and 20.98%, respectively. Regarding size classification of fruiting body, the model demonstrated better accuracy in the side and bot-tom views, with percentages of 64.38% and 63.65%, respectively. In terms of classification of fruiting bodythe number, the model achieved similar accuracy in the top and bottom views, with percentages of 74.09% and 74.46%, respectively. For shape classification, the model achieved the highest accuracy in the top view, reaching 73.42%. In color classification, the model performed best in the top view, with an accuracy of 75.77%. Regarding the presence of pigmentation, the model achieved the highest accuracy in the side view, reaching 80.29%. Lastly, in assessing the presence of damage, the model achieved the highest accuracy in the bottom view, reaching 90.78%. Comparing the results of single-view experiments, various phenotypic classification tasks achieved the highest accuracy in different views. The top view provides important clues for distinguishing the number, shape, and color of fruiting bodies, but its ability to distinguish size, pigmentation, and damage is relatively weaker compared to other views. However, the side and bottom views provide additional complementary information, leading to improved accuracy in various phenotypic classification tasks based on multi-view approaches.Fig. 10 Different grade tasks predict correct Venn diagrams of image sets from three single views.

Figure 10 presents the Venn diagrams of correctly classified images on the test set for each Fig. 10 presents the Venn diagrams of correctly classified images on the test set for each classification task corresponding to three different single-view. Figure 10a depicts the Venn diagram for the size classification of Auricularia cornea. In this diagram, the images classified correctly under top view, side view, and bottom view collectively occupy only 39% of the union of all correctly classified images, indicating significant differences in classified features provided by different views. Specifically, in the top view and side view sections of Fig. 10a, their intersection is 43%, with 20% of images correctly classified solely under top view and 32% under side view. Regarding size classification, the features provided by top view and bottom view within single views exhibit high similarity, but these two views differ significantly from the features provided by side view. Therefore, the combination of top and side view and bottom and side view provides more comprehensive classification features, resulting in a significant improvement in model accuracy. However, the improvement in model accuracy achieved by combining top and bottom view features is relatively limited. Additionally, Fig. 10b–e demonstrate that there is considerable overlap in classification features provided directly by the three views in number classification. In terms of shape, color, pigmentation, and damage classification, there are significant differences between the features provided by the three single views. Auricularia cornea phenotypes that cannot be accurately classified from a single viewpoint can achieve precise classification by integrating phenotype features from different viewpoints, thereby enhancing model accuracy.

The Venn diagram of single-view and multi-view classification results across three perspectives illustrates that multi-view methods provide a more comprehensive phenotype classification of Auricularia cornea fruiting bodies. The model developed in this study effectively extracts phenotype characteristics of Auricularia cornea fruiting bodies from multiple views.

Comparison between single-task and multi-task

Table 5 in this study further demonstrates the impact of single-task and multi-task classification on the accuracy of identifying and categorizing Auricularia cornea fruiting body phenotypes. Single-task and multi-task classification models were trained under various single-view and multi-view scenarios. The Exact Match Ratio of the multi-task models consistently surpassed that of the single-task models across all views. The multi-task models achieved notable Exact Match Ratios of 94.03% and 94.58% under the top and side view and the bottom and side view, respectively. Table 5 The exact match ratio of single-task grading and multi-task grading under different views is presented.

Different view	Accuracy of each indicator	
Top view (%)	Side view (%)	Bottom view (%)	Top and side view (%)	Bottom and side view (%)	Top and Bottom view (%)	
 Exact match ratio	 Single-task	20.80	13.56	20.98	28.93	30.02	21.52	
 Multi-task	24.59	18.44	24.59	94.03	94.58	28.02	

Figure 11 presents the Venn diagram of image sets accurately predicted by the multi-task model across six phenotype recognition tasks under the top and side view, bottom and side view, as well as top and bottom view. It is observed that the intersection of accurately predicted image sets from the top and side view, and the bottom and side view, constitutes 95% of their union. The Exact Match Ratio between these two views differs by only 0.%. Conversely, the model trained under the top and bottom view alone correctly predicts only 1% of the images independently when compared to the other two multi-view settings. Therefore, a combination of either top or bottom view with the side view enables accurate phenotype identification and classification of Auricularia cornea fruiting bodies.Figure 11 Use this study multi-task grading model to predict the correct set of images in venn diagrams from three distinct multi-view perspectives.

The six phenotype classification tasks of Auricularia cornea fruiting bodies are not independent, they exhibit significant correlations. When a single phenotype is classified in isolation, the model tends to focus solely on the features closely associated with that phenotype, thereby reducing its overall accuracy. This issue is exacerbated when the dataset size is small, and the model complexity is low. Multi-task learning addresses this problem by guiding the model to extract varied phenotypic features from the images for each task. Consequently, the trained model becomes attuned to a broader range of phenotypic characteristics. Multi-task learning enhances the model’s generalization ability, mitigates overfitting, and ultimately improves overall performance.

Results of the multi-view multi-label fast model

To improve the performance of our recognition system, we used 5-fold cross-validation to train the model. It aids in statistical measurements, such as classification accuracy for each phenotype and the absolute accuracy of correctly identifying all phenotypes. The experimental results, as shown in Table 6, indicate that our proposed method achieved an average absolute accuracy of 94.66% (±0.63%). Specifically, the model attained 100% accuracy for the number, pigmentation, and damage phenotypes; over 98% accuracy for shape and color; and 96.29% (±0.46%) average accuracy for size, which was the most challenging to train. Table 6 Validity assessment measures with 5-fold cross validation.

Fold	Exact match ratio (%)	Accuracy of each indicator	
Size (%)	Number (%)	Shape (%)	Colour (%)	Pigmentation (%)	Damage (%)	
0	94.03	95.84	100.00	99.28	98.46	100.00	100.00	
1	94.48	95.93	100.00	98.92	99.19	100.00	100.00	
2	94.98	96.75	100.00	99.01	99.19	100.00	100.00	
3	94.66	96.47	100.00	98.82	99.10	100.00	100.00	
4	95.20	96.47	100.00	99.73	99.00	100.00	100.00	
Average	94.66	96.29	100.00	99.15	98.99	100.00	100.00	

Figure 12 illustrates the training process over 5-fold cross-validation, with the vertical axis depicting absolute accuracy. As shown in the figure, our model exhibits stability across all five training runs, consistently maintaining stable performance after approximately 250 epochs.Figure 12 The exact match ratio for Auricularia cornea phenotype recognition and classification for 5-fold cross validation.

The results from the 5-fold cross-validation experiments confirm that our model is highly reliable, without issues of overfitting or underfitting.

Comparison with different lightweight stem CNN

In addition to the multi-view and multi-task fast network, this study conducted various backbone contrast experiments to demonstrate the superiority of our lightweight CNN. In our approach, the model architecture consists of two parts: a Lightweight Fast CNN serving as the backbone network and a class-specific attention classifier for multi-task classification. We replaced the backbone network with other classical lightweight CNN models, including MobileNet, EfficientNet, and ShuffleNet, to assess their performance.

The experimental results are shown in Table 7. Leveraging the class-specific attention classifier proposed in this study, the Exact Match Ratio of EfficientNet reached 96.02%, while MobileNet and ShuffleNet achieved Exact Match Ratios of over 93%. However, their models require more space and time to execute. Fortunately, the model proposed in this study maintains excellent performance with less runtime space and time. The Exact Match Ratio reached 94.66%. The model parameters of this study amount to 4.039 million, with a total size of 73.19 MB. It only takes 11.9ms to predict the phenotype of an Auricularia cornea fruiting body. Table 7 Exact match ratio of comparative experiments on different backbone networks.

Different backbone	Exact match ratio (%)	Speed (ms)	FLOPs	Params (M)	Model size (MB)	Precision (%)	Recall (%)	Macro-F1 (%)	
MobileNet	93.4	16.11	676.227 M	2.247	166.95	98.53	98.4	98.45	
Efficientnet	96.02	21.18	861.742 M	4.183	198.01	99.36	98.65	98.99	
Shufflenet	93.67	18.57	1.212 G	5.381	120.4	98.65	97.99	98.3	
Our method	94.66	11.9	978.368 M	4.039	73.19	99.41	99.47	99.44	

Discussion

The identification and classification of crop phenotypic traits is a critical issue in agriculture. Utilizing machine vision to achieve non-destructive and rapid phenotypic recognition and classification of agricultural products is a significant solution. Edible mushrooms, as an economically important crop, pose a challenge for accurate phenotypic classification due to their complex fruiting body morphology. Auricularia cornea, in particular, exhibits such complexity that single-view images are insufficient for comprehensive phenotype extraction. Therefore, this study collected images of Auricularia cornea from three perspectives and meticulously annotated six phenotypic traits. These images were then used as the dataset for training and testing the proposed model.

In the study conducted by Ref20–23, a traditional machine learning algorithm based on image processing was used to identify different phenotypic features. However, the implementation of these algorithms is challenging as each phenotypic feature requires a separate design. When analyzing multiple phenotypic features, the image analysis needs to be performed multiple times, which is time-consuming and difficult to applied to other phenotypic recognition and analysis tasks. In contrast, the study by Ref24–27 utilized deep learning models that typically rely on single-view images to classify the overall features of crops, resulting in an inability to accurately classify multiple distinct phenotypes simultaneously. This study addresses these challenges by designing a multi-view input, multi-task learning network for phenotypic identification and classification using deep learning methods.The model of this study was validated on the multi view and multi label Auricularia cornea dataset constructed in this study. From Tables 4 and 5, it is evident that using two views as model input achieves an Exact Match Ratio of 94.66%, with a parameter count of 4.039 million and a total model size of 73.19 MB. The average inference time per image is 11.9 milliseconds. This method meets the requirements for phenotypic classification of Auricularia cornea fruiting bodies and provides a meaningful reference for non-destructive, rapid phenotypic classification of other mushroom.

To ensure the practicality of the model, the backbone network of this study’s model is constructed based on PConv, achieving model lightweighting. Compared to traditional convolution, the structural design of the backbone network based on PConv reduces the model parameters from 9.36 to 4.039 M, decreases the model size from 96.53 to 73.19 MB, and lowers the average inference time per image from 19.36 ms to 11.9 ms. Consequently, the classification efficiency is enhanced, meeting real-time requirements, reducing hardware requirements for deployment devices, and expanding application scenarios. Typically, reducing the model parameters may weaken the model’s feature extraction capability, often necessitating enhancements in attention mechanisms, loss functions, etc., to overcome this issue28. Therefore, in this study, the lightweight backbone network is combined with channel attention, and a class-specific attention multi-task classifier is constructed. With a slight increase in parameters, it markedly enhances recognition accuracy and substantially reduces the model size. As demonstrated in Table 7, our algorithm achieves the smallest model size and the fastest classification speed among other mainstream lightweight algorithms while maintaining high recognition accuracy. This renders our model an effective solution for multi-view multi-task classification.

In summary, the multi-view, multi-task rapid network developed in this study provides a robust tool for crop phenotypic identification and classification. This method enables agricultural professionals to perform non-destructive phenotypic recognition and classification of agricultural products with both speed and accuracy, with low hardware requirements.

Conclusions

In this study, we collected a dataset of Auricularia cornea featuring three views and six labels and proposed a method for phenotypic feature extraction. Compared to other current phenotypic recognition and classification methods, our model combines information from two different views of Auricularia cornea fruiting bodies and accurately classifies six phenotypes in a single recognition process. Our experiments demonstrated that phenotypic information varies across different views, and our model effectively integrates this information. Multi-task learning, as opposed to single-task learning, captures more phenotypic information from the fruiting bodies, resulting in superior generalization performance. By leveraging multi-view data and multi-task learning, we significantly improved the classification accuracy of each phenotype. Our method achieved a classification accuracy of 94.66% and a prediction speed of 11.9 ms. Compared to other representative lightweight deep learning models, our model requires lower hardware specifications and achieves the highest speed while maintaining high accuracy. Therefore, we conclude that our method can enhance the accuracy and efficiency of Auricularia cornea quality grading, reduce costs, and serve as an automated phenotypic analysis tool for breeding programs. In future work, we plan to utilize our findings to design new automated methods for the recognition and classification of other phenotypes in mushroom phenotyping, such as mycelium and spores.

Author contributions

Conceptualization, Y.X.; Data curation, Y.X. and Y.L.; Formal analysis, Y.X. and H.L.; Investigation, Y.X. and S.Q.; Methodology, Y.X. and Z.L.; Project administration, Z.L., L.Z.; Resources, Y.X., S.Q. and H.L.; Software, Y.X. and Y.L.; Validation, Y.X.; Visualization, L.W., L.Z.; Writing—original draft, Y.X.; Writing—review & editing, Y.X., L.Z., L.W. and Z.L. All authors reviewed the manuscript.

Funding

This research was funded by the Scientific Research Project of the Department of Education of Jilin Province, grant number JJKH20210331KJ; Jilin Province Science and Technology Development Plan Project, grant number: 20240305025YY; Scientific Research Project of the national key research and de-velopment program of China(2023YFD1301802); the national key research and development program of China(2023YFD1201600).

Data availibility

The datasets generated during and/or analysed during the current study are not publicly available, because they are currently being utilized for ongoing research projects but are available from the corresponding author on reasonable request.

Competing interests

The authors declare no competing interests.

Informed consent

Informed consent was obtained from all subjects involved in the study.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Li Y The status, opportunities and challenges of edible fungi industry in China: Develop with chinese characteristics, realize the dream of powerful mushroom industrial country J. Fungal Res. 2018 16 125 131
Li, Y. et al. The status, opportunities and challenges of edible fungi industry in China: Develop with chinese characteristics, realize the dream of powerful mushroom industrial country. J. Fungal Res. 16, 125–131 (2018).
2. Royse, D. J., Baars, J. & Tan, Q. Current overview of mushroom production in the world. Edible and Medicinal Mushrooms: Technology and Applications, 5–13 (2017).
3. Yuan X Research progress on mushroom phenotyping Mycosystema 2021 40 721 742
Yuan, X. et al. Research progress on mushroom phenotyping. Mycosystema 40, 721–742 (2021).
4. Tsang YP An intelligent model for assuring food quality in managing a multi-temperature food distribution centre Food Control 2018 90 81 97 10.1016/j.foodcont.2018.02.030
Tsang, Y. P. et al. An intelligent model for assuring food quality in managing a multi-temperature food distribution centre. Food Control 90, 81–97 (2018).10.1016/j.foodcont.2018.02.030
5. Wang F An automatic sorting system for fresh white button mushrooms based on image processing Comput. Electron. Agric. 2018 151 416 425 10.1016/j.compag.2018.06.022
Wang, F. et al. An automatic sorting system for fresh white button mushrooms based on image processing. Comput. Electron. Agric. 151, 416–425 (2018).10.1016/j.compag.2018.06.022
6. Liu Q Fang M Li Y Gao M Deep learning based research on quality classification of shiitake mushrooms Lwt 2022 168 113902 10.1016/j.lwt.2022.113902
Liu, Q., Fang, M., Li, Y. & Gao, M. Deep learning based research on quality classification of shiitake mushrooms. Lwt 168, 113902 (2022).10.1016/j.lwt.2022.113902
7. Wu Y A size-grading method of antler mushrooms using yolov5 and pspnet Agronomy 2022 12 2601 10.3390/agronomy12112601
Wu, Y. et al. A size-grading method of antler mushrooms using yolov5 and pspnet. Agronomy 12, 2601 (2022).10.3390/agronomy12112601
8. Zhu X Zhu K Liu P Zhang Y Jiang H A special robot for precise grading and metering of mushrooms based on yolov5 Appl. Sci. 2023 13 10104 10.3390/app131810104
Zhu, X., Zhu, K., Liu, P., Zhang, Y. & Jiang, H. A special robot for precise grading and metering of mushrooms based on yolov5. Appl. Sci. 13, 10104 (2023).10.3390/app131810104
9. Xu, Y. et al. Method for the classification of black fungus quality using mics-cotnet. Trans. Chin. Soc. Agric. Eng. 39, 5 (2023).
10. Shi X Chai X Yang C Xia X Sun T Vision-based apple quality grading with multi-view spatial network Comput. Electron. Agric. 2022 195 106793 10.1016/j.compag.2022.106793
Shi, X., Chai, X., Yang, C., Xia, X. & Sun, T. Vision-based apple quality grading with multi-view spatial network. Comput. Electron. Agric. 195, 106793 (2022).10.1016/j.compag.2022.106793
11. Chen, Y., Wu, Y., Cheng, J. & Tao, D. A deep multi-view learning method for rice grading. In 2019 IEEE International Conference on Real-time Computing and Robotics (RCAR), 726–730 (IEEE, 2019).
12. Chen W Liu M Zhao C Li X Wang Y Mtd-yolo: Multi-task deep convolutional neural network for cherry tomato fruit bunch maturity detection Comput. Electron. Agric. 2024 216 108533 10.1016/j.compag.2023.108533
Chen, W., Liu, M., Zhao, C., Li, X. & Wang, Y. Mtd-yolo: Multi-task deep convolutional neural network for cherry tomato fruit bunch maturity detection. Comput. Electron. Agric. 216, 108533 (2024).10.1016/j.compag.2023.108533
13. Wang D Wang J Ren Z Li W Dhbp: A dual-stream hierarchical bilinear pooling model for plant disease multi-task classification Comput. Electron. Agric. 2022 195 106788 10.1016/j.compag.2022.106788
Wang, D., Wang, J., Ren, Z. & Li, W. Dhbp: A dual-stream hierarchical bilinear pooling model for plant disease multi-task classification. Comput. Electron. Agric. 195, 106788 (2022).10.1016/j.compag.2022.106788
14. Zhu, K. & Wu, J. Residual attention: A simple but effective method for multi-label recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 184–193 (2021).
15. Ren Z Breeding of a new white Auricularia cornea ehrenb. white wood ear Mol. Plant Breed. 2018 16 954 959
Ren, Z. et al. Breeding of a new white Auricularia cornea ehrenb. white wood ear. Mol. Plant Breed. 16, 954–959 (2018).
16. Chen, J. et al. Run, don’t walk: Chasing higher flops for faster neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12021–12031 (2023).
17. Hu, J., Shen, L. & Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7132–7141 (2018).
18. Crawshaw, M. Multi-task learning with deep neural networks: A survey. http://arxiv.org/abs/2009.09796 (2020).
19. Kendall, A., Gal, Y. & Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7482–7491 (2018).
20. Huang X Jiang S Chen Q Zhao J Identification of defect pleurotus geesteranus based on computer vision Trans. Chin. Soc. Agric. Eng. 2010 26 350 354
Huang, X., Jiang, S., Chen, Q. & Zhao, J. Identification of defect pleurotus geesteranus based on computer vision. Trans. Chin. Soc. Agric. Eng. 26, 350–354 (2010).
21. Chen H-H Ting C-H The development of a machine vision system for shiitake grading J. Food Qual. 2004 27 352 365 10.1111/j.1745-4557.2004.00642.x
Chen, H.-H. & Ting, C.-H. The development of a machine vision system for shiitake grading. J. Food Qual. 27, 352–365 (2004).10.1111/j.1745-4557.2004.00642.x
22. Hwang H Development of on-line automatic grading and internet based real time production management system for shiitake Jpn. J. Food Eng. 2005 6 1 7 10.11301/jsfe2000.6.1
Hwang, H. Development of on-line automatic grading and internet based real time production management system for shiitake. Jpn. J. Food Eng. 6, 1–7 (2005).10.11301/jsfe2000.6.1
23. Chen H Xia Q Zuo T Tan H Bian Y Determination of shiitake mushroom grading based on machine vision Trans. Chin. Soc. Agric. Mach. 2014 45 281
Chen, H., Xia, Q., Zuo, T., Tan, H. & Bian, Y. Determination of shiitake mushroom grading based on machine vision. Trans. Chin. Soc. Agric. Mach. 45, 281 (2014).
24. Zuo, Y. & Zhao, M. Sa-efficientnet: Quality grading model of stropharia rugoso-annulate. In 2022 International Conference on Computer Engineering and Artificial Intelligence (ICCEAI), 358–362 (IEEE, 2022).
25. Li T Quality grading algorithm of oudemansiella raphanipes based on transfer learning and mobilenetv2 Horticulturae 2022 8 1119 10.3390/horticulturae8121119
Li, T. et al. Quality grading algorithm of oudemansiella raphanipes based on transfer learning and mobilenetv2. Horticulturae 8, 1119 (2022).10.3390/horticulturae8121119
26. Bakkouri, I. & Afdel, K. Convolutional neural-adaptive networks for melanoma recognition. In Image and Signal Processing: 8th International Conference, ICISP 2018, Cherbourg, France, July 2-4, 2018, Proceedings 8, 453–460 (Springer, 2018).
27. Bakkouri I Bakkouri S 2mgas-net: Multi-level multi-scale gated attentional squeezed network for polyp segmentation Signal Image Video Process. 2024 18 5377 5386 10.1007/s11760-024-03240-y
Bakkouri, I. & Bakkouri, S. 2mgas-net: Multi-level multi-scale gated attentional squeezed network for polyp segmentation. Signal Image Video Process. 18, 5377–5386 (2024).10.1007/s11760-024-03240-y
28. Ye, Y. et al. Channel pruning via optimal thresholding. In Neural Information Processing: 27th International Conference, ICONIP 2020, Bangkok, Thailand, November 18–22, 2020, Proceedings, Part V 27, 508–516 (Springer, 2020).
