
==== Front
J Bone Oncol
J Bone Oncol
Journal of Bone Oncology
2212-1366
2212-1374
Elsevier

S2212-1374(24)00109-X
10.1016/j.jbo.2024.100629
100629
VSI: MI Orthopedics
Radiographic imaging and diagnosis of spinal bone tumors: AlexNet and ResNet for the classification of tumor malignancy
Guo Chengquan a
Chen Yan a1⁎
Li Jianjun lij2046@126.com
a1⁎
a Department of Orthopedic Surgery, Shengjing Hospital of China Medical University, Shenyang, Liaoning 110000 China
⁎ Corresponding authors. lij2046@126.com
1 These authors contributed equally to this work and should be considered as co-corresponding authors.

18 8 2024
10 2024
18 8 2024
48 1006291 6 2024
8 8 2024
11 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Highlights

• CNNs innovate bone tumor diagnosis.

• AlexNet, ResNet enhance spinal tumor accuracy.

• 1532 radiological cases from 580 patients enhance research credibility.

• Enhanced diagnosis boosts personalized treatment outcomes.

• Enhanced diagnosis boosts personalized treatment outcomes.

Objective

This study aims to explore the application of radiographic imaging and image recognition algorithms, particularly AlexNet and ResNet, in classifying malignancies for spinal bone tumors.

Methods

We selected a cohort of 580 patients diagnosed with primary spinal osseous tumors who underwent treatment at our hospital between January 2016 and December 2023, whereby 1532 images (679 images of benign tumors, 853 images of malignant tumors) were extracted from this imaging dataset. Training and validation follow a ratio of 2:1. All patients underwent X-ray examinations as part of their diagnostic workup. This study employed convolutional neural networks (CNNs) to categorize spinal bone tumor images according to their malignancy. AlexNet and ResNet models were employed for this classification task. These models were fine-tuned through training, which involved the utilization of a database of bone tumor images representing different categories.

Results

Through rigorous experimentation, the performance of AlexNet and ResNet in classifying spinal bone tumor malignancy was extensively evaluated. The models were subjected to an extensive dataset of bone tumor images, and the following results were observed. AlexNet: This model exhibited commendable efficiency during training, with each epoch taking an average of 3 s. Its classification accuracy was found to be approximately 95.6 %. ResNet: The ResNet model showed remarkable accuracy in image classification. After an extended training period, it achieved a striking 96.2 % accuracy rate, signifying its proficiency in distinguishing the malignancy of spinal bone tumors. However, these results illustrate the clear advantage of AlexNet in terms of proficiency despite a lower classification accuracy. The robust performance of the ResNet model is auspicious when accuracy is more favored in the context of diagnosing spinal bone tumor malignancy, albeit at the cost of longer training times, with each epoch taking an average of 32 s.

Conclusion

Integrating deep learning and CNN-based image recognition technology offers a promising solution for qualitatively classifying bone tumors. This research underscores the potential of these models in enhancing the diagnosis and treatment processes for patients, benefiting both patients and medical professionals alike. The study highlights the significance of selecting appropriate models, such as ResNet, to improve accuracy in image recognition tasks.

Keywords

Spinal bone tumors
Radiographic imaging
AlexNet
ResNet
Convolutional neural networks
Deep learning image classification
==== Body
pmc1 Introduction

1.1 Review of spinal bone tumors and their causation

Primary bone tumors typically refer to abnormal growth originating within bone tissue, encompassing benign bone tumor-like lesions, benign bone tumors, and malignant bone tumors in a broad sense. Bone tumors are biologically classified based on cell or matrix origin [1]. The main categories include osteogenic tumors (derived from bone cells), chondrogenic tumors (derived from cartilage cells), fibrous tumors (derived from fibrous tissue), myogenic tumors (derived from muscle tissue), adipogenic tumors (derived from adipose tissue), vascular tumors (derived from vascular tissue), and tumors of undetermined origin. Subsequently, they are further classified as benign, intermediate, or malignant based on their differentiation characteristics and biological properties.

The spine, as the central axis of the human body, is divided into cervical, thoracic, lumbar, sacral, and coccygeal segments [2]. Most tumors that occur in the spine are malignant, with the majority being metastatic bone tumors, followed by primary malignant tumors, intermediate bone tumors, and benign tumors. Metastatic bone tumors are caused by cancer cells spreading to the bones from other parts of the body. They account for the vast majority of spinal tumors because when cancer cells spread through the bloodstream or lymphatic system in the body, they often initially target bone tissue. Bone tumors are not common in clinical practice, but they encompass a wide range of conditions. Patients with primary malignant spinal tumors typically present with persistent pain unrelated to activity and not relieved by rest, along with neurological symptoms such as weakness, spasms, sensory deficits, and potentially urinary or bowel incontinence due to compression of the spinal cord or nerve roots. Other signs may include palpable masses or swelling, abnormal sensations, and systemic symptoms like weight loss, low-grade fever, and fatigue.

Primary malignant tumors in the spine are relatively rare, with approximately 80 % of adult spinal tumors being malignant. Patients with primary malignant tumors typically present with pain, which is unrelated to their activity level and is not relieved at rest. Patients may experience weakness, spasms, corresponding sensory deficits, and even urinary and bowel incontinence when the lesion affects nerve roots. Malignant primary spinal cord tumors, such as ependymomas, lymphomas, and Ewing sarcomas, can also lead to systemic symptoms, including weight loss, low-grade fever, and fatigue [3].

The etiology of bone tumors needs to be better understood [4]. Potential factors contributing to the occurrence of bone tumors include genetic factors, history of radiation therapy, bone trauma, growth sites, and other factors such as chronic bone marrow inflammation, exposure to chemicals, and immune system abnormalities. Malignant bone tumors commonly manifest as dull pain in the affected bone area. Initially, the pain may be sporadic, but as the condition progresses, it intensifies and gradually becomes persistent, sometimes affecting sleep, especially with nocturnal pain.

1.2 Detection of spinal bone tumors based on diagnosis and prognosis methods

The symptoms of benign bone tumors are typically subtle, including pain, swelling or a lump, fractures, abnormal bone shape, and restricted movement. They are typically detected through imaging tests such as X-rays, CT scans, MRIs, and bone scans and sometimes require a biopsy for confirmation [5]. Other manifestations may include pathologic fractures, localized swelling, fever, night sweats, growth and development impairment, etc. The initial diagnosis of bone tumors is typically made through epidemiological characteristics, medical history, physical examination, and imaging studies. However, due to the complexity of their classification, the initial diagnosis can be challenging.

Bone tumors of an unclear nature often require a biopsy for macroscopic and histological assessment to guide subsequent treatment plans. A biopsy is a crucial step in obtaining tumor tissue samples. Through microscopic examination of histological features, the nature of the tumor, as well as its degree of differentiation and other pathological characteristics, can be determined.

Treatment options for bone tumors include surgery, radiation therapy, chemotherapy, targeted therapy, and radioisotope therapy. Treatment choice depends on factors such as the tumor's type, size, location, the patient's health status, the stage and extent of tumor spread, and individual factors.

In cases of poorly differentiated tumors, genetic testing and other auxiliary methods may be necessary for precise subtyping. The treatment approaches for bone tumors vary greatly depending on their nature. For benign bone tumors that do not involve surrounding tissues, regular follow-up is the primary approach, while large tumors may require routine excision.

Intermediate bone tumors usually necessitate a one-stage conventional or extended excision based on biopsy results or a two-stage surgical approach following biopsy. In the case of malignant tumors without metastasis, a one-stage extended excision is usually required, along with potential prosthetic implantation, followed by adjuvant radiotherapy and chemotherapy. Palliative treatment is considered if metastasis is confirmed.

The prognosis varies significantly for different types of bone tumors. Benign bone tumors generally have a favorable prognosis, with only a few cases potentially undergoing malignant transformation [6]. Non-metastatic intermediate and malignant bone tumors can often be cured with extensive excision. In cases of metastatic malignant bone tumors, the 5-year survival rate ranges from 23 % to 61 %. The prognosis highly depends on tumor biology, size, location, and other factors.

The diagnosis of bone tumors is the primary step in managing patients and holds significant value in guiding treatment plans. Early and accurate diagnosis allows for timely intervention in the disease process, alleviating disease progression and reducing the chances of tumor spread and metastasis. It also enhances the overall survival rate of patients, alleviates their physical and emotional distress, and reduces the economic burden.

Misdiagnosis or missed diagnosis may lead to patients missing the optimal treatment window, significantly affecting their quality of life and prognosis. The diagnosis of bone tumors relies on several criteria [7]: 1) Epidemiological characteristics, such as age, gender, and affected site; 2) Symptoms, including pain, swelling, pathologic fractures, etc.; 3) Physical signs, such as tenderness, venous congestion, elevated skin temperature, etc.; 4) Imaging examinations, which may reveal lesion locations on X-rays, lytic changes, infiltrative growth, periosteal reactions, sclerosis areas on CT scans, soft tissue infiltration, high signal intensity on MRI scans, increased radiotracer uptake on nuclear bone scans; 5) Macroscopic morphology of the tumor, describing its shape, color, texture, size, etc; 6) Histopathology of tumor tissue, examining cell atypia, vascular invasion, matrix distribution, etc; 7) Additional techniques, such as immunohistochemistry, electron microscopy, special staining, flow cytometry, karyotype analysis, protein/gene testing, etc.

However, conventional methods like medical history and imaging studies have their limitations. Epidemiological indicators are only effective in differentiating a few specific tumors, as most bone tumors generally exhibit a trend of onset at a young age, a preference for males, and non-specific affected sites. The symptoms and signs of tumors are often similar and lack specificity, with benign tumors frequently growing covertly and malignant tumors commonly presenting with pain and swelling.

Imaging evaluation applies only to a few tumors, with limited diagnostic value for most tumors, such as sporadic malignant tumors that lack typical imaging features. Therefore, relying solely on imaging for tumor diagnosis is not highly reliable. Histopathological diagnosis is considered the gold standard for determining the biological nature of tumors. In clinical practice, a lesion biopsy is performed for histopathological examination. A multidisciplinary approach involving orthopedic, radiological, and pathological experts is often employed to diagnose the tumor qualitatively. Therefore, precise and reliable histopathological diagnosis is paramount for formulating scientifically sound treatment plans for bone tumors.

Bone tumors are a type of neoplasm that predominantly affects the bones and their associated tissues, including nerves, blood vessels, and bone marrow. In clinical practice, bone tumors are not very common and represent only a fraction of all neoplastic diseases. Benign tumors and tumor-like lesions are more commonly found in the bones of the extremities, while primary tumors in the spine are relatively rare.

Due to the intricate anatomy of the spine and the overlapping structures [8], plain radiography has limited capabilities for observation. CT and MRI, as cross-sectional imaging techniques, can obtain multi-planar images, making them significant for tumor localization and qualitative diagnosis. In this study, we conducted a retrospective analysis of clinical data from 580 patients with bone tumors, aiming to explore the imaging characteristics of different diagnostic methods in spinal bone tumors. Deep learning and CNN-based image recognition can significantly enhance patients' diagnosis and treatment processes by clinically classifying bone tumors.

1.3 Automatic tumor diagnostics techniques based on artificial intelligence

As science and technology continue to advance, image recognition technology has made remarkable strides and is progressively finding applications in various facets of human life [9]. Image recognition technology represents a critical domain within artificial intelligence, aiming to replicate the human image recognition process through computer-based techniques [10]. It achieves the goal of image recognition by performing operations such as feature extraction and classification to gain a deeper understanding of the content of images.

Medical image recognition is one of the significant applications in the field of medical imaging. It leverages computer vision and deep learning technologies to assist doctors in identifying and analyzing anomalies, organ structures, and abnormalities within medical images. However, radiographic images used for training neural networks are influenced by variations in radiation intensity, beam angle, patient positioning, and image magnification. These variations can affect the consistency and quality of morphological data, thereby influencing the training effectiveness of deep neural networks. In recent years, medical image recognition algorithms based on CNN have achieved remarkable success, leading to revolutionary advancements in medical imaging [11].

With the assistance of deep learning, accurately and efficiently categorizing primary bone tumors at the histopathological level based on their degree of invasiveness [12] as benign, intermediate, or malignant can not only overcome the limitations of pathologists' experiential constraints and subjective biases but also allow for quantitative assessments of the biological nature of tumors. This, in turn, facilitates the optimization of subsequent treatment plans for the disease. In this study, we utilized artificial intelligence deep learning theory as a foundation, established a comprehensive database of radiographic images for all categories of bone tumors, employed various neural network frameworks to construct effective deep learning models, and compared their performance with radiologists of different expertise levels [13]. The study validated the feasibility and effectiveness of using deep learning for qualitative classification in primary bone tumor histopathology.

This research provides a potential solution to the existing challenges in artificial bone tumor histopathological diagnosis. It offers a crucial theoretical basis for the future application of deep learning in histopathological diagnosis of all categories of bone tumors [14]. Examples of the different categories of these tumors are shown in Fig. 1.Fig. 1 Image classification and labeling of tumors in different states.

2 Methods

2.1 Principles of image recognition

To let the model recognize a tumor in a spinal X-ray scan is to let the model recognize whether there are specific features in this image, such as tumor texture and structure, specific shapes of lesions and edge features, etc. Then, based on these features, the model can conclude whether or not this spinal radiograph is abnormal.

2.2 CNN

The CNN is a deep learning architecture that relies on convolutional operations to extract image features. It has found extensive applications in computer vision research, including image classification, image retrieval, object detection, image segmentation, and image feature transfer. Compared to conventional neural networks, CNNs offer the advantages of cost-effectiveness and high classification accuracy.

2.2.1 The features of a CNN

CNNs employ convolutional kernels, or feature filters, to extract features from images. These kernels slide over the input image to generate a new feature map, capturing a portion of the image's features. The feature map is then divided into segments using a maximum pooling operation, where the maximum value from each segment is selected and organized in a maximum pooling matrix, highlighting the most salient features.

Subsequently, the pooled data is flattened to create one-dimensional data. These extracted feature maps are combined to produce a probability distribution for classifying predicted images through the Softmax layer. This process allows CNNs to effectively identify and classify objects or patterns within images.

2.2.2 CNN structure

Convolutional neural networks consist of several vital layers: input, convolutional, pooling, and output. These layers work coordinated to process and analyze image data, as illustrated in Fig. 2.1) Convolutional layer: The convolutional layer is a critical component in a convolutional neural network. It involves the use of a sliding window, often referred to as a convolutional kernel, which moves across a feature map to extract image features. During this process, the convolutional kernel performs a dot product with the portion of the feature map it covers, and the results are mapped to the convolutional layer. In this specific case, both the depth of the convolutional kernel and the depth of the input feature matrix are 3. Furthermore, the depth of the output feature matrix matches the number of convolutional kernels, which is 2.

2) Pooling Layer: After the convolution process, the image will contain multiple channels. The pooling layer is responsible for downsampling the features produced by the convolutional layer. Its primary goals are to eliminate redundant information in the image, maintain the same number of channels, reduce the image's dimensions, and mitigate overfitting of the data. In practice, pooling can be categorized into two primary types: maximum pooling and average pooling.

3) Output Layer: The output layer consists of a fully connected layer and an activation function. The fully connected layer extracts the feature set and calculates the probability that the image belongs to a particular class. A fully connected layer means that every neuron is connected to all neurons in the preceding and succeeding layers, ensuring that the final output is based on comprehensive image information. In the output stage, the activation function can be either Sigmoid, returning values between 0 and 1 to indicate the probability of classifying the image as “yes,” or the Softmax function, providing probabilities for each category.

Fig. 2 CNN structure.

2.3 Algorithm based on CNN

2.3.1 AlexNet model

In 2012, Alex Krizhevshy introduced the AlexNet deep learning model, comprising 11 convolutional neural networks. This model includes 5 convolutional layers, 3 pooling layers, and 3 fully connected layers. The structural arrangement of the AlexNet model is illustrated in Fig. 3. Notably, the size of the convolutional kernels progressively decreases from 11 to 5 to 3, and the feature values are halved through the application of maximum pooling for downsampling. The image features have been comprehensively extracted when the fifth convolutional layer is reached. Ultimately, the final classification result is derived by combining two fully connected layers followed by a Softmax layer.Fig. 3 AlexNet Model Structure.

The size calculation of the feature matrix after convolution or pooling is shown in Equation (1), where W is the height and width of the input feature matrix, F is the size of the convolution or pooling kernel, and P represents the number of zeros added around the matrix. For example, padding [1], [2] represents adding 1 column of zeros to the left, 2 columns of zeros to the right, 1 column of zeros above, and 2 columns of zeros below the feature matrix. S is the step size of convolution or pooling. If the step size S is greater than 1, the calculation of the output feature matrix is Equation (2). The central convolution and pooling steps are as follows:(1) N=W-F+2PS+1

(2) N=W-FS+1

1) In the Conv1 stage, the input is [227,227,3], the convolutional kernel size is 11, padding [1], [2], the step size is 4, and the number of convolutional kernels is 48 × 2 = 96, according to formula 2, N=(227-11)/4+1=55, the output after convolution [55,55,96].

2) In the Maxpool1 stage, the input is [55,55,96], the pooling kernel size is 3, the padding is 0, and the step length is 2. According to formula 2, N=(55-3)/2+1=27, and the output is [27,27,96], it can be seen that the maximum pooling downsampling operation only changes the height and width of the feature matrix and does not change the depth.

3) In the Conv2 stage, the input is [27,27,96], the convolutional kernel size is 5, the padding is [2], [2], the step size is 1, and the number of convolutional kernels is 128 × 2 = 256, according to formula 1, N=(27-5+4)/1+1=27, the output after convolution is [27,27,256].

4) In the Maxpool2 stage, the input is [27,27,256], the size of the pooling kernel is 3, the padding is 0, and the step length is 2. According to formula 2, N=(27-3)/2+1=13, and the output is [13,13,256]. In the Maxpool2 stage, the input is [27,27,256], the size of the pooling kernel is 3, the padding is 0, and the step length is 2. According to formula 2,N=(27-3)/2+1=13, and the output is [13,13,256].

5) In the conv3 stage, the input is [13,13256], the convolutional kernel size is 3, the padding is [1], [1], the step length is 1, and the number of convolutional kernels is 192 × 2 = 384, according to formula 1, N=(13-3+2)/1+1=13, the output after convolution is [13,13,384].

6) In the Conv4 stage, the input is [13,13,384], the convolutional kernel size is 3, the padding is [1], [1], the step length is 1, and the number of convolutional kernels is 192 × 2 = 384, according to formula 1, N=(13-3+2)/1+1=13, output [13,13,384].

7) Conv5 stage, input [13,13,384], convolution kernel size 3, padding [1], [1], step length stripe 1, number of convolution kernels 128 × 2 = 256, using formula 1, N=(13-3+2)/1+1=13, convolutional output [13,13,256].

8) In the Maxpool3 stage, input [13,13,256], pooling kernel size 3, padding 0, step length stripe 2, and output [6,6,256] are obtained from formula 2, where N=(13-3)/2+1=6.

9) The AlexNet model has made significant progress compared to previous models mainly due to the following five characteristics. Firstly, GPU was used for accelerated network training for the first time, significantly improving the training speed. A GPU can speed up by 20 to 50 times compared to a CPU. Secondly, the ReLU activation function was used to simplify the calculation of derivatives and effectively solve the gradient vanishing problem in deep networks. Thirdly, LRN local response normalization was used to make the values with more significant responses relatively larger, improving the model's accuracy. Fourthly, maximum overlap pooling means that the pooling operation overlaps on some pixels. In the pooling operation, the step size S is smaller than the pooling kernel F to avoid overfitting. Fifth, in the first two layers of the fully connected layer, Dropout is used to randomly inactivate a portion of neurons, indirectly reducing network training parameters and reducing overfitting during model training.

2.3.2 ResNet model

Ever since the inception of the AlexNet model [15], the application of deep learning convolutional neural networks in computer vision research has grown progressively prevalent. Deep learning entails using deep neural networks to extract features and data mining [16]. The depth of the neural network corresponds to the complexity of the features extracted from the image. With the increase in network depth and model size, Microsoft Research Asia introduced the deep residual network ResNet, which extended its depth to 152 layers in 2015. ResNet uses residual structures and avoids Dropout, which helps mitigate gradient vanishing, explosion, and degradation issues.

Fig. 4 illustrates the network structure formed by straightforwardly stacking convolutional layers and pooling layers. It is evident that the 48-layer network exhibits higher training and testing set errors compared to the 10-layer network. Therefore, there are better solutions than solely stacking networks through convolutional layers and employing maximum pooling downsampling layers. There are two primary reasons for this. Firstly, as the layers increase in depth, the issue of vanishing or exploding gradients becomes more pronounced. Assuming that the error gradient for each layer is less than 1, during the backpropagation process, it gets multiplied by an error gradient of less than 1 as it propagates forward through each layer.Fig. 4 Forward propagation without and with Dropout.

As the network's depth increases, more coefficients get multiplied by values less than 1. Consequently, they gradually approach 0, causing the gradient to become smaller. Conversely, if the gradient for each layer is greater than 1, during backpropagation, each layer it passes through multiplies it by a number greater than 1. As the number of layers deepens, the gradient becomes larger, potentially leading to a gradient explosion. The second issue is degradation, which implies that even after addressing problems related to vanishing and exploding gradients, there is still a phenomenon where the performance of deep layers could be more effective than that of shallow layers.

The ResNet model employs the following two methods to address the issues of gradient vanishing, explosion, and degradation:1) Dropout is discarded in favor of Batch Normalization to expedite the training process. BN accelerates network convergence and enhances accuracy by ensuring that each feature matrix dimension adheres to a distribution with a mean of 0 and a variance of 1. Throughout the training phase, the moving average technique is employed to track the mean and variance. Upon the completion of training, the statistical mean and variance closely approximate the mean and variance of the entire training dataset. Larger batch sizes result in an even closer approximation to the distribution of the entire dataset. This approach accelerates network convergence and enhances accuracy.

2) Propose residual structures to address degradation issues. Residual is the deviation between the predicted value and the actual value. Fig. 5A shows the residual structure used for networks with fewer layers in ResNet, and Fig. 5B proposes a residual structure for networks with more layers in ResNet.Fig. 5 ResNet network configuration. (A) ResNet-34 Residual Analysis; (B) ResNet-50, ResNet-101, ResNet-152 Residual Analysis; and (C) First layer residual structure.

The main pathway depicted on the left side of Fig. 5A involves the feature matrix passing through two 3x3 convolutional layers, resulting in an output termed the main branch. In contrast, the curved line on the right side represents a connection between the input and the output, signifying an identity mapping that preserves the input without any modification. This is referred to as the shortcut branch. The feature matrix, obtained after it passes through a sequence of convolutional layers within the main branch, is combined with the input feature matrix. Following this combination, a non-linear ReLU activation function is applied.

In Fig. 5B, the residual structure follows a specific pattern. It begins with a 1x1 convolutional layer, proceeds to a 3x3 convolutional layer, and finally passes through another 1x1 convolutional layer. The final 1x1 convolutional layer maintains the height and width of the feature matrix while reducing the depth from 256 to 64, essentially reducing dimensionality. With the use of 256 convolutional kernels in the final convolutional layer, the output and input feature matrices maintain consistent dimensions. In other words, the input and output feature matrices have a depth of 256.

Upon careful calculation, it is determined that the residual structure illustrated in Fig. 5A requires approximately 1,170,648 parameters. In contrast, the residual structure in Fig. 5B involves approximately 69,632 parameters. As a result, ResNet utilizes a sequence of stacked residual structures, with an increase in the number of such structures leading to the conservation of more parameters. The ResNet architecture, illustrated in Table 1, maintains a consistent structure across different layer configurations, whether 18, 34, 50, 101, or 152 layers. In the context of the 34-layer version, it initiates with a sequence of layers. The architecture commences with a 7x7 Convolutional layer, which produces 64 output channels and employs a stride of 2. Following this, there are three 3x3 Maximum pooling layers, each with a stride of 2. Subsequently, the architecture includes three conv2, four conv3, six conv4, and three conv5 residual structures, denoted as conv2_x, conv3_x, conv4_x, and conv5_x, respectively. (See Fig. 5C).Table 1 ResNet Network Structure. The number of FLOPs (Floating Point Operations) for each configuration is also listed. Based on our work's ResNet variant and dataset, specific parameter values (such as kernel sizes and number of output channels) must be provided.

layer name	output size	18-layer	34-layer	50-layer	101-layer	152-layer	
conv1	112 × 112	7 × 7,64, stride 2	


	
conv2_x	56 × 56	3 × 3 max pool, stride 2	
3×3,643×3,64×2	3×3,643×3,64×3	1×1,643×3,641×1,256×3	1×1,643×3,641×1,256×3	1×1,643×3,641×1,256×3	


	
conv3_x	28 × 28	3×3,1283×3,128×2	3×3,1283×3,128×4	1×1,1283×3,1281×1,512×4	1×1,1283×3,1281×1,512×4	1×1,1283×3,1281×1,512×8	


	
conv4_x	14 × 14	3×3,2563×3,256×2	3×3,2563×3,256×6	1×1,2563×3,2561×1,1024×6	1×1,2563×3,2561×1,1024×23	1×1,2563×3,2561×1,1024×36	


	
conv5_x	7 × 7	3×3,5123×3,512×2	3×3,5123×3,512×3	1×1,5123×3,5121×1,2048×3	1×1,5123×3,5121×1,2048×3	1×1,5123×3,5121×1,2048×3	


	
	1 × 1	Average pool, 1000-d fc, softmax	


	
FLOPs	1.8 × 109	3.6 × 109	3.8 × 109	7.6 × 109	11.3 × 109	

The first layer of each residual structure accommodates varying input and output dimensions, ensuring flexibility across these components. For example, in conv3_x, the initial input dimensions are [56, 56, 64], while the output dimensions become [28, 28, 128], thus maintaining congruity between the input and output of the primary and shortcut branches. Within the shortcut branch, a 1x1 convolutional kernel is utilized for dimensionality reduction; this results in a halving of the feature matrix's height and width while ensuring the depth aligns with the needs of the subsequent residual structure. Finally, the network employs average pooling for downsampling, followed by fully connected layers and a softmax layer to transform the output into a probability distribution.

2.4 Data acquisition and model training

2.4.1 General data

This study included 580 patients treated in our hospital with primary spinal bone tumors from January 2016 to December 2023. All patients were diagnosed by pathology and underwent X-ray examination [17]. The general information of the patients is shown in Table 2. The Ethics Committee of Shengjing Hospital approved this study, and informed consent was obtained from the patients or their legal guardians.Table 2 General data of the patients.

Parameters	Subject (n = 580)	
Patients with benign tumors (n, %)	295 (50.862 %)	
Patients with malignant tumors (n, %)	285 (49.138 %)	
Age (years)	19–84	
Average age (years)	27.63 ± 7.82	
Chest and dorsal discomfort	453 (78.103 %)	
Restricted range of motion	63 (9.138 %)	
Sweating irregularities	127 (21.897 %)	
Localized inflammation	36 (6.207 %)	
Radiating discomfort	121 (20.862 %)	
Elevated body temperature	136 (23.448 %)	

In total, 1,532 images were extracted from this imaging dataset. Among these, 679 images correspond to benign bone tumors, while the remaining 853 images represent malignant tumors. Two experienced professional radiologists labeled these images; their annotations are considered the ground truth data.

The dataset was subsequently divided into a training set and a validation set, following a ratio of 2:1 [18]. In the training model, the predefined parameters are the number and size of convolutional kernels, the initial value of weights in the convolutional kernel, the padding, and the step size, and the weights and deviation values in the convolutional kernel are obtained through training, and the classification and probability of the image are predicted according to the parameters obtained from the training. The training of the ResNet network is based on the method of transfer learning, and it is trained using the parameters that others have trained before, which can be quickly and easily compared even when the data set is small. The ResNet network is trained using the pre-trained parameters of others so that it can be trained quickly and produce satisfactory results even when the data set is small. Experiments were conducted to train the models, AlexNet and ResNet, respectively, and the same lung image was used for prediction. Experiments were conducted to compare each model's loss and correct rates in the training and validation sets and to predict the probability of correctly categorizing the image [19]. The number of iteration epochs used by the models in this experiment is 10, and the number of iterations can be increased appropriately if better training results are desired. The number of iterations determines the duration of the training process and the accuracy of the models in predicting the categories.

3 Results

3.1 Training of algorithms based on CNN

The training and test errors versus iteration of AlexNet and ResNet are performed. Fig. 6 shows that the left side figure shows forward propagation for AlexNet, and the correct side figure shows forward propagation for ResNet.Fig. 6 Comparison of error rates between 48-layer and 10-layer on training and testing sets. Two models, AlexNet and ResNet, are evaluated for performance.

3.2 Comparison results of algorithms based on CNN

Through experiments, we found that the models are trained under the same experimental conditions, with 10 iterations. The average time of each iteration is examined, and the probability of predicting the category is shown in Table 3. Through experiments, we found that the ResNet model has a slightly higher accuracy in predicting the category under the same conditions, the AlexNet model takes the significantly least time to train, and the iteration time is about 3 s for an epoch. The ResNet model has a terrible training time at about 32 s for an epoch but a slightly better prediction accuracy. Wholistically, AlexNet is more proficient in diagnosing bone tumors than ResNet.Table 3 Results of model comparison experiments.

	AlexNet	ResNet	
Time each epoch took (Unit seconds)	3	32	
Accuracy (Predicting category probabilities)	95.6 %	96.2 %	

4 Discussion

In this study, we utilized a limited set of radiographic resources to train various deep neural networks with classical architectures. Our findings indicate that the models generated are practical in distinguishing spinal bone tumor types based on their invasiveness at the histopathological level. Among the networks evaluated, AlexNet outperformed others in discriminating at the image patch level and surpassed pathologists with intermediate and junior titles in diagnostic performance at the slide level, reaching a level comparable to pathologists with senior titles. Furthermore, we observed that neural network models could efficiently extract specific visual regions within spinal bone tumor types and rely on these features to identify regions for correct classification and diagnostic predictions [20].

Research on artificial intelligence for diagnosing bone tumors primarily focuses on interpreting bone tumor radiological images. Bao et al. integrated clinical information, epidemiological data, and manually selected radiological features from X-ray images of clinical bone tumor patients (a total of 18 learnable features) and built a machine-learning model for the preliminary classification of a wide range of bone tumor types (29 in total) using a Bayesian algorithm. Depending on the sample size, the Bayesian machine learning model achieved diagnostic accuracy rates ranging from 44 % to 62 % in the Top-1 classification for differentiating bone tumor types based on clinical information and radiological features. This study was the first to explore the feasibility of using machine learning for the initial differentiation of bone tumors with a relatively large sample size, providing a theoretical basis and reference for future research in bone tumor classification using AI technology [21]. However, the main limitation of the above Bayesian model is that all features in the machine learning model need human judgments (rather than being directly extracted from the original data). Different evaluators' professional experience, subjective interpretation, and real-time working conditions can influence the judgment results of the same feature in the same case, which can negatively affect the quality of the labeled training set samples and the training results of the model.

Yu et al. integrated radiological X-ray information from multiple medical centers of bone tumor patients. They used convolutional neural network algorithms to build a deep neural network classification model solely based on regular bone tumor X-ray data to predict the invasiveness of bone tumors. The neural network model ultimately achieved a ROC AUC of 0.877 in binary classification (benign vs. non-benign) and a prediction performance level of 0.560 in ternary classification (benign, intermediate, and malignant), which not only outperformed junior radiologists but also reached a level of discrimination comparable to radiologists specializing in musculoskeletal radiology. However, for bone tumors, radiological image data for AI model training are relatively scarce because a single bone tumor lesion can only produce a few X-ray images taken from different angles, which are low-density two-dimensional morphological data and can only show the bony regions that can be visualized with radiographic imaging (without displaying other non-bony regions like cartilage and fibrous tissue). Typically, to train a deep learning classification model with a low overfitting level and good performance containing millions of parameters, the amount of data for each type in the training set needs to be much larger than the number of X-ray images of benign, intermediate, and malignant bone tumors (reaching thousands to tens of thousands of samples for each type) in the above studies. Therefore, more than a limited number of original bone tumor radiographic data may be required for deep models to accurately capture/extract discriminative features from the morphological information, thus limiting the model's generalization capability.

Furthermore, the pixel data of radiological images of bone tumors used for training neural networks may be subject to interference from radiation intensity, irradiation angles, patient positions, and film magnification ratios, causing variations in the morphological data of the same case under different conditions and ultimately affecting the training effectiveness of deep neural networks. In this research, convolutional neural network models demonstrated good learning curves in distinguishing bone tumor morphology information. AlexNet and ResNet achieved relatively good discrimination performance in the dataset provided in this study [22].

Research on using artificial intelligence (AI) to diagnose bone tumors primarily focuses on interpreting bone tumor imaging studies. Bao et al. integrated clinical information, epidemiological data, and selected imaging features (18 learnable features) from clinical bone tumor patients. They used a Bayesian algorithm to construct a machine-learning model capable of preliminary differentiation of a wide range of bone tumor types (29). This Bayesian machine learning model achieved a top diagnosis accuracy ranging from 44 % to 62 %, depending on the sample size. This study is the first to explore the feasibility of using machine learning with a relatively large sample size for preliminary differentiation of bone tumor types based on imaging studies, providing a theoretical basis and reference for subsequent research on bone tumor classification and differentiation using AI technology.

However, the major limitation of the Bayesian model mentioned above is that all features in the machine learning model need to be manually selected (rather than extracted directly from raw data). The differences in expertise, subjective interpretation, and real-time working conditions of different evaluators may affect the interpretation of the same feature in the same case, negatively impacting the quality of the training set samples and the model training results. Yu et al. integrated imaging data from bone tumor patients across multiple medical centers and used convolutional neural network algorithms to build a deep neural network classification model based solely on regular bone tumor plain X-ray data to predict the aggressiveness of bone tumors.

In this study, the radiographic information available for training AI models is relatively limited. This is because only a few X-ray images can be obtained for a given bone tumor lesion, taken from different angles, and these images provide sparse two-dimensional morphological information. Additionally, they only capture the bony regions that can be visualized under radiographic exposure (excluding soft tissue, cartilage, and other non-bony areas). Training a deep learning classification model with relatively low overfitting and high performance, which contains millions of parameters, requires much larger data for each type than the number of plain X-ray images available for benign, intermediate, and malignant bone tumors in the studies above. Thus, using a limited number of raw bone tumor X-ray images may need more data for deep models to capture and extract discriminative features from morphological information accurately, ultimately limiting the model's generalization ability. Furthermore, the pixel data from bone tumor plain X-rays used for training neural networks may be affected by variations in radiation intensity, irradiation angles, patient positioning, and image magnification, which can cause morphological data from the same case to change under different circumstances, ultimately affecting the training efficacy of deep neural networks.

AlexNet exhibited relatively better differentiation performance in the dataset provided. This study used over four hundred thousand images as training data and the best-performing neural network demonstrated relatively superior performance in patient-level bone tumor diagnosis based on the CNN models trained using bone tumor radiographic information, surpassing junior and mid-level pathology doctors and achieving a level of differentiation comparable to senior-level pathology doctors [23].

5 Conclusion

In summary, this study has investigated the utility of radiographic imaging for diagnosing spinal bone tumors. It has assessed the effectiveness of image recognition algorithms, specifically AlexNet, in classifying the malignancy of these tumors. The principal findings of this research are as follows:

Firstly, bone tumors affecting the spine, while relatively uncommon, exhibit a diverse range of conditions. Due to the complex anatomical structures within the spine, patients often present with pain in various regions, such as the neck, chest, and lower back, as well as signs of spinal cord and nerve root compression. In such cases, imaging examinations are pivotal in diagnosing bone tumors. Secondly, this study employed AlexNet to classify spinal bone tumor images and compared them against ResNet. The results of these models' training and comparison revealed that the ResNet model excelled in image classification accuracy, achieving a remarkable 96.2 % accuracy rate. While AlexNet exhibited faster training times, its accuracy slightly lagged. Finally, integrating deep learning and convolutional neural network-based image recognition technology holds great promise for the qualitative classification of bone tumors. With the training of deep learning models, it becomes possible to accurately categorize bone tumors as benign, intermediate, or malignant. This approach overcomes the constraints of subjective judgment and limited experience commonly associated with pathologists, optimizing the treatment plans for bone tumors, improving patients' survival rates, and reducing the physical and emotional burdens they endure.

In summary, this research underscores the potential of deep learning in bone tumor image classification, particularly in selecting appropriate models such as ResNet to achieve high accuracy in image recognition tasks. As technology advances, integrating these models into clinical practice promises to significantly enhance patients' diagnosis and treatment processes, benefiting both patients and medical professionals alike.

Ethical approval

The Ethics Committee of Shengjing Hospital approved this study, and the patients or guardians signed the informed consent.

Funding

Fund number: 81971829. National Natural Science Foundation of China project: Study the mechanism of CD98-mediated Glutamine metabolism-regulating Macrophage polarization in delayed fracture healing after Splenectomy.

CRediT authorship contribution statement

Chengquan Guo: Writing – original draft, Data curation. Yan Chen: Visualization, Validation, Supervision. Jianjun Li: Writing – review & editing, Project administration, Funding acquisition.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
==== Refs
References

1 Boriani S. Bandiera S. Casadei R. Boriani L. Donthineni R. Chondrosarcoma of the mobile spine: report on 22 cases Spine 42 5 2017 292 298
2 Demura S. Kawahara N. Murakami H. Nambu K. Kato S. Assessment of the risk of adjacent vertebral body fracture after posterolateral fusion with massive bone autografts in patients with lumbar spinal stenosis using a decision tree analysis Eur. Spine J. 25 2 2016 596 602 26153679
3 Chang K.L. Chen Y.C. Yang C.Y. Primary malignant tumors of the spine: a 42-year nationwide population-based epidemiological study World Neurosurg. 122 2019 e771 e779
4 R.L. Siegel, K.D. Miller, A. Jemal, Cancer statistics, 2020. CA Cancer J. Clin., 70(1) (2020) 7-30.
5 Weil A.G. Bone tumors: a practical guide to imaging Am. J. Med. 2020
6 Damron T.A. Ward W.G. Staging and surgical treatment of primary and metastatic bone tumors Clin. Orthop. Relat. Res. 474 3 2016 120 126 26280681
7 Kilpatrick S.E. Dickey I.D. A classification of primary bone neoplasms: a review Semin. Diagn. Pathol. 16 3 2019 186 200
8 Cai L. Zhang C. Wei L. Lin W. Xu L. Radiological and clinical features of spinal bone tumors J. Orthop. Surg. Res. 15 1 2020 1 7 31900192
9 Krizhevshy A. Sutskever I. Hinton G.E. ImageNet classification with deep convolutional neural networks Commun. ACM 60 6 2017 84 90
10 Esteva A. Kuprel B. Novoa R.A. Ko J. Swetter S.M. Blau H.M. Thrun S. Dermatologist-level classification of skin cancer with deep neural networks Nature 542 7639 2017 115 118 28117445
11 Litjens G. Kooi T. Bejnordi B.E. Setio A.A.A. Ciompi F. Ghafoorian M. Ginneken B. A survey on deep learning in medical image analysis Med. Image Anal. 42 2017 60 88 28778026
12 Wang S. Kang Y. Varma M. Liu M. Deep learning for primary bone tumor classification Front. Genet. 11 2020 154 32194630
13 Zhang J. Xie Y. Zhang W. Yang J. Primary bone tumor classification based on deep learning J. Vis. Commun. Image Represent. 60 2019 289 295
14 Lecun Y. Bengio Y. Hinton G. Deep learning Nature 521 7553 2015 436 444 26017442
15 He K. Zhang X. Ren S. Sun J. Deep residual learning for image recognition Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016 770 778
16 K. Simonyan, A. Zisserman, Intense convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).
17 Wang H. Guo W. Zhu J. Yu K. Wang Y. Ye J. Xu S. The radiological evaluation of spinal tumors Cancer Imaging 20 1 2020 1 14
18 Ronneberger O. Fischer P. Brox T. U-Net: Convolutional networks for biomedical image segmentation International Conference on Medical Image Computing and Computer-Assisted Intervention 2015 234 241
19 Deng J. Dong W. Socher R. Li L.J. Li K. Fei-Fei L. Imagenet large scale visual recognition challenge Int. J. Comput. Vis. 115 3 2009 211 252
20 Jamaludin A. Kadir T. Zisserman A. SpineNet: automated classification and evidence visualization in spinal MRIs Med. Image Anal. 41 2017 63 73 28756059
21 Nadeem M.W. Goh H.G. Ali A. Bone age assessment empowered with deep learning: a survey, open research challenges, and future directions Diagnostics 10 10 2020 781 33022947
22 Tawalbeh S. Alquran H. Alsalatie M. Deep feature engineering in colposcopy image recognition: a comparative study Bioengineering 10 1 2023 105 36671677
23 Tao Y. Huang X. Tan Y. Qualitative histopathological classification of primary bone tumors using deep learning: a pilot study Front. Oncol. 11 2021 735739
