
==== Front
PLoS One
PLoS One
plos
PLOS ONE
1932-6203
Public Library of Science San Francisco, CA USA

10.1371/journal.pone.0307912
PONE-D-23-40540
Research Article
Biology and life sciences
Cell biology
Chromosome biology
Chromatin
Chromatin modification
DNA methylation
Biology and life sciences
Genetics
Epigenetics
Chromatin
Chromatin modification
DNA methylation
Biology and life sciences
Genetics
Gene expression
Chromatin
Chromatin modification
DNA methylation
Biology and life sciences
Genetics
DNA
DNA modification
DNA methylation
Biology and life sciences
Biochemistry
Nucleic acids
DNA
DNA modification
DNA methylation
Biology and life sciences
Genetics
Epigenetics
DNA modification
DNA methylation
Biology and life sciences
Genetics
Gene expression
DNA modification
DNA methylation
Medicine and Health Sciences
Oncology
Cancers and Neoplasms
Genitourinary Tract Tumors
Prostate Cancer
Medicine and Health Sciences
Urology
Prostate Diseases
Prostate Cancer
Medicine and Health Sciences
Oncology
Cancers and Neoplasms
Medicine and Health Sciences
Nephrology
Renal Cancer
Computer and Information Sciences
Neural Networks
Biology and Life Sciences
Neuroscience
Neural Networks
Medicine and Health Sciences
Health Care
Health Care Policy
Treatment Guidelines
Medicine and Health Sciences
Diagnostic Medicine
Cancer Detection and Diagnosis
Medicine and Health Sciences
Oncology
Cancer Detection and Diagnosis
Engineering and Technology
Management Engineering
Decision Analysis
Decision Trees
Research and Analysis Methods
Decision Analysis
Decision Trees
Diagnostic classification based on DNA methylation profiles using sequential machine learning approaches
DNA methylation and machine learning approaches
https://orcid.org/0000-0003-2501-5201
Wojewodzic Marcin W. Data curation Formal analysis Funding acquisition Investigation Methodology Project administration Resources Software Writing – original draft Writing – review & editing 1 2 3 *
Lavender Jan P. Data curation Formal analysis Investigation Software Visualization Writing – original draft 3
1 Cancer Registry of Norway, Norwegian Institute of Public Health, Oslo, Norway
2 Chemical Toxicology, Norwegian Institute of Public Health, Oslo, Norway
3 University of Birmingham, Birmingham, United Kingdom
Reismann Marc Editor
Charité Universitätsmedizin Berlin CVK: Charite Universitatsmedizin Berlin - Campus Virchow-Klinikum, GERMANY
Competing Interests: The authors have declared that no competing interests exist.

* E-mail: Marcin.Wojewodzic@kreftregisteret.no
6 9 2024
2024
19 9 e03079124 12 2023
10 7 2024
© 2024 Wojewodzic, Lavender
2024
Wojewodzic, Lavender
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Aberrant methylation patterns in human DNA have great potential for the discovery of novel diagnostic and disease progression biomarkers. In this paper we used machine learning algorithms to identify promising methylation sites for diagnosing cancerous tissue and to classify patients based on methylation values at these sites. We used genome-wide DNA methylation patterns from both cancerous and normal tissue samples, obtained from the Genomic Data Commons consortium and trialled our methods on three types of urological cancer. A decision tree was used to identify the methylation sites most useful for diagnosis. The identified locations were then used to train a neural network to classify samples as either cancerous or non-cancerous. Using this two-step approach we found strong indicative biomarker panels for each of the three cancer types. These methods could likely be translated to other cancers and improved by using non-invasive liquid methods such as blood instead of biopsy tissue.

The author(s) received no specific funding for this work. Data Availability* The source code is available in GitHub repository: https://github.com/bazyliszek/methAI and in Zenodo: https://doi.org/10.5281/zenodo.11530987 * All example datasets were used in previously published studies and are available. We used publicly available and anonymous data from The Cancer Genome Atlas and Harmonized Cancer Datasets in the Genomic Data Commons Data Portal (https://www.cancer.gov/tcga, https://portal.gdc.cancer.gov/) investigating kidney, prostate and bladder tissues (dataset https://portal.gdc.cancer.gov/legacyarchive/search/f) assayed in multiple, independent experiments using the Illumina 450k microarray. The clinical metadata and the manifest file of the samples were downloaded through the GDC legacy archive on 2019-10-19 and 2019-11-09. We used the Genomic Data Commons (GDC) data download tool (https://gdc.cancer.gov/access-data/gdc-data-transfer-tool) using the manifest files deposited on our GitHub.
Data Availability

* The source code is available in GitHub repository: https://github.com/bazyliszek/methAI and in Zenodo: https://doi.org/10.5281/zenodo.11530987 * All example datasets were used in previously published studies and are available. We used publicly available and anonymous data from The Cancer Genome Atlas and Harmonized Cancer Datasets in the Genomic Data Commons Data Portal (https://www.cancer.gov/tcga, https://portal.gdc.cancer.gov/) investigating kidney, prostate and bladder tissues (dataset https://portal.gdc.cancer.gov/legacyarchive/search/f) assayed in multiple, independent experiments using the Illumina 450k microarray. The clinical metadata and the manifest file of the samples were downloaded through the GDC legacy archive on 2019-10-19 and 2019-11-09. We used the Genomic Data Commons (GDC) data download tool (https://gdc.cancer.gov/access-data/gdc-data-transfer-tool) using the manifest files deposited on our GitHub.
==== Body
pmcIntroduction

Epigenetics is the study of biochemical modifications, heritable through cell division, carrying information independent of DNA sequence [1]. This information has direct consequences for cellular phenotype and cell differentiation, including cancer formation. Although the biochemical modifications can vary for DNA, methylation signatures seem to be most optimal in terms of tradeoffs between variation and stability to provide data which can be used for disease diagnosis. This, together with recent breakthroughs and the decreasing cost of sequencing technology, is a probable reason that methylation is the branch of epigenetics that has progressed furthest in the last decade.

More importantly, clear evidence has emerged in recent years that aberrant DNA methylation plays a key role in many diseases, including cancer [2–4]. As a key epigenetic modification, this biochemical process can modulate gene expression to influence the cell differentiation which can possibly lead to cancer [2]. For cancer, aberrant methylation patterns seem to be a promising source of early diagnostic biomarkers. Mayeux and colleagues divide the major types of biomarkers into: biomarkers of exposure (used in risk prediction) and biomarkers of disease used in the screening, diagnosis and monitoring of disease progression) [5]. Early diagnosis biomarkers are of particular interest as it has been shown that, for many cancer types, early detection is strongly correlated with the patient’s chance of survival [6].

Finding methylation patterns in human DNA that are specific to different environmental exposures (e.g. cancerogenic substances) and clinical diagnoses would be a powerful tool to develop cost-efficient predictive assays which could guide clinicians and save lives. Traditionally, finding biomarkers from methylation data is done by applying linear models, as these are effective at uncovering the main differences between two groups. This approach is still very powerful whenever two heterogeneous groups are being compared. More importantly, the linear models are also effective when the sample size is low. These models can also account for covariates, like age or smoking, that are two of the most influential factors on the human methylome [7]. However, diagnosis is a classification problem that can also be approached by machine learning algorithms when the sample size allows.

With the advent of modern technology, decreasing sequencing costs, and the digitalization of international cancer biobanks, we are rapidly approaching the point where AI algorithms may guide oncologists in cancer diagnosis. Such classification approaches are already emerging with well developed machine learning approaches starting to play a leading role in areas such as the medical imaging of ‘melanoma’ or ‘prostate grade cancer’ [8, 9]. The development of AI algorithms even happens in open rooms with AI competitions launched among non experts, at least in the field of cancer [10]. This is significantly driven by the quantity of data accumulated, and as such, sequencing methods have only recently started to gain much attention due to the limited number of samples available in public repositories as well as regulatory issues with sharing such data. With new initiatives that enable genomic and clinical data sharing across federated networks, such as the Beacon protocol (used as a model for the federated discovery and sharing of genomic data) [11, 12], work in this area is now in progress and will open the door to new machine learning approaches.

One of the places with a relatively high number of genomic samples available is the Genomic Data Commons data portal (GDC, https://portal.gdc.cancer.gov/). GDC represents a rich source of harmonized cancer datasets available for immediate ethical use for cancer biology and biomarkers discoveries. Most of the data are publicly available and anonymised, so therefore there is no possibility to connect individual patients with this data. GDC contains transcriptomic and genomic data as well as epigenomics information about cancers isolated from patients. Most of the epigenomics information in GDC is related to methylation profiles measured on the ‘450k platform’ [13]. For this platform, 450 000 of the 28 million CpG sites (locations in the genetic code where a cytosine is immediately followed by a guanine) present in the human genome are measured and a value between 0 to 1 is reported for each CpG site. As the distribution of the methylation is binomial for human DNA a threshold of 0.3 can be used to classify CpGs as either methylated (> = 0.3) or unmethylated (<0.3) [14].

Artificial intelligence (AI) can operate on such data [15] although further development of this field is still required. As recently examples, a deep learning model for regression of genome-wide DNA methylation was implemented for genomic data [16]. The authors proposed a deep learning method for prediction of the genome-wide DNA methylation, in which the Methylation Regression was implemented by Convolutional Neural Networks (MRCNN). Through minimizing the continuous loss function, experiments show that their model was convergent and more precise than the previous state-of-art method (DeepCpG [17]) according to the results of the evaluation. Also, integrative analysis identifies potential DNA methylation biomarkers for ‘pan-cancer’ diagnosis and prognosis [18]. Unfortunately, this research used the ‘ghost probe’ ‘cg203000343’ for classification which contaminates their data. In another recent example, machine learning methods and DNA methylation data were used to distinguish primary lung squamous cells carcinomas from head and neck metastases [19]. Other recent studies have also shown it is possible to use DNA methylation to predict disease outcome [20, 21]. All this should open methodological possibilities of using methylation data for prognosis. However, owing to the security of the genomic data it is desired to use a minimal number of genomic locations for classification purposes. Using a minimal set of genomic locations also allows for the identification of novel biomarkers.

The primary goal of our work was to develop machine learning algorithms for the diagnostic classification of DNA methylation profiles. We used data obtained from the GDC platform relating to three urological cancers; prostate, bladder and kidney. These were chosen as they each had a relatively large number of epigenetic data points available as for both cancerous tissue and normal (adjacent tissue). To develop the algorithm, we used sequential methods for feature selection and then deep learning. These methods are proposed for further independent validation in laboratories that store prostate, bladder, and kidney tissue samples and have sufficient resources.

Results and discussion

In our AI analysis, cancer samples downloaded from GDC were more numerous than normal samples and we chose not to perform a data augmentation strategy. Using Principal Component Analysis (PCA) dimensionality reduction to reduce the data from approximately 400 000 dimensions to just two dimensions we found that the cancer samples can be visually distinguished from the normal samples with relative ease for all tissue types (Fig 1). Given that a method using PCA on 400 000 data points would be impractical in a clinical environment, this was primarily done with the intent of visualizing the data to better understand how to proceed. In particular, we wanted to see if normal samples clustered with each other, or if there were any biases in the data that needed further corrections. We also checked for samples that had been accidentally swapped during deposition to GDC. Using this visualization, we observed that the cancer data points had noticeably higher variation in methylation spread.

10.1371/journal.pone.0307912.g001 Fig 1 Principle component 1 and 2 for cancer (red) and normal (blue) samples for the three urinary cancer types: a) bladder, b) kidney, c) prostate. Data comes from training and validation sets.

Using the decision tree approach we were able to select a low number CpG sites based on the criterias described in the method section while minimising the loss of any information useful for distinguishing between normal and cancer signals (Table 1, Fig 2). The CpGs were gene annotated and we reported their identifiers, locations (on human genome v38), and separation score (Table 1). The CpG sites to be used for neural network classification were further refined by removing those in which the median value for the cancer and normal groups both fell on the same side of the cutoff point. Performed gene enrichment analysis, did not show any direct relationship to cancer pathways.

10.1371/journal.pone.0307912.g002 Fig 2 Boxplot visualisation of CpG sites (on hg38 genome) methylation per group cancer (red) and normal (blue) picked by the decision tree for each tissue type and further used by the neural network.

Data for A. bladder, B. kidney and C. prostate tissue is presented. The threshold of 0.3, defining methylated or unmethylated was marked with a dashed line. Data comes from the training and validation sets.

10.1371/journal.pone.0307912.t001 Table 1 CpG sites selected to distinguish between normal and cancer tissues using feature selection.

Name of CpGs as well as location is given, as well as gene annotation and position in genomic features. The order of CpGs is related to significance given in the neural network. CpGs with negative separation scores were removed.

Type	CpG	Chr	Position	Gene annotation	Location	Separation Score	
Bladder	cg04049981	chr17	50817533	-	-	2,64	
 	cg04349084	chr8	23745164	RP11-175E9.1	-	2,949	
 	cg04449953	chr8	143404769	-	N_Shelf	3,191	
 	cg22746681	chr10	50152871	-	Island	1,356	
 	cg06667961	chr6	31559916	-	-	-16,968	
 	cg16086237	chr12	130442711	RIMBP2	-	3,293	
 	cg18997129	chr7	143408757	EPHA1;EPHA1-AS1;	-	1,622	
 	cg19098932	chr1	2413713	PEX10	S_Shore	1,235	
 	cg19900821	chr1	162389343	-	-	0,862	
 	cg20696143	chr12	132405765	-	N_Shore	2,453	
 	cg21720373	chr7	56487789	-	S_Shelf	3,171	
Kidney	cg00221185	chr6	30684491	AL662797.1;PPP1R18	N_Shelf	3,092	
 	cg00375457	chr11	5596043	AC015691.13;HBG2;TRIM6;TRIM6-TRIM34	-	-9,741	
 	cg01904393	chr12	51421299	SLC4A8	N_Shelf	3,216	
 	cg02578087	chr3	8629675	SSUH2	-	3,164	
 	cg03272310	chr17	41548202	-	N_Shore	6,51	
 	cg03564506	chr7	98878278	TRRAP	N_Shore	-136,373	
 	cg03852551	chr6	167182751	TCP10L2	-	3,643	
 	cg06621027	chr2	79952633	CTNNA2	-	1,845	
 	cg14204586	chr1	155962067	ARHGEF2	-	8,381	
 	cg22274117	chr6	16713382	ATXN1	-	12,399	
Prostate	cg00183173	chr6	170024144	-	Island	-50,587	
 	cg00333364	chr11	8683255	AC091053.1;RPL27A;SNORA45A	S_Shore	-2,998	
 	cg00344260	chr12	44876623	NELL2	Island	-48,392	
 	cg00958578	chr16	53715	POLR3K;SNRNP25	Island	-102,875	
 	cg02049405	chr6	30127488	-	Island	1,255	
 	cg08206623	chr11	2886104	CDKN1C	S_Shore	0,064	
 	cg09691340	chr17	44325514	SLC25A39	Island	2,712	
 	cg09808235	chr1	2611149	MMEL1	-	1,787	
 	cg11417025	chr7	16465964	SOSTDC1	-	1,792	
 	cg12265829	chr14	24334816	ADCY4; RP11-934B9.3	Island	1,914	
 	cg15267232	chr10	8055726	GATA3	Island	2,158	
 	cg22621867	chr3	51956285	GPR62	Island	1,658	

The final CpGs used in the neural network were subjected to clustering using the Ward clustering method and visualized (Fig 3). Bladder and kidney cancer types were generally characterized by CpGs that were hypomethylated (low methylation), while for prostate we identified hypermethylated (high methylation) sites defining cancer. We then performed ROC curve analysis on our testing set and found our neural network produced very accurate results (Fig 4).

10.1371/journal.pone.0307912.g003 Fig 3 Cluster maps for A. bladder, B. kidney and C. prostate cancer. Colour scale represents methylation values. Samples are clustered by the Cancer and Normal samples and chosen CpGs. Data originates from the training and validation sets but the testing set is excluded.

10.1371/journal.pone.0307912.g004 Fig 4 Receiver operating characteristic (ROC curves) for A. bladder, B. kidney and C. prostate cancer. First neural layer had 7 and second 4 neurons. Area under the curve (AUC) is marked. Data originates from the testing set only.

Conclusions

As novel and cost-effective genomic methods emerge and computational methods become more sophisticated, the monitoring of changes in epigenetic signatures of human DNA could be a powerful tool for finding prognostic cancer biomarkers as well as for many screening programs. Ideally, many future screening programs could be based on the epigenomic tests although this would require a very high sensitivity and specificity of these tests [15].

Traditional classification methods (such as linear models based on methylation data differences between groups) focus on areas of high variance within the methylation levels. However, they may miss interactions between points that could offer insight into the underlying biology behind developing tumours. AI methods may be more capable of identifying these interactions, and selecting CpG sites that play a role in the development of cancer in conjunction with another CpG location, but that do not necessarily display a high variance in methylation levels between cancerous and non-cancerous tissues themselves. Our approach also paves the way for the development of more AI methods that can work on methylation data, combining it with other data types, to reveal insights for which linear methods may not be adequate.

However, there is no doubt that to achieve such a goal, investment in a large number of samples with measured methylation statuses is still desired. Currently, genomic samples are scarce in public databases and their availability is hampered by data privacy issues, however the GDC platform already offers a rich source of anonymised methylation data. We used this in our work for machine learning approaches, to demonstrate the viability of our AI approach in the context of urological cancers.

Using normalized methylation values in data provided by the GDC community we have been able to identify unique signatures for three cancer types, and distinguish these from their respective normal tissues. As we were looking for major differences between samples we did not use any metadata related to survival or age of the patient. This approach is desired as it decreases the possibility of back identification of any given patients which could compromise the future use of this technique in diagnostics. Importantly, we were still able to get a clear signal for these cancers without compromising personal information.

For this study, we divided our data into methylated and unmethylated groups prior to processing, based on a threshold of 0.3. This resulted in a clear separation of the cancer and normal tissues for many CpG sites (quantified by their separation values in Table 1). Using this approach, CpGs with a high separation value are likely to be useful for designing future targeted approaches. We also performed gene enrichment analysis, which did not show any direct relationship to any known cancer pathways, suggesting that the CpG sites we found do not share a common function, including cancer.

We have also noticed that other feature engineering approaches (such as PCA) used for neural networks perform well, however many of them would require the whole methylation profile to be obtained from future patients and would dramatically increase the computational resources required. Although we believe this is possible, we have demonstrated that only few CpG sites are needed to obtain a high specificity and sensitivity when classifying patients.

In our study, a computationally demanding step was using the decision tree for feature selection but dimensionality reduction was a necessary step for successfully training the neural network. We used a simple AI architecture, where only a few neurons were used in each layer of the net, however, this was enough to train the model efficiently as shown by the validation set. As we kept the testing set completely separate until all other changes were finished we do not believe we have any overfitting and believe that the accuracy it demonstrated is representative of how our model would perform on similar external data.

There is a risk, however, that the lack of an independent, external set of the samples from other resources could have introduced bias into our results. The GDC data we used helps offset this somewhat given that it comes from multiple independent projects that have been deposited on GDC over several years. There is still a risk, however, that unaccounted for discrepancies in the way different projects handled this initial data (whether in the acquisition, handling, or processing stages) could have introduced biases into our sets. It would be beneficial to further test the conclusions we have reached on independently collected data. Unfortunately, we do not have access to any external data sources, however by freely releasing our algorithms we allow other users who have access to such data to test our approach.

The found biomarkers are related directly to the tissue where cancer was present, and compared with adjacent normal tissue. This creates the possibility that there are field effects in this study and therefore further validation is needed. This could be done in an independent cohort. These studies were done in tissue but could be further converted to work based on blood which would be beneficial for the development of screening tests.

Our study used data from tissues samples and thus biopsies would be necessary to apply our classification algorithm, which is not an option in screening tests. As such, this approach could be improved upon but transitioning to human material that can both be unintrusively collected and has a high chance of containing information about developing cancer. Possible options include blood, urine or spit. The first approach could be done, for instance, using large prediagnostic blood biobanks, such as the Janus Serum Biobank [22]. Samples in the blood would also need deconvolution methods to further account for different cell compositions in the blood samples.

Our study only focused on three urological cancer types, however, it is possible that pre-diagnostic screening methods could be developed for other cancers using similar methods. Further studies to explore this possibility would also be beneficial.

Finally, our results provide a demonstration of the fundamental power of AI approaches and of the usefulness of methylation signatures for diagnoses, and will hopefully pave the way towards the development of a better understanding of how epigenetics plays a role in cancers and in the design of better and more cost-effective screening methods for urological cancers without compromising patient data safety.

Material and methods

Our code was written in Python3, using pickle, matplotlib, pylab, scipy, numpy, seaborn, and pandas as external libraries. The program was run using University of Birmingham’s BlueBEAR HPC service, which provides a High Performance Computing service to the University’s research community (http://www.birmingham.ac.uk/bear). BlueBEAR HPC uses a Slurm files system for job submission.

Data was downloaded from Genomic Data Commons (GDC) Data Portal (see data availability for manifest files and our GitHub repository, github.com/bazyliszek/methAI). All methylation data were already normalized across samples by the GDC consortium and therefore were ready for further processing. Shortly, the GDC portal performs automatic data normalization with each major release of the data. Their Methylation Array Harmonization Workflow (MAHW) uses raw methylation array data for the HumanMethylation 450k platform, measuring CpG methylation as beta values, calculated from array intensities (Level 2 data) as Beta = M/(M+U). This differs from the Methylation Liftover Pipeline in that the raw methylation array data is used instead of submitted methylation beta values, and the data is processed through the SeSAMe software [23, 24]. Additionally, the analysis results from the MAHW are of higher quality than results from the Methylation Liftover Pipeline. In our work we used only HumanMethylation 450k data. SeSAMe software can remove non-detection artifacts that would otherwise survive existing pipelines, such as Y-chromosome probes. It also corrects detection failures caused by germline and somatic deletions in other DNA methylation array software by calculating the significance of signals in the methylation arrays [23]. Correcting for these artifacts allows SeSAMe to improve upon the detection calling and quality control of methylation data.

Nevertheless, to further confirm that the normalization between datasets was done correctly by GDC we plotted a histogram of the two datasets for each tissue type, and calculated the first four moments for the distributions between cancer and normal samples for each tissue type (S1 Fig) to check that there were no noticeable differences between the two groups.

The downloaded data contained a large number of missing (NA) values. This created a problem as the neural net requires numerical inputs and the way these values were handled could adversely affect the decision tree’s performance as well. On inspection, we discovered that the majority of these could be attributed to non-existent CpG sites (location was always specified as *) in the data files. However, even after removing these, there were still some NA values. Throwing out all rows or columns with NA values would have reduced our dataset too much to be a valid option. Instead, we first discarded all CpG sites in which NA values made up more than 15% of the data. We then discarded all patients in which NA values made up more than 15% of the data. CpG sites were removed first as we had more of them available so removing a CpG site caused less loss of information in our dataset than removing a patient. Still, there were a few patients with a sufficiently high ratio of NA values that it was still beneficial to remove them from the study.

After pruning out the CpG sites and patients with the most NA values, we were left with 439 patients in the bladder set, 484 in the kidney set, and 542 in the prostate set. Among these, a small number of NA values still remained in the data. We explored increasing the cutoff of NA values needed to exclude a patient or CpG site from the study to further reduce these, but rejected that out of concern for diminishing our dataset too much. We also considered replacing all NA values with a designated number (such as -1) so that the neural net would be able to distinguish them, but rejected that as well as it may have biased the neural net towards treating NA values similarly to low values, and vis versa. In the end, we settled on replacing all remaining NA values with the arithmetic mean value for that CpG site in the relevant tissue as this was considered the option the meant they were least likely to have a significant impact on our results.

The methylation values were then converted to binary by using a cutoff point of 0.3 as previously suggested [14]. CpG sites with methylation values higher than this were treated as methylated, while those with lower values were classified as unmethylated. This also has a biological meaning related to gene regulation.

After the preprocessing stage, the data was randomly divided into three sets. 70% of the data was assigned to the training set (307 for bladder, 339 for kidney, 379 for prostate), 20% was assigned to the validation set (88 for bladder, 97 for kidney, 108 for prostate), and the remaining 10% to the testing set (44 for bladder, 48 for kidney, 55 for prostate). The validation set was withheld and used later to optimize and compare different approaches, while the testing set was put aside and not used for any purpose other than calculating the final values in the results section. Values in the testing set were also excluded from any data visualization.

The ~400 000 dimensions (CpG sites) in our dataset was far more than a neural network could reasonably be expected to train on in a feasible amount of time. To resolve this we applied feature selection techniques to reduce the dataset to a more manageable number of dimensions. Initially we tried using Principal Component Analysis (PCA) to reduce the dimensions from ~400 000 to 10. This worked well and the initial tests with the neural network showed that it could classify the patients extremely accurately using the PCA data set. However, as the PCA set was constructed with input from 400 000 CpG sites, any test based on this procedure would require all 400 000 CpG sites to be measured for any patient. The cost of doing this means that this approach would be impractical in most clinical settings. Instead we attempted to identify those CpG sites most relevant for classifying the data. For this, we used a decision tree to identify approximately 10 of the most relevant CpG sites for dividing the data for each tissue type. These points were then used to train the neural network. The decision to focus on around 10 CpG sites was heavily based on the trade off between the cost associated with using a higher number if the test was to be used in a clinical setting, weighed against the loss of accuracy that reducing the number of CpG sites would cause. Analyzing 10 CpG sites would be relatively inexpensive compared with a 400k array, while still providing enough information for the neural network to classify the cases.

The optimal CpG locations to give the neural network would contain a list of locations that show a high variance in methylation status between the cancerous and normal groups (primary separating locations), and locations that help it classify the data in cases which are not separated by those high variance points (secondary separating locations). In order to try to get both of these, the decision tree was run to a max depth of two, with the first dividing CpG site being chosen as the one that best separated the two groups and the second and third sites (if applicable) being the ones that were best able to correct for any mistakes the first point made. After each run the sites found were added to the list and all subsequent runs of the decision tree were prevented from dividing on those locations in future. This forced the subsequent runs to find a variety of primary and secondary separating locations, increasing the number of features that the neural network could use for classification. At the end of the process, this resulted in a list of at least 10 CpG sites for each tissue type (actual number was slightly variable as the decision tree did not always return the same amount of sites, however it would keep running until it had at least 10, this resulted in 10 to 12 values being chosen at the end. We considered cutting it off at 10, however the purpose of using a decision tree to select the values instead of linear models was that the effectiveness of the CpG sites it could select may be highly dependent on each other being included.). When training the decision tree we used entropy as a measure of the confidence the system had in classifying a data point at a given position within the tree, to determine which features it would select. The Entropy at a given node was calculated accordingly: [25], Entropy=−Normal*log(Normal)−Cancer*log(Cancer) (1)

Where ‘Normal’ represents the fraction of the data points in that node that were taken from adjacent apparently non-cancerous tissue, and ‘Cancer’ is the fraction of the data points in that node that were taken from visibly cancerous tissue. This equation creates a curve that is at its minimum (0) when the data is fully divided and maximum (1) when the data is evenly mixed.

When training the decision tree, we found that a depth of 2 was sufficient to classify the data in >90% of cases. Running at a max depth of 2 gave us a mix of locations with high variance between the cancer and normal tissue types, and locations that were good at fine tuning the details and correcting those data points that the first locations failed to classify correctly.

The CpG lists from the decision tree were further narrowed down by excluding any in which the median methylation for each dataset fell on the same side of the cutoff point. This was done to reduce chances of the neural network basing a decision on unreliable or outlying data. The separation score (a rough estimate of how distant the two medians of the two groups were from each other’s side of the cutoff point) was calculated using the formula: SeparationScore=−CancerMedian−cutoffCancerSD*NormalMedian−cutoffNormalSD (2)

Negative values occur when both groups have a median on the same side of the cutoff point and were excluded from the rest of the study. The values were recorded (Table 1). The final list of CpG locations was used to train three neural networks for each tissue type. A sigmoid activation function (S(x)) was chosen to calculate the output of each neuron, due to the ease of calculating the gradient for back propagation and the fact that it gives bounded outputs for all potential inputs.

S(x)=11+e−x (3)

Where e is Euler’s number and x is the sum of all weighted inputs for a given neuron. The first neural network had 7 neurons in the first hidden layer and 4 in the second hidden layer. The second neural network had 10 neurons in both hidden layers. The third neural network had three hidden layers containing 5, 4, and 3 neurons respectively. All three architectures also used a learning rate function that decreased slowly over time. These architectures and the initial learning rate were chosen based on our experience training other neural nets on similar data sets, with the intent to use the validation set to optimize these parameters. However, this turned out to be unnecessary as all three structures reached the point where they were able to perfectly, or near perfectly, classify all patients in the validation set on all tissue types. This resulted in an Area Under the Curve of, or very close to, 1 for all structures over all tissue types during Receiver operating characteristic (ROC) analysis. As such, no further adjustments were needed. As our ROC curves produced such high values for each tissue type a larger testing set is likely needed to get a more accurate assessment on the neural nets actual performance. All we can say from the current results is that it is likely very high for both bladder and kidney tissue. To check if the CpGs used had any known biological functions (for instance known cancer pathway) we have used webtool Enrichr was used to find out if there is any association between genes that were in approximation of the found CpGs and their functions (https://maayanlab.cloud/Enrichr/) [26].

Supporting information

S1 Fig (DOCX)

The computations described in this paper were performed using the University of Birmingham’s BlueBEAR HPC service, which provides a High Performance Computing service to the University’s research community. Additional computations were performed on resources provided by Sigma2—the National Infrastructure for High Performance Computing and Data Storage in Norway (project number NN8039K). We thank the GDC help desk for the data and for assistance given related to retrieving the data from their servers.

10.1371/journal.pone.0307912.r001
Decision Letter 0
Reismann Marc Academic Editor
© 2024 Marc Reismann
2024
Marc Reismann
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version0
13 May 2024

PONE-D-23-40540Diagnostic classification based on DNA methylation profiles using sequential machine learning approaches.PLOS ONE

Dear Dr. Marcin,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. The manuscript has great potential against the background of current attempts to find discriminatory markers for distinguishment of malignant entities. I find the approach regarding the definition of biomarkers or particularly biomarker signatures based on the application of machine learning methods to DNA methylation data intriguing. However, the reviewers have expressed some concerns regarding particularly methodological aspects. Some issues have to be explained in more detail. Although I decided for “major revision” I am very confident that a satisfactory based on the reviewers´ comments is possible.

Please submit your revised manuscript by Jun 27 2024 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Marc Reismann, MD, PhD

Academic Editor

PLOS ONE

Journal Requirements:

1. When submitting your revision, we need you to address these additional requirements.

Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at 

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and 

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Did you know that depositing data in a repository is associated with up to a 25% citation advantage (https://doi.org/10.1371/journal.pone.0230416)? If you’ve not already done so, consider depositing your raw data in a repository to ensure your work is read, appreciated and cited by the largest possible audience. You’ll also earn an Accessible Data icon on your published paper if you deposit your data in any participating repository (https://plos.org/open-science/open-data/#accessible-data).

3. Please note that PLOS ONE has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, all author-generated code must be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

4. When completing the data availability statement of the submission form, you indicated that you will make your data available on acceptance. We strongly recommend all authors decide on a data sharing plan before acceptance, as the process can be lengthy and hold up publication timelines. Please note that, though access restrictions are acceptable now, your entire data will need to be made freely accessible if your manuscript is accepted for publication. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If you are unable to adhere to our open data policy, please kindly revise your statement to explain your reasoning and we will seek the editor's input on an exemption. Please be assured that, once you have provided your new statement, the assessment of your exemption will not hold up the peer review process.

5. We note that you have included the phrase “data not shown” in your manuscript. Unfortunately, this does not meet our data sharing requirements. PLOS does not permit references to inaccessible data. We require that authors provide all relevant data within the paper, Supporting Information files, or in an acceptable, public repository. Please add a citation to support this phrase or upload the data that corresponds with these findings to a stable repository (such as Figshare or Dryad) and provide and URLs, DOIs, or accession numbers that may be used to access these data. Or, if the data are not a core part of the research being presented in your study, we ask that you remove the phrase that refers to these data.

6. Please include your full ethics statement in the ‘Methods’ section of your manuscript file. In your statement, please include the full name of the IRB or ethics committee who approved or waived your study, as well as whether or not you obtained informed written or verbal consent. If consent was waived for your study, please include this information in your statement as well. 

7. We notice that your supplementary figure is included in the manuscript file. Please remove and upload it with the file type 'Supporting Information'. Please ensure that each Supporting Information file has a legend listed in the manuscript after the references list.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Partly

Reviewer #2: Partly

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: N/A

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The core concept this manuscript attempts to address is discrimination between cancer and adjacent normal tissue by methylation profiling, in this instance via Illumina 450K array. 3 urological cancer datasets are used as an example. The authors use a machine learning model to extract CpG sites that differentiate between cancer and normal in each tumour type, and use these to construct a classification model.

This manuscript appears mostly technically sound in terms of the implementation of the authors chosen methods, although as with any bioinformatics analysis or machine learning model it is impossible to definitively assess methodological rigour without viewing the underlying code (which is not public/viewable at the time of review). Nevertheless, the descriptions given in the methods section are sufficient for an adequate representation of the rationale underlying this particular implementation, if not replication. Machine learning-based classification of tumours is a well established approach (e.g., the MNP CNS tumour classification system) and though the model in this manuscript uses a different algorithm, the principle of classification via methylation is sound. The manuscript is written in an intelligible manner with a few grammatical mistakes but nothing that should impede a proper understanding of the written language.

I have some technical/rationale queries for the authors after reading the manuscript:

1. I understand data was previously normalised across samples by GDC before retrieval by the authors. Did this normalisation include removal of CpGs associated with the X and Y chromosomes, probes with known mismatches, and probes adjacent to SNPs? If not, did the authors remove these CpGs as part of their preprocessing? This is not clear from the methods.

2. Regarding principal component analyses and dimensionality reduction of the data: The separation of cancer and normal samples here is a little lackluster compared to what I usually see in cancer vs normal analyses. Feeding the entire methylation array into a PCA is a rather curious approach. Many of the probes on the array will demonstrate no or very little variance between cancer and normal tissue, but will introduce significant noise to the final resolving power of the analysis. I find that selecting a subset of ~10,000 probes with high variability across the cohort (e.g., ranking probes by median absolute deviation) dramatically improves separation of biological groups when performing PCA/TSNE while capturing almost all biological variability. Similar subsetting steps prior to PCA are a fairly common step in most discrimination/class discovery analyses using methylation data. I mention this primarily as it may improve segregation on the PCA analyses the authors performed.

3. If the goal is to identify a minimal selection of CpGs that can act as markers to reliably discriminate between cancer and normal, why not perform a simple differential methylation analysis of cancer v normal for each tumour type? This would give a list of differentially methylated CpGs. Sites with a consistently large difference between cancer and normal tissue could then be chosen as standalone markers for validation or combined and incorporated into a classification model. It seems this approach would have more biological grounding, and would likely reflect major underlying biological changes between tumour/normal. While the machine learning approach to feature selection is sophisticated, I wonder about the added value of using a rather convoluted method to achieve a relatively simple goal (binary classification).

Reviewer #2: This is an interesting manuscript that investigates DNA methylation patterns in three type of cancer using archived data from the Genomic Data Commons. A few minor issues in the approach should be addressed prior to publication.

Under Methods:

Add specific detail in the first paragraph. The manuscript currently states, “We used ‘Python’ as the primary programming language for this project with libraries such as ‘pickle’ and ‘matplotlib’.” Please include all programming languages and libraries used.

Explain how data are determined to be “NA” in Illumina 450k data. Beyond the absence of a specific CpG, why were specific measurements designated NA?

In the third paragraph, it is state, “We explored a few options for how to deal with these, and settled on replacing all remaining NA values with the average value for that CpG site in the relevant tissue as this was considered the option least likely to have an impact on our results.” Specify the “average” used (was this the arithmetic mean? Geometric mean? Median?). Describe all the options that were explored. In the Results, provide the data that were used to make an objective determination that the average would be used.

In paragraph 6, the statement “… the initial tests with the neural network showed that it could classify the patients extremely accurately using the PCA data set (data not shown)” is not very useful. Please provide the data and specify the accuracy of the classification.

Also in paragraph 6, it is stated that “… any test based on this procedure would require all 400 000 CpG sites to be measured for any patient. This would be impractical and unreliable in clinical settings …” Please provide an explanation as to why the approach would be unreliable.

Also in paragraph 6, it is stated that “… we used a decision tree to identify approximately 10 of the most relevant CpG sites …”. Please explain why the target value of 10 was chosen. Was this a subjective choice or based on prior data or experience?

In paragraph 7, it is stated that the “actual number was slightly variable ...” please provide the actual range of values.

In paragraph 8, the explanation of Entropy as “a measure of how mixed the data is calculated accordingly” is confusing. Please provide a more specific explanation.

In paragraph 9, the statement, “we found that a depth of 2 was almost always sufficient to classify the data” should be replaced with actual classification data. What does “almost always sufficient” objectively mean?

In paragraph 12, consider replacing the term “chosen based on intuition.” Were initial learning rates selected based on prior experience with similar systems?

Finally, a specific statement about the limitations of the available data sets and analyses that could affect performance of the approach is needed. For example, variability in biological sample acquisition, storage, or processing could within and between studies could impact the analysis. Also, contributions from biological (e.g. grade of cancer) versus technical (e.g., low probe reliability) sources of variance could affect interpretation.

**********

6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

10.1371/journal.pone.0307912.r002
Author response to Decision Letter 0
Submission Version1
11 Jun 2024

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Partly

Reviewer #2: Partly

________________________________________

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: N/A

Reviewer #2: Yes

________________________________________

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

________________________________________

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

________________________________________

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The core concept this manuscript attempts to address is discrimination between cancer and adjacent normal tissue by methylation profiling, in this instance via Illumina 450K array. 3 urological cancer datasets are used as an example. The authors use a machine learning model to extract CpG sites that differentiate between cancer and normal in each tumour type, and use these to construct a classification model.

This manuscript appears mostly technically sound in terms of the implementation of the authors chosen methods, although as with any bioinformatics analysis or machine learning model it is impossible to definitively assess methodological rigour without viewing the underlying code (which is not public/viewable at the time of review). Nevertheless, the descriptions given in the methods section are sufficient for an adequate representation of the rationale underlying this particular implementation, if not replication. Machine learning-based classification of tumours is a well established approach (e.g., the MNP CNS tumour classification system) and though the model in this manuscript uses a different algorithm, the principle of classification via methylation is sound. The manuscript is written in an intelligible manner with a few grammatical mistakes but nothing that should impede a proper understanding of the written language.

I have some technical/rationale queries for the authors after reading the manuscript:

** MW: Thank you for your constructive comments. We have now opened the code in GitHub according to FAIR principle and the code can be now accessed under the following link: https://github.com/bazyliszek/methAI as well in Zenodo. We made a change in the main text “The source code is available in GitHub repository (https://github.com/bazyliszek/methAI), and in Zenodo under DOI: 10.5281/zenodo.11530987” and can be cited Wojewodzic, M. W., & Lavender, J. P. (2024). Diagnostic classification based on DNA methylation profiles using sequential machine learning approaches (Version v0.1.0-alpha-2) [Computer software]. https://doi.org/10.5281/zenodo.11530987

1. I understand data was previously normalised across samples by GDC before retrieval by the authors. Did this normalisation include removal of CpGs associated with the X and Y chromosomes, probes with known mismatches, and probes adjacent to SNPs? If not, did the authors remove these CpGs as part of their preprocessing? This is not clear from the methods.

*** MW: Thank you for your comments. We have now improved this in our description of the method, as we understand this could be beneficial for the reader to better understand our data sources and how the GDC processed the data prior to our usage.

We have consulted the author of the GDC pipeline to better understand the process by which probes are normalised and how probes with known mismatches or those adjacent to SNPs/repeats are discarded, as described in this paper https://academic.oup.com/nar/article/45/4/e22/2290930.

We have described some of these steps in our manuscript.

Shortly, GDC portal automatically normalised data across samples at every major release of the data. The Methylation Array Harmonization Workflow, used by GDC, uses raw methylation array data from different generations of Illumina Infinium DNA methylation arrays: Human Methylation 27 (HM27), HumanMethylation 450 (HM450) and EPIC platforms, to measure the level of methylation at known CpG sites as beta values, calculated from array intensities (Level 2 data) as Beta = M/(M+U). This differs from the Methylation Liftover Pipeline in that the raw methylation array data is used instead of submitted methylation beta values, and the data is processed through the software package SeSAMe [1]. In our work we used only HumanMethylation 450 (HM450) data.

SeSAMe effectively removes non-detection artefacts that survive existing pipelines, including Y-chromosome probes. SeSAMe corrects to detection failures that occur in other DNA methylation array software commonly due to germline and somatic deletions by utilizing a novel way to calculate the significance of detected signals in methylation arrays. By correcting for these artifacts as well as other improvements to DNA methylation data processing, SeSAMe improves upon detection calling and quality control of processed DNA methylation data. SeSAMe output files include: two Masked Methylation Array IDAT files, one for each color channel, that contains channel data from a raw methylation array after masking potential genotyping information; and a subsequent Methylation Beta Value TXT file derived from the two Masked Methylation Array IDAT files, that displays the calculated methylation beta value for CpG sites.

We also included following publication in the paper.

[1]. Zhou, Wanding, Triche Timothy J., Laird Peter W. and Shen Hui. "SeSAMe: Reducing artifactual detection of DNA methylation by Infinium BeadChips in genomic deletions." Nucleic Acids Research. (2018): doi: 10.1093/nar/gky691

[2]. Zhou, Wanding, Laird Peter L., and Hui Shen. "Comprehensive characterization, annotation and innovative use of Infinium DNA methylation BeadChip probes." Nucleic Acids Research. (2016): doi: 10.1093/nar/gkw967

2. Regarding principal component analyses and dimensionality reduction of the data: The separation of cancer and normal samples here is a little lackluster compared to what I usually see in cancer vs normal analyses. Feeding the entire methylation array into a PCA is a rather curious approach. Many of the probes on the array will demonstrate no or very little variance between cancer and normal tissue, but will introduce significant noise to the final resolving power of the analysis. I find that selecting a subset of ~10,000 probes with high variability across the cohort (e.g., ranking probes by median absolute deviation) dramatically improves separation of biological groups when performing PCA/TSNE while capturing almost all biological variability. Similar subsetting steps prior to PCA are a fairly common step in most discrimination/class discovery analyses using methylation data. I mention this primarily as it may improve segregation on the PCA analyses the authors performed.

*** MW: Thank you for this comment. We agree with the reviewer, however at this stage of analysis we just wanted to visualise how the data clustered before performing feature selection later. In particular, we wanted to see if normal samples clustered with each other, or if there were any biases in the data that needed further corrections. This would also identify samples that were accidentally swapped during deposition to GDC. As we can notice there is a separation between groups on PC1 axes.

Initially we tried using Principal Component Analysis (PCA) to reduce the dimensions from ~400 000 to 10. However, because these 10 features were generated from 400 000, using them to classify a patient would require analysing those 400 000 CpG sites for the patient. This would be impractical in most real-world settings. As such, we instead attempted to find a small number of CpG sites that could be used without PCA.

We have now implemented this into our manuscript in the results section.

3. If the goal is to identify a minimal selection of CpGs that can act as markers to reliably discriminate between cancer and normal, why not perform a simple differential methylation analysis of cancer v normal for each tumour type? This would give a list of differentially methylated CpGs. Sites with a consistently large difference between cancer and normal tissue could then be chosen as standalone markers for validation or combined and incorporated into a classification model. It seems this approach would have more biological grounding, and would likely reflect major underlying biological changes between tumour/normal. While the machine learning approach to feature selection is sophisticated, I wonder about the added value of using a rather convoluted method to achieve a relatively simple goal (binary classification).

*** MW: Thank you for this comment. Yes, we are aware of that and in fact we have discussed this in the introduction: “Traditionally, finding biomarkers from methylation data is done by applying linear models, as these are effective at uncovering the main differences between two groups. However, diagnosis is a classification problem that can also be approached by machine learning algorithms when the sample size allows”.

Yet, to account for this comment we included also following text in the discussion: “Traditional classification methods (such as linear models based on methylation data differences between groups) focus on areas of high variance within the methylation levels. However, they may miss interactions between points that could offer insight into the underlying biology behind developing tumours. AI methods may be more capable of identifying these interactions, and selecting CpG sites that play a role in the development of cancer in conjunction with another CpG location, but that do not necessarily display a high variance in methylation levels between cancerous and non-cancerous tissues themselves.”

Our approach also paves the way for the development of more AI methods that can work on methylation data and reveal insights for which linear methods may not be adequate.

Reviewer #2: This is an interesting manuscript that investigates DNA methylation patterns in three type of cancer using archived data from the Genomic Data Commons. A few minor issues in the approach should be addressed prior to publication.

*** MW: Thank you very much for the interest in our paper and the feedback we got.

Under Methods:

Add specific detail in the first paragraph. The manuscript currently states, “We used ‘Python’ as the primary programming language for this project with libraries such as ‘pickle’ and ‘matplotlib’.” Please include all programming languages and libraries used.

*** MW: Thank you. We have not implemented versions of the packages. We added missing packages we used in the text: pylab, scipy, numpy, seaborn, pandas.

Explain how data are determined to be “NA” in Illumina 450k data. Beyond the absence of a specific CpG, why were specific measurements designated NA?

*** MW: The NA values originate from the Illumina platform and are associated with low quality probes, or chromosome Y. We have now mentioned this in the method section and how preprocessing is done by GDC.

In the third paragraph, it is state, “We explored a few options for how to deal with these, and settled on replacing all remaining NA values with the average value for that CpG site in the relevant tissue as this was considered the option least likely to have an impact on our results.” Specify the “average” used (was this the arithmetic mean? Geometric mean? Median?). Describe all the options that were explored. In the Results, provide the data that were used to make an objective determination that the average would be used.

*** MW: We have used the arithmetic mean of all non-NA data points at that CpG site. We have also described in detail which alternative methods we explored and rejected (the same paragraph).

In paragraph 6, the statement “… the initial tests with the neural network showed that it could classify the patients extremely accurately using the PCA data set (data not shown)” is not very useful. Please provide the data and specify the accuracy of the classification.

Also in paragraph 6, it is stated that “… any test based on this procedure would require all 400 000 CpG sites to be measured for any patient. This would be impractical and unreliable in clinical settings …” Please provide an explanation as to why the approach would be unreliable.

*** MW: Unfortunately, we no longer have the data showing the accuracy of the 400 000 classification approach given that this method was rejected as unusable for other reasons. We have expanded a little upon the rationale behind this rejection.

The 400 000 approach was rejected as unreliable on the basis that such a large quantity of data being collected would increase the rate of errors being introduced to the set. However, in responding to recommendations to the other reviewer we removed the statement that this would be unreliable and focused on the impracticality and expense of attempting to use a 400 000 CpG method in a clinical setting. As such, it no longer refers to the approach as unreliable.

Also in paragraph 6, it is stated that “… we used a decision tree to identify approximately 10 of the most relevant CpG sites …”. Please explain why the target value of 10 was chosen. Was this a subjective choice or based on prior data or experience?

*** MW: The decision to focus on around 10 CpG sites was heavily based on the trade-off between the cost associated with using a higher number if the test was to be used in a clinical setting, weighed against the loss of accuracy that reducing the number of CpG sites would cause. Analysing 10 CpG sites is relatively inexpensive but involves very little information loss. We have implemented this in the paper now.

In paragraph 7, it is stated that the “actual number was slightly variable ...” please provide the actual range of values.

*** MW: Thank you. We have added the actual number of return values (10 to 12). We have now implemented this to the article (paragraph 7) and added justification.

In paragraph 8, the explanation of Entropy as “a measure of how mixed the data is calculated accordingly” is confusing. Please provide a more specific explanation.

MW: We have improved this further. "Entropy is a measure of the confidence with which the system can classify a data point at a given position in the tree" to the paper. (Paragraph 8)

In paragraph 9, the statement, “we found that a depth of 2 was almost always sufficient to classify the data” should be replaced with actual classification data. What does “almost always sufficient” objectively mean?

*** MW: Depth of 2 was sufficient for classification for more

Attachment Submitted filename: PLOS_ONE_PONE-D-23-40540_Response_to_Reviewers.docx

10.1371/journal.pone.0307912.r003
Decision Letter 1
Reismann Marc Academic Editor
© 2024 Marc Reismann
2024
Marc Reismann
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version1
11 Jul 2024

Diagnostic classification based on DNA methylation profiles using sequential machine learning approaches.

PONE-D-23-40540R1

Dear Dr. Marcin,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. If you have any questions relating to publication charges, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Marc Reismann, MD, PhD

Academic Editor

PLOS ONE

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. If the authors have adequately addressed your comments raised in a previous round of review and you feel that this manuscript is now acceptable for publication, you may indicate that here to bypass the “Comments to the Author” section, enter your conflict of interest statement in the “Confidential to Editor” section, and submit your "Accept" recommendation.

Reviewer #1: All comments have been addressed

Reviewer #2: All comments have been addressed

**********

2. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

3. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

6. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The authors have made suitable efforts to engage with the previous review comments and have made modifications to the manuscript to add clarity and detail where suggested. The authors have defended their analysis and, while I find it a slightly curious approach, there is no indication that this is unsound or inappropriate. My main minor suggestion to improve the manuscript would be to add further descriptive prose to the results section. At present, after the first paragraph, it is a little lackluster and doesn't give much indication of the results (e.g., "Performed gene enrichment analysis, did not show any direct relationship to cancer pathways."; "...performed ROC curve analysis on our testing set and found our neural network produced very accurate results."). While the associated figures and tables add some context, description in the text would be useful to highlight specific details e.g, accuracy metrics etc.

Reviewer #2: The authors have addressed my concerns. The manuscript has been updated and is now suitable for publication.

**********

7. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

**********

10.1371/journal.pone.0307912.r004
Acceptance letter
Reismann Marc Academic Editor
© 2024 Marc Reismann
2024
Marc Reismann
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
16 Jul 2024

PONE-D-23-40540R1

PLOS ONE

Dear Dr. Marcin,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

If revisions are needed, the production department will contact you directly to resolve them. If no revisions are needed, you will receive an email when the publication date has been set. At this time, we do not offer pre-publication proofs to authors during production of the accepted work. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few weeks to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Marc Reismann

Academic Editor

PLOS ONE
==== Refs
References

1 Waddington CH . A Manual of Embryology. Nature. 1940. pp. 728–729. doi: 10.1038/146728a0
2 Turner JD , Kirschner SA , Molitor AM , Evdokimov K , Muller CP . Epigenetics. International Encyclopedia of the Social & Behavioral Sciences. 2015. pp. 839–847. doi: 10.1016/b978-0-08-097086-8.14142-x
3 Wang J , Han X , Sun Y . DNA methylation signatures in circulating cell-free DNA as biomarkers for the early detection of cancer. Science China Life Sciences. 2017. pp. 356–362. doi: 10.1007/s11427-016-0253-7 28063009
4 Li J , Xing X , Li D , Zhang B , Mutch DG , Hagemann IS , et al . Whole-Genome DNA Methylation Profiling Identifies Epigenetic Signatures of Uterine Carcinosarcoma. Neoplasia. 2017. pp. 100–111. doi: 10.1016/j.neo.2016.12.009 28088687
5 Mayeux R. Biomarkers: Potential uses and limitations. Neurotherapeutics. 2004. pp. 182–188. doi: 10.1602/neurorx.1.2.182 15717018
6 Xie Y , Meng W-Y , Li R-Z , Wang Y-W , Qian X , Chan C , et al . Early lung cancer diagnostic biomarker discovery by machine learning methods. Transl Oncol. 2020;14 : 100907. doi: 10.1016/j.tranon.2020.100907 33217646
7 Yang Y , Gao X , Just AC , Colicino E , Wang C , Coull BA , et al . Smoking-Related DNA Methylation is Associated with DNA Methylation Phenotypic Age Acceleration: The Veterans Affairs Normative Aging Study. Int J Environ Res Public Health. 2019;16. doi: 10.3390/ijerph16132356 31861365
8 Jain V , Chatterjee JM . Machine Learning with Health Care Perspective: Machine Learning and Healthcare. Springer Nature; 2020.
9 Tolkach Y , Dohmgörgen T , Toma M , Kristiansen G . High-accuracy prostate cancer pathology using deep learning. Nature Machine Intelligence. 2020. pp. 411–418. doi: 10.1038/s42256-020-0200-7
10 Kaggle Competitions in the Classroom: Retrospectives and Recommendations. Volume 47, Number 4, August 2020. 2020. doi: 10.1287/orms.2020.04.13
11 Canakoglu A , Pinoli P , Gulino A , Nanni L , Masseroli M , Ceri S . Federated sharing and processing of genomic datasets for tertiary data analysis. Briefings in Bioinformatics. 2020. doi: 10.1093/bib/bbaa091 34020536
12 Rieke N , Hancox J , Li W , Milletarì F , Roth HR , Albarqouni S , et al . The future of digital health with federated learning. NPJ Digit Med. 2020;3 : 119. doi: 10.1038/s41746-020-00323-1 33015372
13 Morris T , Lowe R . Report on the Infinium 450k Methylation Array Analysis Workshop. Epigenetics. 2012. pp. 961–962. doi: 10.4161/epi.20941 22722984
14 Warden CD , Lee H , Tompkins JD , Li X , Wang C , Riggs AD , et al . COHCAP: an integrative genomic pipeline for single-nucleotide resolution DNA methylation analysis. Nucleic Acids Res. 2019;47 : 8335–8336. doi: 10.1093/nar/gkz663 31340028
15 Topol E. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again. Hachette UK; 2019.
16 Tian Q , Zou J , Tang J , Fang Y , Yu Z , Fan S . MRCNN: a deep learning model for regression of genome-wide DNA methylation. BMC Genomics. 2019;20 : 192. doi: 10.1186/s12864-019-5488-5 30967120
17 Angermueller C , Lee HJ , Reik W , Stegle O . DeepCpG: accurate prediction of single-cell DNA methylation states using deep learning. Genome Biol. 2017;18 : 67. doi: 10.1186/s13059-017-1189-z 28395661
18 Ding W , Chen G , Shi T . Integrative analysis identifies potential DNA methylation biomarkers for pan-cancer diagnosis and prognosis. Epigenetics. 2019;14 : 67–80. doi: 10.1080/15592294.2019.1568178 30696380
19 Jurmeister P , Bockmayr M , Seegerer P , Bockmayr T , Treue D , Montavon G , et al . Machine learning analysis of DNA methylation profiles distinguishes primary lung squamous cell carcinomas from head and neck metastases. Sci Transl Med. 2019;11 . doi: 10.1126/scitranslmed.aaw8513 31511427
20 Karagod VV, Sinha K. A novel machine learning framework for phenotype prediction based on genome-wide DNA methylation data. 2017 International Joint Conference on Neural Networks (IJCNN). 2017. doi: 10.1109/ijcnn.2017.7966050
21 Lian H , Han Y-P , Zhang Y-C , Zhao Y , Yan S , Li Q-F , et al . Integrative analysis of gene expression and DNA methylation through one-class logistic regression machine learning identifies stemness features in medulloblastoma. Mol Oncol. 2019;13 : 2227–2245. doi: 10.1002/1878-0261.12557 31385424
22 Rounge TB , Lauritzen M , Erlandsen SE , Langseth H , Holmen OL , Gislefoss RE . Ultralow amounts of DNA from long-term archived serum samples produce quality genotypes. Eur J Hum Genet. 2020;28 : 521–524. doi: 10.1038/s41431-019-0543-x 31719661
23 Zhou W , Triche TJ Jr , Laird PW , Shen H . SeSAMe: reducing artifactual detection of DNA methylation by Infinium BeadChips in genomic deletions. Nucleic Acids Res. 2018;46 : e123. doi: 10.1093/nar/gky691 30085201
24 Zhou W , Laird PW , Shen H . Comprehensive characterization, annotation and innovative use of Infinium DNA methylation BeadChip probes. Nucleic Acids Res. 2017;45 : e22. doi: 10.1093/nar/gkw967 27924034
25 Russell SJ , Russell SJ , Norvig P , Davis E . Artificial Intelligence: A Modern Approach. Prentice Hall; 2010.
26 Kuleshov MV , Jones MR , Rouillard AD , Fernandez NF , Duan Q , Wang Z , et al . Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 2016;44 : W90–7. doi: 10.1093/nar/gkw377 27141961
