
==== Front
PLoS Genet
PLoS Genet
plos
PLOS Genetics
1553-7390
1553-7404
Public Library of Science San Francisco, CA USA

39186782
10.1371/journal.pgen.1011377
PGENETICS-D-24-00407
Research Article
Biology and Life Sciences
Genetics
Mutagenesis
Biology and Life Sciences
Genetics
Gene Identification and Analysis
Genetic Screens
Biology and life sciences
Biochemistry
Proteins
DNA-binding proteins
Biology and Life Sciences
Molecular Biology
Molecular Biology Techniques
Gene Mapping
Research and Analysis Methods
Molecular Biology Techniques
Gene Mapping
Research and Analysis Methods
Animal Studies
Experimental Organism Systems
Model Organisms
Caenorhabditis Elegans
Research and Analysis Methods
Model Organisms
Caenorhabditis Elegans
Research and Analysis Methods
Animal Studies
Experimental Organism Systems
Animal Models
Caenorhabditis Elegans
Biology and Life Sciences
Organisms
Eukaryota
Animals
Invertebrates
Nematoda
Caenorhabditis
Caenorhabditis Elegans
Biology and Life Sciences
Zoology
Animals
Invertebrates
Nematoda
Caenorhabditis
Caenorhabditis Elegans
Biology and Life Sciences
Genetics
Gene Types
Suppressor Genes
Biology and Life Sciences
Genetics
Genomics
Animal Genomics
Invertebrate Genomics
Biology and Life Sciences
Genetics
Heredity
Genetic Suppression
A machine learning enhanced EMS mutagenesis probability map for efficient identification of causal mutations in Caenorhabditis elegans
Machine learning assists causal mutation identification
https://orcid.org/0009-0001-2746-613X
Guo Zhengyang Conceptualization Data curation Formal analysis Methodology Supervision Validation Visualization Writing – original draft Writing – review & editing
Wang Shimin Investigation
https://orcid.org/0009-0004-6696-5179
Wang Yang Formal analysis Visualization
Wang Zi Investigation
https://orcid.org/0000-0003-1512-7824
Ou Guangshuo Conceptualization Funding acquisition Supervision Writing – original draft Writing – review & editing *
Tsinghua-Peking Center for Life Sciences, Beijing Frontier Research Center for Biological Structure, McGovern Institute for Brain Research, State Key Laboratory of Membrane Biology, School of Life Sciences and MOE Key Laboratory for Protein Science, Tsinghua University, Beijing, China
Xu X. Z. Shawn Editor
University of Michigan, UNITED STATES OF AMERICA
The authors have declared that no competing interests exist.

* E-mail: guangshuoou@tsinghua.edu.cn
26 8 2024
8 2024
20 8 e101137712 4 2024
27 7 2024
© 2024 Guo et al
2024
Guo et al
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Chemical mutagenesis-driven forward genetic screens are pivotal in unveiling gene functions, yet identifying causal mutations behind phenotypes remains laborious, hindering their high-throughput application. Here, we reveal a non-uniform mutation rate caused by Ethyl Methane Sulfonate (EMS) mutagenesis in the C. elegans genome, indicating that mutation frequency is influenced by proximate sequence context and chromatin status. Leveraging these factors, we developed a machine learning enhanced pipeline to create a comprehensive EMS mutagenesis probability map for the C. elegans genome. This map operates on the principle that causative mutations are enriched in genetic screens targeting specific phenotypes among random mutations. Applying this map to Whole Genome Sequencing (WGS) data of genetic suppressors that rescue a C. elegans ciliary kinesin mutant, we successfully pinpointed causal mutations without generating recombinant inbred lines. This method can be adapted in other species, offering a scalable approach for identifying causal genes and revitalizing the effectiveness of forward genetic screens.

Author summary

Exploring gene functions through chemical mutagenesis-driven genetic screens is pivotal, yet the cumbersome task of identifying causative mutations remains a bottleneck, limiting their high-throughput potential. In this investigation, we uncovered a non-uniform mutation pattern induced by Ethyl Methane Sulfonate (EMS) mutagenesis in the C. elegans genome, highlighting the influence of proximate sequence context and chromatin status on mutation frequency. Leveraging these insights, we engineered a machine learning enhanced pipeline to construct a comprehensive EMS mutagenesis probability map for the C. elegans genome. This map operates on the principle that causative mutations are selectively enriched in genetic screens targeting specific phenotypes amid the backdrop of random mutations.

Applying this mapping tool to Whole Genome Sequencing (WGS) data derived from genetic suppressors rescuing a C. elegans ciliary kinesin mutant, we achieved precise identification of causal mutations without resorting to the conventional generation of recombinant inbred lines. Our work not only advances understanding of mutation dynamics but also revitalizes the efficacy of forward genetic screens, contributing to the refinement of genetic exploration methodologies with implications for various organisms.

http://dx.doi.org/10.13039/501100001809 National Natural Science Foundation of China 31991190, 31730052, 31861143042, 31525015, 31561130153 https://orcid.org/0000-0003-1512-7824
Ou Guangshuo http://dx.doi.org/10.13039/501100012165 Key Technologies Research and Development Program 2017YFA0503501, 2019YFA0508401 https://orcid.org/0000-0003-1512-7824
Ou Guangshuo This work was supported by the National Natural Science Foundation of China Grants to GO (31991190, 31730052, 31861143042, 31525015, 31561130153) and National Key R&D Program of China Grants (2017YFA0503501, 2019YFA0508401). The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. PLOS Publication Stagevor-update-to-uncorrected-proof
Publication Update2024-09-06
Data AvailabilityMutation data of MMP dataset were directly downloaded from the MMP homepage (http://genome.sfu.ca/mmp/mmp_mut_strains_data_Mar14.txt).Python scripts and the model are accessible on GihHub at https://github.com/young55775/Genetorch-developing. C. elegans CHIP-chip and CHIP-seq data can be obtained from the modENCODE project (https://compbio.med.harvard.edu/modencode/webpage/Chromatin.v0.6.html) and visualized in Wormbase Genome Browser (JBrowse (wormbase.org)). ModEncode data can also be downloaded from GEO Accession Viewer (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi) with the accession number listed in S1 Table. Code availability: https://github.com/young55775/EMSForest-A-Machine-Learning-Enhanced-EMS-Mutagenesis-Probability-Map.
Data Availability

Mutation data of MMP dataset were directly downloaded from the MMP homepage (http://genome.sfu.ca/mmp/mmp_mut_strains_data_Mar14.txt).Python scripts and the model are accessible on GihHub at https://github.com/young55775/Genetorch-developing. C. elegans CHIP-chip and CHIP-seq data can be obtained from the modENCODE project (https://compbio.med.harvard.edu/modencode/webpage/Chromatin.v0.6.html) and visualized in Wormbase Genome Browser (JBrowse (wormbase.org)). ModEncode data can also be downloaded from GEO Accession Viewer (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi) with the accession number listed in S1 Table. Code availability: https://github.com/young55775/EMSForest-A-Machine-Learning-Enhanced-EMS-Mutagenesis-Probability-Map.
==== Body
pmcIntroduction

Forward genetic screens have been instrumental in demystifying the molecular mechanisms underpinning myriad biological processes across a diverse array of organisms. The procedure commences with the creation of a genetically varied population, often achieved via the application of chemical mutagens such as Ethyl Methane Sulfonate (EMS) [1–3] or through irradiation [4,5], thereby seeding the genome with a variety of mutations. This random mutagenesis generates a pool of organisms each harboring a distinct set of genetic aberrations. Subsequent stages involve diligent phenotypic screening, identifying responsible genes, and decoding biological significance. This strategy has been pivotal in unveiling gene functions across multiple biological spectra, spanning developmental processes [6,7], signal transduction [8], and disease mechanisms [9]. Genes unearthed through these screens often illuminate deeper molecular mechanisms and are frequently employed to glean insights into analogous processes in more complex organisms, thereby contributing profoundly to our understanding of biological systems and phenomena.

Once a desirable mutant phenotype is identified through chemical mutagenesis, a key challenge lies in determining which among potentially thousands of induced mutations is responsible for the observed phenotype [10]. Identifying the causal mutation necessitates a substantial investment of labor and time, often demanding rigorous mapping and validation work to confirm the genetic change responsible for the observed phenotype [11–13]. These limitations, especially when contrasted with RNAi [14,15] or CRISPR-Cas9-based reverse genetics [16] methodologies, have led to the gradual obsolescence of EMS chemical mutagenesis and screens in the model organisms. RNAi, with its ability to knock down genes transiently and its amenability to high-throughput formats [17], enables the systematic analysis of gene function across the genome, offering a powerful tool for functional genomics [18]. With CRISPR-Cas9, not only can specific genes be targeted, but specific types of mutations (e.g., point mutations, deletions, insertions) can be introduced [19], offering a level of precision and control that is simply unattainable with chemical mutagenesis [20,21].

On the other hand, chemical mutagenesis has unique advantages in the generation of a broad spectrum of alleles [22], including hypomorphic, hypermorphic, and neomorphic alleles, providing a rich reservoir for investigating gene function in a depth unattainable through complete gene knockdown or knockout strategies like RNAi or CRISPR-Cas9. Thus, forward genetic screens facilitate the exploration of nuanced gene functions [23], interactions [24] and pathway dynamics [25], often unveiling surprising connections and uncharted biological landscapes that are not immediately evident through targeted gene perturbation approaches [26]. Furthermore, this strategy can illuminate the functionality of genes of unknown function or open reading frames that might be overlooked with hypothesis-driven reverse genetics [27]. This stochastic, phenotype-driven approach allows for the serendipitous discovery of novel genetic interactions and pathways without preconceived notions about gene function, offering the potential to uncover entirely new facets of biology.

Whole-genome sequencing (WGS) has emerged as a formidable instrument for pinpointing causal mutations derived from genetic screens [28]. In the context of Caenorhabditis elegans (C. elegans), several strategies have been devised to reduce the number of candidate causal mutations. One prevalent method involves mating mutants, which are in the N2 background, with the polymorphic Hawaiian strain (CB4856), while alternative strategies exploit DNA variants inherent to the initial background or introduced through mutagenesis [29]. For the identification of suppressor mutations, the Sibling Subtraction Method (SSM) has been developed, which, by excluding genetic variants present in both mutants and their non-mutant siblings, markedly diminishes the roster of candidates [10]. However, these strategies, given their reliance on genetic crosses, do not lend themselves to high-throughput applications. Given that the recent advancement of AlphaMissense has predicted millions of pathogenic variants [30], the incorporation of homologous ones into model organisms for genetic suppressor studies may enhance our comprehension of rescue mechanisms and forge paths towards the proposal of innovative therapeutic strategies. Consequently, there is a pressing demand for the development of scalable methods for identifying causal mutations.

Machine learning, widely used as a statistical method, is a powerful tool for detecting hidden relationships between variables, fitting predictive models to data, and identifying informative groupings within data [31]. It proves particularly valuable when datasets are too complex for straightforward human analysis, enabling the capture of underlying relationships between variables. With rapid advancements in this field, successful attempts have been made to link phenotypes to potential causal variants based on sequencing data [32] or structural information [30]. Most existing tools predict the impact of a variant based on sequence conservation or clinical datasets, such as SIFT [33] or SnpEff [34]. However, when dealing with genetic screening involving highly effective mutagens, pathogenic alleles are not always the causal ones, and the genotype/phenotype relationship can be complicated by different genetic backgrounds. Furthermore, there are many genes whose functions are not fully understood and are not necessarily disease-causing but still worthy of study. A new model for genetic screening is needed to detect causal mutations independent of their known functions to predict the significance of mutated genes in chemical screening experiments.

In this study, we systematically analyzed C. elegans WGS data derived from 737 EMS-mutagenized worms, furnished by the Million Mutation Project (MMP) [22], alongside chromatin immunoprecipitation sequencing (ChIP-seq) datasets of DNA-binding proteins from the modENCODE project [35]. Contrary to expectations, we discovered that EMS-induced mutations were not uniformly distributed across the C. elegans genome. Instead, a genome-wide EMS mutagenesis bias was uncovered, which correlates with both adjacent sequence context and chromatin structure, facilitating the development of a machine learning based model for predicting EMS-induced mutagenesis probabilities. This model enabled the generation of a genome-wide EMS mutagenesis probability map, which predicts the expected frequency of random mutations for each nucleotide within the C. elegans genome. Consequently, our method enhances our capability to discern causative mutations—enriched through genetic screening for a specific phenotype—from random mutations. By employing this pipeline to analyze WGS data of genetic suppressors that rescue a ciliary kinesin harboring a missense mutation in C. elegans, we identified the causal mutations without the requisite establishment of recombinant inbred lines, thus enhancing the efficiency of high-throughput forward genetics.

Result

A non-uniform pattern of EMS-induced mutations across the C. elegans genome

We examined the distribution of a cumulative 265,169 Single Nucleotide Variants (SNVs) and insertion/deletions (InDels) across the 737 EMS-mutagenized genomes from the Million Mutation Project (Fig 1A). A conspicuous non-uniform pattern emerged, characterized by ‘hot spots’—regions exhibiting a higher-than-average SNV density—and ‘cold spots,’ or regions manifesting a lower density. Given that EMS predominantly introduces an alkyl group at N7-guanine and O6-guanine, consequently inducing transitions from ‘G/C’ to ‘A/T,’ we postulated that mutation frequency might be modulated by the genomic concentration of ‘G/C’ nucleotides. Nonetheless, our correlation analysis revealed that the ‘G/C’ pair distribution only partially elucidated the mutation frequency (Fig 1A and 1B). Moreover, we discerned variations in mutation frequency amongst different chromosomes (Fig 1C), with chromosome X exhibiting an elevated mutation frequency and chromosome IV manifesting a diminished frequency, neither of which could be attributed to the ‘G/C’ distribution. These observations parallel the findings from two additional WGS datasets procured from our recent Artificial Evolution Summer School, in which undergraduates perform forward genetic screens for animals exhibiting dumpy (Dpy) or uncoordinated movement (Unc) phenotypes (S1A–S1F Fig). These results imply that EMS-induced variations are not uniformly distributed across the genome, potentially due to disparate local chromosomal features.

10.1371/journal.pgen.1011377.g001 Fig 1 Uneven distribution of EMS-induced SNV.

A) Scatter and line plots representing mutations in the MMP dataset. The Scatter plot shows the mutation number on each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. The Line plot represents the number of ‘C/G’ base pairs in each 100,000 bp region. B) Scatter plot representing the relationship between ‘C/G’ base pair contents and the number of mutations from the MMP data of each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. Pearson correlation was used to evaluate the association between them. C) Kernel density estimation plot of the number of mutations on each chromosome.

The adjacent sequence context influences the probability of EMS mutagenesis

We investigated the distribution of the four nucleotides both upstream (3’ +4/+3/+2/+1) and downstream (5’ -4/-3/-2/-1) of mutated bases across the genome (Fig 2A and 2B). Should EMS-induced mutations transpire indiscriminately, uninfluenced by adjacent bases, an asymmetry-free distribution would be anticipated. However, our chi-square test illuminated that the positions (+2/+1) and (-1/-2) adjacent to the ‘C/G’ nucleotide wield a significant impact on the efficacy of EMS mutagenesis. Examining the MMP dataset, we discerned that certain 5-base sequences exhibit a markedly heightened susceptibility to mutation. For example, sequences such as ‘AGGGG’ and ‘CCCCT’ are roughly tenfold more predisposed to mutation relative to sequences like ‘GTCGA’ and ‘TCGAC’ (Fig 2C). The substantive impact on mutagenesis probability is not confined to the immediate neighbors of nucleotides. Even while maintaining the +1/0/-1 position constant, the +2 and -2 positions continue to sway mutagenesis probability (Fig 2D and 2E), suggesting a viable avenue for predicting biases in mutation frequency. Utilizing this information, the base mutation frequency can be determined by calculating the mutation rate of each pattern in all the mutation events in the Million Mutation Project. Hereafter, we refer to this base mutation frequency as P0, which is determined by: P0=C′C0. C′ represents the count of a specific five-base pattern underwent mutation in the MMP dataset and C0 is the number of the same pattern in the reference genome.

10.1371/journal.pgen.1011377.g002 Fig 2 Flanking sequences affect EMS mutagenesis efficiency.

A) The distribution patterns of flanking sequences of each kind of base (A, T, C, G) in the genome DNA. The number with ‘+’ or ‘-’ represents the base left (5’)/right (3’) to the position 0. These distributions were set as expectation values that if the mutagen selects randomly on a genome, it will cause mutations with the same distribution in flanking sequences. B) The distribution patterns of flanking sequences for each kind of mutated base in MMP dataset. Note that these distribution patterns are different from the corresponding distributions in (A). Chi-square test: * p<0.05, ** p<0.01, ***p<0.001, ****p<0.000. C) Heatmap recording the mutation rate of each 5-base pattern. X-axis was sorted by the first 3 bases and Y-axis containing the next two. D) Bar plots showing the variation of C mutagenesis rate when -1 and +1 positions are fixed, as shown around the inner circle. E) Bar plots showing the variation of G mutagenesis rate when -1 and +1 positions are fixed, as shown around the inner circle.

Multiple DNA-binding protein patterns show correlation with the mutation events

Employing the adjacent sequence context, we crafted a graphical representation that elucidates the frequency of EMS-induced mutations. To exemplify the method, we showcase the map in relation to mutation events across a broad range on chromosome V, spanning from position 1,800,000 to 3,450,000 (Fig 3A and 3B). We incorporated a representative gene, ttn-1, situated on chromosome V between positions 6,120,909 and 6,202,632 (S2A and S2B Fig). In both scenarios, the maps reveal expansive genomic segments characterized by a notable diminution in EMS-induced mutations, the ‘silent region’ of which belies the predictions formulated on adjacent sequence context.

10.1371/journal.pgen.1011377.g003 Fig 3 Mutation numbers in the MMP data shows association with the DNA-binding protein features.

Take the MMP data and DNA-binding protein features on Chr V:1800000~3450000 as an example. A) Line plot showing the prediction made from the Flanking sequence preferences of EMS mutagenesis (P0). B) The mutations observed in the MMP dataset. C) Raw ChIP-chip signal data of UNC-62 binding in young adult worms. D) Raw ChIP-chip signal data of UNC-39 binding in worm embryo. E) Raw ChIP-seq signal data of histone H3 in L3 larval worms. F) Raw ChIP-seq signal data of histone H4 in L3 larval worms. G) Line plot showing the prediction made by a Random Forest regressor trained by DNA-binding protein data.

We posited that a unique chromatin state might elucidate the emergence of these ‘silent regions’. Tightly supercoiled DNA, potentially having circumscribed interactions with extraneous substances [36], along with particular DNA-binding proteins, might confer protection to their target sequences, rendering them less vulnerable to chemicals that engage with DNA [37]. Consequently, we probed the C. elegans ChIP-seq datasets of DNA-binding proteins and histone modifications from the modENCODE project (Figs 3C–3F and S2C–S2F). On chromosome V, a silent region extends over 200kb in length (Fig 3B). Within this expanse, the binding activity of the transcription factor UNC-62 and UNC-39 manifests irregularly, as evidenced by the ChIP-seq signal (Fig 3C and 3D). The absence of ChIP-seq signals in this area might denote a distinctive chromatin state that is less amenable to external substances. Analogously, the ‘silent region’ of the gene ttn-1 also undergoes a comparable albeit less marked depletion of transcription factor binding (S2C and S2D Fig). Nonetheless, a distinct histone binding pattern is evident in this region (S2E and S2F Fig). While histone binding shows a slight increase across the preceding 200 kb ‘silent region’ on chromosome V (Fig 3E and 3F), particularly in H3, the profile of transcription factor binding provides a clearer boundary of this area. This suggests that transcription factor binding may be a more effective feature for capturing the relationship between abnormal mutation frequency and DNA-binding proteins in this case. Probing mutations and the attributes of DNA-binding proteins suggests that regions with elevated or reduced mutation probabilities may originate from direct causative factors or mere coincidence. Nevertheless, it seems judicious to elucidate the binding patterns of these proteins and explore their association with mutation probability.

Using ChIP-seq dataset collections from the WormBase, we next examined the potential correlation between DNA-binding protein features and mutagenesis probability with machine learning techniques.

Nevertheless, it seems judicious to elucidate the binding patterns of these proteins and explore their association with mutation probability. Employing ChIP-seq dataset compilations from WormBase, we subsequently scrutinized the potential correlation between DNA-binding protein characteristics and mutagenesis probability through the lens of machine learning methodologies.

Random Forest regressor (Rfr) modeling to predict the mutation frequency

Employing ChIP-seq dataset compilations from modENCODE and WormBase [38], we examined the potential correlation between DNA-binding protein characteristics and mutagenesis probability through machine learning methods, which can identify complex chromatin feature patterns and learn their relationship with mutation events

A decision tree is a simple yet powerful prediction method [39]. It splits data into subsets in order to create groups with similar values of the target feature. Therefore, the predictor can make predictions from new observations with the relationship it learned from the already exist data. Random Forests are a combination of tree predictors [40]. Each tree is trained on a randomly selected subset of the data and features, capturing slightly different information, and reducing the risk of overfitting, thus improving overall performance.

We developed a regression model using Random Forest regressor. The underlying hypothesis is that, despite the sparsity of mutation events in the Million Mutation Project (200–350 mutations per 100,000 bp, Fig 1A), the mutation probability for base pairs with the same ’property’ is well-represented in the dataset. The random forest algorithm’s task is to categorize base pairs into different groups based on DNA binding protein patterns and flanking sequences, enabling accurate prediction of each group’s average mutation expectation (Fig 4A). Nucleotides with similar flanking sequences and DNA-binding protein attributes are grouped into the same output node of each tree. The mean mutagenesis frequency for that node then serves as the predicted mutagenesis frequency for any nucleotide with similar characteristics assigned to that node during predictions.

10.1371/journal.pgen.1011377.g004 Fig 4 Random Forest modeling and validation.

A) Schematic diagram of the Random Forest regressor (RFr) model. 600 trees were modeled and each tree randomly used at most 70% percent of the 24 features (≤16 features). The smallest leaf size is set to be 30000 to avoid overfitting. In this model, each decision tree in the forest will put bases with similar properties (flanking sequences and DNA-binding protein patterns) together into an output node, and use the average mutation rate in each node as output. Consequently, the forest will take the average output of every 600 trees as the final output, with which the mutation rate of this kind of base can be predicted. B) Prediction made only by P0. The worm genome was divided into 300,000 bp blocks and each dot represents the expectation of mutation number made by P0 and the actual mutation counts in MMP dataset. C) The performance of RFr on the test set, which contains 20% of the whole dataset. The test set was sorted by the prediction and was divided into 100,000 bp blocks. Each dot represents the expectation of mutation number made by RFr and versus actual mutations counts in MMP dataset in each 100,000 bp long blocks. D) Scatter plot of the overall performance of RFr on the whole genome. The genome was divided into 300,000 bp blocks. Each dot represents the expectation of mutation number made by RFr and the actual mutations counts in MMP dataset in each 300,000 bp long blocks. E) Scatter plot of the overall performance of RFr on 19639 genes of C. elegans. Each dot represents the expectation of mutation number made by RFr and the actual mutations counts in MMP dataset of each gene. F) Representative EMS mutagenesis map of gene atm-1. Followed by the mutation map in the MMP dataset and dumpy screening. G) Representative EMS mutagenesis map of gene hum-7. Followed by the mutation map in the MMP dataset and dumpy screening.

To create our model, we used a dataset containing 24 DNA-binding protein datasets from modENCODE, which includes information about histone binding, modifications, transcription factor interactions, and other epigenetic modifications (S1 Table). We constructed 600 decision trees, each based on the mentioned method, with each tree trained on a randomly chosen subset of up to 16 protein features. To create our model, we used a dataset containing 24 DNA-binding protein datasets from modENCODE, which includes information about histone binding, modifications, transcription factor interactions, and other epigenetic modifications (S1 Table). each trained on a randomly chosen subset of up to 16 protein features, based on the method described above.

The ultimate model output can be succinctly represented as P=α600∑i=1600fi(x,P0) (see also Materials and Methods), which involves averaging the predictions from each individual tree to provide a more precise overall prediction. In this equation, the variable α factor in experimental variations stemming from fluctuations in EMS concentrations, developmental stages, ambient temperatures, and other potential elements that may impact EMS mutagenesis effectiveness.

To validate the model’s performance, our initial validation assessed predictions grounded solely on P0, demonstrating congruent patterns across every 300 kb genomic block. These patterns were juxtaposed with the discerned fluctuations in mutation occurrences within the MMP dataset (Fig 4B). Upon training utilizing a randomly selected 80% of the data, the model’s performance was appraised against the test set. The test set was categorized based on prediction results and subsequently partitioned into 100kb blocks. Anticipated mutation tallies were then contrasted against actual mutation instances documented in the MMP dataset (Fig 4C). Further, a thorough evaluation juxtaposed the anticipated mutation tallies within 300kb genomic blocks (Fig 4D) and individual genes, as per the WBcel235 annotation (Fig 4E), against actual mutation tallies from the MMP dataset. Collectively, these comparisons refine the rudimentary P0 map, facilitating a meticulous prediction of EMS mutagenesis probability at individual nucleotides (Figs 3G, 4F, 4G and S2G).

Subsequently, we selected two representative genes and created EMS mutagenesis maps, highlighting variations in mutagenesis frequency across the entire gene. These maps were then compared to mutations observed in the MMP dataset and sequencing data from the dumpy screening mentioned earlier (Fig 4F and 4G). In regions identified as ‘silent region’ based on our predictions, neither screening approach revealed mutations. Conversely, due to the limited number of SNVs in the MMP dataset, numerous nucleotides remained unmutated in this dataset. Notably, the continuous absence of observed mutations on the actual mutation occurrence map did not diminish the model’s ability to predict the potential for mutagenesis: In the dumpy screening, we observed mutations in regions where the MMP dataset displayed minimal mutations (Fig 4F and 4G). This suggests that our model recognized the specific characteristics of nucleotides in these regions and transferred the knowledge that similar nucleotides were prone to mutation in other genomic regions in order to make accurate predictions.

The EMS mutagenesis probability map facilitates the identification of causal mutations

We employed the EMS mutagenesis probability map to discern the causal mutation within a genetic suppressor screen, amending defects instigated by a missense mutation in the ciliary kinesin OSM-3. This kinesin drives intraflagellar transport, a process pivotal for the construction of olfactory cilia in C. elegans sensory neurons. These segments harbor an abundance of G protein-coupled receptors [41], enabling the organism to perceive environmental stimuli, inclusive of various odorant molecules. A functional impairment of the OSM-3 kinesin culminates in a specific loss of distal ciliary segments [42]; concomitantly, the delocalization of GPCRs or the prevention of odorant-receptor interaction causes animal behavioral defects, such as an incapacitation in executing osmotic avoidance (Osm).

Our preceding genetic screens isolated the E251K missense mutation within the motor domain of OSM-3 kinesin. This mutation parallels a pathogenic variant, E253K, identified in the KIF1A kinesin [43], mutations within which give rise to a spectrum of neurological disorders collectively recognized as KIF1A-Associated Neurological Disorder (KAND). The E253K mutation in KIF1A is hypothesized to disturb the flexibility of switch I in the motor domain, thereby suppressing γ-phosphate release and subsequent ATP binding and hydrolysis within the motor domain [44,45]. An in vitro single-molecule motility assay revealed that E253K induces a stringent binding to the microtubule (MT) yet precludes the engagement in processive motion of KIF1A. Consistently, C. elegans harboring the E251K mutation in OSM-3 disrupts their distal ciliary segments with full penetrance. Whereas wild-type animals exploit their distal ciliary segments to uptake the fluorescent dye DiI from the culture medium [46], all examined OSM-3 E251K mutant animals failed to do so (N > 500), manifesting a dye filling defect (Dyf) phenotype unable to take up dye from the environment, which is detectable under a fluorescence stereoscope.

Utilizing the Dyf phenotype as an efficacious readout probing ciliary defects, we executed genetic suppressor screens aimed at restoring dye-filling capacity and recuperating ciliary distal segments. Our screens isolated 38 independent suppressors. We conducted whole-genome sequencing of all suppressors and employed the EMS mutagenesis probability map to analyze the WGS data. Initially, background mutations, defined as those shared across a wide range of samples, were identified, and removed. Subsequently, we evaluated the effectiveness of EMS mutagenesis (Materials and Methods). The remaining mutations were compared to the adjusted mutagenesis map (Fig 5A). To assess the enrichment effect on each gene by the genetic screening, we calculated the fold change and p-value for each gene showing mutations in the background-removed mutation pool (Fig 5A).

10.1371/journal.pgen.1011377.g005 Fig 5 Predict target genes with the EMS mutagenesis map.

A) Schematic diagram of the analysis pipeline. After genetic screening of the EMS mutated worms, those with a target phenotype were sequenced and mutation information was pooled. Then it could be compared with the EMS mutagenesis map based on the RFr (see Materials and Methods). Finally, candidate genes’ mutation profiles were analyzed by the loss of function impact of InDel, stop-gain, splicing, missense mutations. B) Volcano plot showing the result of an OSM-3 (E251K, cas22599, n = 38) suppressor screening comparing to the EMS mutagenesis map. There were 5 genes showing significant difference between mutation expectation and actual mutation number. C) Bar plot showing the mutation properties of these five genes. D) Taking all the mutations inducing a high impact on protein into consideration, each of the 38 samples has one high-impact mutation in one of the three genes. E) Ciliary defects in the OSM-3 (E251K, cas22599) mutant animals rescued by a K04F10.2 mutated allele (Arg506*, cas23441). K04F10.2 gDNA tagged with Scarlet overexpression in cas23441 exhibits shorter cilia. Arrowheads indicate the ciliary base and transition zone, and arrows indicate junctions between the middle and distal segments. F) Cilium length (mean ± SD) in each group, n = 30 to 50. *** P<0.001 one-way ANOVA. G) Volcano plot showing the result of a dumpy screening (N = 240) from EMS mutated N2 wildtype strain comparing to the EMS mutagenesis map. Dumpy-related or body wall related genes are labeled. H) Bar plot showing the mutation properties of genes labeled in (G).

Our analysis unveiled the top five candidate genes, which demonstrated a markedly biased EMS mutagenesis rates in our suppressor screen but did not exhibit any enriched mutation frequency in our Dpy or Unc screens (Fig 5B). Among these candidates, two ciliary kinases, DYF-5 and DYF-18, have been recognized for their efficacy in rescuing ciliary defects in osm-3 mutant animals. We obtained 10 and 16 mutant alleles for dyf-5 and dyf-18, respectively. The majority of these mutant alleles induce missense mutations, introduce stop codons, or cause splicing mutations within the coding region. These results underscore that the EMS mutagenesis probability map facilitates the identification of causal mutations in genes known to participate in this process. The remaining three genes are not well characterized, and it is unclear which one or ones might be implicated in the regulatory mechanisms of OSM-3.

Through categorizing the mutational characteristics of three genes, we observed that two of them harbored mutations—other than missense, introduced stop, or splicing mutations—which are unlikely to disrupt the function of the gene products. In stark contrast, 12 genetic suppressors encompass various loss-of-function mutations within the coding region of the K04F10.2 gene. Noteworthily, given the substantial size of the initial group subjected to screening, the concurrence of two suppressor alleles within the same strain emerges as exceptionally rare. Notably, the mutations with a significant impact did not coincide across the 38 suppressor strains (Fig 5D), suggesting that the 38 genetic suppressors induce defects in three genes: dyf-5, dyf-18, and K04F10.2. Given that we have already procured 12 distinct alleles of K04F10.2, and generating additional mutants of this gene may not yield further insight, we endeavored to conduct transgenic rescue experiments to further determine whether K04F10.2 operates as the suppressor gene. To this end, we introduced the genomic DNA of K04F10.2, tagged with the red fluorescent protein Scarlet and controlled by its 2kb endogenous promoter, into OSM-3 E251K; K04F10.2 double mutant animals (Fig 5E and 5F). The IFT-dynein heavy chain CHE-3, which traverses the entire length of cilia, was marked with green fluorescence, and used as a ciliary marker to ascertain ciliary length (Fig 5E) [41]. As anticipated, GFP fluorescence was observed along cilia measuring 8.09±0.86 μm. In mutants harboring the OSM-3 E251K mutation, the distal ciliary segment was absent, resulting in cilia of the shorter length (4.40±0.93 μm). While OSM-3 E251K;K04F10.2 double mutants restored ciliary length to 7.54±0.60 μm, introducing Scarlet-tagged wild-type genomic DNA of K04F10.2 into the double mutant reduced ciliary length to 6.22±0.81 μm, similar to that of the OSM-3 E251K single mutant. These results indicate that K04F10.2 mutations act as suppressors for OSM-3 E251K. Henceforth, we designate this gene as Joubert syndrome homolog 26 (jbts-26) due to its inferred role in primary cilia. These results demonstrated that the EMS mutagenesis probability map is instrumental in isolating causal mutations in genes previously unlinked to specific processes. 

Discussion

In conclusion, this research delineates the non-uniform mutation rate instigated by EMS mutagenesis across the C. elegans genome, underscoring that both proximate sequence context and chromatin status are ostensibly correlated with mutation frequency. Leveraging these factors, we deployed a Machine Learning assisted pipeline to sculpt a genome-wide EMS mutagenesis probability map. Operating under the premise that causative mutations will achieve enrichment through genetic screens pertinent to a specific phenotype amidst random mutations, we utilized the map to scrutinize Whole Genome Sequencing (WGS) data derived from C. elegans forward genetic screens. The findings illuminate that the map expediently facilitates the discernment of genes known for their instrumental roles in these processes. Notably, the map also prognosticated a novel gene entwined in each regulation, a prediction substantiated through transgenic rescue experiments. This holistic approach to causal gene identification eschews labor intensive genetic crossing and demonstrates that bioinformatic analysis of WGS data via the EMS mutagenesis probability map can present a potent conduit for directly interfacing phenotype with genotype. Our model performs well when multiple alleles per gene are obtained but is less effective when screens yield only one or two alleles per gene. For challenging forward genetic screens such as manual fluorescence-based screens under high-power microscopy, we recommend using the sibling subtraction method. Machine learning approaches rely on statistical effects to achieve desired accuracy, assuming a sufficient number of candidates are screened—typically, we recommend 30 or more samples per experiment. Our tool was originally designed for large-scale causal gene identification. For small datasets, sibling subtraction remains the most reliable method due to its manageable workload when dealing with only a few alleles.

Given the affordability of WGS, WGS data are usually available before sibling subtraction experiments. Analyzing one hundred WGS datasets using the EMS probability map takes less than one minute, and only several seconds for 10 to 20 WGS datasets. Our model significantly enhances the likelihood of identifying potential causal mutations for genes that have low baseline mutation probabilities, irrespective of the number of isolated mutant alleles. Therefore, investing an additional minute in computation could potentially expedite mutant cloning efforts.

While the EMS probability map cannot be universally applied, especially in challenging screens that yield a limited number of mutant alleles, it proves valuable in many forward genetic screens capable of efficiently generating dozens of mutants with manageable effort. Nonetheless, it’s crucial to employ multiple complementary approaches to enhance forward genetic screens. Their strength lies in their capacity to uncover novel mutations that might otherwise go unnoticed, and ability to unveil unknown aspects of biology. Thus, the EMS probability map can be a useful addition to the forward genetic toolkit.

The Random Forest model elucidated the relationship between various DNA-binding protein patterns and mutation frequency. DNA-binding proteins such as histones and transcription factors are commonly employed to assess the accessibility of genomic DNA, from which the topological arrangement of nucleosomes can be inferred [47]. However, this relationship becomes more intricate due to the challenge of precisely uncovering how DNA-binding protein patterns affect EMS effectiveness. For example, the occupancy of transcription factors on chromatin may reduce accessibility to EMS, while regions lacking ChIP-seq signals likely indicate a tightly coiled DNA topology with nucleosomes, making them less accessible to other molecules. The machine learning model established the relationship between EMS mutation frequency and these features by leveraging these correlations, resulting in a comprehensive model capable of handling diverse patterns. Further experiments are necessary to decipher these patterns and gain a deeper understanding of the actual topology and conditions of these regions.

We summarize the mutation counts from various datasets: In the MMP project, 737 WGS datasets of EMS mutagenized worms revealed a total of 265,169 mutations, averaging 360 mutations per worm, or approximately 0.36 mutations per 100,000 base pairs (bp). The DPY dataset, consisting of 240 worms, showed 65,171 mutations, averaging 272 mutations per worm, or approximately 0.27 mutations per 100,000 bp. Lastly, the UNC dataset, with 118 worms, yielded 34,730 mutations, averaging 293 mutations per worm, or approximately 0.30 mutations per 100,000 bp. While the MMP project generates a slightly higher mutation rate (360 mutations per worm) compared to our DPY (272 mutations per worm) or UNC (293 mutations per worm) screens, all mutation rates fall within a similar range.

The elucidation of jbts-26 as a suppressor gene mitigating ciliary defects induced by the OSM-3 E251K mutation inaugurates avenues for mechanistic investigations. Previous research has illuminated that the abrogation of ciliary kinases DYF-5 or DYF-18 permits an alternative ciliary heterotrimeric kinesin-2 to ectopically enter the ciliary distal region [48], thereby substituting the functionality absent in OSM-3. In essence, these two ciliary kinases may not exert their influence directly upon OSM-3. Rather, they curtail a concurrent ciliary transport pathway, and their loss enables heterotrimeric kinesin-2 to compensate for the OSM-3 deficit. Consequently, in the singular mutants of dyf-5 or dyf-18, the organisms exhibit aberrantly elongated cilia [48]. In contrast, jbts-26 might orchestrate its function through a mechanism disparate from these two ciliary kinases. We and others have not been able to discern any overt ciliary defects in the single mutant defective in jbts-26. A pioneering systematic exploration of ciliary genes has annotated jbts-26 as a conserved putative binding associate of a microtubule-severing protein, namely katanin, hinting at a role in microtubule regulation [49]. Nonetheless, the underlying mechanisms remain mysterious. Considering that the E251K mutation in OSM-3 is analogous to the E253K mutation in KIF1A, future studies will ascertain whether the inhibition of jbts-26 can ameliorate neuronal defects induced by KIF1A E253K, potentially unveiling a novel therapeutic target for intervening in KIF1A E253K-associated neurological disorders.

Recent endeavors to conduct genetic suppressor screens for various missense mutations within OSM-3 have been undertaken, inclusive of published suppressors of a motor hinge mutation, OSM-3 G444E [44,45,50]. Nonetheless, no mutations of jbts-26 were unveiled as suppressors for the ciliary defects attributed to OSM-3 G444E, positing that jbts-26 might exert its suppressive effects on the E251K mutation in OSM-3 in a residue-specific modality. This observation accentuates the imperative of functional residuomics: each residue may exhibit intrinsic uniqueness, and mutations at each individual residue may be ameliorated via divergent, distinct mechanisms.

Given the ubiquitous chemical principle underpinning EMS’s capacity to induce genetic mutations across diverse species, it is conceivable that the adjacent sequence context and chromatin status might also engender a non-uniform distribution of EMS-induced alterations throughout all genomes. Congruent with our findings, the efficacy of EMS mutagenesis in rice (Huanghuazhan) is correlated with its flanking sequence and chromatin status [51]. Consequently, it may not be startling to observe that systematic analyses of published WGS data across species divulge a pervasive bias in EMS mutation rates at genome-wide levels. Identifying EMS-induced causal mutations in species characterized by complexities in genome size, ploidy number, and life cycle surpassing those of C. elegans poses a substantially more formidable challenge. Therefore, formulating analogous EMS mutagenesis probability maps for these species could markedly expedite forward genetics endeavors therein.

Materials and methods

WGS data from the Million Mutation project

WGS data from the Million Mutation project can be accessed via the official webpage of Simon Fraser University (http://genome.sfu.ca/mmp/). For our study, we utilized mutation data from a total of 737 strains isolated after EMS mutagenesis was used to train and test the model. Raw sequencing data in fastq format are available from Home—SRA—NCBI (nih.gov). The average sequencing depth of these data were about 15x fold and 265169 variations were found in this EMS mutation dataset.

DNA-binding protein data

C. elegans ChIP-chip and ChIP-seq data can be obtained from the modENCODE project (data.modencode.org) and visualized in WormBase Genome Browser (JBrowse (wormbase.org)).

P0 calculation from the MMP dataset

Flanking sequence preferences of EMS mutagenesis (referred to as P0) were calculated with the MMP dataset and C. elegans genome WBcel235 (Caenorhabditis_elegans - Ensembl Genomes 57) as P0=C′C0

C′ represents the count of a specific five-base pattern underwent mutation in the MMP dataset and C0 is the number of the same pattern in the reference genome.

Preparation of DNA-binding protein data and Random Forest regressor modeling

The raw ChIP-chip and ChIP-seq data were preprocessed to better reflect the actual distribution of DNA-binding protein on the genome. A python script was employed to apply a moving average with the window size of 100 bp to smooth the data across each chromosome. The smoothed DNA-binding protein dataset was then utilized to train and validate a random forest regressor after randomly spilt the data into a 80% training set and a 20% test set.

The random forest comprised a total of 600 decision trees, each trained on the provided training dataset. To prevent overfitting, we implemented feature selection by randomly considering up to 70% of the available features for each tree. Additionally, to maintain model integrity, we set a minimum leaf size of 30,000 instances. To produce the final prediction, the output from each individual tree was aggregated. This ensemble approach allowed us to generate a comprehensive mutation probability assessment for each type of nucleotide.

During validation, the test set, comprising both mutation information and additional features, was sorted based on their predicted mutation rates. Subsequently, the test set was divided into 100kb blocks. To validate the model’s predictions, we computed the mutation expectation and the actual mutation count in the MMP dataset. By calculating the likelihood of mutation for each individual base pair, we were able to determine the mutation rate for each block, or for any genomic sequence length, using the formula: Pseq=∑i∈seqPi.

This approach allowed us to assess the model’s predictive accuracy across different genomic regions.

During a genetic screening involving the potential mutation of millions of base pairs, the number of variations for each nucleotide can be reasonably modeled to follow a binomial distribution. Consequently, the expected mutation count can be calculated as: E(seq)=∑n∈seqE(n)=n*∑n∈seqPn.

This formula is applicable within a population of n strains, and it allows us to estimate the anticipated number of mutations across the sequence under investigation.

The EMS effectiveness can be subject to variation due to factor such as changes in temperature, the age of the worms, minor temporal differences, and alterations in mutagen concentration during the EMS mutagenesis process. To assess this effectiveness, we compare the average number of mutation events per strain in different batches of EMS mutation experiments. We took the mutation frequency of the Million Mutation Project as a baseline. We denote the mutagen effectiveness of the MMP strains as α0 and the efficiency in another experiment as α1, the random forest regression model takes the form of: P=α0f(x,P0)

where x represents the DNA-binding protein features used by the model to make predictions and adjust the initial prediction of P0. This model can be applied to another dataset as: P′seq=α1∑i∈seqf(xi,P0)=N¯1N¯0α0∑i∈seqf(xi,P0)

Here, N¯0 and N¯1 represent the average number of mutation events that occurred in the MMP strains and the batch to be analyzed, respectively. In this way, this model can be used on screening with altered EMS effectiveness.

Model evaluation and feature importance

To evaluate the accuracy of this regression model, we compare the expectation and the actual mutation count within different regions of the C. elegans genome, resulting in the determination of the coefficient of determination (R2), given by: R2=1−MSEVar

Additionally, we assess the importance of each feature used in making predictions through permutation scores. In this analysis, each feature is successively replaced with random noise sharing the same value distribution as the original data. The model, initially trained on the original dataset, is then employed to make predictions using the dataset in which a feature has been replaced with noise. The permutation importance of feature j is subsequently calculated as: ij=R2−1n∑i=1nRi,j2

Where feature j undergoes n shuffling iterations.

The mutation and feature importance data (S3 Fig), have been normalized to a range of 0 to 1 using the formula: V=v−min(v)max(v)−min(v)

Worm culture

C. elegans were maintained under a consistent temperature of 20°C according to the standard method. Nematode growth medium (NGM) with Escherichia Coli OP50 seeded on it was used to cultivate these worms. C. elegans strains used in this study are listed in S2 Table.

EMS mutagenesis

Worms synchronized at the late L4 stage were carefully collected with 4 mL M9 buffer (S1 Table). Subsequently, these collected worms were placed in 50 mM EMS buffer at room temperature with continuous rotation for a duration of 4 hours. Following this treatment, the worms underwent a thorough washing process with M9 buffer and were then cultured under standard conditions. Approximately 20 hours later, the adult worms were subjected to a bleaching procedure to isolate their eggs (referred to as F1). These eggs were subsequently distributed across approximately 100 separate 9 cm NGM plates, with an average of 50 to 100 eggs placed on each plate. Any adult worms displaying the desired phenotype were meticulously collected and placed in individual culture settings. After a careful examination of their offspring, these individuals were subjected to sequencing using an Illumina next-generation sequencer.

WGS data analyze

Mutation data of MMP dataset were directly downloaded from the MMP homepage (http://genome.sfu.ca/mmp/mmp_mut_strains_data_Mar14.txt). Raw reads obtained from the next-generation WGS were assessed for duplication and quality with FastQC and were trimmed using Trim_galore (version 0.4.4) to remove the adaptor sequence and low-quality reads. After that, clean reads were aligned to the reference genome (WBcel235) using BWA-MEM2 (version 2.2) with default parameters. In all these sequencing results, > 20x average coverage is ensured. Variations were detected using freebayes (version 1.3.6) and annotated using SnpEff. Finally, a filter was set to ensure the variation quality and limit false positive rate: only sequence depth >5 and allele frequency > 0.8 variations were taken into further analyze.

Strain construction

To perform the rescue experiment, we synthesized the pK04F10.2::K04F10.2::Scarlet DNA fragment using SOEing PCR [47] and subsequently created the transgenic strain through microinjection.

Supporting information

S1 Fig Uneven distribution of EMS-induced SNV in screening of dumpy and uncoordinated phenotype.

A) Scatter and line plots representing mutations in the Dumpy screening dataset (n = 240). The Scatter plot shows the mutation number on each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. The Line plot represents the number of ‘C/G’ base pairs in each 100,000 bp regions. B) Scatter plot representing the relationship between ‘C/G’ base pair contents and the number of mutations from the Uncoordinated screening data of each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. Pearson correlation was used to evaluate the association between them. C) Kernel density estimation plot of the number of mutations on each chromosome in uncoordinated screening dataset. D) Scatter and line plots representing mutations in the Uncoordinated screening dataset (n = 118). The Scatter plot shows the mutation number on each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. The Line plot represents the number of ‘C/G’ base pairs in each 100,000 bp regions. E) Scatter plot representing the relationship between ‘C/G’ base pair contents and the number of mutations from the dumpy screening data of each chromosome. Each dot represents the number of mutated bases within every 100,000 bp counted from the first base pair of each chromosome. Pearson correlation was used to evaluate the association between them. F) Kernel density estimation plot of the number of mutations on each chromosome in dumpy screening dataset.

(TIF)

S2 Fig Mutation numbers in the MMP data shown association with the DNA-binding protein features.

Take the MMP data and DNA-binding protein features on Chr V:6120909~6202632 (ttn-1) as an example. A) Line plots showing the prediction made from the Flanking sequence preferences of EMS mutagenesis (P0). B) The mutations observed in the MMP dataset. C) Raw ChIP-chip signal data of UNC-62 binding in young adult worms. D) Raw ChIP-chip signal data of UNC-39 binding in worm embryo. E) Raw ChIP-seq signal data of histone H3 in L3 larval worms. F) Raw ChIP-seq signal data of histone H4 in L3 larval worms. G) Line plots showing the prediction made by a Random Forest regressor trained by DNA-binding protein data.

(TIF)

S3 Fig Feature importance analysis of the Random Forest regressor.

A-F) Permutation importance (see Materials and Methods) of each feature used to model the Random Forest regressor. The importance is shown in the heatmap alongside the chromosome. Mutation rate is normalized to 0~1 and is shown in the bar plot attaching to the right side of the heatmap.

(TIF)

S1 Table Features for Random Forest training.

(DOCX)

S2 Table C. elegans Strains in this study.

(DOCX)

S3 Table PCR products for C. elegans transgenesis.

(DOCX)

10.1371/journal.pgen.1011377.r001
Decision Letter 0
Xu X.Z. Shawn Guest Editor
Zhu Xiaofeng Section Editor
© 2024 Xu, Zhu
2024
Xu, Zhu
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version0
11 Jul 2024

Dear Dr. Ou,

Thank you very much for submitting your Research Article entitled 'A Machine Learning Enhanced EMS Mutagenesis Probability Map for Efficient Identification of Causal Mutations in Caenorhabditis elegans' to PLOS Genetics. My apologies for the delay. One reviewer was unable to provide comments, and we had to find a new reviewer. 

The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important topic but identified some concerns that we ask you address in a revised manuscript.

We therefore ask you to modify the manuscript according to the review recommendations. Your revisions should address the specific points made by each reviewer. In particular, please revise the text related to Machine Learning in such a way that it can be better appreciated by the general audience. 

In addition we ask that you:

1) Provide a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript.

2) Upload a Striking Image with a corresponding caption to accompany your manuscript if one is available (either a new image or an existing one from within your manuscript). If this image is judged to be suitable, it may be featured on our website. Images should ideally be high resolution, eye-catching, single panel square images. For examples, please browse our archive. If your image is from someone other than yourself, please ensure that the artist has read and agreed to the terms and conditions of the Creative Commons Attribution License. Note: we cannot publish copyrighted images.

We hope to receive your revised manuscript within the next 30 days. If you anticipate any delay in its return, we would ask you to let us know the expected resubmission date by email to plosgenetics@plos.org.

If present, accompanying reviewer attachments should be included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist.

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org.

Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process.

To resubmit, log into your Editorial Manager account and select the option 'Revise Submission' in the 'Submissions Needing Revision' folder.

Please let us know if you have any questions while making these revisions.

Yours sincerely,

Shawn Xu

Guest Editor

PLOS Genetics

Xiaofeng Zhu

Section Editor

PLOS Genetics

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: Chemical-mutagenesis forward-genetic screening is a powerful tool in the realm of genetic screening. Forward genetic techniques offer advantages over reverse genetic techniques (RNAi and CRISPR/Cas9) in that it can reveal novel genes and lead to the discovery of unexpected connections and mechanisms. However, identifying the responsible mutation for a given phenotype is laborious and can require a significant time investment.

The authors of this study outline the development of a machine learning tool that can be used to identify causative mutations generated through EMS chemical mutagenesis. They show that the distribution of EMS-induced mutations is non-uniform throughout the genome, both within and across chromosomes; specific nucleotide sequences and different chromatin statuses are biased towards different mutation frequency. This is a key observation that is used to train the Random Forest regressor (Rfr) machine learning tool which can predict what genes give rise to an observed phenotype following mutagenesis. The Rfr model is used to identify the suppressor gene K04F10.2 after EMS-mutagenesis, which can restore ciliary function in OSM-3 mutant C. elegans. Taken together, the conclusions reached in the study are very promising. However, I have a number of concerns about the figures and the explanation of key takeaways as outlined below.

Major points

1) Introduction. The discussion of genetic screening techniques is adequate. Given that the study intersects with the field of machine learning, there should also be a priming outline of machine learning practices that already exist in the field. Add more background for strictly biological researchers who may not be familiar with machine learning techniques and probability mapping.

2) Line 201. Formulation of P0 should also be inserted into the main text since many future panels rely on an understanding of this calculation.

3) Figure 3E/F. It is stated that a distinctive histone binding pattern is not seen in the ‘silent region’ of figure 3 as is the case in the ttn-1 gene. Yet there does seem to be a slight increase in histone binding activity in this region compared to the entire chromosome sequence depicted, particularly in H3. The text does not elaborate on panels E/F specifically, and the conclusion that this pattern is “not observed” is not entirely convincing.

4) Figure 3E/F. There are conflicting results for how histone modification affects mutation frequency in the ‘silent region’ versus the ttn-1 gene. Is it known how “alternative chromatin states” (line 226) affect mutagenesis? Are there experiments that can elucidate the binding patterns at play here? Figure 3 as a whole requires a more detailed explanation with greater emphasis on the rhetorical flow of logic.

5) Figure 5E/F, lines 436-440. The results given in these panels are unclear and appear incongruous with the written text. Did OSM-3 E251K; K04F10.2 double mutants receive the Scarlet-tagged protein or wild type K04F10.2? Is the succeeding condition with Scarlet protein meant to show a diminished capacity of K04F10.2-tagged Scarlet to rescue OSM-3?

Minor points

1) Line 456. typo in “laborintensive”

2) Lines 229-241. Nearly identical paragraphs repeated back-to-back.

3) Line 257. Missing reference for WormBase.

4) Figure 5 legend. Font inconsistency.

5) Figure S1. Misplaced label in panel B.

Reviewer #2: The manuscript by Guo et al presents an innovative and potentially game-changing strategy for chemically-induced forward genetic screening. In the past, mutagenesis-based genetic screens were standard in model organism fields. With technological advances including bioinformatics and genomic engineering, reverse genetic strategies have become the standard to ascertain gene function. Due to the difficulty in identifying causal lesions in mutagenesis screens and ease of knock-down/knockout strategies, forward screens are falling out of favor. However, as authors point out, the power of forward genetic screens is their unbiased nature, their ability to identify novel mutations that may be otherwise missed, and their power to reveal unknown biology. Here, authors use a strategy that capitalizes on the numerous community generated resources (WormBase, MillionMutations project, CHIP-seq, modENCODE) to identify causal mutations in WGS in two proof-of-principle EMS-based screens. In addition to this, authors show that mutation frequency is influenced by nearby sequences and chromatin status. This manuscript will appeal to the readership of PLOS Genetics, with some modification.

As written, the reader needs to be expert in C. elegans, in statistics/machine learning, and in cilia biology. This reviewer is versed in two, but not machine learning and struggled with this section of the manuscript. For example, it would help to more thoroughly explain how the Random Forest Regressor works and why this model was chosen. The same would be true for the non-worm or non-cilia reader in those section. Authors must make the manuscript accessible to the broad readership of PLOS Genetics, so that the impact and usefulness of this powerful strategy can be appreciated and employed by model organism geneticists.

Second, while the data is convincing that this strategy works, how can this be applied in a practical sense? It would be great if authors could work with WormBase/Alliance of Genome Resources to develop a user-friendly interface. I realize this is beyond the scope of this manuscript, but hope this is a future direction.

A few minor things:

Lines 195, 197: typo isare

Lines 230-242 are redundant/garbled

Throughout manuscript: check C. elegans nomenclature for italicizing gene names

line 369 define “back door” (example of needing to be a cilia/kinesin aficionado)

suggestion – request K04F10.2 be named jbts-26

line 434 – Fig. 5E

lines 465, 468: I think kinesin-II is older nomenclature. Heterotrimeric kinesin-2?

Maureen Barr

Reviewer #3: The comments to the authors are provided as an attachment.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: Yes

Reviewer #2: None

Reviewer #3: Yes

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: Yes: Rui Xiao

Reviewer #2: No

Reviewer #3: No

Attachment Submitted filename: comments to authors.docx

10.1371/journal.pgen.1011377.r002
Author response to Decision Letter 0
Submission Version1
22 Jul 2024

Attachment Submitted filename: 20240718 PLOS Genetics response.docx

10.1371/journal.pgen.1011377.r003
Decision Letter 1
Xu X.Z. Shawn Guest Editor
Zhu Xiaofeng Section Editor
© 2024 Xu, Zhu
2024
Xu, Zhu
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version1
27 Jul 2024

Dear Dr Oh,

We are pleased to inform you that your manuscript entitled "A Machine Learning Enhanced EMS Mutagenesis Probability Map for Efficient Identification of Causal Mutations in Caenorhabditis elegans" has been editorially accepted for publication in PLOS Genetics. Congratulations!

Before your submission can be formally accepted and sent to production you will need to complete our formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Please note: the accept date on your published article will reflect the date of this provisional acceptance, but your manuscript will not be scheduled for publication until the required changes have been made.

Once your paper is formally accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you’ve already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosgenetics@plos.org.

In the meantime, please log into Editorial Manager at https://www.editorialmanager.com/pgenetics/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production and billing process. Note that PLOS requires an ORCID iD for all corresponding authors. Therefore, please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field.  This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager.

If you have a press-related query, or would like to know about making your underlying data available (as you will be aware, this is required for publication), please see the end of this email. If your institution or institutions have a press office, please notify them about your upcoming article at this point, to enable them to help maximise its impact. Inform journal staff as soon as possible if you are preparing a press release for your article and need a publication date.

Thank you again for supporting open-access publishing; we are looking forward to publishing your work in PLOS Genetics!

Yours sincerely,

Shawn Xu

Guest Editor

PLOS Genetics

Xiaofeng Zhu

Section Editor

PLOS Genetics

www.plosgenetics.org

Twitter: @PLOSGenetics

----------------------------------------------------

Comments from the reviewers (if applicable):

----------------------------------------------------

Data Deposition

If you have submitted a Research Article or Front Matter that has associated data that are not suitable for deposition in a subject-specific public repository (such as GenBank or ArrayExpress), one way to make that data available is to deposit it in the Dryad Digital Repository. As you may recall, we ask all authors to agree to make data available; this is one way to achieve that. A full list of recommended repositories can be found on our website.

The following link will take you to the Dryad record for your article, so you won't have to re‐enter its bibliographic information, and can upload your files directly: 

http://datadryad.org/submit?journalID=pgenetics&manu=PGENETICS-D-24-00407R1

More information about depositing data in Dryad is available at http://www.datadryad.org/depositing. If you experience any difficulties in submitting your data, please contact help@datadryad.org for support.

Additionally, please be aware that our data availability policy requires that all numerical data underlying display items are included with the submission, and you will need to provide this before we can formally accept your manuscript, if not already present.

----------------------------------------------------

Press Queries

If you or your institution will be preparing press materials for this manuscript, or if you need to know your paper's publication date for media purposes, please inform the journal staff as soon as possible so that your submission can be scheduled accordingly. Your manuscript will remain under a strict press embargo until the publication date and time. This means an early version of your manuscript will not be published ahead of your final version. PLOS Genetics may also choose to issue a press release for your article. If there's anything the journal should know or you'd like more information, please get in touch via plosgenetics@plos.org.

10.1371/journal.pgen.1011377.r004
Acceptance letter
Xu X.Z. Shawn Guest Editor
Zhu Xiaofeng Section Editor
© 2024 Xu, Zhu
2024
Xu, Zhu
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
20 Aug 2024

PGENETICS-D-24-00407R1

A Machine Learning Enhanced EMS Mutagenesis Probability Map for Efficient Identification of Causal Mutations in Caenorhabditis elegans

Dear Dr Ou,

We are pleased to inform you that your manuscript entitled "A Machine Learning Enhanced EMS Mutagenesis Probability Map for Efficient Identification of Causal Mutations in Caenorhabditis elegans " has been formally accepted for publication in PLOS Genetics! Your manuscript is now with our production department and you will be notified of the publication date in due course.

The corresponding author will soon be receiving a typeset proof for review, to ensure errors have not been introduced during production. Please review the PDF proof of your manuscript carefully, as this is the last chance to correct any errors. Please note that major changes, or those which affect the scientific understanding of the work, will likely cause delays to the publication date of your manuscript.

Soon after your final files are uploaded, unless you have opted out or your manuscript is a front-matter piece, the early version of your manuscript will be published online. The date of the early version will be your article's publication date. The final article will be published to the same URL, and all versions of the paper will be accessible to readers.

Thank you again for supporting PLOS Genetics and open-access publishing. We are looking forward to publishing your work!

With kind regards,

Judit Kozma

PLOS Genetics

On behalf of:

The PLOS Genetics Team

Carlyle House, Carlyle Road, Cambridge CB4 3DN | United Kingdom

plosgenetics@plos.org | +44 (0) 1223-442823

plosgenetics.org | Twitter: @PLOSGenetics
==== Refs
References

1 Brenner S. The genetics of Caenorhabditis elegans. Genetics. 1974;77 (1 ):71–94. doi: 10.1093/genetics/77.1.71 4366476
2 Sega GA . A review of the genetic effects of ethyl methanesulfonate. Mutat Res. 1984;134 (2–3 ):113–42. doi: 10.1016/0165-1110(84)90007-1 6390190
3 St Johnston D. The art and design of genetic screens: Drosophila melanogaster. Nat Rev Genet. 2002;3 (3 ):176–88. doi: 10.1038/nrg751 11972155
4 Puck TT , Harvey WF . Gamma-ray mutagenesis measurement in mammalian cells. Mutat Res. 1995;329 (2 ):173–81. doi: 10.1016/0027-5107(95)00028-h 7603499
5 Russell WL , Russell LB , Kelly EM . Radiation dose rate and mutation frequency. Science. 1958;128 (3338 ):1546–50. doi: 10.1126/science.128.3338.1546 13615306
6 Nüsslein-Volhard C , Wieschaus E . Mutations affecting segment number and polarity in Drosophila. Nature. 1980;287 (5785 ):795–801. doi: 10.1038/287795a0 6776413
7 Nüsslein-Volhard C , Wieschaus E , Kluding H . Mutations affecting the pattern of the larval cuticle inDrosophila mel anogaster: I. Zygotic loci on the second chromosome. Wilehm Roux Arch Dev Biol. 1984;193 (5 ):267–82. doi: 10.1007/BF00848156 28305337
8 Iwanami N , Sikora K , Richter AS , Mönnich M , Guerri L , Soza-Ried C , et al . Forward Genetic Screens in Zebrafish Identify Pre-mRNA-Processing Path ways Regulating Early T Cell Development. Cell Rep. 2016;17 (9 ):2259–70. doi: 10.1016/j.celrep.2016.11.003 27880902
9 Zhu Z , Li D , Jia Z , Zhang W , Chen Y , Zhao R , et al . Global histone H2B degradation regulates insulin/IGF signaling-mediate d nutrient stress. EMBO J. 2023;42 (19 ):e113328. doi: 10.15252/embj.2022113328 37641865
10 Joseph BB , Blouin NA , Fay DS . Use of a Sibling Subtraction Method for Identifying Causal Mutations i n Caenorhabditis elegans by Whole-Genome Sequencing. G3 (Bethesda). 2018;8 (2 ):669–78. doi: 10.1534/g3.117.300135 29237702
11 Doitsidou M , Jarriault S , Poole RJ . Next-Generation Sequencing-Based Approaches for Mutation Mapping and I dentification in Caenorhabditis elegans. Genetics. 2016;204 (2 ):451–74. doi: 10.1534/genetics.115.186197 27729495
12 Schneeberger K. Using next-generation sequencing to isolate mutant genes from forward genetic screens. Nat Rev Genet. 2014;15 (10 ):662–76. doi: 10.1038/nrg3745 25139187
13 Serrano M , Kombrink E , Meesters C . Considerations for designing chemical screening strategies in plant bi ology. Front Plant Sci. 2015;6 :131. doi: 10.3389/fpls.2015.00131 25904921
14 Fire A , Xu S , Montgomery MK , Kostas SA , Driver SE , Mello CC . Potent and specific genetic interference by double-stranded RNA in Cae norhabditis elegans. Nature. 1998;391 (6669 ):806–11. doi: 10.1038/35888 9486653
15 Gao M , Monian P , Pan Q , Zhang W , Xiang J , Jiang X . Ferroptosis is an autophagic cell death process. Cell Res. 2016;26 (9 ):1021–32. doi: 10.1038/cr.2016.95 27514700
16 Shalem O , Sanjana NE , Zhang F . High-throughput functional genomics using CRISPR-Cas9. Nat Rev Genet. 2015;16 (5 ):299–311. doi: 10.1038/nrg3899 25854182
17 Simpson KJ , Davis GM , Boag PR . Comparative high-throughput RNAi screening methodologies in C. elegans and mammalian cells. N Biotechnol. 2012;29 (4 ):459–70. doi: 10.1016/j.nbt.2012.01.003 22306616
18 Gudmunds E , Wheat CW , Khila A , Husby A . Functional genomic tools for emerging model species. Trends Ecol Evol. 2022;37 (12 ):1104–15. doi: 10.1016/j.tree.2022.07.004 35914975
19 Doudna JA , Charpentier E . Genome editing. The new frontier of genome engineering with CRISPR-Cas 9. Science. 2014;346 (6213 ):1258096. doi: 10.1126/science.1258096 25430774
20 Hsu PD , Lander ES , Zhang F . Development and applications of CRISPR-Cas9 for genome engineering. Cell. 2014;157 (6 ):1262–78. doi: 10.1016/j.cell.2014.05.010 24906146
21 Wang T , Wei JJ , Sabatini DM , Lander ES . Genetic screens in human cells using the CRISPR-Cas9 system. Science. 2014;343 (6166 ):80–4. doi: 10.1126/science.1246981 24336569
22 Thompson O , Edgley M , Strasbourger P , Flibotte S , Ewing B , Adair R , et al . The million mutation project: a new approach to genetics in Caenorhabd itis elegans. Genome Res. 2013;23 (10 ):1749–62. doi: 10.1101/gr.157651.113 23800452
23 Baer J , Taylor I , Walker JC . Disrupting ER-associated protein degradation suppresses the abscission defect of a weak hae hsl2 mutant in Arabidopsis. J Exp Bot. 2016;67 (18 ):5473–84. doi: 10.1093/jxb/erw313 27566817
24 Glass F , Härtel B , Zehrmann A , Verbitskiy D , Takenaka M . MEF13 Requires MORF3 and MORF8 for RNA Editing at Eight Targets in Mit ochondrial mRNAs in Arabidopsis thaliana. Mol Plant. 2015;8 (10 ):1466–77. doi: 10.1016/j.molp.2015.05.008 26048647
25 Mishra A , Singh A , Sharma M , Kumar P , Roy J . Development of EMS-induced mutation population for amylose and resista nt starch variation in bread wheat (Triticum aestivum) and identificat ion of candidate genes responsible for amylose variation. BMC Plant Biol. 2016;16 (1 ):217. doi: 10.1186/s12870-016-0896-z 27716051
26 Li D , Liu Y , Yi P , Zhu Z , Li W , Zhang QC , et al . RNA editing restricts hyperactive ciliary kinases. Science. 2021;373 (6558 ):984–91. doi: 10.1126/science.abd8971 34446600
27 Ruan B , Hua Z , Zhao J , Zhang B , Ren D , Liu C , et al . OsACL-A2 negatively regulates cell death and disease resistance in ric e. Plant Biotechnol J. 2019;17 (7 ):1344–56. doi: 10.1111/pbi.13058 30582769
28 Sarin S , Prabhu S , O’Meara MM , Pe’er I , Hobert O . Caenorhabditis elegans mutant allele identification by whole-genome se quencing. Nat Methods. 2008;5 (10 ):865–7. doi: 10.1038/nmeth.1249 18677319
29 Wicks SR , Yeh RT , Gish WR , Waterston RH , Plasterk RH . Rapid gene mapping in Caenorhabditis elegans using a high density poly morphism map. Nat Genet. 2001;28 (2 ):160–4. doi: 10.1038/88878 11381264
30 Cheng J , Novati G , Pan J , Bycroft C , Žemgulytė A , Applebaum T , et al . Accurate proteome-wide missense variant effect prediction with AlphaMi ssense. Science. 2023;381 (6664 ):eadg7492. doi: 10.1126/science.adg7492 37733863
31 Greener JG , Kandathil SM , Moffat L , Jones DT . A guide to machine learning for biologists. Nature Reviews Molecular Cell Biology. 2022;23 (1 ):40–55. doi: 10.1038/s41580-021-00407-0 34518686
32 Mountjoy E , Schmidt EM , Carmona M , Schwartzentruber J , Peat G , Miranda A , et al . An open approach to systematically prioritize causal variants and genes at all published human GWAS trait-associated loci. Nat Genet. 2021;53 (11 ):1527–33. doi: 10.1038/s41588-021-00945-5 34711957
33 Ng PC , Henikoff S . SIFT: predicting amino acid changes that affect protein function. Nucleic Acids Research. 2003;31 (13 ):3812–4. doi: 10.1093/nar/gkg509 12824425
34 Cingolani P , Platts A , Wang le L , Coon M , Nguyen T , Wang L , et al . A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin). 2012;6 (2 ):80–92. doi: 10.4161/fly.19695 22728672
35 Celniker SE , Dillon LAL , Gerstein MB , Gunsalus KC , Henikoff S , Karpen GH , et al . Unlocking the secrets of the genome. Nature. 2009;459 (7249 ):927–30. doi: 10.1038/459927a 19536255
36 Bowater RP , Bohálová N , Brázda V . Interaction of Proteins with Inverted Repeats and Cruciform Structures in Nucleic Acids. Int J Mol Sci. 2022;23 (11 ):6171. doi: 10.3390/ijms23116171 35682854
37 Arnold AR , Barton JK . DNA protection by the bacterial ferritin Dps via DNA charge transport. J Am Chem Soc. 2013;135 (42 ):15726–9. doi: 10.1021/ja408760w 24117127
38 Stein L , Sternberg P , Durbin R , Thierry-Mieg J , Spieth J . WormBase: network access to the genome and biology of Caenorhabditis elegans. Nucleic Acids Research. 2001;29 (1 ):82–6. doi: 10.1093/nar/29.1.82 11125056
39 Krzywinski M , Altman N . Classification and regression trees. Nature Methods. 2017;14 (8 ):757–8.
40 Breiman L. Random Forests. Mach Learn. 2001;45 (1 ):5–32.
41 Pan X , Ou G , Civelekoglu-Scholey G , Blacque OE , Endres NF , Tao L , et al . Mechanism of transport of IFT particles in C. elegans cilia by the con certed action of kinesin-II and OSM-3 motors. J Cell Biol. 2006;174 (7 ):1035–45. doi: 10.1083/jcb.200606003 17000880
42 Ou G , Blacque OE , Snow JJ , Leroux MR , Scholey JM . Functional coordination of intraflagellar transport motors. Nature. 2005;436 (7050 ):583–7. doi: 10.1038/nature03818 16049494
43 Themistocleous AC , Baskozos G , Blesneac I , Comini M , Megy K , Chong S , et al . Investigating genotype-phenotype relationship of extreme neuropathic pain disorders in a UK national cohort. Brain Commun. 2023;5 (2 ):fcad037. doi: 10.1093/braincomms/fcad037 36895957
44 Nitta R , Kikkawa M , Okada Y , Hirokawa N . KIF1A alternately uses two loops to bind microtubules. Science. 2004;305 (5684 ):678–83. doi: 10.1126/science.1096621 15286375
45 Yun M , Zhang X , Park CG , Park HW , Endow SA . A structural pathway for activation of the kinesin motor ATPase. EMBO J. 2001;20 (11 ):2611–8. doi: 10.1093/emboj/20.11.2611 11387196
46 Hedgecock EM , Culotti JG , Thomson JN , Perkins LA . Axonal guidance mutants of Caenorhabditis elegans identified by fillin g sensory neurons with fluorescein dyes. Dev Biol. 1985;111 (1 ):158–70. doi: 10.1016/0012-1606(85)90443-9 3928418
47 Klemm SL , Shipony Z , Greenleaf WJ . Chromatin accessibility and the regulatory epigenome. Nature Reviews Genetics. 2019;20 (4 ):207–20. doi: 10.1038/s41576-018-0089-8 30675018
48 Yi P , Xie C , Ou G . The kinases male germ cell-associated kinase and cell cycle-related ki nase regulate kinesin-2 motility in Caenorhabditis elegans neuronal ci lia. Traffic. 2018;19 (7 ):522–35. doi: 10.1111/tra.12572 29655266
49 Sanders AAWM , de Vrieze E , Alazami AM Alzahrani F , Malarkey EB , Sorusch N , et al . KIAA0556 is a novel ciliary basal body component mutated in Joubert sy ndrome. Genome Biol. 2015;16 (1 ):293.26714646
50 Imanishi M , Endres NF , Gennerich A , Vale RD . Autoinhibition regulates the motility of the C. elegans intraflagellar transport motor OSM-3. J Cell Biol. 2006;174 (7 ):931–7. doi: 10.1083/jcb.200605179 17000874
51 Yan W , Deng XW , Yang C , Tang X . The Genome-Wide EMS Mutagenesis Bias Correlates With Sequence Context and Chromatin Structure in Rice. Front Plant Sci. 2021;12 :579675. doi: 10.3389/fpls.2021.579675 33841451
