
==== Front
Mol Ther Nucleic Acids
Mol Ther Nucleic Acids
Molecular Therapy. Nucleic Acids
2162-2531
American Society of Gene & Cell Therapy

S2162-2531(24)00187-2
10.1016/j.omtn.2024.102300
102300
Commentary
How well does the adaptive feature representation learning approach identify human mRNA N4-acetylcytidine sites?
Kumar Rahul 1
Wang Yanfeng yfwang0703@163.com
2∗
Dhanda Sandeep Kumar sandeep.dhanda@stjude.org
3∗∗
1 Department of Biotechnology, Indian Institute of Technology, Hyderabad, India
2 Beidahuang Industry Group General Hospital, Harbin 150001, China
3 Department of Oncology, St. Jude Children’s Research Hospital, Memphis, TN 38103, USA
∗ Corresponding author: Yanfeng Wang, Beidahuang Industry Group General Hospital, Harbin 150001, China. yfwang0703@163.com
∗∗ Corresponding author: Sandeep Kumar Dhanda, Department of Oncology, St. Jude Children’s Research Hospital, Memphis, TN 38103, USA. sandeep.dhanda@stjude.org
29 8 2024
10 9 2024
29 8 2024
35 3 102300© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
==== Body
pmcMain text

RNA molecules undergo extensive modifications by enzymes known as RNA modification enzymes. Over 170 different types of these post-transcriptional modifications have been identified,1 and it is expected that this number will continue to grow. The field of RNA modification research has experienced a resurgence, driven by mounting evidence highlighting its crucial role in gene expression regulation. Among these modifications is N4-acetylcytidine (ac4C), a highly conserved modification found across all domains of life. It involves the addition of an acetyl group to the fourth nitrogen atom of cytidine and is primarily located within coding sequences of the transcriptome. Originally identified at the wobble position 34 of bacterial tRNAMet, ac4C has since been discovered in eukaryotic tRNAs and 18S rRNA. ac4C plays a crucial role in enhancing translation efficiency by promoting precise codon recognition during protein synthesis.2 Conventional methods for identifying these modification sites genome-wide are often costly, slow, and labor intensive. Nonetheless, computational techniques may provide a swift, efficient, and cost-effective alternative to traditional experimental methods. In this issue of Molecular Therapy Nucleic Acids, Pham et al.3 developed a robust computational predictor to identify ac4C sites using an adaptive feature representation learning framework. Notably, they employed a wide array of feature descriptors, classifiers, and feature optimization techniques to select the most informative baseline models for building their ac4C predictor (ac4C-AFL). However, the study has limitations: the training dataset may not fully represent all possible ac4C modification sites, and the number of negative samples (non-ac4C sites) used was limited.

In 2023, Su et al.4 constructed a reliable benchmarking dataset based on the acRIP-seq data.2 They partitioned these data into 80% data for developing the prediction model and 20% of the samples for evaluating the model transferability. Notably, this is a redundancy-reduced dataset, and primary RNA sequence fragments of 201 base pairs. Utilizing these data samples, Su et al. developed an iRNA-ac4C, a single classifier-based model with optimal features from Kmer, nucleotide chemical properties, and accumulated nucleotide frequency. In this issue, Pham et al. utilized the same dataset and fine-tuned the sequence length to focus on the optimal neighboring nucleotides around the modification sites. They employed a wide array of 16 feature encoding methods, encompassing aspects such as sequence composition, physicochemical characteristics, positional information, and NLP-based representations. The authors implemented a two-step feature selection approach. First, they ranked the features using their novel ensemble feature importance scoring (EFIS) techniques. Then, they applied a sequential forward search to determine the optimal subset of features for each of the 16 different encodings. Notably, this exhaustive feature selection approach significantly reduced the feature dimension in the optimal sets, improving prediction performance compared to that using the original dimensions. Subsequently, they trained the optimal features using 11 different classifiers, including machine learning and deep learning classifiers, resulting in 176 baseline models. Finally, a two-step feature selection technique was employed to identify the most effective baseline models, combining their predictions and training them with an SVM classifier to create the final ac4C-AFL prediction model. This model is freely accessible as a web server at https://balalab-skku.org/ac4C-AFL.

Furthermore, Pham et al. analyzed the feature distribution of the 110-dimensional (110D) optimal probabilistic features along with the top five feature descriptors using a t-distributed stochastic neighbor embedding (t-SNE) plot. While individual feature descriptors demonstrated some separation between positive and negative samples in t-SNE plots, the combined 110D feature vector exhibited far superior separation, with minimal overlap. This highlights the effectiveness of adaptive feature representation learning in distinguishing ac4C from non-ac4C samples. Consequently, this approach offers a significant improvement in discriminating these two classes and holds potential for identifying other post-transcriptional modifications.

Comprehensive training and independent validation of ac4C-AFL achieved decent performances. Overall, ac4C-AFL significantly outperformed the existing best predictor, iRNA-ac4C. Despite its consistent performance between the training and independent datasets, there is still room for improvement, as the current accuracy is ∼83%. The reliance on handcrafted features, which may inadvertently remove important information during the optimization process, could be a contributing factor. Exploring deep learning algorithms and diverse architecures5 might be a promising avenue for improving performance.

The rapid advancement of experimental techniques for identifying various RNA epigenetic modifications necessitates the development of sequence-based computational methods, like the one proposed by Pham et al., to decipher their biological functions. In their study, Pham et al. employed a systematic approach and derived optimal baseline models based on adaptive feature representation learning, which effectively differentiates ac4C from non-ac4C sites. The web server they provide will be a valuable resource for the research community to predict potential ac4C sites before conducting experiments. However, further research is needed to fully understand the precise mechanisms underlying modified mRNA formation and their impact on biological processes.

Acknowledgments

The authors did not receive any funding to complete this manuscript.

Declaration of interests

The authors declare no competing interests.
==== Refs
References

1 Zhang Y. Lu L. Li X. Detection technologies for RNA modifications Exp. Mol. Med. 54 2022 1601 1616 36266445
2 Arango D. Sturgill D. Alhusaini N. Dillman A.A. Sweet T.J. Hanson G. Hosogane M. Sinclair W.R. Nanan K.K. Mandler M.D. Acetylation of Cytidine in mRNA Promotes Translation Efficiency Cell 175 2018 1872 1886.e24 30449621
3 Pham N.T. Terrance A.T. Jeon Y.J. Rakkiyappan R. Manavalan B. ac4C-AFL: A high-precision identification of human mRNA N4-acetylcytidine sites based on adaptive feature representation learning Mol. Ther. Nucleic Acids 35 2024 102192
4 Su W. Xie X.Q. Liu X.W. Gao D. Ma C.Y. Zulfiqar H. Yang H. Lin H. Yu X.L. Li Y.W. iRNA-ac4C: A novel computational method for effectively detecting N4-acetylcytidine sites in human mRNA Int. J. Biol. Macromol. 227 2023 1174 1181 36470433
5 Khamparia A. Singh K.M. A systematic review on deep learning architectures and applications Expet Syst. 36 2019 e12400
