
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae200
bbae200
Problem Solving Protocol
AcademicSubjects/SCI01060
TransAC4C—a novel interpretable architecture for multi-species identification of N4-acetylcytidine sites in RNA with single-base resolution
Liu Ruijie Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, 430000, China

Zhang Yuanpeng Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, 430000, China

Wang Qi Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, 430000, China

Zhang Xiaoping Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, 430000, China
Shenzhen Huazhong University of Science and Technology Research Institute, Shenzhen, 518000, China

Corresponding authors. Xiaoping Zhang, Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, No. 1277 Jiefang Avenue, Wuhan, 430000, China. E-mail: xzhang@hust.edu.cn; Qi Wang, Department of Urology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, No. 1277 Jiefang Avenue, Wuhan, 43000, China. Tel.: 027-85726114; E-mail: q_wang@hust.edu.cn
5 2024
02 5 2024
02 5 2024
25 3 bbae20008 11 2023
08 4 2024
15 4 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

N4-acetylcytidine (ac4C) is a modification found in ribonucleic acid (RNA) related to diseases. Expensive and labor-intensive methods hindered the exploration of ac4C mechanisms and the development of specific anti-ac4C drugs. Therefore, an advanced prediction model for ac4C in RNA is urgently needed. Despite the construction of various prediction models, several limitations exist: (1) insufficient resolution at base level for ac4C sites; (2) lack of information on species other than Homo sapiens; (3) lack of information on RNA other than mRNA; and (4) lack of interpretation for each prediction. In light of these limitations, we have reconstructed the previous benchmark dataset and introduced a new dataset including balanced RNA sequences from multiple species and RNA types, while also providing base-level resolution for ac4C sites. Additionally, we have proposed a novel transformer-based architecture and pipeline for predicting ac4C sites, allowing for highly accurate predictions, visually interpretable results and no restrictions on the length of input RNA sequences. Statistically, our work has improved the accuracy of predicting specific ac4C sites in multiple species from less than 40% to around 85%, achieving a high AUC > 0.9. These results significantly surpass the performance of all existing models.

N4-acetylcytidine
RNA
Transformer
deep learning
multi-species
National Key Scientific Instrument Development Project 81927807 Wuhan Science and Technology Plan Application Foundation Frontier Project 2020020601012247 Science, Technology and Innovation Commission of Shenzhen Municipality 10.13039/501100010877 JCYJ20190809102415054 Shenzhen Medical Research Funds B2302054
==== Body
pmcINTRODUCTION

Ribonucleic acid (RNA) modification is a frequently observed biological phenomenon in RNA sequences, encompassing m6A, m7G and m5C [1]. Within this array of RNA modifications, N4-acetylcytidine (ac4C) has been detected in rRNA and tRNA, and recent findings have revealed the presence of ac4C sites in mRNA [2–4]. The ac4C modifications have been demonstrated to play a role in gene expression regulation, the efficiency of the translation process and RNA stability [5, 6]. Furthermore, it is widely acknowledged that ac4C modifications of RNA are closely linked to cancer, neural disorders and dysfunction in the cardiovascular system [7–9]. Therapeutic approaches targeting ac4C mediators such as NAT10 have also undergone experimental validation. By inhibiting NAT10 using Remodelin, Ma et al. [10] demonstrated a significant reduction in the metastasis of hepatocellular carcinoma cells, and Zi et al. [11] likewise observed a greater likelihood of apoptosis in cancer cells in acute myelogenous leukemia. However, a global inhibition of ac4C via its mediator NAT10 necessitates a high dosage, thereby posing a substantial risk of side effects. Consequently, it is of utmost importance to accurately identify ac4C sites across RNA to facilitate the recognition of disease-related mechanisms and enable further advancements in drug discovery.

One of the traditional most widely used experimental methods to identify ac4C sites is acetylated RNA immunoprecipitation (acRIP-seq). This method involves co-incubating acetylated RNA-specific antibodies with randomly interrupted RNA fragments and subsequently selecting the fragments with acetylation modifications for sequencing [6]. Although acRIP-seq can indicate the level of acetylation in specific regions of the target sequence, it cannot identify ac4C sites at the base-level resolution. In recent times, a more complex method known as N4-acetylcytidine sequencing (ac4C-seq) has been developed to achieve the identification of ac4C sites at the single base level [12]. However, this method’s sensitivity in identifying ac4C sites relies heavily on a high sequencing depth. The expensive and labor-intensive methods of identifying ac4C sites have hindered progress in understanding the molecular mechanisms and human diseases associated with RNA ac4C. Consequently, it is crucial to be able to predict ac4C sites on RNA, as it can enhance the effectiveness of identifying these sites.

The self-attention mechanism has been proposed to comprehend the relationships in sequence. This mechanism has proven to achieve commendable performance in the domains of text classification and translations as the model excels in capturing the connections between the segments of the sequence while demonstrating proficient interpretative capabilities by illustrating the associations among individual tokens [13, 14]. Among the various models that utilize the self-attention mechanism, the Transformer stands out as one of the most important ones. It has served as the foundation for constructing numerous large language models [15, 16]. Nevertheless, the Transformer model often encounters difficulties in small-scale settings where overfitting is a concern [17–19]. Additionally, it does not perform optimally in specific downstream tasks that involve sequential analysis [19]. Consequently, it is imperative to meticulously design Transformer-based structures.

With the aid of artificial intelligence, certain models have been developed in order to forecast ac4C sites. Initially, there were the PACES and XG-ac4C models, which employed a range of biophysical characteristics and traditional machine learning techniques, yet they only achieved a sensitivity level below 60% [20, 21]. Subsequently, DeepAc4C implemented the convolutional neural network (CNN), along with the physicochemical patterns of the sequence and the distributed representation information of the sequence, to construct the prediction model [22]. Additionally, two models that incorporate the attention mechanism, namely, EMDL-ac4C and LSA-ac4C, have been constructed in parallel [23, 24]. They achieves better performance than former models needing multiple human-defined sequence features.

Although these models achieve a comparably good level of performance, there are still six unresolved problems. (1) All the models were constructed using a dataset that focused solely on the length of mRNA ranging from 400 to 415 nt, thus lacking information on shorter or longer mRNA sequences. (2) Furthermore, all of these models were built on a dataset that was verified using acRIP-seq, which does not provide information on ac4C sites at the single-base level. (3) Another limitation is that these models exclusively consider ac4C sites of mRNA in Homo sapiens, neglecting any potential modifications in other species. (4) It is worth noting that ac4c modifications occur in various types of RNA, yet the aforementioned models solely focus on mRNA. (5) Additionally, while XG-ac4C and DeepAc4C models consider the interpretation of the biophysical characteristics that constitute the features used for prediction, they lack a positional base-level interpretation. (6) Finally, there is still room for improvement in the predictive ability of these models.

To address these issues, we have undertaken the task of reconstructing the previous dataset on acRIP-seq in order to establish a more uniform distribution of RNA length ranging from 215 to 415 nt. Additionally, we have constructed a novel dataset from the data verified by the ac4C-seq, which encompasses three distinct species: H. sapiens, yeast and archaea. In this dataset, the targeted sequence is set at 21 nt. Notably, the new dataset incorporates various RNA types, including mRNA, tRNA, rRNA, ncRNA, RNase P RNA, SRP RNA and snoRNA, thereby facilitating a comprehensive analysis of ac4C on RNA.

Due to the benefit of the processing relationships between words in the sequence, the powerful, interpretable and simple module—transformer encoder—was introduced, thereby also showcasing promise for the interpretation of prediction results. Moreover, as the transformer encoder may struggle in specific tasks without numerous data, we also have carefully incorporated convolutional neural networks and Bi-directional Long Short-Term Memory networks to accurately predict the ac4C sites for both the acRIP-seq data and the ac4C-seq data, referred to as transAC4C-415 nt and transAC4C-21nt, respectively. Importantly, these models enable a positional base-level interpretation of the model based on the corresponding attention weight. Furthermore, by integrating these two models, we have proposed a new pipeline for ac4C, which enables the analysis of any available RNA sequences, thereby enhancing the feasibility of ac4C investigations (Figure 1).

Figure 1 The scheme of the article. (A) The datasets used in this paper are verified by two methods: acRIP-seq and ac4C-seq. The former contains only information on H. sapiens, while the latter contains information on H. sapiens, yeast and archaea. (B) The structure of the model used in this paper: the input was split into 3mer words and embedded. After that, the matrix is processed by the transformer encoder, one Bi-LSTM, four Conv1d and three FC layers. (C) In this paper, the prediction performance and the interpretation of predictions were evaluated. (D) The pipeline in this paper; given a batch of sequences, our model can make exact predictions of every site, determining the overall extent of the sequence undergoing ac4C and giving the following interpretation.

MATERIALS AND METHODS

Data collection

The utilized sequences in this manuscript are the corresponding transcriptional DNA sequences within the genome, thus encompassing four bases, namely, A (adenine), T (thymine), C (cytosine) and G (guanine).

The human ac4C acRIP-seq data utilized in this investigation were initially obtained from the experiments performed by Arango et al. [6]. Zhao et al. selected the original data based on the criterion that sequences with a minimum of five CXX motif repeats in the ac4C peaks were considered positive samples, while sequences with a minimum of five CXX motif repeats in the non-ac4C peaks were considered negative samples. Wang et al. then processed this data using CD-HIT with a threshold of 0.4. This processing yielded 1148 positive samples and 5439 negative samples for the training dataset, as well as 467 positive samples and 2151 negative samples for the independent test dataset [20, 22]. To create balanced sets, Wang constructed 10 training sets and 10 testing sets with ratios of 1148:1148 and 467:467, respectively. Additionally, ensemble learning was employed by integrating 10 models from the test dataset.

Based on Wang’s data, a sequence truncation technique was applied as dividing a longer sequence into two slices with at least five consecutive CXX motifs (Figure 2A). From Wang’s files, a total of 4930 unique negative sequences and 1148 positive sequences were collected for the 10 train sets. Following this, any negative sequences longer than 400 nt were truncated to fall within the range of [300,400]. Furthermore, 20% of the newly truncated sequences were randomly selected and retained. Similarly, negative sequences longer than 300 nt were truncated to fall within the [300,215] range, and 15% of the resulting truncated sequences were randomly selected and retained. Additionally, 25% of negative sequences longer than 400 were sampled and kept to balance the distribution of sequence length. In the case of positive sequences within the train set, they underwent the same procedure, with a ratio of 100% and 25% for each step. Consequently, we have successfully achieved a balanced train set, ensuring an equal number of positive and negative sequences. Furthermore, the length distribution of the sequences has also been balanced. As a result, there is no need to utilize computational ensemble learning, which is difficult to interpret and computational-costly (Supplementary Table 1).

Figure 2 The evaluation of transAC4C-415 nt on the test set. (A) Scheme of balancing the distribution of sequence length: the long sequence is divided into two new sequences with at least five positive CXX motifs and then sampled. (B) The average performance on the 10 test sets. (C) The AUC for transAC4C-415 nt in the 10 test sets. (D) The AUC and ACC of trans-415 nt by combining K bases into a word, k = 1, 2, 3, 4.

For the 10 balanced test sets, negative sequences longer than 400 are truncated into the [300,400] range, and 100% of the new truncated sequences were randomly kept. Subsequently, negative sequences longer than 300 are truncated into the [300,215] range, and 25% of the new truncated sequences were randomly kept. The same procedure is applied to positive sequences in the train set. As a result, we have successfully obtained 10 sets of balanced tests, which allows for a fair comparison with existing models (Supplementary Table 1).

For the ac4C-seq data, which attains a resolution of a single base of ac4C, the sources of Sas-Chen et al. and Liu et al. are utilized [12, 25, 26]. These sources encompass sequences that possess cytidine at the core and are 21 nt in length. The former encompasses various species across three major classifications, namely, H. sapiens, yeast (Saccharomyces cerevisiae) and archaea (Thermococcus kodakarensis, Pyrococcus furiosus, Thermococcus sp. AM4 and Saccharolobus solfataricus). It comprises mRNA, tRNA, rRNA, ncRNA, RNase P RNA, SRP RNA and snoRNA. On the other hand, the latter concentrates solely on H. sapiens. The data from the latter contain two kinds of sequences: one from WT hESCs, another from the NAT-10 knockdown hESCs. We use all WT sites excluding any site presented in the NAT10-KD data to obtain a high-confident set. As a result, the data from these two papers are consolidated as positive sequences. Conversely, negative sequences are derived from the aforementioned acRIP-seq, whereby cytidine is situated at the core and is also 21 nt in length. Furthermore, for the positive sequences of yeast and archaea, they all possess a CCG at the core. Consequently, the corresponding negative sequences in their train sets and test sets also contain CCG at the core to prevent the data leakage problem. Consequently, three train sets were constructed for H. sapiens, yeast and archaea, while there are six test sets for the individual species. The ratio of the number of the train set to the test set is 8:2 (Supplementary Table 2).

For the construction of the unbalanced datasets, as there are 10 independent test sets for transAC4C-415 nt, we retain the positive sequences and combine all the unique negative sequences in the 10 datasets, constructing a dataset (pos to neg = 1:7.5). By evaluating the dataset, we find the ratio of negative sequences in the following region ([215,300], (300,350], (350,400], (400, 415]) is approximately 2.8:2.4:2.5:1. Therefore, we generate sequences in the (400, 415] region by dividing the sequence into two pieces but still in (400, 415]. Taking them together, the imbalanced dataset with pos to neg = 1:10 and the ratio of sequences in the following region ([215,300], (300,350], (350,400], (400, 415]) approximately as 0.93:0.8:0.84:1 is now constructed. The imbalanced dataset for transAC4C-21nt was constructed by retaining the positive sequences of the test set and generating negative sequences from sources of sequences previously identified by acRIP-seq as not undergoing ac4C modification with cytidine at the center for human and CCG at center for species other than human as most of positive sequences for other species have CCG in the middle. The length is set at 21 bp and is checked without intersection with previous sequences, and the ratio of pos to neg is also 1:10 in the imbalanced dataset.

The architecture of transAC4C

The transAC4C consists of four parts: the transformer encoder, one Bidirectional Long Short-Term Memory (Bilstm) layer, four 1D Convolutional (Conv1d) layers and three fully connected (FC) Layers. The rationale for this architecture is that the transformer encoder and the Bi-LSTM enable the model to learn the contextual information well, and the Conv1d layers enable the model to extract important information from the input, thus leading to the model’s universality. The FC layers were used to connect the features of the input to the final output, the classification in this paper.

The transformer encoder

The initial proposal for the transformer encoder was put forward by Vaswani et al. [27]. This novel architecture brings about a revolution in tasks based on sequences through the utilization of two key components: Multi-Head Self-Attention and Position-wise Fully Connected Feed-Forward networks. The former employs parallel processing, thus facilitating simultaneous focus on various segments of input to capture context. The latter carries out an independent analysis of attention outputs, thereby identifying intricate patterns within the data. Both of these modules incorporate layer normalization and residual connections as a means of preventing the vanishing gradient problem. The process of calculating the attention weight, as described in the paper, is illustrated as follows:

(1) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}\left\{\begin{array}{c}Q=X{W}^Q\\ {}K=X{W}^K\\ {}V=X{W}^V\end{array}\right.\end{align*}\end{document}

(2) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}attention = softmax\left(\frac{Q{K}^T}{\sqrt{d_k}}\right)V\end{align*}\end{document}

where X refers the input matrix and WQ, WK and WV refer to the three transformation layers in the module. Three matrices Q(Query), K(Key) and V(Value) were first calculated as formula (1). Then, the output of this module, the attention matrix, was calculated as formula (2), where dK refers to the dimension of K and the attention weight is \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $softmax\left(\frac{Q{K}^T}{\sqrt{d_k}}\right)\ with\ three\ dimensions$\end{document}: [Batch size, input length, input length], and to visualize it by sequence length, they are all averaged by its second dimension as [Batch size, input length].

The Bi-LSTM module

Bi-LSTM embodies a noteworthy progression in the design of recurrent neural networks (RNNs) [28–30]. It excels in capturing intricate patterns within sequential data by simultaneously processing input sequences in both forward and backward directions. This distinctive bidirectional approach empowers the model to comprehensively comprehend the context of each time step, extracting insights from both preceding and subsequent information. Unlike conventional LSTMs, the invaluable capability of Bi-LSTM to grasp long-term dependencies renders it highly applicable in diverse domains, particularly in natural language processing tasks such as machine translation and sentiment analysis [31, 32]. The aptitude of Bi-LSTM to comprehend the sequential nature of data endows it with immense power as a tool in the realm of deep learning.

The Conv1d module

Conv1d serves as a fundamental component in the realm of deep learning to manipulate sequential data [33]. Its function entails the application of a sliding mechanism, which has been acquired through learning, over the input data to extract local patterns and features. In the present study, each Conv1d layer is accompanied by an activation layer, a normalization layer and a dropout layer, with the dropout layer having a ratio of 0.5 for transAC4C-415 nt and 0.1 for transAC4C-21 nt [34–36]. The strides and padding values for all Conv1d layers are 3 and 1, respectively. Furthermore, the kernel size varies across different Conv1d layers, with values of 7, 3, 3 and 3, while the in/out channels are represented by (128,64), (64,32), (32,16) and (16,32). Lastly, a max pooling layer is implemented after the final Conv1d layer [37].

The FC module

The FC layers, otherwise known as linear layers, perform linear calculations using computing neurons [38]. To prevent sudden losses of computing neurons, three FC layers were utilized. The respective in/out channels for transAC4C-415 nt are (1600,1024), (1024,64) and (64,2). For transAC4C-21nt, however, only one FC layer was applied. This FC layer serves as the final output of Conv1d, which is flattened with a dimension of 24. Consequently, the in/out channels for transAC4C-21 nt are (24,2). The output of the FC layers is then processed through the Softmax layer to achieve the classification ability.

The training schedule

The loss function used in this paper is termed cross entropy and the Adam optimizer with weight decay at 1e-2 was used [39, 40]. The learning rate was set at 5e-6, and the batch size was set at 64. Ten percent of the train set was randomly split as the validation set. The epoch was set at 100 and to prevent overfitting, the early-stopping strategy with patience equal to five was applied, which means that during the training process, if the average accuracy of the validation set did not improve consecutively in five epochs, the training process would stop and the corresponding model would be kept as the final model [41].

In silico mutagenesis

To prove that the rationale as attention weights is explainable in our model, in silico mutagenesis was conducted. For every base in the targets sequence, it would be replaced to other three types, and the greatest difference in the score is recorded as the importance the base contributes to the prediction.

The evaluation metrics

Six common judgments have been tried in this study including accuracy (ACC), sensitivity (SN), specificity (SP), area under the receiver operating characteristic curve (AUC), area under the precision–recall curve (AUPR) and the Matthews correlation coefficient (MCC). They are calculated as:

(3) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}ACC=\frac{TP+ TN}{TP+ TN+ FP+ FN}\end{align*}\end{document}

(4) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}Sn=\frac{TP}{TP+ FN}\end{align*}\end{document}

(5) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}SP=\frac{TN}{TN+ FP}\end{align*}\end{document}

(6) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}MCC=\frac{TP\times TN- FP\times FN}{\sqrt{\left( TP+ FP\right)\times \left( TP+ FN\right)\times \left( TN+ FN\right)\times \left( FN+ FP\right)}}\end{align*}\end{document}

(7) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}AUC={\int}_0^1 TPR\left({FPR}^{-1}(x)\right) dx\end{align*}\end{document}

(8) \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} \begin{align*}AUPR={\int}_0^1 Precision\left({Recall}^{-1}(x)\right) dx\end{align*}\end{document}

where TP, TN, FP and FN refer to true positive, true negative, false positive and false negative, respectively. The AUC was calculated as integral of TPR under every FPR by x. The x is the threshold of determining samples as positive or negative. Recall is the same as the TPR, which means the true-positive rate, and the precision is the ratio of true positives out of all that have been predicted as positives.

RESULTS

Construction of transAC4C-415 nt model on the datasets by acRIP-seq

Former models were constructed based on the datasets by acRIP-seq, where over 80% of the sequences are above 400 nt. Therefore, their datasets and models shed rare light on shorter mRNA, which is common in nature. Due to this limitation, a sequence truncation technique was applied by dividing a longer sequence into two slices with at least five consecutive CXX motifs as the previous study considered sequences with at least five CXX motifs in the ac4C peak area as the sequence undergoing ac4C, and they are sampled to keep the dataset balanced (Figure 2A) [20]. As a result, the acRIP-seq datasets presented in this paper exhibit a more balanced distribution of mRNA, both in terms of their positive and negative status, as well as their varying lengths.

Subsequently, we developed a prediction model for identifying ac4C sites in mRNA. The model consists of a transformer encoder, one Bi-LSTM layer, four 1D convolutional layers and three fully connected layers (Figure 1B). This model, named transAC4C-415 nt, is designed to handle RNA sequences of up to 415 nt in length. The model demonstrates strong performance, with an AUC value exceeding 0.85 across the training, validation and test sets (Figure 2B and C, Supplementary Table 3). To assess the optimality of this model for the given dataset, ablation experiments were conducted. The results indicate that reducing the Bi-LSTM and transformer encoder significantly diminishes the prediction ability of transAC4C-415 nt on the training set, suggesting an underfitting scenario (Table 1). Conversely, reducing the number of Conv1d layers leads to a decrease in the prediction ability of transAC4C-415 nt on the validation and test sets, indicating an overfitting situation (Table 1). Thus, the best accuracy (ACC) and AUC values on the validation and test sets can only be achieved when using the transAC4C-415 nt model with one Bi-LSTM layer, four Conv1d layers, the transformer encoder and three fully connected layers in combination (Table 1). Moreover, the former model incorporates a channel-mechanism module called SENet, which was originally used for computer vision tasks [42]. To further prove the importance of the transformer encoder in our architecture, the replacement experiment has been conducted. The result shows that replacing the transformer encoder with SENet greatly decreases the performance in the 10 test sets, with 3.48% averaged AUC loss and 3.71% averaged ACC loss, thus further demonstrating the superiority of the transformer in sequence analysis (Supplementary Figure 1). Additionally, the optimal choice of word length, measured in bases, was also considered. The results demonstrate that using three bases as a word yields the highest accuracy and AUC values, which is consistent with the fact that a codon consists of three bases (Figure 2D).

Table 1 The result of the ablation experiment for the transAC4C-415 nt

Ablation	trainACC	trainAUC	validACC	validAUC	testACC	testAUC	
transAC4C	0.7856	0.8699	0.7883	0.8576	0.7797	0.8537	
No CNN	0.8779	0.9456	0.7721	0.8432	0.7371	0.8110	
CNN*1	0.9275	0.9756	0.7796	0.8540	0.7411	0.8101	
CNN*2	0.8029	0.8773	0.7634	0.8383	0.7594	0.8298	
CNN*3	0.8060	0.8815	0.7597	0.8373	0.7594	0.8325	
No Bi-LSTM	0.7632	0.8380	0.7409	0.8169	0.7608	0.8294	
SLSTM	0.7584	0.8357	0.7634	0.8241	0.7557	0.8310	
No Transformer Encoder	0.7650	0.8319	0.7385	0.8072	0.7554	0.8245	
FC*2	0.7285	0.7993	0.7123	0.7837	0.7141	0.7955	
FC*1	0.7255	0.8093	0.7061	0.7893	0.7156	0.7963	
The best scores are shown in bold.

Comparison of transAC4C-415 nt with the state-of-the-art model

Subsequently, the evaluation of transAC4C-415 nt was conducted to demonstrate that our model achieved the highest level of performance compared to the state-of-the-art (SOTA) models. Three SOTA models, namely, DeepAc4C, EMDL-ac4C and LSA-ac4C, were taken into consideration. These models were constructed using the same dataset as our model, and both DeepAc4C and EMDL-ac4C utilized the same criterion as transAC4C, which required a minimum of five CXX motifs in the ac4C peak region as the sequence undergoing ac4C, and could accept mRNA shorter than 416 nt as input. However, LSA-ac4C assumed that the cytidine closest to the acetylation peak was the cytidine undergoing ac4C and only accepted mRNA sequences of 201 nt with cytidine at the center. Therefore, we compared the performance of DeepAc4C and EMDL-ac4C on our test sets, as well as the performance of DeepAc4C, EMDL-ac4C and LSA-ac4C on three independent sets that were verified by ac4C-seq, which can achieve single-base resolution of ac4C.

Compared to DeepAc4C and EMDL-ac4C, our method achieved consistently much better ACC, AUC and MCC (Table 2). To prevent the problem that truncation of sequence can lead to false classifications, the performances on the original datasets, which were used inDeepAc4C and EMDL-ac4C were also verified and the performances were close (Supplementary Figure 2).

Table 2 The comparison of performance on the test sets between the transAC4C-415 nt to the existing SOTA models

Comparison with the SOTA model	Independent test set from acRIP-seq	Archaea	Human	
ACC	AUC	MCC	ACC	ACC	
transAC4C-415 nt	0.7797	0.8537	0.5599	0.5092	0.4095	
EMDL-ac4C	0.736	0.8259	0.4957	0.3148	0.0109	
DeepAc4C	0.7472	0.82136	0.53587	0.3796	0.1055	
LSA-ac4C	–	–	–	0.1522	0.3242	
The best scores are shown in bold.

Since single-base identification of ac4C is of utmost importance, the models’ performance on the verified single-base resolution of ac4C by ac4C-seq was subsequently examined. Despite not being specifically trained to identify ac4C at a single base, transAC4C-415 nt still outperforms all the models by a significant margin (Table 2). This demonstrates that our model exhibits the highest level of performance and indicates that only our model can properly learn the data.

Construction of transAC4C-21 nt model on the datasets by ac4C-seq

Although our previous model for transAC4C-415 nt has demonstrated superior performance, it is still limited in three aspects: the exclusion of other species, the omission of other types of RNA and the lack of exact prediction of ac4C sites. To address these limitations, a new model, named transAC4C-21 nt, was developed. This new model was trained on newly verified datasets obtained from ac4C-seq, which encompass three major species: H. sapiens, yeast (S. cerevisiae) and archaea (T. kodakarensis, P. furiosus, Thermococcus sp. AM4 and S. solfataricus). The datasets include various types of RNA such as mRNA, tRNA, rRNA, ncRNA, RNase P RNA, SRP RNA and snoRNA. Furthermore, the input sequence for the transAC4C-21 nt model is set at 21 nt with cytidine positioned at the center. This design ensures that most RNA molecules or fragments can be analyzed with single-base resolution.

Subsequently, the performance of the transAC4C-21 nt model was evaluated using an independent test set with cross-validation and the test set includes species such as H. sapiens, T. kodakarensis, P. furiosus, Thermococcus sp. AM4, S. cerevisiae and S. solfataricus. As the dataset is limited in size, cross-validation was used. The transAC4C-21 nt exhibits stably good performance both in 5-fold and 10-fold experiments (Table 3). To enable the more powerful and generalized prediction of ac4C sites in RNA, ensemble learning with soft voting was applied. The transAC4C-21 nt model achieved the best performance by integrating the 10 models in the 10-fold cross-validation (Table 3). For the independent test set, it achieves an AUC of approximately 0.9 and an ACC of 0.8385, 0.8086, 0.8298, 0.8241, 0.8333 and 0.9117, respectively. These results demonstrate the strong predictive capability of the transAC4C-21 nt model (Figures 3 and 4). To further prove the idea that Bi-LSTM promotes the learning of the model and Conv1d prevents overfitting in the test set, ablation experiments were conducted. Based on the averaged performance on the individual test sets by the ensemble learning, we found that the result is similar to the result of ablation experiments in transAC4C-415 nt. Reducing the Bi-LSTM and transformer encoder leads to lower predictive performance on the training set, and reducing the number of Conv1d layers leads to lower predictive performance on the validation and test sets, therefore, demonstrating our model structures are ideal (Table 4).

Table 3 The result of the cross-validation experiments for the transAC4C-21 nt

	10-fold cross-validation	5-fold cross-validation	
TransAC4C-21 nt	Valid AUC	Test AUC
no ensemble	Test AUC ensemble	Valid AUC	Test AUC
no ensemble	Test AUC ensemble	
S. cerevisiae	0.9452	0.9066	0.9308	0.9215	0.9045	0.9308	
H. sapiens	0.8743	0.8870	0.9035	0.8543	0.8829	0.8968	
Thermococcus sp. AM4	0.8869	0.8728	0.8875	0.8726	0.8598	0.8836	
T. kodakarensis	0.8869	0.8761	0.8929	0.8726	0.8707	0.8834	
P. furiosus	0.8869	0.9090	0.9244	0.8726	0.9029	0.9212	
S. cerevisiae	0.8869	0.8827	0.9012	0.8726	0.8765	0.8642	

Figure 3 The AUC of transAC4C-21 nt on the test set, respectively. Three transAC4C-21 nt were trained on H. sapiens, yeast (S. cerevisiae) and archaea (T. kodakarensis, P. furiosus, Thermococcus sp. AM4 and S. solfataricus), and they are tested on its test set, respectively.

Figure 4 The performance of transAC4C-21 nt on the test set, respectively. Three transAC4C-21 nt were trained on H. sapiens, yeast (S. cerevisiae) and archaea (T. kodakarensis, P. furiosus, Thermococcus sp. AM4 and S. solfataricus), and they are tested on its test set, respectively.

Table 4 The result of the ablation experiment for the transAC4C-21 nt

Ablation experiments	trainACC	trainAUC	validACC	validAUC	testACC	testAUC	
transAC4C-21 nt	0.9231	0.9775	0.851	0.912	0.841	0.9067	
No CNN	0.8721	0.9387	0.8163	0.8754	0.8331	0.9013	
CNN*1	0.9569	0.9893	0.8271	0.8854	0.8323	0.8978	
CNN*2	0.9243	0.9778	0.8323	0.8912	0.8385	0.9028	
CNN*3	0.9055	0.9689	0.8478	0.9027	0.838	0.9027	
No Bi-LSTM	0.9164	0.9658	0.8401	0.8912	0.833	0.8925	
SLSTM	0.9015	0.9643	0.8351	0.8834	0.8281	0.8812	
No Transformer	0.9011	0.9531	0.8473	0.9037	0.8324	0.9025	
The best scores are shown in bold.

Performance of transAC4C-21 nt model cross-species and cross-RNA types

Moreover, the per RNA category performance of transAC4C-21 nt was evaluated to prove its ability to predict ac4C sites across RNA types. The result shows that the model trained on H. sapiens and yeast both achieves AUC > 0.87 and ACC > 0.83 cross RNA types (Figure 5A). The model trained on archaea works best on tRNA and achieves AUC > 0.85 across RNA types despite a marginal loss in mRNA. Furthermore, as our dataset lacks information on other important species, such as Mus musculus and Drosophila melanogaster, the ability for the transAC4C-21 nt to cross-species predict ac4C sites is needed. However, the evaluation of the results on a wider range of species revealed that the transAC4C-21 nt model lost performance with AUC values ranging from 0.54 to 0.88, which can only achieve moderate prediction ability (Figure 5B). As the datasets of three species have different compositions to RNA types, the ability to predict ac4C sites cross-species in per RNA category was further explored. The result shows that the model trained on H. sapiens works well with an AUC > 0.78 in mRNA types across species that contain positive mRNA samples and the model trained on archaea works well with an AUC >0.88 in tRNA across species (Figure 5C and D). With the commendable performance observed within each distinct species, the transAC4C-21 nt demonstrates suitability for ac4C prediction at specific sites for humans, archaea and yeast in individual RNA types. TransAC4C-21 nt also potentially possesses the capability to identify specific sites in both mRNA and tRNA across various species.

Figure 5 Performance of transAC4C-21 nt model cross RNA types and cross-species. Three transAC4C-21 nt were trained on H. sapiens, yeast (S. cerevisiae) and archaea (T. kodakarensis, P. furiosus, Thermococcus sp. AM4 and S. solfataricus). (A) The performance of the model per RNA category. (B) The overall performance (AUC) of the model (X axis) on the corresponding test set (Y axis) is recorded. (C) The performance (AUC) of the model (X axis) on the tRNA of the corresponding test set (Y axis) is recorded. (D) The overall performance (AUC) of the model (X axis) on the mRNA of the corresponding test set (Y axis) is recorded. NA means no positive samples in the corresponding species.

Positional base-level interpretation of the transAC4C model

Benefitting from the transformer encoder, the model assigned an attention weight to each word comprising three bases, with the word’s contribution to the prediction being more crucial as the weight values increase. Additionally, each word encompasses information on three bases.

In the case of transAC4C-415, the model’s evaluation of the ac4C sequence that was correctly classified was subsequently conducted. The result demonstrated that the model consistently assigned a low attention weight to the padding sequence, indicating that the model effectively learned the data (Figure 6A). Furthermore, the plot illustrates that the peak of attention weight not only emerges in the region with a minimum of five CXX repeats. Consequently, we computed the top five words for each correctly classified ac4C sequence in the test set, and the top five words are GCG, GGG, CAG, AGC and CTG. This implies that the three bases with high GC content significantly contribute to the ac4C of RNA.

Figure 6 The interpretation of the model transAC4C-415 nt and transAC4C-21 nt. (A) The interpretation of transAC4C-415 nt on one of the correctly classified positive samples. (B–D) The interpretation of transAC4C-21 nt on one of the correctly classified positive samples from human, archaea and yeast. (E) The comparison of the percentages of top bases of the targeted sequence overlapped in the top 3-mer in the targeted sequence between the two strategies (attention weights and in silico mutagenesis) and random choosing. Top3 bases and top1 3-mer for transAC4C-21 nt and top15 bases and top5 3-mer for transAC4C-415 nt were analyzed. (F) The comparison of the percentages of top bases of the targeted sequence overlapped in the top 3-mer in the targeted sequence between the two strategies (attention weights and in silico mutagenesis) and random choosing. Top6 bases and top2 3-mer for transAC4C-21 nt and top15 bases and top5 3-mer for transAC4C-415 nt were analyzed. The paired t-test was used. **: P < 0.01, ****: P < 0.0001

In the case of transAC4C-21 nt, the model’s evaluation of the correctly classified ac4C sequence was also performed. The interpretation of the transAC4C-21 nt was performed similarly to transAC4C-21 nt (Figure 6B–D). For the four species of archaea, GGT was consistently interpreted as important. For humans and yeast, CCG and AXA were interpreted as important (Supplementary Table 4). Moreover, sequences on the 5’ direction from the cytidine undergoing ac4C also have more attention weights, therefore, contributing more to the model. To further verify the interpretation result by the attention weights, in silico mutagenesis was applied, which iteratively changes a single base of the targeted sequence to get the positional importance of the base. We observe overlaps in four identical samples of peaks in the per-prediction positional importance results by these two methods (Supplementary Figure 3). The statistical result shows that the result by attention weight significantly greatly overlaps with random situations (Figure 6E and F), which further demonstrates that using attention weights as the interpretation is reasonable.

The pipeline for evaluating RNA sequences undergoing ac4C by transAC4C

Based on the transAC4C-21 nt, our method can identify underlying ac4C sites in target RNA by iteratively dividing RNA into 21 nt slices. Based on the padding mechanism, transAC4C-21 nt can also accept slices shorter than 21 nt as long as the cytidine is in the center (Figure 7A). Furthermore, for researchers who are solely interested in determining which RNA is undergoing ac4C, the transAC4C-415 nt methodology can predict this as long as the mRNA is not longer than 415 nt (Figure 7B).

Figure 7 The pipeline evaluating the potential of sequences undergoing ac4C. (A) For those interested in which site of the sequence is acylated, the target sequence is sliced and predicted by transAC4C-21 nt. (B) For those interested in whether the whole sequence has ac4C sites, if the sequence is no longer than 415 nt, the whole sequence is predicted by ransAC4C-415 nt. (C, D) For situations interested in whether the whole sequence would undergo the ac4C modification, if the sequence is longer than 415nt, the sequence would be sliced and evaluated by transAC4C-21 nt (C) and transAC4C-415n (D) and assigned with an ac4C score defined as the average probability of every slice.

Moreover, for researchers who are interested in ac4C RNA longer than 415 nt, the ac4C score of each sequence can be determined by utilizing both the transAC4C-21 nt and transAC4C-415 nt model. This is achieved by calculating the probabilities of each RNA slice undergoing ac4Cs and then averaging them (Figure 7C). To compare the ac4C score obtained from these two models, we calculated the ac4C score for the mRNA of all the genes in the KEGG fatty acid metabolism pathway. It was recently discovered that only the ELOVL6, ACSL1, ACSL3, ACSL4, ACADSB and ACAT1 genes exhibit high ac4C modifications, and the mRNA of genes in the KEGG fatty acid metabolism pathway is longer than 415 nt [43]. Therefore, the mRNA of these six genes was considered as identified ac4C mRNA, while the mRNA of the remaining 51 genes was considered as not-identified ac4C mRNA. This allowed us to construct a new test set comprising of long mRNA sequences. The results demonstrate that the identified ac4C mRNA displays a significantly higher ac4C score when calculated using the transAC4C-21 nt model (Figure 8A). However, there is no significant difference in the ac4C score calculated by transAC4C-415 nt, even though it successfully identified that the mRNA of ELOVL6 and ACSL3 contains ac4C slices (Figure 8B). Furthermore, we evaluated the performance of the DeepAc4C model on the mRNA of genes in the KEGG fatty acid metabolism pathway. This model sacrificed the input of positional information by flattening the matrix on the feature dimension rather than the length dimension, allowing for no limitation on the length of input RNA. The results reveal that the DeepAc4C model performed the worst, as it failed to identify any of the mRNA of ELOVL6, ACSL1, ACSL3, ACSL4, ACADSB and ACAT1 as ac4C mRNA (Figure 8C).

Figure 8 The performance of tranAC4C on long sequences. (A) The dot plot of ac4C scores for identified ac4C mRNA and not-identified ac4C mRNA by transAC4C-21 nt. (B) The dot plot of ac4C scores for identified ac4C mRNA and not identified ac4C mRNA by transAC4C-415 nt. (C) The dot plot of ac4C scores for identified ac4C mRNA and not-identified ac4C mRNA byDeepAc4C. (D) The AUC of classifying identified ac4C mRNA and not-identified ac4C mRNA by ac4C scores calculated from the transAC4C-21 nt. The unpaired t-test was used as ns: P > 0.05; **: P < 0.01. (E) The precision–recall curve of the transAC4C-415 and transAC4C-21 nt on the unbalanced test set. (F) The performance of transAC4C-415 nt compared to DeepAc4C and EMDL-ac4C on the unbalanced test set.

To provide a more direct demonstration of the prediction performance of the transAC4C-21 nt model on the genes in the KEGG fatty acid metabolism pathway, we plotted the receiver operating characteristic curve. The results exhibit that by assigning each sequence with the ac4C score calculated using the transAC4C-21 nt methodology, the identified ac4C mRNA and not-identified ac4C mRNA were effectively distinguished, with an AUC greater than 0.93 (Figure 8D). Therefore, it is recommended to calculate the ac4C score using the transAC4C-21 nt model when selecting potential ac4C RNA longer than 415 nt. Additionally, since the lowest ac4C score among ELOVL6, ACSL1, ACSL3, ACSL4, ACADSB and ACAT1 is 0.410, it is suggested to use this value as the cut-off criterion for distinguishing between ac4C RNA longer than 415 nt and non-ac4C RNA longer than 415 nt.

Moreover, to gauge the model’s true performance in the world scenario, performance on the unbalanced datasets with the ratio of positive samples to negative samples as 1:10 was evaluated. The result shows that the model can achieve AUPR of 0.37 for transAC4C-415 nt and AUPR from 0.53 to 0.91 for transAC4C-21 nt (Figure 8E), which is at least three times bigger than the floor of the AUPR in the dataset with the ratio of positive samples to negative samples as 1:10. Moreover, the comparison to the former models on the unbalanced dataset was performed. Despite the fact that EMDL-ac4C has greater MCC, transAC4C-415 nt consistently has greater AUC and AUPR, demonstrating better performance in discriminating positive samples and negative samples (Figure 8F). Therefore, we believed that our trained model has good feasibility and superior performance in a real-world scenario.

DISCUSSION

Our model possesses numerous advantages when compared to the existing SOTA methods (Table 5) [22–24]. The transAC4C comprises two distinct models: transAC4C-21 nt and transAC4C-415 nt. These models were individually trained on the acRIP-seq data and ac4C-seq data. By utilizing the aforementioned pipeline, our transAC4C model enables the utilization of any RNA inputs. This allows for the identification of potential RNA sequences undergoing ac4C or the precise prediction of ac4C sites within target RNA sequences. Additionally, the training data encompass information from various RNA types and species. Consequently, this results in the construction of the most comprehensive model available at present. The careful design of the model ensures that the sequence length of inputs processed by the transformer remains consistent with the sequence length of the initial input matrix. As a result, this facilitates the positional interpretation and base-level interpretation of every prediction (Table 5).

Table 5 The overall comparison between the transAC4C to the existing SOTA models

	EMDL-ac4C	DeepAC4C	LSA-ac4C	transAC4C	
Input sequence	Length ≤ 415 nt	No limitation	Length at 201 nt	No limitation	
RNA types	mRNA	mRNA	mRNA	mRNA, tRNA, rRNA, ncRNA, RNase P RNA and SRP RNA	
Single-base resolution	No	No	Yes, but sensitivity below 40%	Yes, and sensitivity higher than 80%	
Species	H. sapiens	H. sapiens	H. sapiens	H. sapiens, S. cerevisiae,  
T. kodakarensis, P. furiosus, Thermococcus sp. AM4 and S. solfataricus	
Universality	No, failed in other test sets	No, failed in other test sets	No, failed in other test sets	Yes, validated cross different test sets	
Interpretation	No	Importance of the selected biophysical features of all the prediction	No	Positional base-level interpretation of every prediction	

In both of our models, we have observed a significant decrease in performance when removing the transformer encoder. Particularly for input lengths ranging from 215 to 415 nt, the decrease is most pronounced. This suggests that the transformer encoder is particularly crucial for processing longer sequence outputs. Additionally, we have observed that attention weights are widely distributed across positions, which is in line with the transformer’s ability to capture long-range dependencies in sequence analysis. We therefore believe that, in addition to its interpretability, the transformer encoder enhances the understanding of long-range associations in target sequences. This observation highlights the importance of the transformer architecture in biological sequential analysis, as it effectively models and comprehends biological sequential data such as nucleotides and peptides.

The transAC4C aids in the exploration of the nature of ac4C in RNAs through its interpretation of the prediction. Our investigation revealed that the identification of RNA segments heavily relies on the presence of high GC-content words in the transAC4C-415 nt model. Furthermore, in the transAC4C-21 nt model, the occurrence of high GC-content words, particularly the CCG word, persists. Moreover, the transAC4C-21 nt also imparts valuable insights into the positional interpretation. Notably, the model places significant emphasis on the beginning and end of the sequence, suggesting that the nucleotide composition in these positions plays a crucial role in ac4C modifications in RNA.

However, there still exist some limitations in the paper. The number of sequences analyzed in the paper is not big enough due to the fact that the newly proposed technique ac4C-seq relies on sequencing with very high sequencing depth; therefore, the identified ac4C sequences with single-base resolution are not many available now [12]. This also constrains the paper from taking other important species including M. musculus and D. melanogaster. Moreover, the model developed within this study lacks a superior ability to identify ac4C sites across various species, as well as different tissue types, which is the potential for future advancements in ac4C sites prediction. Furthermore, despite the fact that the model attains a light-weighed explanation in terms of prediction through the use of attention weights and in silico mutagenesis, it’s not theoretically proven to have fully coincided with the contribution of the variables. Therefore, more powerful techniques for explanation need to be proposed, which is also an urgent need in the current field.

Key Points

TransAC4C was created using new comprehensive datasets, enabling multi-species ac4C prediction in RNA with single-base resolution.

TransAC4C is a highly effective model that significantly enhances the accuracy of ac4C prediction.

The new pipeline, which is based on TransAC4C, allows for the prediction of ac4C in any RNA sequence without limitations.

TransAC4C is the most interpretable model, providing positional interpretation for each prediction.

Supplementary Material

Supplementary_Figures_bbae200

Supplementary_Table_1_bbae200

Supplementary_Table_2_bbae200

Supplementary_Table_3_bbae200

Supplementary_Table_4_bbae200

Supplementary_Table_5_bbae200

ACKNOWLEDGEMENTS

The computation is completed in the HPC Platform of Huazhong University of Science and Technology.

FUNDING

This study was supported by the National Key Scientific Instrument Development Project (81927807), the Wuhan Science and Technology Plan Application Foundation Frontier Project (2020020601012247), the Science, Technology and Innovation Commission of Shenzhen Municipality (JCYJ20190809102415054) and the Shenzhen Medical Research Funds (B2302054).

DATA AVAILABILITY

The source code and data are available from https://github.com/BioLJ/TransAC4C. The detailed information of data is shown in Supplementary Tables 1 and 2. The detailed results presented in the Results section are shown in Supplementary Tables 3–5.

Author Biographies

Ruijie Liu is an MD candidate at Huazhong University of Science and Technology. His research interest is applications of artificial intelligence into biology and clinical medicine.

Yuanpeng Zhang is an MD candidate at Huazhong University of Science and Technology. His research interests are sequence analysis and mechanism insights in gene networks.

Qi Wang is a post-doctor of Huazhong University of Science and Technology. His research area includes molecular biology and application of bioinformatics.

Xiaoping Zhang is a professor at Huazhong University of Science and Technology. His research area focuses on cancer biology and bioinformatics.
==== Refs
References

1. Qiu  L, Jing  Q, Li  Y, Han  J. RNA modification: mechanisms and therapeutic targets. Mol Biomed  2023;4 (1 ):25.37612540
2. Ito  S, Akamatsu  Y, Noma  A, et al.  A single acetylation of 18 S rRNA is essential for biogenesis of the small ribosomal subunit in Saccharomyces cerevisiae. J Biol Chem  2014;289 (38 ):26201–12.25086048
3. Wei  W, Zhang  S, Han  H, et al.  NAT10-mediated ac4C tRNA modification promotes EGFR mRNA translation and gefitinib resistance in cancer. Cell Rep  2023;42 (7 ):112810.37463108
4. Yang  Z, Wilkinson  E, Cui  YH, et al.  NAT10 regulates the repair of UVB-induced DNA damage and tumorigenicity. Toxicol Appl Pharmacol  2023;477 :116688.37716414
5. Yan  Q, Zhou  J, Wang  Z, et al.  NAT10-dependent N4-acetylcytidine modification mediates PAN RNA stability, KSHV reactivation, and IFI16-related inflammasome activation. Nat Commun  2023;14 (1 ):6327.37816771
6. Arango  D, Sturgill  D, Alhusaini  N, et al.  Acetylation of cytidine in mRNA promotes translation efficiency. Cell  2018;175 (7 ):e1872–1886.e24.
7. Chen  X, Hao  Y, Liu  Y, et al.  NAT10/ac4C/FOXP1 promotes malignant progression and facilitates immunosuppression by reprogramming glycolytic metabolism in cervical cancer. Adv Sci (Weinh)  2023;10 (32 ):e2302705.37818745
8. Wang  C, Hou  X, Guan  Q, et al.  RNA modification in cardiovascular disease: implications for therapeutic interventions. Signal Transduct Target Ther  2023;8 (1 ):412.37884527
9. Luo  J, Cao  J, Chen  C, Xie  H. Emerging role of RNA acetylation modification ac4C in diseases: current advances and future challenges. Biochem Pharmacol  2023;213 :115628.37247745
10. Ma  R, Chen  J, Jiang  S, et al.  Up regulation of NAT10 promotes metastasis of hepatocellular carcinoma cells through epithelial-to-mesenchymal transition. Am J Transl Res  2016;8 (10 ):4215–23.27830005
11. Zi  J, Han  Q, Gu  S, et al.  Targeting NAT10 induces apoptosis associated with enhancing endoplasmic reticulum stress in acute myeloid Leukemia cells. Front Oncol  2020;10 :598107.33425753
12. Thalalla Gamage  S, Sas-Chen  A, Schwartz  S, Meier  JL. Quantitative nucleotide resolution profiling of RNA cytidine acetylation by ac4C-seq. Nat Protoc  2021;16 (4 ):2286–307.33772246
13. Xie  J, Hou  Y, Wang  Y, et al.  Chinese text classification based on attention mechanism and feature-enhanced fusion neural network. Comput Secur  2020;102 (3 ):683–700.
14. Jia  Y . Attention mechanism in machine translation. J Phys Conf Ser  2019;1314 (1 ):012186  IOP Publishing.
15. Floridi  L, Chiriatti  M. GPT-3: its nature, scope, limits, and consequences. Mind Mach  2020;30 (4 ):681–94.
16. Devlin  J, Chang  M-W, Lee  K, et al.  Bert: pre-training of deep bidirectional transformers for language understanding. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Naacl Hlt 2019) 2019;1 :4171–86.
17. Shao  R, Bi  X-J. Transformers meet small datasets. IEEE Access  2022;10 :118454–64.
18. Zhang  W, Vaidya  I. Mixup training leads to reduced overfitting and improved calibration for the transformer architecture, arXiv preprint. arXiv  2021;2102.11402.
19. Zeng  A, Chen  M, Zhang  L, et al.  Are transformers effective for time series forecasting?  Proceedings of the AAAI conference on artificial intelligence  2023;37 :11121–8.
20. Zhao  W, Zhou  Y, Cui  Q, Zhou  Y. PACES: prediction of N4-acetylcytidine (ac4C) modification sites in mRNA. Sci Rep  2019;9 (1 ):11112.31366994
21. Alam  W, Tayara  H, Chong  KT. XG-ac4C: identification of N4-acetylcytidine (ac4C) in mRNA using eXtreme gradient boosting with electron-ion interaction pseudopotentials. Sci Rep  2020;10 (1 ):20942.33262392
22. Wang  C, Ju  Y, Zou  Q, Lin  C. DeepAc4C: a convolutional neural network model with hybrid features composed of physicochemical patterns and distributed representation information for identification of N4-acetylcytidine in mRNA. Bioinformatics  2021;38 (1 ):52–7.34427581
23. Jia  J, Wei  Z, Cao  X. EMDL-ac4C: identifying N4-acetylcytidine based on ensemble two-branch residual connection DenseNet and attention. Front Genet  2023;14 :1232038.37519885
24. Lai  FL, Gao  F. LSA-ac4C: a hybrid neural network incorporating double-layer LSTM and self-attention mechanism for the prediction of N4-acetylcytidine sites in human mRNA. Int J Biol Macromol  2023;253 (Pt 3 ):126837.37709212
25. Liu  R, Wubulikasimu  Z, Cai  R, et al.  NAT10-mediated N4-acetylcytidine mRNA modification regulates self-renewal in human embryonic stem cells. Nucleic Acids Res  2023;51 (16 ):8514–31.37497776
26. Sas-Chen  A, Thomas  JM, Matzov  D, et al.  Dynamic RNA acetylation revealed by quantitative cross-evolutionary mapping. Nature  2020;583 (7817 ):638–43.32555463
27. Vaswani  A, Shazeer  N, Parmar  N, et al.  Attention is all you need. Adv Neural Inf Process Syst  2017;30 1–11.
28. Elman  JL . Finding structure in time. Cogn Sci  1990;14 (2 ):179–211.
29. Hochreiter  S, Schmidhuber  J. Long short-term memory. Neural Comput  1997;9 (8 ):1735–80.9377276
30. Graves  A, Graves  A. In: Janusz Kacprzyk (ed). Long Short-Term Memory. Supervised Sequence Labelling with Recurrent Neural Networks. Berlin, Heidelberg: Springer, 2012, 37–45.
31. Hameed  Z, Garcia-Zapirain  B. Sentiment classification using a single-layered BiLSTM model. IEEE Access  2020;8 :73992–4001.
32. Hameed  Z, Garcia-Zapirain  B, Ruiz  IO. A computationally efficient BiLSTM based approach for the binary sentiment classification. In: 2019 IEEE International Symposium on Signal Processing and Information Technology (ISSPIT). Ajman, United Arab Emirates, Khaled Assaleh: IEEE, 2019, 1–4.
33. Kim  Y . Convolutional neural networks for sentence classification. https://uwspace.uwaterloo.ca/bitstream/handle/10012/9592/Chen_Yahui.pdf?sequence=3&isAllowed=y (24 April 2024 last accessed).
34. Rasamoelina  AD, Adjailia  F, Sinčák  P. A review of activation function for artificial neural network. In: 2020 IEEE 18th World Symposium on Applied Machine Intelligence and Informatics (SAMI), Herl'any, Slovakia. Levente Kovács: IEEE, 2020, 281–6.
35. Bjorck  N, Gomes  CP, Selman  B, et al.  Understanding batch normalization. Adv Neural Inf Proces Syst  2018;31 :7705–16.
36. Baldi  P, Sadowski  PJ. In: Burges CJ (eds). Understanding Dropout, Advances in Neural Information Processing Systems. NY, United States: Curran Associates Inc, 2013, 26.
37. Zhou  P, Qi  Z, Zheng  S, et al.  Text classification improved by integrating bidirectional LSTM with two-dimensional max pooling, arXiv preprint. arXiv  2016;1611.06639 .
38. Albrecht  MR, Driessen  B, Kavun  EB, et al.  Block ciphers–focus on the linear layer (feat. PRIDE)  In: Advances in Cryptology–CRYPTO 2014: 34th Annual Cryptology Conference. Santa Barbara, CA, USA: Springer, 2014  Proceedings, David Hutchison Part I 34. 2014, 57–76.
39. De Boer  P-T, Kroese  DP, Mannor  S, et al.  A tutorial on the cross-entropy method. Ann Oper Res  2005;134 :19–67.
40. Kingma  DP, Ba  J. Adam: a method for stochastic optimization, arXiv preprint. arXiv  2014;1412.6980 .
41. Prechelt  L . Early stopping-but when? Neural Networks: Tricks of the trade. Springer, Heidelberg, David Hutchison, 2002, 55–69.
42. Hu  J, Shen  L, Sun  G. Squeeze-and-excitation networks. Proceedings of the IEEE conference on computer vision and pattern recognition. 2018;1 :7132–41.
43. Dalhat  MH, Mohammed  MRS, Alkhatabi  HA, et al.  NAT10: an RNA cytidine transferase regulates fatty acid metabolism in cancer cells. Clin Transl Med  2022;12 (9 ):e1045.36149760
