==== Front Res Sq ResearchSquare Research Square American Journal Experts 37398207 10.21203/rs.3.rs-2963881/v1 10.21203/rs.3.rs-2963881 preprint 1 Article CELL-E: A Text-To-Image Transformer for Protein Localization Prediction Khwaja Emaad 12* http://orcid.org/0000-0002-0734-9868 Song Yun S. 234 http://orcid.org/0000-0003-1704-4141 Huang Bo 456* 1 UC Berkeley - UCSF Joint Graduate Program in Bioengineering, CA, USA. 2 Computer Science Division, UC Berkeley, Berkeley, 94720, CA, USA. 3 Department of Statistics, UC Berkeley, Berkeley, 94720, CA, USA. 4 Chan Zuckerberg Biohub - San Francisco, San Francisco, 94158, CA, USA. 5 Department of Pharmaceutical Chemistry, UCSF, San Francisco, 94143, CA, USA. 6 Department of Biochemistry and Biophysics, UCSF, San Francisco, 94143, CA, USA. Contributing authors: yss@berkeley.edu; 8 Author information E.K. played a key role in the advancement of the approach, carrying out the majority of the coding, designing and conducting a significant number of the experiments, and producing an initial version of the manuscript. The remaining authors also offered consistent input on all aspects of the project, assessed the code, and helped with the final draft of the manuscript. * Corresponding author(s): emaad@berkeley.edu; bo.huang@ucsf.edu; 02 6 2023 rs.3.rs-2963881https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use. nihpp-rs2963881v1.pdf Accurately predicting cellular activities of proteins based on their primary amino acid sequences would greatly improve our understanding of the proteome. In this paper, we present CELL-E, a text-to-image transformer model that generates 2D probability density images describing the spatial distribution of proteins within cells. Given an amino acid sequence and a reference image for cell or nucleus morphology, CELL-E predicts a more refined representation of protein localization, as opposed to previous in silico methods that rely on pre-defined, discrete class annotations of protein localization to subcellular compartments. text-to-image synthesis transformers single-cell imaging generative models ==== Body pmc1 Introduction In recent years, advancements in sequencing technologies have allowed for the comprehensive cataloging of proteins and their amino acid sequences across a wide range of organisms [1]. Despite this progress, the exact functions and cellular dynamics of many proteins remain unclear. In order to gain a deeper understanding of these proteins, researchers have sought ways to predict their properties, including structure, interactions, subcellular localization, and trafficking patterns, from their amino acid sequences. This type of computational analysis has the potential to shed light on the “dark matters” of the proteome and enable large-scale screening before expensive experimental validation. These tools have numerous applications in biomedical research, such as drug design and therapeutic target discovery [2]. In this study, our focus is on predicting subcellular localization of proteins from their amino acid sequences, which serves as the spatial context for their cellular functions. The localization of a protein to a specific subcellular compartment can be driven by either active transport or passive diffusion in conjunction with specific protein-protein interactions, often involving localization “signals” in the amino acid sequence [3–5]. In many cases, however, the exact mechanisms for sequence recognition and trafficking are not yet fully understood [6]. For example, there is ongoing debate about the mechanism behind the import of proteins via the nuclear localization sequence (NLS) [7]. Given these challenges, machine learning utilizing existing knowledge of protein localization has become a particularly useful tool. Although computational prediction of protein subcellular localization from primary amino acid sequences is an active area of research, most works train the model with class annotation of subcellular compartments (e.g., nucleus, plasma membrane, endoplasmic reticulum, etc.) [8] which are available from databases such as UniProt [9]. This approach has two major limitations. First, many proteins are present in different and variable amounts across multiple subcellular compartments. Second, protein localization could be highly heterogeneous and dynamic depending on the cell type and cell state (including cell cycle state). Neither of these two aspects have been captured by existing discrete class annotations. Consequently, machine-learning-based protein localization prediction still has limited applications. Furthermore, to assist mechanistic discoveries, it is highly desirable for the machine learning models to be explainable. To investigate the relationship between sequence and subcellular localization, we present CELL-E, a text-to-image transformer model which predicts the probability of protein localization on a per-pixel level from a given amino acid sequence and a conditional reference image for the cell or nucleus morphology and location (Fig. 1). It relies on transfer learning via amino acid embeddings from a pre-trained protein language model and two quantized image encoders trained from a live-cell imaging dataset. By generating a two-dimensional probability density function (2D PDF) atop the reference image, CELL-E naturally accounts for multi-compartment localization and the cell type/state information implicitly encoded by the cell morphology. We demonstrate the capability of CELL-E to predict localization of proteins, identify changes in localization due to mutations, and uncover sequence features correlated with the specification of subcellular protein localization. 2 Results 2.1 The CELL-E Model CELL-E (Fig. 2) is inspired by the text-to-natural-image generation model of DALL-E [10] (See Section S.2.1 for a review of relevant work). Similar to DALL-E, our model autoregressively learns text and image tokens as a single stream of data. On the other hand, While the goal of general text-to-image models is to produce images with high perceptual strength, they do not necessarily aim for quantitative accuracy [10–12]. Therefore, CELL-E was designed with the following considerations: 1. Transfer learning. Training CELL-E requires a library of cellular images and corresponding morphological reference images for a large number of proteins. For this purpose, we utilized the recently established OpenCell library[13], which contains a library of 1,311 CRISPR-edited HEK293T human cell lines, each having one target protein fluorescently tagged and imaged by confocal microscopy with accompanying DNA staining as the reference for nuclei morphology. The high image quality and consistency makes OpenCell a good choice as the training and validation dataset (See Section S.3.1 for more information). Still, data availability in this domain remains a large obstacle. For example, DALL-E was trained on 250 million text-images pairs [10], orders of magnitude larger than the OpenCell dataset. We utilize transfer learning by incorporating frozen embeddings from a pre-trained protein language model as the input representation of the amino acid text sequence. This approach reduces the number of learned paramaters, thereby alleviating the burden for CELL-E to also learn the amino acid sequence space. This allows training to be concentrated on the relationship between sequence and image tokens. We evaluated multiple protein language models (see Supplementary Information and Table S1) and eventually chose the BERT-based model from Rao et. al. [14], which we refer to as the TAPE model, for subsequent work. 2. Morphological reference. In our initial efforts, we found that a transformer using just the amino acid tokens and protein image tokens is capable of generating cell-like images from the amino acid sequence alone (Fig. S3). However, quantifying protein localization information in the generated images is challenging. Furthermore, an estimation of a single snapshot of protein localization is not necessarily a quantifiable indication of global behavior. Therefore, in addition to amino acid tokens and protein image tokens, we add a 3rd embedding space to include tokens representing the overall cell morphology from a reference image. The reference image provides the model with information regarding the localization of subcellular structures and compartments. Moreover, cell morphology implicitly provides the cell type and cell state context for CELL-E predictions. 3. Image model. Instead of the Vector Quantized Varational Autoencoder (VQVAE) previously used to analyze OpenCell imaging data [15], we chose to use Vector Quantized Generative Adversarial Network (VQGAN) [16] which produces images with comparatively higher spatial frequency. To simplify the task of the protein image VQGAN, we let it predict per-pixel binary representations of protein localization (i.e., a thresholded image). This allows us to use the marginal probabilities predicted for each image token from CELL-E to create a weighted sum on the image tokens. This latent space linear combination is then used to generate a continuous 2D probability density function of protein localization, which resembles a gray-scale image (Fig. 7). We note that the same model can also be trained to output gray-scale images directly (See Supplementary Notes S.2.2 and Fig. S4). 2.2 Performance Evaluation Fig. 3 and Fig. S2 show the CELL-E predictions for several proteins in the validation dataset. High similarities can be seen between the predictions and the ground truth. Even though the reference images only depict the nuclei, which is a limitation of the OpenCell training data, CELL-E can reasonably paint the shape of the cell for cytoplasmic proteins. Interestingly, the case of Mitogen-Activated Protein Kinase 9 (MAPK9) contains a cell in metaphase (top row of Fig. 3). CELL-E correctly predicts the round shape of its distribution around the mitotic chromosomes instead of the more expanded distribution for the adjacent interphase cell. This result suggests that CELL-E can indeed capture cell state information from the morphological reference images. We used several metrics to evaluate the reconstruction performance of CELL-E, summarized in Table S2. Among the metrics, nucleus proportion accuracy measures how close the estimated proportion of pixel intensity within the nucleus is to the ground truth thresholded image. We believe this is the most relevant metric as it is not obscured by small spatial variations and nucleus boundaries can be obtained from the reference images. Description of other metrics and more information on the evaluation procedure can be found in Section S.3.7. Using these metrics, we performed ablations studies to optimize our model architecture and choice of protein language embedding (see Section S.2.3, Fig. S5 and Table S1). While CELL-E is not specifically trained as a discrete localization classifier, we also performed naive comparison between CELL-E model and 1D protein localization classifiers MuLoc [17] and Subcons [18] specifically trained with annotated protein localizations. We focused on nuclear classification using a simple classification criteria on CELL-E output (see Section S.3.7), and the results are summarized in Table 1. We observed a relatively high degree of accuracy from this method compared to the task-specific models. CELL-E was a close second for validation set proteins despite not seeing localization annotations during training. 2.3 Analysis of NLS using CELL-E As a first test to show that CELL-E can recognize specific, functional sequence features, we let it predict the images for Green Fluorescent Protein (GFP), which is non-native to human and does not contain known localization signals, as well as GFP appended with two commonly used NLS’s KRPAATKKAGQAKKKK from nucleoplasmin [19] and PAAKRVKLD from N-Myc [20]) that drive nuclear localization of a protein. We also appended a randomly generated sequence as a control. A randomly chosen nuclear image from the OpenCell dataset was used as the morphological reference. CELL-E does not localization of GFP (or random sequence + GFP) to a specific subcellular compartment with high confidence, whereas the two NLS-GFP fusions are clearly predicted to be localized within the nucleus (Fig. 4). Therefore, CELL-E has the potential to perform computational insertion screenings for the functional sufficiency of putative localization sequence features. Next, we examined whether CELL-E can identify NLS in a protein by computationally performing truncation/deletion studies. For this purpose, we chose DNA Topoisomerase I (TOP1), whose N-terminal intrinsically disordered region (amino acid (aa) 1–199) is essential for its nuclear localization [21]. An experimental study generated a series of deletion mutants for this region and imaged the subcellular localization in HeLa cells when fused to eGFP [22]. To computationally reproduce this study, we fed the exact sequences of the deletion mutants to CELL-E. As shown in Fig. 5, the predictions were largely consistent with the experimental data, recapturing the inability for aa 1–67 to drive nuclear localization despite containing a putative NLS, as well as the sufficiency of aa 148–199 as an NLS. Lastly, we demonstrate a more direct approach than computational insertion or deletion studies to identify putative sequence features responsible for protein localization. Specifically, we split the generated image patches into two groups, one with the target protein being present and the other being absent based on the average pixel intensity within the 16 × 16 image patch. Then, we calculated the difference of attention weights for each amino acid token to contribute to the two groups. Fig. 6 highlights the amino acids with higher weights for the “present” group. The highlighted amino acids include the three putative NLSs (Motifs II, III, and IV) in the experimentally verified aa 148–199 range, as well as part of the new aa 117–146 NLS identified in [22]. On the other hand, the putative NLS (Motif I) in the experimentally invalidated aa 1–69 range are not activated. The attention map also suggest that aa 89–107 (KIKKE) could be another NLS in this protein. We must point out that the calculation of attention map was simply based on a protein being “present” or “not present” in image patches and did not specify “nuclear localization” at all. Therefore, it should be capable of serving as a general approach to discover putative sequence features driving protein localization to a variety of subcellular compartments. 3 Discussion CELL-E’s performance seems to be currently limited by the scope of the OpenCell dataset, which only accounts for a handful of proteins within a single cell type and imaging modality. As the OpenCell project is an active development, we expect stronger performance as more data become available. The availability of brightfield (e.g., phase-contrast) images as the morphological reference will also likely improve the prediction of cytoplasmic protein localization compared to using nuclei images. Furthermore, the utility of the model comes in terms of linking embedding spaces of dependent data. One could imagine follow up experiments where rather than images being the prediction, other signatures such as protein mass spec could be predicted. Additionally, other sources of information, such as structural embeddings could be incorporated to bolster CELL-E’s capabilities. 4 Methods We use a multi-phase training approach similar to DALL-E, but our model also uses pre-trained language-model input embeddings for the amino acid text sequences via TAPE: Phase 1 A Vector Quantized-Generative Adversarial Network (VQGAN) [16] is trained to represent a single channel 256 × 256nucleus image as a grid comprised of 16 × 16 image tokens (Fig. S7), each of which could be one of 512 tokens. Phase 2 A similar VQGAN is trained on images corresponding to binarized versions of protein images. These tokens represent the spatial distribution of the protein (Fig. S9). Phase 3 The VQGAN image tokens are concatenated to 1000 amino acid tokens for the autoregressive transformer which models a joint distribution over the amino acids, nucleus image, and protein threshold image tokens. 4.1 Model Specifics The optimization problem is modelled as maximizing the evidence lower bound (ELBO) [23, 24] on a joint likelihood distribution over protein threshold images u, nucleus images x, amino acids y, and tokens z for the protein threshold image: Theorem 1. pθ,ψ(u,x,y,z)=pθ(u∣x,y,z)pψ(x,y,z) This is bounded by: Theorem 2. lnpθ,ψ(u,x,y)≥Ez~qϕ(z∣u)[lnpθ(u∣x,y,z)]−KL(qϕ(x,y,z∣u),pψ(x,y,z)) where qϕ is the distribution 16 × 16 image tokens from the VQGAN corresponding to the threshold protein image u,pθ is the distribution over protein threshold generated by the VQGAN given the image tokens, and pψ indicates the joint distribution over the amino acid, nucleus, and protein threshold tokens within the transformer. 4.2 Nucleus Image Encoder Training both image VQGANs maximizes ELBO with respect to ϕ and θ. The VQGAN improves upon existing quantized autoencoders by introducing a learned discriminator borrowed from GAN architectures [16]. The Nucleus Image Encoder is a VQGAN which represents 256 × 256 nucleus reference images as 256 16 × 16 image patches. The VQGAN codebook size was set to n=512 image patches. Further details can be found in Section S.3.4. 4.3 Protein Threshold Image Encoder The protein threshold image encoder learns a dimension reduced representation of a discrete binary PDF of per-pixel protein location, represented as an image image. We adopt a VQGAN architecture identical to the Nucleus VQGAN. The VQGAN serves to approximate the total set of binarized image patches. While in theory a discrete lookup of each pixel arrangement is possible, this would require ~ 1.16 × 1077 entries, which is computationally infeasible. Furthermore, some distributions of pixels might be so improbable that having a discrete entry would be a waste of space. Protein images are binarized with respect to a mean-threshold, via: u‾i,j=1,ui,j≥μ,0,ui,j<μ, ∀ pixels u∈ image U of size i×j, where μ is the mean pixel intensity in the image (Fig. S8). The 16 × 16 image patches learned within the VQGAN codebook therefore correspond to local protein distributions. In Section 4.6, we detail how a weighted sum over these binarized image patches is used to determine a final probability density map. Hyperparameters and other training details can be found in Section S.3.5. 4.4 Amino Acid Embedding For language transformers, it is necessary to learn both input embedding representations of a text vector as well as attention weights between embeddings [25]. In practice, this creates a need for very large datasets [26]. The OpenCell dataset contains 1,311 proteins, while the human body is estimated to contain upwards of 80,000 unique proteins [27]. It is unlikely that such a small slice could account for the large degrees of variability found in nature. In order to overcome this obstacle, we opted for a transfer learning strategy, where fixed amino acid embeddings from a pre-trained language model exposed to a much larger dataset were utilized. We found the strongest performance came from TAPE embeddings [14]. Utilizing pre-trained embeddings had the two-fold benefit of giving our model a larger degree of protein sequence context, as well as reducing the number of trained model parameters, which allowed us to scale the depth of our network. We tried training using random initialization for amino acid embeddings (See Section S.2.3), however, we noted overfitting on the validation set image reconstruction and high loss on validation sequences. We also experimented with other types of protein embeddings, including UniRep [28] and ESM1-b [29]. 4.5 CELL-E Transformer The transformer pϕ utilizes an input comprised of amino acid tokens, a 256 × 256 nucleus image crop, and the 256 × 256 corresponding protein image threshold crop. In this phase, ϕ and θ are fixed, and a prior over all tokens is learned by maximizing ELBO with respect to ϕ. It is a decoder-only model [30]. The model is trained on a concatenated sequence of text tokens, nucleus image tokens, and protein threshold image tokens, in order. Within the CELL-E transformer, image token embeddings were cast into the same dimensionality as the language model embedding to in order to maintain the larger protein context information, however the embeddings corresponding to the image tokens within this dimension are learned (See Fig. 2). 4.6 Probability Density Maps When generating images, the model is provided with the amino acid sequence and nucleus image. The transformer autoregressively predicts the protein-threshold image. In order to select a token, the model outputs logits which contain probability values corresponding to the codebook identity of the next token. The image patch vi is selected by filtering for the top 25% of tokens and applying top-k sampling with gumbel noise [31]. Ordinarily, the final image is generated by converting the predicted codebook indices of the protein threshold image to the VQGANs decoder. However, to generate the probability density map v‾, we include the full range of probability values corresponding to image patches, pvi, obtained from the output logits. The values are clipped between 0 and 1 and multiplied by the embedding weights within the VQGAN’s decoder, wi: Theorem 3. v‾=w⋅p(v)=∑i=1nwipvi This output is normalized and displayed as a heatmap (Fig. 7). Supplementary Material Supplement 1 7 Acknowledgements B.H. is supported by the National Institutes of Health (R01GM131641). Y.S.S. and B.H. are Chan Zuckerberg Biohub - San Francisco Investigators. Y.S.S. is supported by NIH grant R35-GM134922. Fig. 1 Given an amino acid sequence and a reference nucleus image, CELL-E makes a prediction of protein localization with respect to the nucleus as a 2D probability density function, shown as heatmap, with color indicating relative confidence for each pixel. Fig. 2 Graphical depiction of CELL-E. Solid lines correspond to pre-trained components. Gray dashed lines are learned in Phase 1 and 2 (Reference Image and Protein Threshold VQGANs). Black dashed lines correspond to components learned in Phase 3. A start token is prepended to the sequence and the final protein image token is removed. The amino acid sequence embedding from the model is preserved, and embedding spaces for the image tokens are cast in the same depth and concatenated with the amino acid sequence embedding. The transformer is tasked with reproducing the original sequence of tokens (e.g., the input sequence with start token shifted to the right one position). Fig. 3 Prediction results of several types of proteins from the validation set, unseen to the model during training. The nucleus channel is depicted in grayscale, and the protein channel is shown as an overlay in red (Fig. S1 for clarification). The thresholded image (Column 2) is designated “Ground Truth” because those are the types of images exposed to the model during training. The predicted probability map is obtained from a weighted sum of potential image patches and normalized to 1. Fig. 4 Predicted localization of GFP and modified-GFP sequences. Fig. 5 CELL-E’s predicted localization (images) of eGFP fusions from [22] and corresponding localization annotations (table) from the original paper. In the table on the right hand side, green indicates agreement between CELL-E and experimental results, while red indicates disagreement. aa 1–199 contains the entire N-terminus region. aa 1–146 only contains Motifs I and V. aa 1–67 only contains Motif-I. aa 148–199 contains Motif II, III, IV and V. Fig. 6 Attention weights for significant tokens when patches containing a large percentage of protein are selected (bottom-right figure). Previous computationally identified putative NLSs are boxed in black (top figure). These are aa 59–65 (Motif I, KKHKEKE), aa 150–156 (Motif II, KKIKTED), aa 174–180 (Motif III, KKPKNKD), and aa 192–198 (Motif IV, KKKPKKE). Additionally, the new NLS identified in Mo et. al.[22], Motif V (aa 117–146 ), is highlighted. Fig. 7 Simplified example of probability map calculation. Each circle corresponds to an image token within the quantized VQGAN embedding space. Each PDF patch (yellow) is obtained as a weighted sum over all protein threshold image VQGAN codebook vectors. Table 1 Nuclear Localization Prediction Accuracy Train Validation VQGAN 0.99 ± 0.08 0.99 ± 0.09 CELL-E 0.89 ± 0.31 0.72 ± 0.45 MuLoc 0.71 ± 0.45 0.79 ± 0.41 Subcons 0.43 ± 0.49 0.69 ± 0.46 VQGAN indicates the accuracy evaluated on the ground truth threshold image passed through the VQGAN image encoder. As CELL-E selects tokens from this VQGAN to produce its outputs, these values represent the best possible performance for our model. 6 Code availability Our model is a heavily modified version of an open source text-to-image transformer [32], available via the MIT license (Copyright (c) 2021 Phil Wang). Our code is available at https://github.com/BoHuangLab/Protein-Localization-Transformer via the MIT license (Copyright (c) 2022 Emaad Khwaja, Yun Song, & Bo Huang). ==== Refs References [1] Hu T. , Chitnis N. , Monos D. & Dinh A. Next-generation sequencing technologies: An overview. Human Immunology 82 , 801–811 (2021). URL https://www.sciencedirect.com/science/article/pii/S0198885921000628.33745759 [2] Palma C.-A. , Cecchini M. & Samorì P. Predicting self-assembly: from empirism to determinism. Chemical Society Reviews 41 , 3713–3730 (2012). URL https://pubs.rsc.org/en/content/articlelanding/2012/cs/c2cs15302e. Publisher: The Royal Society of Chemistry.22430648 [3] Chacinska A. , Koehler C. M. , Milenkovic D. , Lithgow T. & Pfanner N. Importing Mitochondrial Proteins: Machineries and Mechanisms. Cell 138 , 628–644 (2009). URL https://www.sciencedirect.com/science/article/pii/S0092867409009672.19703392 [4] Imai K. & Nakai K. Prediction of subcellular locations of proteins: where to proceed? Proteomics 10 , 3970–3983 (2010).21080490 [5] Ahmed H. R. & Glasgow J. Sokolova M . & van Beek P . (eds) A Novel Particle Swarm-Based Approach for 3D Motif Matching and Protein Structure Classification. (eds Sokolova M . & van Beek P .) Advances in Artificial Intelligence, Lecture Notes in Computer Science, 1–12 (Springer International Publishing, Cham, 2014). [6] Gardy J. L. & Brinkman F. S. L. Methods for predicting bacterial protein subcellular localization. Nature Reviews Microbiology 4 , 741–751 (2006). URL https://www.nature.com/articles/nrmicro1494. Bandiera_abtest: a Cg_type: Nature Research Journals Number: 10 Primary_atype: Reviews Publisher: Nature Publishing Group.16964270 [7] Lu J. Types of nuclear localization signals and mechanisms of protein import into the nucleus. Cell Communication and Signaling 19 , 60 (2021). URL 10.1186/s12964-021-00741-y.34022911 [8] Almagro Armenteros J. J. , Sønderby C. K. , Sønderby S. K. , Nielsen H. & Winther O. DeepLoc: prediction of protein subcellular localization using deep learning. Bioinformatics 33 , 3387–3395 (2017). URL 10.1093/bioinformatics/btx431.29036616 [9] The UniProt Consortium. UniProt: the universal protein knowledgebase. Nucleic acids research 45 , D158–D169 (2017). cPlae: England. 27899622 [10] Ramesh A. Zero-Shot Text-to-Image Generation. arXiv:2102.12092 [cs] (2021). URL http://arxiv.org/abs/2102.12092. ArXiv: 02.2112092. [11] Ding M. CogView: Mastering Text-to-Image Generation via Transformers. arXiv:2105.13290 [cs] (2021). URL http://arxiv.org/abs/2105.13290. ArXiv: 2105.13290. [12] Ramesh A. , Dhariwal P. , Nichol A. , Chu C. & Chen M. Hierarchical Text-Conditional Image Generation with CLIP Latents (2022). URL http://arxiv.org/abs/2204.06125. ArXiv:2204.06125 [cs]. [13] Cho N. H. OpenCell: Endogenous tagging for the cartography of human cellular organization. Science (New York, N.Y.) 375 , eabi6983 (2022). Place: United States. 35271311 [14] Rao R. Evaluating Protein Transfer Learning with TAPE. arXiv:1906.08230 [cs, q-bio, stat] (2019). URL http://arxiv.org/abs/1906.08230. ArXiv: 1906.08230. [15] Kobayashi H. , Cheveralls K. C. , Leonetti M. D. & Royer L. A. Self-Supervised Deep Learning Encodes High-Resolution Features of Protein Subcellular Localization. preprint, Cell Biology (2021). URL 10.1101/2021.03.29.437595. [16] Esser P. , Rombach R. & Ommer B. Taming Transformers for High-Resolution Image Synthesis. arXiv:2012.09841 [cs] (2021). URL http://arxiv.org/abs/2012.09841. ArXiv: 2012.09841. [17] Jiang Y. , Wang D. , Wang W. & Xu D. Computational methods for protein localization prediction. Computational and Structural Biotechnology Journal 19 , 5834–5844 (2021). URL https://www.sciencedirect.com/science/article/pii/S2001037021004451.34765098 [18] Salvatore M. , Warholm P. , Shu N. , Basile W. & Elofsson A. SubCons: a new ensemble method for improved human subcellular localization predictions. Bioinformatics 33 , 2464–2470 (2017). URL 10.1093/bioinformatics/btx219.28407043 [19] Dingwall C. , Robbins J. , Dilworth S. M. , Roberts B. & Richardson W. D. The Nucleoplasmin Nuclear Location Sequence Is Larger and MoreComplex than That of S¥−40 Large T Antigen. The Journal of Cell Biology 107 , 9 (1988).2839524 [20] Ray M. , Tang R. , Jiang Z. & Rotello V. M. Quantitative Tracking of Protein Trafficking to the Nucleus Using Cytosolic Protein Delivery by Nanoparticle-Stabilized Nanocapsules. Bioconjugate Chemistry 26 , 1004–1007 (2015). URL 10.1021/acs.bioconjchem.5b00141. Publisher: American Chemical Society.26011555 [21] Alsner J. , Svejstrup J. Q. , Kjeldsen E. , Sørensen B. S. & Westergaard O. Identification of an N-terminal domain of eukaryotic DNA topoisomerase I dispensable for catalytic activity but essential for in vivo function. The Journal of Biological Chemistry 267 , 12408–12411 (1992).1319995 [22] Mo Y.-Y. , Wang C. & Beck W. T. A Novel Nuclear Localization Signal in Human DNA Topoisomerase I*. Journal of Biological Chemistry 275 , 41107–41113 (2000). URL https://www.sciencedirect.com/science/article/pii/S0021925819556435.11016921 [23] Kingma D. P. & Welling M. Auto-Encoding Variational Bayes. arXiv:1312.6114 [cs, stat] (2014). URL http://arxiv.org/abs/1312.6114. ArXiv: 1312.6114. [24] Rezende D. J. , Mohamed S. & Wierstra D. Stochastic Backpropagation and Approximate Inference in Deep Generative Models, 1278–1286 (PMLR, 2014). URL https://proceedings.mlr.press/v32/rezende14.html. ISSN: 1938-7228. [25] Vaswani A. et al. Guyon I . (eds) Attention is All you Need. Advances in Neural Information Processing Systems, Vol. 30 (Curran Associates, Inc., 2017). URL https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. [26] Popel M. & Bojar O. Training Tips for the Transformer Model. The Prague Bulletin of Mathematical Linguistics 110 , 43–70 (2018). URL http://content.sciendo.com/view/journals/pralin/110/1/article-p43.xml. [27] Schuler G. D. A gene map of the human genome. Science (New York, N.Y.) 274 , 540–546 (1996).8849440 [28] Bepler T. & Berger B. Learning the protein language: Evolution, structure, and function. Cell Systems 12 , 654–669.e3 (2021). URL https://linkinghub.elsevier.com/retrieve/pii/S2405471221002039.34139171 [29] Rives A. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences 118 , e2016239118 (2021). URL 10.1073/pnas.2016239118. _eprint: 10.1073/pnas.2016239118. [30] Liu P. J. Generating Wikipedia by Summarizing Long Sequences (2023). URL https://openreview.net/forum?id=Hyg0vbWC-. [31] Jang E. , Gu S. & Poole B. Categorical Reparameterization with Gumbel-Softmax. arXiv:1611.01144 [cs, stat] (2017). URL http://arxiv.org/abs/1611.01144. ArXiv: 1611.01144. [32] Wang P. DALL-E in Pytorch (2022). URL https://github.com/lucidrains/DALLE-pytorch. Original-date: 2021–01-05T20:35:16Z. [33] Vig J. BERTology Meets Biology: Interpreting Attention in Protein Language Models (2021). URL http://arxiv.org/abs/2006.15222. ArXiv:2006.15222 [cs, q-bio] version: 3. [34] Zaheer M. Big Bird: Transformers for Longer Sequences (2021). URL http://arxiv.org/abs/2007.14062. ArXiv:2007.14062 [cs, stat] version: 2. [35] Elnaggar A. ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE transactions on pattern analysis and machine intelligence 44 , 7112–7127 (2022).34232869 [36] Wang Y. A High Efficient Biological Language Model for Predicting Protein–Protein Interactions. Cells 8 , 122 (2019). URL https://www.mdpi.com/2073-4409/8/2/122. Number: 2 Publisher: Multidisciplinary Digital Publishing Institute.30717470 [37] Steinegger M. , Mirdita M. & Söding, J. Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold. Nature Methods 16 , 603–606 (2019). URL 10.1038/s41592-019-0437-4.31235882 [38] Suzek B. E. , Wang Y. , Huang H. , McGarvey P. B. & Wu C. H. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics 31 , 926–932 (2015). URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4375400/.25398609 [39] Mistry J. Pfam: The protein families database in 2021. Nucleic Acids Research 49 , D412–D419 (2021). URL 10.1093/nar/gkaa913.33125078 [40] Berman H. M. The Protein Data Bank. Nucleic Acids Research 28 , 235–242 (2000). URL 10.1093/nar/28.1.235.10592235 [41] Alley E. C. , Khimulya G. , Biswas S. , AlQuraishi M. & Church G. M. Unified rational protein engineering with sequence-based deep representation learning. Nature Methods 16 , 1315–1322 (2019). URL https://www.nature.com/articles/s41592-019-0598-1. Number: 12 Publisher: Nature Publishing Group.31636460 [42] Devlin J. , Chang M.-W. , Lee K. & Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs] (2019). URL http://arxiv.org/abs/1810.04805. ArXiv: 1810.04805. [43] Pan G. , Sun C. , Liao Z. & Tang J. in Machine and Deep LearningDeep learning (DL) for Prediction of Subcellular Localization (ed.Cecconi D .) Proteomics Data Analysis Methods in Molecular Biology, 249–261 (Springer US, New York, NY, 2021). URL 10.1007/978-1-0716-1641-315. [44] Zou J. A primer on deep learning in genomics. Nature Genetics 51 , 12–18 (2019). URL https://www.nature.com/articles/s41588-018-0295-5. Bandiera_abtest: a Cg_type: Nature Research Journals Number: 1 Primary_atype: Reviews Publisher: Nature Publishing Group Subject term: Computational biology and bioinformatics;Genetics;Genome informatics Subject_term_id: computational-biology-and-bioinformatics;genetics;genome-informatics.30478442 [45] Yun K. , Huyen A. & Lu T. Deep Neural Networks for Pattern Recognition. arXiv:1809.09645 [cs] (2018). URL http://arxiv.org/abs/1809.09645. ArXiv: 1809.09645. [46] Pang L. , Wang J. , Zhao L. , Wang C. & Zhan H. A Novel Protein Subcellular Localization Method With CNN-XGBoost Model for Alzheimer’s Disease. Frontiers in Genetics 9 , 751 (2019). URL 10.3389/fgene.2018.00751.30713552 [47] Yang W.-Y. , Lu B.-L. & Yang Y. A Comparative Study on Feature Extraction from Protein Sequences for Subcellular Localization Prediction, 1–8 (2006). [48] Hager K. M. , Striepen B. , Tilney L. G. & Roos D. S. The nuclear envelope serves as an intermediary between the ER and Golgi complex in the intracellular parasite Toxoplasma gondii. Journal of Cell Science 112 (Pt 16 ), 2631–2638 (1999).10413671 [49] Mim C. & Unger V. M. Membrane curvature and its generation by BAR proteins. Trends in Biochemical Sciences 37 , 526–533 (2012). URL https://www.sciencedirect.com/science/article/pii/S0968000412001387.23058040 [50] Ewing G. W. pH is a Neurally Regulated Physiological System. Increased Acidity Alters Protein Conformation and Cell Morphology and is a Significant Factor in the Onset of Diabetes and Other Common Pathologies. The Open Systems Biology Journal 5 (2012). URL https://benthamopen.com/ABSTRACT/TOSYSBJ-5-1. [51] Martorana A. Probing Protein Conformation in Cells by EPR Distance Measurements using Gd3+ Spin Labeling. Journal of the American Chemical Society 136 , 13458–13465 (2014). URL 10.1021/ja5079392. Publisher: American Chemical Society.25163412 [52] Lou H.-Y. , Zhao W. , Zeng Y. & Cui B. The Role of Membrane Curvature in Nanoscale Topography-Induced Intracellular Signaling. Accounts of Chemical Research 51 , 1046–1053 (2018). URL 10.1021/acs.accounts.7b00594. Publisher: American Chemical Society.29648779 [53] Ohno M. , Karagiannis P. & Taniguchi Y. Protein Expression Analyses at the Single Cell Level. Molecules 19 , 13932–13947 (2014). URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6270791/.25197931 [54] Grün D . Revealing dynamics of gene expression variability in cell state space. Nature Methods 17 , 45–49 (2020). URL https://www.nature.com/articles/s41592-019-0632-3. Number: 1 Publisher: Nature Publishing Group.31740822 [55] Kotliar D. Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-Seq. eLife 8 , e43803 (2019). URL 10.7554/eLife.43803. Publisher: eLife Sciences Publications, Ltd.31282856 [56] Goodfellow I. Ghahramani Z ., Welling M ., Cortes C ., Lawrence N . & Weinberger K. Q . (eds) Generative Adversarial Nets. (eds Ghahramani Z ., Welling M ., Cortes C ., Lawrence N . & Weinberger K. Q .) Advances in Neural Information Processing Systems, Vol. 27 (Curran Associates, Inc., 2014). URL https://proceedings.neurips.cc/paper_files/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf. [57] Mansimov E. , Parisotto E. , Ba J. L. & Salakhutdinov R. Gene rating Images from Captions with Attention. arXiv:1511.02793 [cs] (2016). URL http://arxiv.org/abs/1511.02793. ArXiv: 1511.02793. [58] Reed S. Balcan M. F . & Weinberger K. Q . (eds) Generative Adversarial Text to Image Synthesis. (eds Balcan M. F . & Weinberger K . Q.) Proceedings of The 33rd International Conference on Machine Learning, Vol. 48 of Proceedings of Machine Learning Research, 1060–1069 (PMLR, New York, New York, USA, 2016). URL https://proceedings.mlr.press/v48/reed16.html. [59] Xu T. AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks. arXiv:1711.10485 [cs] (2017). URL http://arxiv.org/abs/1711.10485. ArXiv: 1711.10485. [60] Osorio D. , Rondó-Villarreal P. & Torres Sáez R. Peptides: A Package for Data Mining of Antimicrobial Peptides. The R Journal 7 , 4–14 (2015). [61] Kidera A. , Konishi Y. , Oka M. , Ooi T. & Scheraga H. A. Statistical analysis of the physical properties of the 20 naturally occurring amino acids. Journal of Protein Chemistry 4 , 23–55 (1985). URL 10.1007/BF01025492. [62] Sandberg M. , Eriksson L. , Jonsson J. , Sjöström M. & Wold S. New chemical descriptors relevant for the design of biologically active peptides. A multivariate characterization of 87 amino acids. Journal of medicinal chemistry 41 , 2481–2491 (1998). URL 10.1021/jm9700575.9651153 [63] Cruciani G. Peptide studies by means of principal properties of amino acids derived from MIF descriptors. Journal of Chemometrics 18 , 146–155 (2004). URL 10.1002/cem.856. _eprint: 10.1002/cem.856. [64] Liang G. & Li Z. Factor Analysis Scale of Generalized Amino Acid Information as the Source of a New Set of Descriptors for Elucidating the Structure and Activity Relationships of Cationic Antimicrobial Peptides. QSAR & Combinatorial Science 26 , 754–763 (2007). URL 10.1002/qsar.200630145. _eprint: 10.1002/qsar.200630145. [65] Tian F. , Zhou P. & Li Z. T-scale as a novel vector of topological descriptors for amino acids and its application in QSARs of peptides. Journal of Molecular Structure 830 , 106–115 (2007). URL http://www.sciencedirect.com/science/article/pii/S0022286006314. [66] Mei H. , Liao Z. H. , Zhou Y. & Li S. Z. A new set of amino acid descriptors and its application in peptide QSARs. Biopolymers 80 , 775–786 (2005).15895431 [67] van Westen G. J. Benchmarking of protein descriptor sets in proteochemometric modeling (part 1): comparative study of 13 amino acid descriptor sets. Journal of Cheminformatics 5 , 41 (2013). URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3848949/.24059694 [68] Yang L. ST-scale as a novel amino acid descriptor and its application in QSAM of peptides and analogues. Amino Acids 38 , 805–816 (2010).19373543 [69] Georgiev A. G. Interpretable numerical descriptors of amino acid space. Journal of Computational Biology: A Journal of Computational Molecular Cell Biology 16 , 703–723 (2009).19432540 [70] Zaliani A. & Gancia E. MS-WHIM Scores for Amino Acids: A New 3D-Description for Peptide QSAR and QSPR Studies. J. Chem. Inf. Comput. Sci. (1999). [71] Wallis J. , Miller T. , Lerner C. & Kleerup E. Three-dimensional display in nuclear medicine. IEEE Transactions on Medical Imaging 8, 297–230 (1989). Conference Name: IEEE Transactions on Medical Imaging. [72] Thul P. J. A subcellular map of the human proteome. Science 356 , eaal3321 (2017). URL 10.1126/science.aal3321. Publisher: American Association for the Advancement of Science.28495876 [73] Schnell U. , Dijk F. , Sjollema K. A. & Giepmans B. N. G. Immunolabeling artifacts and the need for live-cell imaging. Nature Methods 9 , 152–158 (2012). URL 10.1038/nmeth.1855.22290187 [74] Walsh I. , Pollastri G. & Tosatto S. C. E. Correct machine learning on protein sequences: a peer-reviewing perspective. Briefings in Bioinformatics 17 , 831–840 (2016). URL 10.1093/bib/bbv082.26411473 [75] Fu L. , Niu B. , Zhu Z. , Wu S. & Li W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics 28 , 3150–3152 (2012). URL 10.1093/bioinformatics/bts565.23060610 [76] Su J. , Lu Y. , Pan S. , Wen B. & Liu Y. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864 [cs] (2021). URL http://arxiv.org/abs/2104.09864. ArXiv: 2104.09864. [77] Bo P. Improve the Transformer self-attention mechanism with just a few lines of code (almost no increase in computation). URL https://zhuanlan-zhihu-com.translate.goog/p/191393788?_x_tr_sl=en&_x_tr_tl=zh-CN&_x_tr_hl=en&_x_tr_pto=wapp. [78] Child R. , Gray S. , Radford A. & Sutskever I. Generating Long Sequences with Sparse Transformers (2019). URL http://arxiv.org/abs/1904.10509. ArXiv:1904.10509 [cs, stat]. [79] Stringer C. , Wang T. , Michaelos M. & Pachitariu M. Cellpose: a generalist algorithm for cellular segmentation. Nature Methods 18 , 100–106 (2021). URL https://www.nature.com/articles/s41592-020-01018-x. Number: 1 Publisher: Nature Publishing Group.33318659 [80] Abnar S. & Zuidema W. Quantifying Attention Flow in Transformers, 4190–4197 (Association for Computational Linguistics, Online, 2020). URL https://aclanthology.org/2020.acl-main.385.