
==== Front
J Phys Chem B
J Phys Chem B
jp
jpcbfk
The Journal of Physical Chemistry. B
1520-6106
1520-5207
American Chemical Society

39185763
10.1021/acs.jpcb.4c02461
Article
Protein Classes Predicted by Molecular Surface Chemical Features: Machine Learning-Assisted Classification of Cytosol and Secreted Proteins
Hu Guanghao †
Moon Jooa †
https://orcid.org/0000-0002-4065-1807
Hayashi Tomohiro *†‡
† Department of Materials Science and Engineering, School of Materials Science and Chemical Technology, Tokyo Institute of Technology, 4259 Nagatsuta-cho, Midori-ku, Yokohama-shi, Kanagawa-ken 226-8502, Japan
‡ The Institute for Solid State Physics, The University of Tokyo, 5-1-5, Kashiwanoha, Kashiwa, Chiba 277-0882, Japan
* E-mail: tomo@mac.titech.ac.jp.
26 08 2024
05 09 2024
128 35 84238436
26 04 2024
14 08 2024
13 08 2024
© 2024 The Authors. Published by American Chemical Society
2024
The Authors
https://creativecommons.org/licenses/by/4.0/ Permits the broadest form of re-use including for commercial purposes, provided that author attribution and integrity are maintained (https://creativecommons.org/licenses/by/4.0/).

Chemical structures of protein surfaces govern intermolecular interaction, and protein functions include specific molecular recognition, transport, self-assembly, etc. Therefore, the relationship between the chemical structure and protein functions provides insights into the understanding of the mechanism underlying protein functions and developments of new biomaterials. In this study, we analyze protein surface features, including surface amino acid populations and secondary structure ratios, instead of entire sequences as input for the classifier, intending to provide deeper insights into the determination of protein classes (cytosol or secreted). We employed a random forest-based classifier for the prediction of protein locations. Our training and testing data sets consisting of secreted and cytosol proteins were constructed using filtered information from UniProt and 3D structures from AlphaFold. The classifier achieved a testing accuracy of 93.9% with a feature importance ranking and quantitative boundary values for the top three features. We discuss the significance of these features quantitatively and the hidden rules to determine the protein classes (cytosol or secreted).

Japan Society for the Promotion of Science 10.13039/501100001691 JP22H045302 Ministry of Education, Culture, Sports, Science and Technology 10.13039/501100001700 NA Japan Society for the Promotion of Science 10.13039/501100001691 JP23H04059 document-id-old-9jp4c02461
document-id-new-14jp4c02461
ccc-price
==== Body
pmc1 Introduction

Proteins have evolved strategies to adapt themselves to their subcellular environments, resulting in optimal functioning within the corresponding environmental context.1,2 Consequently, studying protein subcellular localization can reveal the connections between protein properties and functions. The differentiation of protein features in body fluids gains attention, as body fluids serve as the medium for various cellular activities and provide a stable environment for biomaterials to function properly.3 Cell membranes divide the body fluid environment into extra- and intracellular parts. Intracellular and extracellular fluids maintain cellular homeostasis and mediate various biological processes. These fluids exhibit distinct compositional differences that facilitate their respective functions.

One primary difference between intra- and extracellular environments lies in the electrostatic potential, which is influenced by the concentration of ions and charged molecules, resulting in a unique electrical environment for each.4 The concentration of macromolecules, including proteins, nucleic acids, and carbohydrates, differs significantly between intracellular and extracellular fluids, reflecting the diverse functional requirements inside and outside the cell.5 Lastly, protein interactions with membranes and the degrees of molecular crowding vary between intracellular and extracellular environments. Intracellularly, proteins engage with densely packed organelle membranes, participating in signal transduction and vesicle trafficking.6,7 In contrast, extracellular proteins interact with the less crowded plasma membrane, contributing to cell adhesion, communication, and nutrient transport.8,9 These differences in interaction patterns and membrane crowdedness reflect the diverse functions that proteins serve inside and outside the cell.

Protein structures and surface amino acids strongly correlate with environmental conditions considering specific and nonspecific interactions.10 Proteins employ specific amino acid pairings that yield enhanced binding affinities with water molecules, which help prevent nonspecific interactions.11 Factors including shape and size, residue propensity, secondary structure exposure, and hydrogen bonding play crucial roles in protein complex formation.12 Furthermore, the prevalence of various chemical bonds effectively indicates surface antiadhesion properties.13 Prior research has leveraged amino acids as building blocks for modifying biomaterial surfaces, such as in constructing nonfouling self-assembled monolayers11,14−16 and polymer films.17 These successful applications underscore the connection between surface amino acid characteristics and their functional roles.14 Hence, examining surface features contributes to a better understanding of protein behaviors and potential biomaterial designs.

Machine learning has emerged as a powerful tool for managing large-scale data and identifying statistical relationships, showcasing immense potential in uncovering hidden rules from data, particularly in protein research. The continuous updates and publication of protein databases, such as UniProt,18 AlphaFold,19 Gene Ontology,20 and the Protein Data Bank,21 provide a wealth of reliable resources like annotations, protein sequences, and structures, significantly enhancing the application of machine learning by allowing the incorporation of various protein features as input for diverse models.22,23 Protein-based machine learning models are two major types. One is the deep learning models utilizing protein sequences. These models aim to interpret protein behaviors, such as folding and localization, from the most fundamental genetic features.24−31 Though these models gain prediction results with high accuracies, they often need more details of the underlying mechanisms of prediction targets from the aspect of body–body interactions. While researchers have tried to demonstrate results by analyzing the weighting of sequence segments,22,32 the findings have limited universality in biomaterial design since applying them to nonprotein entities proves challenging. The other type of model treats proteins as well-folded entities and explains protein characteristics based on established structures. These models analyze protein features such as protein images and surface features to illustrate protein characters from a more macroscopic perspective,33−38 enhancing the universality of the results and pushing the outcome more to the aspect of applications.

In this study, we employed machine learning to differentiate between secreted (extracellular) and cytosolic (intracellular) proteins by analyzing protein surface amino acid populations and secondary structure ratios. Additionally, we explained the feature importance by model interpretation and referred to general protein surface properties, aiming to establish a guideline for biomaterial design.

2 Methodology

2.1 Data Set of Proteins

We constructed our data set using two primary resources: UniProt and AlphaFold. UniProt provides species tags, descriptions of protein subcellular locations, and references to protein 3D structures. We opted for the predicted structural data from AlphaFold rather than the experimental data from PDB due to PDB’s limited data set size and numerous incomplete structures (e.g., major urinary protein 4, PDB ID: 3KFF, AlphaFold ID: AF-P11590-F1). AlphaFold provides more highly accurate predictions for most proteins than the databases, as researchers believe the system has effectively learned the fundamental principles of protein nature.39,40 Furthermore, AlphaFold evaluates the confidence of each predicted region using a score, wherein lower scores serve as reliable indicators of intrinsically disordered regions.21

Our data set comprises 708 proteins. We initially filtered protein files on UniProt based on the following criteria: (1) a reviewed file; (2) a clear description of the subcellular location, either’secreted’ or’cytosol’; (3) a corresponding 3D structure available on AlphaFold. The protein samples included proteins from humans, mice, and rats to ensure a sufficiently large data set. We tested models at each stage using the new data set to maintain homogeneity across the three data sets. The expanding mouse and rat protein tests demonstrated consistent accuracy, and model performance steadily improved. For each parameter set per training round, we allocated 150 iterations to divide the data set into 70% training data and 30% test data for model training. We balanced secreted and cytosolic proteins in the training data set by separately assigning random states to the protein of each tag to achieve the exact data size (248 cytosol and 248 secreted proteins). The rest of the data set was then used for testing. An extra data set of 106 proteins was also constructed following the same criteria to validate the performance of the best-trained model.

2.2 Extract Surface Residues

We identify surface amino acids by relative solvent-accessible surface area (SASA). Relative SASA is defined as the real SASA divided by the maximum SASA, which refers to the total exposure of the residue of interest when all other residues except for its two neighbors are removed. We first calculated the relative SASA for all residues in the protein. Then, we select residues with a relative SASA higher than or equal to 0.3 as surface residues (Figure 1), and residue information is available in the downstream profiles from the indexes we selected.

Figure 1 Extract surface residues by the relative solvent-accessible surface area.

2.3 Parameters to Describe Surfaces of Protein Molecules

We considered the following factors: (1) surface compositions of the 20 amino acids and groups of amino acids with similar properties; (2) surface populations of functional groups; (3) surface and overall compositions of protein secondary structures; and (4) ratios of surface residues to overall residues.

We identified 18 descriptors by checking the Pearson correlation map and the pruning process during model training (Figure 2). The correlation map was employed to identify redundant features. Pairs with a correlation value above 0.85 were selected, and we removed the pairs with a higher cumulative correlation value in each pair. In the pruning process, we did feature importance analysis and deleted features with an importance distribution of less than 2%.

Figure 2 Pearson correlation map of 34 descriptors.

2.4 Model Selection

Our purposes require the model to reach a high prediction accuracy while maintaining interpretability, and thus, we selected random forest (RF).

We compared the performance of RF with four other commonly used algorithms; each has unique features: Artificial neural networks (ANNs) and logistic regression (LR) operate using gradient descent; K-nearest neighbor (KNN) and support vector machines (SVM) rely on distance calculations; the RF algorithm comprises decision trees. We collected performance data for each algorithm by running 150 random states on training-test splitting, and RF demonstrated superior performance (Figure 3). Besides, RF has four main features: (1) based on fully grown decision trees; (2) bootstrapped data for training each tree; (3) majority voting for the final result; (4) assigning random subsets of features for each decision tree.41,42 These features make RF advantageous in binary classification problems due to its reduced susceptibility to overfitting and robustness.43−46 It also enables feature importance ranking, calculated from the decrease in the Gini index for each feature, which is crucial for identifying predominant features.42 The other four algorithms necessitate feature scaling before training, which strictly requires homogeneity between targets and samples from the training data set during the application, while RF does not. This feature allows RF for more flexible applications and direct correlations with the original values of features.

Figure 3 Comparison between the five algorithms. Test accuracy reveals their performances on test data sets, while probability density value indicates the likelihood of model performance to reach the corresponding accuracy.

3 Results and Discussion

3.1 Comparison Between Surface and Global Features

We compared surface and global features across four categories of residue populations: hydrophilic, hydrophobic, positively charged, and negatively charged residues. The category of hydrophilic residues encompasses positively charged residues (Arg, His, and Lys), negatively charged residues (Asp and Glu), and polar but uncharged residues (Ser, The, Asn, and Gln). The hydrophobic residue category includes Ala, Ile, Leu, Met, Phe, Trp, Tyr, and Val.

Figure 4 shows the population of four collections of residues from secreted and cytosol proteins. Global data generally provide poorer distinctions between secreted and cytosolic proteins, underlining the prominence of surface descriptors over global ones in this binary protein classification problem.

Figure 4 Histograms of the population of four-type residues of secreted (orange) and cytosol (blue) proteins: (a) hydrophilic, (b) hydrophobic, (c) positively charged, and (d) negatively charged. The p values are indicated.

3.2 Model Performance

Confusion matrices and receiver operating characteristic (ROC) curves demonstrate the best-trained model’s capability and robustness to distinguish between positive (secreted proteins) and negative (cytosol proteins) groups (Figure 5a–d). Calculating based on the precision and recall, the model gains an f1 score of 0.935 on the test data set (Table 1). Though it does not reflect the performance on all events, as cytosol and secreted proteins do not hold a definite relationship with negative and positive tags, the f1 score remains as high as 0.906 on the balanced validation data set, indicating the promising performance of the classifier. Besides, the feature importance analysis displayed similar rankings from seven well-trained RF models with an average accuracy of 93.03% (Figure 5b). Among the 18 descriptors, surface compositions of Glu, Cys, and Leu consistently emerged as the leading contributors to the prediction.

Figure 5 (a) Confusion matrix of the best-trained model on the test data set. (b) Confusion matrix of the model on the validation data set. (c) ROC curve of the model on the test data set. (d) ROC curve of the model on the validation data set. (e) Feature importance ranking from the top-seven RF models.

Table 1 Threshold Values of Surface Compositions of Glu, Cys, and Leu to be Recognized as Secreted Proteins

 	precision	recall	
test data set	0.939	0.930	
validation data set	0.906	0.906	

We discovered relationships between features’ contributions and their values by analyzing the components of several proteins’ probabilities for being classified as secreted or cytosolic proteins. Consequently, we plotted the features’ contributions against their values (Figure 6). Horizontal lines at zero contribution separate data points contributing to secreted and cytosolic proteins, while vertical lines effectively divide the two groups and identify boundary values (Table 2). We found a sigmoid-like distribution pattern of surface Glu, indicating a nonmiscible boundary. Data points from surface Glu to surface Leu gradually aggregate toward the boundary, manifesting the descending importance distribution.

Figure 6 Plots of the contribution of surface (a) Glu, (b) Cys, and (c) Leu compositions toward the chance to be recognized as a cytosol protein in the machine learning.

Table 2 Threshold Values of Surface Compositions of Glu, Cys, and Leu to be Recognized as Secreted Proteins

 	surface Glu	surface Cys	surface Leu	
secreted protein	<9.0%	>1.8%	>5.8%	

3.3 Importance Analysis of Surface Glutamic Acid

Glu and Asp are two amino acids with negatively charged side chains. Their structures only differ by an additional carbon on Glu’s side chain, resulting in similar chemical and physical properties. However, their importance varies significantly. Glu is the top-ranked feature, while Asp was removed during feature pruning due to its low importance. Although they share many similarities, the one-carbon difference influences their preferences for secondary structures in long sequences. Alpha-helices form through hydrogen bonds between the side chains of every first and fourth residue.47 While Glu’s structure accommodates this arrangement well, one less carbon in Asp’s structure makes it less favorable for α helix formation.48,49

In this manner, the abundance of surface Glu should correspond to a similar trend in the protein secondary structure. Referring to the correlation map, the alpha-helix and beta-sheet compositions are strongly negatively related. The overall beta-sheet composition plot reveals an expected correlation between the abundance of Glu and deficient surface beta-sheet (Figure 7). Exposing the beta-sheet to the surface promotes aggregation behavior, suggesting that the abundance of Glu could prevent protein aggregation.50

Figure 7 Overall beta-sheet composition distribution of secreted and cytosol proteins.

Nonetheless, the strong positive correlation between overall and surface secondary structure compositions indicates that alpha-helix and beta-sheet structures do not significantly prefer exposure or burial. Thus, although we have rationalized the relationship between secondary structure and the surface composition of Glu, it may only be a minor contributor to Glu’s importance. The primary importance of Glu’s surface composition arises from its contributions to protein surface hydrophilicity and surface charging state.

Glu predominates the population of negatively charged residues for both secreted and cytosolic proteins, where cytosolic proteins have slightly higher Glu compositions (Figure 8a,b). However, such a minor difference still contributes to the discrepancy in protein stability. Glu and Asp differ in their side chain conformational entropy, and substituting Asp with Glu better stabilizes protein conformation from a free energy perspective.51 Additionally, previous research reported that an increase in protein denaturation midpoint originates from such substitutions.52 Considering the overall surface residues, cytosolic proteins have higher Asp and Glu surface compositions but exhibit a more significant increase in Glu. Therefore, Glu generally represents a larger population on cytosolic protein surfaces, promoting specific interactions in a crowded environment.53

Figure 8 Summary of surface negatively charged residue composition of (a) secreted and (b) cytosol proteins. Surface (c) Asp and (d) Glu distributions of secreted and cytosol proteins.

In addition to a greater composition of negatively charged residues on the surface, cytosolic proteins also expose more positively charged residues (Figure 9a,b). This increase in positively charged residues is related to the more negative intracellular electrostatic potential caused by the imbalanced charge separation on the two sides of the cell membrane.54,55 However, the increase in surface negatively charged residues is more substantial. Consequently, considering the populational difference between the two types of charged residues (Figure 9c), the overall charges on cytosolic proteins’ surfaces are more neutral. Surface neutrality helps prevent trappings by negatively charged membranes.56−59 Thus, the abundance of Glu aids in blocking membrane adhesion and maintaining protein mobility.

Figure 9 Histogram of secreted and cytosol proteins’ numbers by compositions of (a) surface positively (+) and (b) surface negatively (−) charged residues. (c) Histogram of secreted and cytosol proteins’ numbers by their (surface (+) charged residue composition-surface (−) charged residue composition).

We also observed an increase in cytosolic proteins’ surface polar residue composition (Figure 10c). Simultaneously, the composition analysis of surface polar residues reveals that uncharged polar residues have a lower weight in cytosolic proteins, while negatively charged residues have a higher weight (Figure 10a,b). However, the surface compositions of uncharged polar residues remain similar in both secreted and cytosolic proteins (Figure 10d), indicating that the negatively charged residue is the primary factor in the shift in polar residue composition. Such hydrophilicity enables the cytosolic protein surface to interact more effectively with aqueous environments and aids in preventing aggregation under crowded cellular conditions.

Figure 10 Summary of surface polar residue compositions of (a) secreted and (b) cytosol proteins. Histogram of secreted and cytosol proteins’ numbers by compositions of (c) surface polar and (d) surface nonpolar residues.

3.4 Importance Analysis of Surface Cysteine

Cysteine, characterized by the large atomic radius of sulfur and the low dissociation energy of the S–H bond, is unique among amino acids due to its redox-active function and exceptional nucleophilicity.60 Its ionization state is highly susceptible to changes in its immediate chemical milieu. Subtle fluctuations in the reduction potential can induce notable differences in the equilibrium between its dithiol groups and disulfide bonds.61,62 Compared to secreted proteins, proteins within the cytosol exhibit a diminished concentration of surface cysteine, as indicated in Figure 11a, and a less pronounced inclination to present cysteine on the surface, as shown in Figure 11b. This observation aligns well with the differing reduction potentials and protein functions inside and outside the cell.

Figure 11 (a) Histogram of secreted and cytosol proteins’ numbers by compositions of surface Cys composition distributions of secreted and cytosol proteins. (b) Histogram of secreted and cytosol proteins’ numbers by the ratio between the numbers of surface Cys and global Cys residues.

Inside the cell, the chemical environment is reducing, with a potential ranging from −220 to 260 mV, while outside the cell, it is oxidizing at approximately −140 mV.63 As such, except in certain hyperthermophilic archaea, intracellular disulfide bonds are scarce and not fundamentally critical for protein structures.64 Surface cysteine on cytosol proteins predominantly serves roles in redox-sensitive regulation and protection against reactive oxygen species (ROS), given that it exists predominantly in the reduced form.65,66 Conversely, in response to the oxidizing extracellular environment, cysteines form intramolecular disulfide bonds to bolster protein structure integrity and are displayed on the surface to stabilize proteins under harsh environmental conditions.60,67 Additionally, surface cysteines aggregate at the active sites of enzymes that catalyze redox processes.68 Consequently, the population difference on surface Cys is both protein function- and chemical environment-related one.

3.5 Importance Analysis of Surface Leucine

Leu is one of the hydrophobic amino acids, closely resembling Ile and Val. They differ in the length and arrangement of their side chain carbons, which, as mentioned earlier, naturally leads to different preferences for secondary structures. Leu is the second-best preferer for α helix; Val and Ile are the best and second-best preferers, respectively.49

Cytosol proteins prefer fewer Leu and Ile residues on the surface, while showing no preference for Val (Figure 12). As mentioned earlier, secondary structural factors are minor contributors to Glu and Asp. The surface compositions of Leu, Val, and Ile further support this idea. Leu and Ile have similar hydrophobicity across various scales, while Val has the weakest hydrophobicity.69−73 If secondary structure preference strongly relates to residues’ positions, we expect more significant trends of Val and Ile than Leu to be located in the core. However, the distribution plots do not reveal that.

Figure 12 Histogram of secreted and cytosol proteins’ numbers by compositions of (a) surface Leu, (b) surface Val, and (c) surface Ile residues.

Leu also demonstrates the predominant population in surface and global hydrophobic compositions (Figure 13a,b,d, and e). Compared to secreted proteins, cytosolic proteins have, on average, 3% fewer surface hydrophobic residues, mainly due to the inward shift of Leu from the surface, as Leu’s overall population remains similar for secreted and cytosolic proteins (Figure 13c). The surface-global ratio further illustrates the trend of cytosolic proteins to hide Leu from the surface (Figure 13f), creating a stable hydrophobic core in cytosolic proteins.

Figure 13 (a) Surface residue composition of hydrophilic (charged and polar) and hydrophobic residues for (a) secreted and (b) cytosol proteins. Compositions of surface hydrophobic residues of (c) secreted and (d) cytosol proteins. Histogram of secreted and cytosol proteins’ numbers by compositions of (e) global and (f) surface Leu residues.

3.6 Double-Tagged Proteins

Besides secreted-only and cytosol-only proteins, some proteins can exist in both environments according to scenarios. Due to the deficiency in data size, we picked 79 double-tagged proteins and compared them with cytosol and secreted proteins on the main collections and most important features. Interestingly, these double-tagged proteins balance out each distinct feature of both kinds (Figure 14). We can roughly tell that their peaks sit between the main peaks of the other two kinds, slightly leaning toward either cytosol or secreted proteins. The comparison reveals that the surface chemical structures of the double-tagged proteins are optimized to accept both intra- and extra-cellular environments.

Figure 14 Comparison among secreted, cytosol, and double-tagged proteins on surface features: (a) Surface polar residue composition. (b) Surface polar uncharged residue composition. (c) Surface positively charged residue composition. (d) Surface negatively charged residue composition. (e) Surface cysteine composition. (f) Surface glutamic acid composition. (g) Surface leucine composition.

4 Conclusions

Using an interpretable RF model and protein surface features, we successfully classified proteins into secreted or cytosol ones. The model provided quantitative references for protein surface amino acid populations.

The model revealed that surface compositions of glutamic acid, cysteine, and leucine were the three most essential features and provided quantitative threshold values for each of them. Further analysis demonstrated the preferences of both secreted and cytosolic proteins for these three features and quantitative boundary values. Interpretations of the features’ importance explained proteins’ strategies to construct compatible surfaces in two different environments. Glutamic acid gains importance by playing a critical role in tuning the protein surface charge and assisting in the construction of a water barrier through polar interactions. The population of surface cysteines relates to the formation of disulfide bonds and, further, antioxidant aspects. The abundance of cysteine on the surfaces of secreted proteins serves as the stabilization strategy in the oxidizing environments and triggers protein functions. Leucine predominates in the change in surface hydrophobic residue composition. In cytosolic proteins, leucine’s inward shifting constructs a more hydrophilic surface and a more hydrophobic core, which is energetically preferred. Generally, protein stability in a crowded environment requires (1) surface neutrality, (2) exposure of hydrophilic residues while hiding hydrophobic residues, and (3) protection from potential oxidative damage.

The model analysis provides the quantitative threshold values of surface amino acid compositions, offering insights for nonfouling surface design. Besides, the results provide a unique perspective on proteins’ in vivo functioning, contributing to a better understanding of protein nature and potentially assisting in protein engineering endeavors. Though surface feature analysis shows promise in this study, other protein features, such as conformational descriptors and intrinsically disordered sections, still require careful examination. Future work can supplement these aspects and evaluate the rules on various materials and platforms by directly utilizing amino acids or implementing alternative modifications to achieve similar effects.

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jpcb.4c02461.Plots showcasing the abilities of other surface features on separating cytosol and secreted proteins, other features include surface-global residue ratio, hydrophobic residues composition, polar residue composition, positively charged residue composition, carbonyl group population, glycine composition, isoleucine composition, lysine composition, serine composition, glutamine composition, arginine composition, asparagine composition, tryptophan composition, and threonine composition, and tables listing proteins included in the analysis (PDF)

Supplementary Material

jp4c02461_si_001.pdf

The authors declare no competing financial interest.

Acknowledgments

We acknowledge Ms. Kazue Taki for her help in arranging this project. This work was supported by JSPS KAKENHI Grant Numbers (JP23H04059 and JP22H04530). It was performed under the Research Program for CORE lab of “Five-star Alliance” in “NJRC Mater. & Dev.” This work was also supported by the Japan−Taiwan Exchange Association.
==== Refs
References

Donnes P. ; Hoglund A. Predicting protein subcellular localization: past, present, and future. Genomics, Proteomics Bioinf. 2004, 2 (4 ), 209–215. 10.1016/S1672-0229(04)02027-3.
Chou K. C. ; Elrod D. W. Protein subcellular location prediction. Protein Eng. 1999, 12 (2 ), 107–118. 10.1093/protein/12.2.107.10195282
Sobczynski D. J. ; Fish M. B. ; Fromen C. A. ; Carasco-Teja M. ; Coleman R. M. ; Eniola-Adefeso O. Drug carrier interaction with blood: a critical aspect for high-efficient vascular-targeted drug delivery systems. Ther Deliv. 2015, 6 (8 ), 915–934. 10.4155/TDE.15.38.26272334
Kumar S. ; Nussinov R. Close-range electrostatic interactions in proteins. ChemBiochem 2002, 3 (7 ), 604–617. 10.1002/1439-7633(20020703)3:7<604::AID-CBIC604>3.0.CO;2-X.12324994
Levy E. D. ; De S. ; Teichmann S. A. Cellular crowding imposes global constraints on the chemistry and evolution of proteomes. Proc. Natl. Acad. Sci. U. S. A. 2012, 109 (50 ), 20461–20466. 10.1073/pnas.1209312109.23184996
Lingwood D. ; Simons K. Lipid rafts as a membrane-organizing principle. Science 2010, 327 (5961 ), 46–50. 10.1126/science.1174621.20044567
Luby-Phelps K. Cytoarchitecture and physical properties of cytoplasm: volume, viscosity, diffusion, intracellular surface area. Int. Rev. Cytol 1999, 192 , 189–221. 10.1016/S0074-7696(08)60527-6.
Hynes R. O. Integrins: bidirectional, allosteric signaling machines. Cell 2002, 110 (6 ), 673–687. 10.1016/s0092-8674(02)00971-6.12297042
Leitinger B. Discoidin domain receptor functions in physiological and pathological conditions. Int. Rev. Cell Mol. Biol. 2014, 310 , 39–87. 10.1016/B978-0-12-800180-6.00002-5.24725424
Andrade M. A. ; O’Donoghue S. I. ; Rost B. Adaptation of protein surfaces to subcellular location. J. Mol. Biol. 1998, 276 (2 ), 517–525. 10.1006/jmbi.1997.1498.9512720
White A. D. ; Nowinski A. K. ; Huang W. ; Keefe A. J. ; Sun F. ; Jiang S. Decoding nonspecific interactions from nature. Chem. Sci. 2012, 3 (12 ), 3488 10.1039/c2sc21135a.
Jones S. ; Thornton J. M. Principles of protein-protein interactions. Proc. Natl. Acad. Sci. U. S. A. 1996, 93 (1 ), 13–20. 10.1073/pnas.93.1.13.8552589
Kwaria R. J. ; Mondarte E. A. Q. ; Tahara H. ; Chang R. ; Hayashi T. Data-Driven Prediction of Protein Adsorption on Self-Assembled Monolayers toward Material Screening and Design. ACS Biomater. Sci. Eng. 2020, 6 (9 ), 4949–4956. 10.1021/acsbiomaterials.0c01008.33455289
Chang R. ; Mondarte E.A.Q. ; Palai D. ; Sekine T. ; Kashiwazaki A. ; Murakami D. ; Tanaka M. ; Hayashi T. ; et al. Protein- and Cell-Resistance of Zwitterionic Peptide-Based Self-Assembled Monolayers: Anti-Biofouling Tests and Surface Force Analysis. Front. Chem. 2021, 9 , 748017 10.3389/fchem.2021.748017.34692644
Chen S. ; Cao Z. ; Jiang S. Ultra-low fouling peptide surfaces derived from natural amino acids. Biomaterials 2009, 30 (29 ), 5892–5896. 10.1016/j.biomaterials.2009.07.001.19631374
Pinazo A. ; Pons R. ; Pérez L. ; Infante M. R. Amino Acids as Raw Material for Biocompatible Surfactants. Ind. Eng. Chem. Res. 2011, 50 (9 ), 4805–4817. 10.1021/ie1014348.
Palai D. ; Tahara H. ; Chikami S. ; Latag G. V. ; Maeda S. ; Komura C. ; Kurioka H. ; Hayashi T. ; et al. Prediction of Serum Adsorption onto Polymer Brush Films by Machine Learning. ACS Biomater. Sci. Eng. 2022, 8 (9 ), 3765–3772. 10.1021/acsbiomaterials.2c00441.35905395
UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res. 2021, 49 (D1 ), D480–D489. 10.1093/nar/gkaa1100.33237286
Jumper J. ; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596 (7873 ), 583–589. 10.1038/s41586-021-03819-2.34265844
The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res. 2021, 49 (D1 ), D325–D334. 10.1093/nar/gkaa1113.33290552
Tunyasuvunakool K. ; et al. Highly accurate protein structure prediction for the human proteome. Nature 2021, 596 (7873 ), 590–596. 10.1038/s41586-021-03828-1.34293799
Jiang Y. ; Wang D. ; Wang W. ; Xu D. Computational methods for protein localization prediction. Comput. Struct. Biotechnol. J. 2021, 19 , 5834–5844. 10.1016/j.csbj.2021.10.023.34765098
Nakai K. ; Wei L. Recent Advances in the Prediction of Subcellular Localization of Proteins and Related Topics. Front. Bioinform. 2022, 2 , 910531 10.3389/fbinf.2022.910531.36304291
Armenteros J. J. A. ; Sonderby C. K. ; Sonderby S. K. ; Nielsen H. ; Winther O. DeepLoc: Prediction of protein subcellular localization using deep learning. Bioinformatics 2017, 33 (21 ), 4049 10.1093/bioinformatics/btx431.29028934
Hua S. ; Sun Z. Support vector machine approach for protein subcellular localization prediction. Bioinformatics 2001, 17 (8 ), 721–728. 10.1093/bioinformatics/17.8.721.11524373
Nakashima H. ; Nishikawa K. Discrimination of intracellular and extracellular proteins using amino acid composition and residue-pair frequencies. J. Mol. Biol. 1994, 238 (1 ), 54–61. 10.1006/jmbi.1994.1267.8145256
Sahu S. S. ; Loaiza C. D. ; Kaundal R. Plant-mSubP: a computational framework for the prediction of single- and multi-target protein subcellular localization using integrated machine-learning approaches. AoB Plants 2020, 12 (3 ), lz068 10.1093/aobpla/plz068.
Thumuluri V. ; Almagro Armenteros J. J. ; Johansen A. R. ; Nielsen H. ; Winther O. DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Res. 2022, 50 (W1 ), W228–W234. 10.1093/nar/gkac278.35489069
Wan S. ; Duan Y. ; Zou Q. HPSLPred: An Ensemble Multi-Label Classifier for Human Protein Subcellular Location Prediction with Imbalanced Source. Proteomics 2017, 17 (17–18 ), 1700262 10.1002/pmic.201700262.
Yu C.-S. ; Chen Y.-C. ; Lu C.-H. ; Hwang J. K. Prediction of protein subcellular localization. Proteins 2006, 64 (3 ), 643–651. 10.1002/prot.21018.16752418
Brandes N. ; Ofer D. ; Peleg Y. ; Rappoport N. ; Linial M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 2022, 38 (8 ), 2102–2110. 10.1093/bioinformatics/btac020.35020807
Almagro Armenteros J. J. ; Salvatore M. ; Emanuelsson O. ; Winther O. ; von Heijne G. ; Elofsson A. ; Nielsen H. Detecting sequence signals in targeting peptides using deep learning Life Sci. Alliance 2019 2 , (5 ), , 10.26508/lsa.201900429.
Kobayashi H. ; Cheveralls K. C. ; Leonetti M. D. ; Royer L. A. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nat. Methods 2022, 19 (8 ), 995–1003. 10.1038/s41592-022-01541-z.35879608
Tahir M. ; Idris A. MD-LBP: An Efficient Computational Model for Protein Subcellular Localization from HeLa Cell Lines Using SVM. Curr. Bioinf. 2020, 15 (3 ), 204–211. 10.2174/1574893614666190723120716.
Xu Y.-Y. ; Yao L.-X. ; Shen H.-B. Bioimage-based protein subcellular location prediction: a comprehensive review. Front. Comput. Sci. 2018, 12 (1 ), 26–39. 10.1007/s11704-016-6309-5.
Xu Y.-Y. ; Shen H.-B. ; Murphy R. F. Learning complex subcellular distribution patterns of proteins via analysis of immunohistochemistry images. Bioinformatics 2020, 36 (6 ), 1908–1914. 10.1093/bioinformatics/btz844.31722369
Yang F. ; Liu Y. ; Wang Y. ; Yin Z. ; Yang Z. MIC_Locator: a novel image-based protein subcellular location multi-label prediction model based on multi-scale monogenic signal representation and intensity encoding strategy. BMC Bioinf. 2019, 20 (1 ), 522 10.1186/s12859-019-3136-3.
Mylonas S. K. ; Axenopoulos A. ; Daras P. DeepSurf: a surface-based deep learning approach for the prediction of ligand binding sites on proteins. Bioinformatics 2021, 37 (12 ), 1681–1690. 10.1093/bioinformatics/btab009.33471069
Jumper J. ; Hassabis D. Protein structure predictions to atomic accuracy with AlphaFold. Nat. Methods 2022, 19 (1 ), 11–12. 10.1038/s41592-021-01362-6.35017726
Pereira J. ; Simpkin A. J. ; Hartmann M. D. ; Rigden D. J. ; Keegan R. M. ; Lupas A. N. High-accuracy protein structure prediction in CASP14. Proteins 2021, 89 (12 ), 1687–1699. 10.1002/prot.26171.34218458
Breiman L. Random Forests. Mach. Learn 2001, 45 (1 ), 5–32. 10.1023/A:1010933404324.
Fawagreh K. ; Gaber M. M. ; Elyan E. Random forests: from early developments to recent advancements. Syst. Sci. Control Eng 2014, 2 (1 ), 602–609. 10.1080/21642583.2014.956265.
Boinee P. ; De Angelis A. ; Foresti G. L. Meta random forests. Int. J. Comput. Intell 2005, 2 (3 ), 138–147.
Couronné R. ; Probst P. ; Boulesteix A.-L. Random forest versus logistic regression: a large-scale benchmark experiment. BMC Bioinf. 2018, 19 (1 ), 270 10.1186/s12859-018-2264-5.
Liaw A. ; Wiener M. Classification and regression by randomForest. R News 2002, 2 (3 ), 18–22.
Robnik-Šikonja M. ″Improving Random Forests″; Springer: Berlin, Heidelberg, 2004; pp. 359 370.
Börner H. G. ; Lutz J. F. 6.15 - Synthetic–Biological Hybrid Polymers: Synthetic Designs, Properties, and Applications. In Polymer Science: A Comprehensive Reference, Matyjaszewski K. ; Möller M. , Eds.; Elsevier: Amsterdam, 2012; pp. 543 586.
Malkov S. ; Zivkovic M. V. ; Beljanski M. V. ; Zaric S. D. ″Correlations of amino acids with secondary structure types: connection with amino acid structure″. arXiv 2005, 10.48550/arXiv.q-bio/050504.
Malkov S. N. ; Zivkovic M. V. ; Beljanski M. V. ; Hall M. B. ; Zaric S. D. A reexamination of the propensities of amino acids towards a particular secondary structure: classification of amino acids based on their chemical structure. J. Mol. Model. 2008, 14 (8 ), 769–775. 10.1007/s00894-008-0313-0.18504624
Bratko D. ; Blanch H. W. Effect of secondary structure on protein aggregation: A replica exchange simulation study. J. Chem. Phys. 2003, 118 (11 ), 5185–5194. 10.1063/1.1546429.
Doig A. J. ; Sternberg M. J. Side-chain conformational entropy in protein folding. Protein Sci. 1995, 4 (11 ), 2247–2251. 10.1002/pro.5560041101.8563620
Lee D. Y. ; Kim K.-A. ; Yu Y. G. ; Kim K.-S. Substitution of aspartic acid with glutamic acid increases the unfolding transition temperature of a protein. Biochem. Biophys. Res. Commun. 2004, 320 (3 ), 900–906. 10.1016/j.bbrc.2004.06.031.15240133
Speer S. L. ; Zheng W. ; Jiang X. ; Chu I.-T. ; Guseman A. J. ; Liu M. ; Pielak G. J. ; Li C. The intracellular environment affects protein-protein interactions. Proc. Natl. Acad. Sci. U. S. A. 2021, 118 (11 ), e2019918118 10.1073/pnas.2019918118.33836588
Chan P. ; Curtis R. A. ; Warwicker J. Soluble expression of proteins correlates with a lack of positively-charged surface. Sci. Rep. 2013, 3 , 3333 10.1038/srep03333.24276756
Honig B. H. ; Hubbell W. L. ; Flewelling R. F. Electrostatic interactions in membranes and proteins. Annu. Rev. Biophys. Biophys. Chem. 1986, 15 , 163–193. 10.1146/annurev.bb.15.060186.001115.2424473
Forest V. ; Cottier M. ; Pourchez J. Electrostatic interactions favor the binding of positive nanoparticles on cells: A reductive theory. Nano Today 2015, 10 (6 ), 677–680. 10.1016/j.nantod.2015.07.002.
Li Z.-L. ; Ding H.-M. ; Ma Y.-q. Interaction of peptides with cell membranes: insights from molecular modeling. J. Phys.: Condens. Matter. 2016, 28 (8 ), 083001 10.1088/0953-8984/28/8/083001.26828575
Ma Y. ; Poole K. ; Goyette J. ; Gaus K. Introducing Membrane Charge and Membrane Potential to T Cell Signaling. Front. Immunol. 2017, 8 , 1513 10.3389/fimmu.2017.01513.29170669
Schavemaker P. E. ; Smigiel W. M. ; Poolman B. Ribosome surface properties may impose limits on the nature of the cytoplasmic proteome. Elife 2017, 6 , e30084 10.7554/eLife.30084.29154755
Pace N. J. ; Weerapana E. Diverse functional roles of reactive cysteines. ACS Chem. Biol. 2013, 8 (2 ), 283–296. 10.1021/cb3005269.23163700
Harris T. K. ; Turner G. J. Structural basis of perturbed pKa values of catalytic groups in enzyme active sites. IUBMB Life 2002, 53 (2 ), 85–98. 10.1080/15216540211468.12049200
Banerjee R. Redox outside the box: linking extracellular redox remodeling with intracellular redox metabolism. J. Biol. Chem. 2012, 287 (7 ), 4397–4402. 10.1074/jbc.R111.287995.22147695
Trivedi M. V. ; Laurence J. S. ; Siahaan T. J. The role of thiols and disulfides on protein stability. Curr. Protein Pept. Sci. 2009, 10 (6 ), 614–625. 10.2174/138920309789630534.19538140
Mallick P. ; Boutz D. R. ; Eisenberg D. ; Yeates T. O. Genomic evidence that the intracellular proteins of archaeal microbes contain disulfide bonds. Proc. Natl. Acad. Sci. U. S. A. 2002, 99 (15 ), 9679–9684. 10.1073/pnas.142310499.12107280
Barford D. The role of cysteine residues as redox-sensitive regulatory switches. Curr. Opin. Struct. Biol. 2004, 14 (6 ), 679–686. 10.1016/j.sbi.2004.09.012.15582391
Requejo R. ; Hurd T. R. ; Costa N. J. ; Murphy M. P. Cysteine residues exposed on protein surfaces are the dominant intramitochondrial thiol and may protect against oxidative damage. Febs J 2010, 277 (6 ), 1465–1480. 10.1111/j.1742-4658.2010.07576.x.20148960
Thornton J. M. Disulphide bridges in globular proteins. J. Mol. Biol. 1981, 151 (2 ), 261–287. 10.1016/0022-2836(81)90515-5.7338898
Prinz W. A. ; Aslund F. ; Holmgren A. ; Beckwith J. The role of the thioredoxin and glutaredoxin pathways in reducing protein disulfide bonds in the Escherichia coli cytoplasm. J. Biol. Chem. 1997, 272 (25 ), 15661–15667. 10.1074/jbc.272.25.15661.9188456
Hessa T. ; Bihlmaier K. ; Lundin C. ; Boekel J. ; Nilsson I. ; White S. H. ; von Heijne G. ; et al. Recognition of transmembrane helices by the endoplasmic reticulum translocon. Nature 2005, 433 (7024 ), 377–381. 10.1038/nature03216.15674282
Kyte J. ; Doolittle R. F. A simple method for displaying the hydropathic character of a protein. J. Mol. Biol. 1982, 157 (1 ), 105–132. 10.1016/0022-2836(82)90515-0.7108955
Moon C. P. ; Fleming K. G. Side-chain hydrophobicity scale derived from transmembrane protein folding into lipid bilayers. Proc. Natl. Acad. Sci. U. S. A. 2011, 108 (25 ), 10174–10177. 10.1073/pnas.1103979108.21606332
Wimley W. C. ; White S. H. Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nat. Struct. Biol 1996, 3 (10 ), 842–848. 10.1038/nsb1096-842.8836100
Zhao G. ; London E. An amino acid ″transmembrane tendency″ scale that approaches the theoretical limit to accuracy for prediction of transmembrane helices: relationship to biological hydrophobicity. Protein Sci 2006, 15 (8 ), 1987–2001. 10.1110/ps.062286306.16877712
