
==== Front
J Chem Inf Model
J Chem Inf Model
ci
jcisd8
Journal of Chemical Information and Modeling
1549-9596
1549-960X
American Chemical Society

39099394
10.1021/acs.jcim.4c00809
Article
Streamlining NMR Chemical Shift Predictions for Intrinsically Disordered Proteins: Design of Ensembles with Dimensionality Reduction and Clustering
https://orcid.org/0000-0002-4523-5570
Bakker Michael J. †
Gaffour Amina †
https://orcid.org/0000-0002-1890-9082
Juhás Martin †‡
Zapletal Vojtěch †
Stošek Jakub †¶
https://orcid.org/0000-0002-3565-5926
Bratholm Lars A. §
https://orcid.org/0000-0003-0844-2666
Pavlíková Přecechtělová Jana *†
† Faculty of Pharmacy in Hradec Králové, Charles University, Akademika Heyrovského 1203/8, 500 05 Hradec Králové, Czech Republic
‡ Department of Chemistry, Faculty of Science, University of Hradec Králové, Rokitanského 62, 500 03 Hradec Králové, Czech Republic
¶ Department of Chemistry, Faculty of Science, Masaryk University, Kotlářská 2, 611 37 Brno, Czech Republic
§ School of Chemistry, University of Bristol, Cantock’s Close, BS8 1TS Bristol, U.K.
* E-mail: precechj@faf.cuni.cz Phone: +420 495 067 488.
05 08 2024
26 08 2024
64 16 65426556
09 05 2024
25 07 2024
24 07 2024
© 2024 The Authors. Published by American Chemical Society
2024
The Authors
https://creativecommons.org/licenses/by/4.0/ Permits the broadest form of re-use including for commercial purposes, provided that author attribution and integrity are maintained (https://creativecommons.org/licenses/by/4.0/).

By merging advanced dimensionality reduction (DR) and clustering algorithm (CA) techniques, our study advances the sampling procedure for predicting NMR chemical shifts (CS) in intrinsically disordered proteins (IDPs), making a significant leap forward in the field of protein analysis/modeling. We enhance NMR CS sampling by generating clustered ensembles that accurately reflect the different properties and phenomena encapsulated by the IDP trajectories. This investigation critically assessed different rapid CS predictors, both neural network (e.g., Sparta+ and ShiftX2) and database-driven (ProCS-15), and highlighted the need for more advanced quantum calculations and the subsequent need for more tractable-sized conformational ensembles. Although neural network CS predictors outperformed ProCS-15 for all atoms, all tools showed poor agreement with HN CSs, and the neural network CS predictors were unable to capture the influence of phosphorylated residues, highly relevant for IDPs. This study also addressed the limitations of using direct clustering with collective variables, such as the widespread implementation of the GROMOS algorithm. Clustered ensembles (CEs) produced by this algorithm showed poor performance with chemical shifts compared to sequential ensembles (SEs) of similar size. Instead, we implement a multiscale DR and CA approach and explore the challenges and limitations of applying these algorithms to obtain more robust and tractable CEs. The novel feature of this investigation is the use of solvent-accessible surface area (SASA) as one of the fingerprints for DR alongside previously investigated α carbon distance/angles or ϕ/ψ dihedral angles. The ensembles produced with SASA tSNE DR produced CEs better aligned with the experimental CS of between 0.17 and 0.36 r2 (0.18–0.26 ppm) depending on the system and replicate. Furthermore, this technique produced CEs with better agreement than traditional SEs in 85.7% of all ensemble sizes. This study investigates the quality of ensembles produced based on different input features, comparing latent spaces produced by linear vs nonlinear DR techniques and a novel integrated silhouette score scanning protocol for tSNE DR.

Univerzita Karlova v Praze 10.13039/100007397 SVV 260 666 GrantovÃ¡ Agentura CeskÃ© Republiky 10.13039/501100001824 19-14886Y Ministerstvo Å kolstvÃ­, MlÃ¡deÅ¾e a TelovÃ½chovy 10.13039/501100001823 e-INFRA CZ (ID:90254) Univerzita Hradec KrÃ¡lovÃ© 10.13039/100018512 2200/04/2024-2026 GrantovÃ¡ Agentura, Univerzita Karlova 10.13039/100007543 344321 document-id-old-9ci4c00809
document-id-new-14ci4c00809
ccc-price
==== Body
pmcIntroduction

The biochemistry dogma suggests that proteins comprise amino acids that fold into 3D structures and determine the function of the protein. A growing field of research on the dark proteome reveals the importance of intrinsically disordered proteins (IDPs), which lack rigid secondary structures, both for their functions and prospective dysfunctions.1−3 This class of proteins comprises a significant portion of the human proteome2,4−10 and plays a crucial role in the regulation of cellular processes including signal transduction,7,8,11 transcription,2,7,8,11,12 and DNA repair.2,5,13 They are also involved in protein–protein interactions,2,4,6−8,11,12 acting as flexible scaffolds for enzymes,2,4,7,9,11,13 and binding to a range of partners with high affinity.4,10,12 In addition, IDPs are also responsible for the regulation of gene expression,5,9,12,14,15 cell differentiation,9,16,17 and the immune response,18,19 as they can adapt and respond rapidly to changes in their environment. These proteins are commonly associated with multiple disorders such as Alzheimer’s5−9,15,19−21 and Parkinson’s5,7,11,19−21 and are of great interest for the investigation of these disorders.

Many IDPs undergo post-translational modification, such as phosphorylation, which involves adding a phosphate group to a protein (Figure S1).5,22 This reversible process mediated by kinase enzymes can elicit diverse changes in the protein’s function, structure, or conformation.23,24 Phosphorylation can also disrupt or create new binding sites, alter protein stability, modulate precise cellular signaling pathways, and enable rapid responses to environmental stimuli.4−9,11,13,15−21,23,24 While extensive research has been conducted on the impact of phosphorylation on IDPs, the exact mechanisms are many times left unexplored. This is partially due to the technical challenges of characterizing and detecting these proteins.25,26

The functional role of IDPs and IDRs in the cell is predominantly based on their intrinsic flexibility and capability to perform various operations.27,28 This makes it particularly challenging to determine their structure using traditional experimental techniques, such as CryoEM, which only derives static structures.29 Utilizing mathematical algorithms, molecular dynamics (MD) can forecast protein behavior over time, providing insight into its dynamics, flexibility, and interactions with other molecules. In combination with NMR chemical shifts (CSs), relaxation data, small-angle scattering (SAXS), circular dichroism (CD), paramagnetic relaxation enhancements (PREs), or any other comparable experimental methods, MD provides a more comprehensive understanding of the structure and behavior of the protein.28,30

NMR CSs can be predicted or computed from the conformations of the proteins and averaged together to compare with experimental data. Quantum calculations can derive them,31,32 although computational resources required to compute chemical shielding for so many residues are daunting. Rapid neural network-derived chemical shift prediction (CSP) tools, such as Sparta+33 or ShiftX2,34 have emerged in recent decades to assist researchers in protein structure characterization. CSP models can be beneficial for ordered proteins, but these tools may be overfit for the concerns of IDPs. Other techniques have been implemented using databases of quantum chemical calculations (ProCS-15) to avoid complications from overfitting and to compute atoms outside the range of the backbone rapidly, unusual atom types (e.g., 31P), or without massive databases for the neural network to train upon.35

While the application of CSP tools is computationally cheap and can be easily implemented on a complete MD simulation, calculating NMR parameters at the quantum mechanics level is significantly more demanding. State-of-the-art protein NMR calculations31,36−38 that build on MD and density functional theory (DFT) typically employ sequential structural ensembles (SEs), see Figure 1a. Biomolecular structures are taken from an MD simulation trajectory at regular time intervals, and the structural ensemble obtained is subject to follow-up NMR calculations.39 The computed NMR parameters are then averaged over the ensemble to mimic the time average observed in the experiment.

Figure 1 Representative schematic comparing conformational ensembles generated (a) sequentially and (b) by clustering.

While such an approach gives reliable predictions, it often requires hundreds of structures40 in the computed ensemble to obtain converged ensemble averages. The number of calculations that must be performed is given by (number of structures in an ensemble) × (number of residues),31,36 leading to a fast consumption of high-performance computing resources. The computational costs become a severe drawback, especially if the DFT calculations are repeated using a different ensemble, model chemistry, or fragment/surrounding size.36 The problem can be tackled by replacing the sequential ensemble with a cluster-based ensemble (CE), see Figure 1b. Cluster analysis (CA)41−43 is applied to an MD simulation trajectory to identify clusters of similar structures, and structures representative of individual clusters constitute an ensemble for NMR calculations.31 The approach has only been used sparsely and is typically employed by clustering based on an RMSD, which works reasonably well for biomolecules with a well-defined secondary and tertiary structure, such as DNA and structured proteins.

Applying RMSD-based clustering to highly flexible intrinsically disordered proteins (IDPs) is problematic.44,45 Additionally, the conformational space of IDPs is immensely high-dimensional, and clustering performed directly on the MD trajectory is inefficient. Therefore, dimensional reduction (DR) techniques45,46 have to be applied before clustering. DR is used in machine learning and data analysis to reduce the number of variables or features in a data set.47,48 It has been pointed out that the complexity of IDP structural data requires the use of nonlinear techniques such as Uniform Manifold Approximation and Projection (UMAP)49 and t-Distributed Stochastic Neighbor Embedding (tSNE).50 A recent study that combines the latter with K-means clustering45 has shown that tSNE is very effective in separating heterogeneous IDP conformations into homogeneous subgroups. The homogeneity of the clusters obtained was assessed in terms of global structural metrics such as RMSD or end-to-end distance. The work provided an unprecedented insight into the latent spaces of IDPs derived from tSNE and their dependence on the hyperparameter setup employed. Two critical aspects of clustering by tSNE were beyond the scope of the work and were not covered by the authors. (1) No quantitative validation of the resulting ensembles against the experimental data has been performed. Consequently, (2) there was no need to determine the cluster centers that could be used as representative structures for validation.

The principal objective of the present study is to derive sufficiently small conformational ensembles of IDPs for reliable, cost-effective NMR calculations and predictions based on MD simulations. In pursuit of this goal, we integrate DR techniques and CAs to design a comprehensive protocol that yields IDP ensembles consistent with experimental NMR data. The development of the protocol builds on (1) the assessment of different tools for the prediction of CSs in IDPs, (2) the evaluation of various DR techniques, both linear and nonlinear, and (3) the analysis of the clustering quality based on internal statistical indicators as well as based on validation against experimental NMR CSs.

To ensure an accurate depiction of the conformational dynamics of the IDPs we are examining, we evaluate the efficacy of different fingerprints for DR, including ϕ/ψ angles, α-carbon distances and angles, and solvent-accessible surface area (SASA), to capture phenomena critical to CS variations. We benchmark the hyperparameters of tSNE to assess their impact on the DR and the correlation of the predicted CSs with the experiment. For this purpose, we also propose a metric enabling the identification of the cluster centers, i.e., structures that represent the conformational characteristics of the clusters in question. The CEs derived from the cluster centers are then rigorously compared with traditional sampling methods to validate their efficacy.

Methodology

Four proteins (hTH1, MAP2C, GB3, and UBIQ) were simulated for this investigation. MAP2C is a microtubule-associated protein involved in stabilizing the cytoskeleton of neural cells. The N-terminal region (<40% of sequence) is disordered but contains a segment with <80% propensity to form an α-helix.51 Phosphorylation in this region seems heavily dependent on Ser184 and Thr220.52 hTH1, human tyrosine hydroxylase I, is involved in the biosynthesis of dopamine53 and dysfunction has been linked to neurological diseases, such as Parkinson’s.54 65 residues of the regulatory domain were determined biologically relevant,55 and the first 53 residues were simulated (Figure 2) for 2 μs, phosphorylated at Ser22 and Ser43. Phosphorylation at Ser43 showed a nominal influence, while phosphorylation at Ser22 encouraged phosphorylation at Ser43, which affects the activity of the protein.56 Ubiquitin is a small regulatory protein found in most tissues of eukaryotic organisms. It plays a key role in various biological processes, including marking proteins for degradation, altering their cellular location, and affecting their activity.57 Its structure has been characterized (1ubq) by NMR spectroscopy.58 GB3 is a cell surface protein found in Streptococcus bacteria. It binds to IgG, a type of antibody with high affinity.59 The structure of GB3 (2oed) has been determined by X-ray crystallography, both alone and in complex with an antibody fragment.60

Figure 2 MAP2C and hTH1 single letter amino acid (AA) sequence codes with the phosphorylation sites identified.

MD Simulations

All molecular dynamics (MD) simulations were performed using GROMACS,61 and the proteins were solvated in an aqueous environment with ionic species Na+ and Cl– at 100 mM. The simulations ran at 298 K, 1 atm pressure (STP) using a 1 fs time step and the Velocity Verlet algorithm, a numerical integration algorithm using the Newtonian equations of motion. Neighbor searching was performed every ten steps. The Particle Mesh Ewald (PME) algorithm was used for electrostatic interactions with a one nm cutoff. A reciprocal grid of 120 × 120 × 120 cells was used with fourth-order B-spline interpolation. A single cutoff of 1.004 nm was used for Van-der-Waals interactions. Temperature coupling was done using the Nose-Hoover algorithm, while pressure coupling was done using the Parrinello–Rahman algorithm. The AMBER99SB-ILDN force field was used, as suggested by a previous investigation51,62−65 and the four-point TIP4P-D water model was implemented due to its success in maintaining disorder in an IDP trajectory.51,66 The trajectory lengths varied depending on the protein and whether the protein was ordered or disordered in nature, and details are provided in Table 1. The starting structures for MAP2C was generated using Flexible-Meccano,67 and ASTEROIDS software68 with prior NMR chemical shift data. The starting structure in hTH1 was selected randomly from an existent 1 μs trajectory, in which the conformation was fully equilibriated. Three replicates were generated for the phosphorylated and nonphosphorylated MAP2C trajectories, each with one using a starting structure containing an α-helix observed experimentally using CD (MAP2C/C and MAP2C/H).

Table 1 Overview of Classical MD Trajectories Obtained for the IDPs and Ordered Proteins (OPs) of Interest and the Naming Schemes Used Throughout the Manuscripta

Protein/xx	Residues	Type	Phosphorylated	Length	
MAP2C/A51	96	IDP	yes	1100 ns	
MAP2C/B51	96	IDP	yes	500 ns	
MAP2C/C51	96	IDP	yes	500 ns	
MAP2C/H51	96	IDP	no	500 ns	
MAP2C/NH151	96	IDP	no	500 ns	
MAP2C/NH251	96	IDP	no	500 ns	
hTH131	53	IDP	yes	2000 ns	
GB3*	56	OP	no	200 ns	
UBIQ*	87	OP	no	40 ns	
a Trajectories that were generated just for this publication are marked with an asterisk (*).

Chemical Shift Prediction Tools

CS predictions were made for both ensemble types by various prediction software. Sparta+33,69 is a database searching program that utilizes protein sequences and structural homology to predict chemical shifts of known conformations. Sparta+ calculates the ϕ, ψ, and χ1 angles from PDB coordinate files, then matches them to the complete database, returning the nearest 20 matches. The program then implements a similarity score, considering minor differences to generate a weighting factor for each of the six nuclei observed (HN, Hα, Cα, Cβ, C′, and N). ShiftX234 follows a similar procedure with different algorithms to derive CSs, showing marginally improved accuracy over Sparta+ in some systems.70 Prosecco71 utilizes a purely neural-network-derived algorithm trained entirely on amino acid sequences and NMR CSs for IDPs to predict the CSs.

The ProCS-15 tool35 is a database–based method that predicts chemical shieldings using values from DFT calculations.72 It operates on models with the same sequence and structural pattern by calculating the chemical shielding based on a sum of three factors. First, chemical shielding is determined by a “naked” H3C–CO-Ala-X-Ala-NH–CH3 tripeptide (where X represents an arbitrary amino acid, see Figure 3) with a defined conformation. Second, the effects of side chains are taken into account. Third, the contributions of hydrogen bonding and the ring-current impact are considered. Since the ProCS-15 method does not use machine learning and instead relies on calculated chemical shielding values, there is a reduced risk of overfitting compared to other methods.

Figure 3 Example of the H3C–CO-Ala-X-Ala-NH–CH3 tripeptide used to compute the backbone contributions to the chemical shielding values. X represents the possible phosphorylated residue (Serine, Threonine, Tyrosine).

Since the original database of chemical shielding used by ProCS-15 does not contain chemical shielding for phosphorylated residues (pSer, pThr, pTyr), this chemical shielding had to be calculated and added to the database. The original FragBuilder code73 was extended to construct phosphorylated residues, and the modified FragBuilder tool was applied to generate input files to optimize the geometry of various tripeptide conformations. Geometry optimizations were performed in implicit water at the PM674 level of theory using the Simple Virtual Screening technique. NMR calculations followed geometry optimizations in implicit water. The calculations used the OPBE functional75−78 and 6-31G(d,p) basis set.79−83 The computed chemical shielding values were compiled into a database along with the corresponding torsion angles. Necessary changes to the source code were introduced into the Phaistos and ProCS-15 source codes, allowing ProCS-15 to identify and process phosphorylated amino acids correctly. Common three-letter codes already defined in phaistos(ref (84))/ProCS-15 were employed: pSer – SEP, pThr — TPO, pTyr – PTR. The phosphate group was approximated as a carboxylate, allowing the use of an already available database of shielding corrections. The source code is available on GitHub.85 The complete procedure and methodology were described in the paper by Larsen et al.35

The built-in GROMACS, dssp tool was utilized to determine the secondary structure. Much of the trajectory analysis, radius of gyration (RG), end-to-end distances (EEDIST), solvent-accessible surface area (SASA), root-squared-mean deviation (RMSD), and root-mean-squared fluctuations (RMSF), were done through the mdtraj(86) python package, statistical analysis (e.g., kurtosis/skew) using stats package, and storage and management of large files utilized the pickle. Finally, to visualize and present the results, we used seaborn and matplotlib libraries.

Although CS prediction tools can be done on massive ensembles (>1000) or even the complete set, more in-depth computationally expensive tools can be made intractable when scaled to these levels. Additionally, these smaller ensembles provide insight into the states that can be divulged from the conformations, although oscillating between states that can interact with surrounding molecules or moieties. The traditional technique for generating these ensembles is sequential, with a median time step between each frame extracted from the trajectory (Figure 4). The Gromos clustering algorithm87 identifies clusters by selecting the structure with the most neighbors in a pool of structures and including its neighbors in the cluster, then repeating the process until all structures have been assigned to a cluster.88 Other clustering techniques implement specific features of the trajectory, such as α carbon distances and angles (Figure 4a), SASA (Figure 4b), the ϕ and ψ angles (Figure 4c) from the backbone. Pure clustering alone may not be sufficient, as it only examines similarities and differences between data points without considering the underlying structure or patterns in the data. In high-dimensional data, multicollinearity is a common issue with highly correlated features. Preprocessing the data while preserving the underlying patterns can help reduce noise and redundancy. Moreover, it allows for easier visualization and interpretation of the data in a lower-dimensional space, allowing for a qualitative assessment.

Figure 4 Features generated from the trajectory using (a) α-carbon distances and angles, (b) atomistic solvent-accessible surface area (SASA), and (c) ϕ and ψ dihedral angles.

Linear dimensionality reduction was implemented by the scikit-learn package for principle component analysis (PCA) and linear discriminant analysis (LDA) methods, deeptime for the time-lagged independent component analysis algorithm (tICA). Nonlinear DR was done using scikit-learn for t-distributed stochastic neighbor embedding (tSNE). From scikit-learn, we also used k-means and agglomerative clustering algorithms and calculated the silhouette score to evaluate the quality of the clusters. In this paper, we explore the integration of DR and CA to obtain intelligently selected frames that still conserve the underlying data in the original trajectories.

Results and Discussion

Conformational Analysis

All trajectories RMSD were computed and plotted in the SI (Figure S2) to ensure the trajectories are not trapped in microstates. Based on RG, the MAP2C/A and MAP2C/C trajectories produce similar averages (3.4 ± 0.6) and among the nonphosphorylated trajectories, MAP2C/H and MAP2C/NH1 have closer RG (2.5 ± 0.5 and 2.8 ± 0.5 nm, respectively) to the experimentally obtained (2.5 ± 0.3 nm) as seen in the statistical analysis (Table 2). MAP2C/A and MAP2C/NH1 have negative kurtosis values from the RG, as the distribution is highly irregular. The distribution (Figure 5a/b) also shows that the MAP2C/A trajectory is bifurcated into expanded (≈4.0 nm) and contracted (≈2.8 nm) states during the trajectory. This is also reflected in their variances, as MAP2C/NH1 and MAP2C/A are distributed over a large range. MAP2C/B has a very low variance and is tightly clustered around the mean. For the hTH1 trajectory, the average RG (1.9 ± 0.4 nm) indicates that the protein is also quite pluriform, and its kurtosis (1.63) and skew (1.27) suggest that the distribution of RG values is skewed toward higher values, but is highly clustered around the mean (Figure S3). This is further supported by the high variance value (0.17), indicating a wide range of RG values in the hTH1 trajectory. The EEDIST plots (Figure 5c/d) show that the MAP2C/A and MAP2C/C are more extended on average than MAP2C/B, and both are more extended than each of the nonphosphorylated trajectories, with MAP2C/NH2 being the most compact.

Figure 5 Computed RG of each frame in the MAP2C trajectories are mapped according to their distribution with a kernel density estimation (KDE) for the phosphorylated (a) and nonphosphorylated (b) trajectories with the averages presented in dotted lines and the experimentally observed51 range shaded in gray for the nonphosphorylated. The end-to-end distance (EEDIST) is shown in box plots from the phosphorylated (c) and nonphosphorylated (d) trajectories.

Table 2 Computed Radius of Gyration (RG) Means (x̅), Standard Deviation (σ), Kurtosis, Skew, and Variance from Each Respective Trajectories

Trajectory	x̅(nm)	σ	Kurtosis	Skew	Variance	
MAP2C/H	2.54	0.46	0.84	1.16	0.21	
MAP2C/NH1	2.95	0.70	–1.04	0.24	0.50	
MAP2C/NH2	2.79	0.46	–0.43	0.27	0.21	
MAP2C/A	3.41	0.61	–1.28	0.03	0.37	
MAP2C/B	2.41	0.31	0.82	0.54	0.10	
MAP2C/C	3.40	0.60	0.54	0.56	0.37	
hTH1	1.91	0.41	1.63	1.27	0.17	
GB3	1.09	0.01	0.51	0.63	0.00	
UBIQ	1.18	0.01	0.31	0.43	0.00	

Among the phosphorylated trajectories, MAP2C/A has the most significant fluctuation for most residues (Figure 6a), whereas MAP2C/B exhibits the least. Among the nonphosphorylated trajectories (Figure 6b), there is significantly less deviation between trajectories. Noticeably, the most significant fluctuation occurs at the end terminals of IDPs, a trend not observed in ordered proteins (Figure S23). It is vital to understand the differences and variations in these properties to cluster, as the primary focus of this investigation is on the generation of NMR chemical shifts. CSs depend on local interactions, either solvent-based or intramolecularly. To investigate the solvent interactions, the overall solvent accessibility was computed for each atom of the trajectory (Figure 6c/d) and showed a distinct divergence in interactions with surrounding solvents as well as intramolecular interactions.

Figure 6 Root-mean-squared fluctuations (RMSF) calculated of the α-Carbon for each residue in the phosphorylated (a) and nonphosphorylated (c) trajectories, with black lines indicating phosphorylation sites, and the computed solvent-accessible surface area (SASA) for the phosphorylated (b) and nonphosphorylated (d) MAP2C trajectories.

The MAP2C/A and C trajectories have the most accessible surface area (≈122 nm2) and MAP2C/B the least (107 nm2), indicating that MAP2C/A and C exhibit more solvent interactions, while MAP2C/B simulates more intramolecular interactions. Among the nonphosphorylated trajectories, the SASAs are quite similar; MAP2C/H has the lowest SASA (114.4 nm2), although the variances (≈60) is nearly doubled in MAP2C/H and NH2 compared to their phosphorylated counterparts (≈30) as seen in Table S1 To compare between the proteins, the average SASA was computed per atom and shown in Table S2, demonstrating that disordered proteins have nearly double SASA per atom compared to ordered proteins. Additionally, from the ordered proteins, low variances in RG and SASA reflect the reduced flexibility in GB3 and UBIQ, as it indicates a relatively tight distribution around the mean.

Secondary Structure Analysis

Based on the DSSPs, approximately 56% of the MAP2C trajectories consisted of random coils, 62.6% in hTH1, 14.7% in GB3, and 20.4% in UBIQ (Table 3) as seen in the time-dependent DSSP (tDDSSP) plots (Figure S5–S8). Well-know stabilizing secondary structures, such as β-sheets and α-helices, appear in much higher quantities in OP versus IDPs. The MAP2C trajectories MAP2C/C and MAP2C/H have significant but transient appearances of α-helices, as seen in the tDDSSP plots. These structures were detected in the experiment and sought after in the simulation.51 Recent interest has emerged in the structural importance and contribution of polyproline type 2 (PPII) helices,28,89,90 as they have been observed in higher proportion in IDPs.91,92 Approximately 18% of the MAP2C trajectories had PPII helices, higher than hTH1 (8.9%), GB3 (0.0%), and UBIQ (2.2%).

Table 3 Secondary Structure Predictions (%) for Each of the Trajectories Using DSSP Implemented by GROMACS

Trajectory	Coils	Bridgesa	Helicesb	PPII	Bend	Turn	Phos. Res.	
MAP2C/H	53.2	2.5	5.3	17.0	18.0	4.0	0.0	
MAP2C/NH1	55.9	1.3	1.0	18.9	19.1	3.8	0.0	
MAP2C/NH2	55.8	2.6	0.3	18.9	18.7	3.7	0.0	
MAP2C/A	56.3	0.7	2.0	19.0	16.4	3.5	2.1	
MAP2C/B	55.2	3.5	0.4	15.3	16.7	6.9	2.1	
MAP2C/C	57.6	0.4	2.3	18.2	16.2	3.2	2.1	
hTH1	62.6	1.8	1.5	8.9	16.4	5.0	3.8	
GB3	14.7	42.3	26.5	0.0	7.1	9.3	0.0	
UBIQ	20.4	33.9	19.8	2.2	4.6	19.1	0.0	
a Combination of isolated and extended β-sheets.

b Combination of 3, 4, and 5-turn α-helices.

In hTH1, these PPII helices appear significant (>30%) in specific regions, Pro5–Pro7, and Ser34–Arg36, which may be stabilizing features of the protein (Figure S9). The PPII helices in the nonphosphorylated MAP2C trajectories seem to form indiscriminately throughout the chain, with solid favoritism around sites Ser222–Arg226, a structure which is maintained upon phosphorylation, with a significant drop in MAP2C/B only (Figure S9). MAP2C/B deviates from the other trajectories additionally with a notable decrease in α-helices structures, which are instead replaced with β-bridges and turns between Arg227 and Ser231. These structures may be significant or an artifact of the trajectory. Additional information about the global properties of the trajectories of interest can be found in Section S1 of the SI.

NMR Chemical Shifts

In addition to global and secondary structure analysis, we seek to evaluate the molecular dynamics for local properties, such as interactions, hydrogen bonding, and NMR CSs. As discussed previously, to aid in calculating CSs, neural network-derived CSP (e.g., Sparta+, ShiftX2) and database-derived CSP (ProCS-15) can produce relatively accurate CSs from a trajectory significantly faster than DFT calculations. A comprehensive comparison of CSPs on sequential ensembles can be found in Section S2 of the SI, and the method of developing these SEs is discussed in Section S3 of the SI. Among the nonphosphorylated MAP2C trajectories, MAP2C/NH2 performs the best, followed by MAP2C/H and MAP2C/NH1 (Table S5). MAP2C/C shows the greatest agreement with experimental, and then MAP2C/A and MAP2C/B, in that order (Figures S17–S18). There were several strong outliers in each of the trajectories discussed and described in full in the SI. Still, the overall trend is that the best-performing atoms according to their R2 values are Cβ, Cα, C′, N, Hα, and HN, although the trend for the RMSE is Hα, HN, Cβ, C′, Cα, and N (Table S5 and S3, respectively).

Performance rankings for CSPs differ between different atoms and proteins, but some trends can be observed. For the intrinsically disordered proteins (hTh1 and MAP2C), Sparta+ and ShiftX2 produced relatively good agreements for Cα, Cβ, and N, although it was not great for Hα and HN. Sparta+ also performed better than ShiftX2 for calculating nonphosphorylated chemical shifts. ShiftX2 was better at predicting C′, and Prosecco best predicted CS based on amino acid sequences. Nevertheless, since Prosecco uses only amino acid sequences as input, it cannot be used to evaluate the performance of MD trajectories. Sparta+ and ShiftX2 ignore proline N atoms due to the challenges in assigning them experimentally, as they require customized heteronuclear NMR experiments.93 Prosecco performs quite well in predicting chemical shifts on the basis of protein sequences. However, its usage is limited because of its inability to describe dynamic behaviors, while Sparta+ and ShiftX2 are well suited for describing dynamic behavior, as they utilize conformation information in the form of PDB files. ProCS-15 performed poorly in predicting chemical shifts for ordered and disordered proteins, as highlighted in Figure 8. It was outperformed by the most commonly used prediction tools, such as Sparta+ and ShiftX2, in terms of both accuracy and linearity for the chemical shift data examined in this study. Furthermore, ProCS-15 showed a higher standard deviation for some residues than other prediction tools, suggesting inconsistent performance for different amino acids (Figure 7). Thus, while it is helpful in specific applications, the general performance of ProCS-15 indicates that it may not be a suitable prediction tool for all protein systems and conditions, and its predictions should be used with caution.

Figure 7 Correlation between experimental and predicted N chemical shifts in hTH1 using (a) Sparta+, (b) ShiftX2, and (c) ProCS-15 with r2 and root-mean squared error in parentheses.

Figure 8 Influence of the phosphorylation of the protein on the 1HN CS by residue plotted for experimental (a), Sparta+ (b), ShiftX2 (c), and ProCS-15 (d) with specific shifted residues marked red (upshifted) and blue (downshifted).

Phosphorylation Influence

The CSs from the phosphorylated and nonphosphorylated proteins of MAP2C were compared, and the CSPs were evaluated for their ability to describe phosphorylation. The MAP2C trajectory was phosphorylated at the Ser184 and Thr220 sites, and the experimental chemical shifts were derived for comparison.51Figure 8 shows the influence of phosphorylation on the HN CSs; the N plots produced similar trends (Figure S21) although the variation is less pronounced. In addition to the phosphorylation sites (Ser184 and Thr220), four residues were deshielded upon phosphorylation (Arg221, Thr219, Leu185, and Ala164) and two were shielded (Glu223 and Glu180). These six residues represent sites where the phosphorylation alters the chemical environment or the phosphate group interacts intramolecularly with the respective HN. Those that were shielded were both negatively charged residues near the phosphorylation sites. Residue Ser183 was not affected by phosphorylation despite its proximity, and the MD trajectory suggests an atomic-level explanation. Due to interactions between the phosphorylated Ser184 and Arg187, the protein appears pinched in many parts, which prevents the phosphate from interacting with its neighbor. The study revealed that recent neural network-based CSPs, such as Sparta+ and ShiftX2, could not incorporate the influence of phosphorylation, which was expected due to their reliance on data sets without phosphorylated CSs. However, the DFT-trained CSP, ProCS-15, distinguished three relevant phosphorylation shifts, indicating potential for this approach, although signficant improvement must be made toward improving its accuracy and precision. This highlights the importance of developing more advanced prediction methods to accurately account for the influence of phosphorylation, which would require smaller ensemble sizes. Furthermore, the limitations of Sparta+ and ShiftX2 highlight the need for higher-level, precise prediction techniques.

GROMOS-Clustered Ensembles

Before we embarked on the derivation of various clustered ensembles, we constructed SEs for comparison, see Section S3. CEs can be generated from different characteristics of the trajectories. To develop a CE, we must first decide on the logical feature based on which clustering will be performed. The Gromos algorithm87 built into GROMACS45,94 employs RMSD, and we implemented the algorithm in this investigation on several trajectories at different RMSD cutoff points. The resultant CEs were averaged and compared to sequential ensembles of the same size (Figure S24–S25), and additional figures and explanations can be found in Section S4 in the SI. Regardless of the ensemble size or trajectory, the CEs produced less agreement with experimental values in all trajectories and atoms tested. Even for the less pluriform, more homogeneous GB3 protein trajectory, the ensembles produced rarely exceeded agreement with those generated sequentially. One reason is that the RMSD matrix used for this method captures the system’s overall behavior rather than focusing on specific regions or interactions. This can be problematic as it ignores local fluctuations and conformational changes, which are crucial for accurately describing the behavior of macromolecules reflected in NMR CSs. As a result, CEs generated using this technique fail to capture the level of detail necessary for further analysis. Better considerations of the possible clustering features of the protein, both local and global, are required to generate more precise CEs, particularly for disordered systems.

Dimensionality Reduction

Multicollinearity (high correlation between features) in high-dimensional data sets can confound clustering results. Multicollinearity arises when two or more predictor variables are highly correlated. This can lead to difficulties in various analytical tasks, including clustering.95 To address this, dimensionality reduction techniques can be applied. These methods reduce the number of features (dimensions) while preserving essential information. Doing so mitigates noise, redundancy, and multicollinearity, leading to more effective clustering. By reducing the dimensions, we also create a more manageable representation of the data, making it easier to visualize and analyze. Each dimensionality reduction yields distinct latent spaces (Figure S26). These spaces emerge from intricate variations in input features, unveiling diverse patterns within the system and singling out specific states that may be of significance biologically. Choosing a suitable DR technique can depend on many factors, outlined in Section S5. For our purposes, the linear method tICA and nonlinear method tSNE were selected to generate the reduced landscapes in the procedure. Several other linear (PCA, LDA) and nonlinear (UMAP) techniques were attempted, although, as explained in Section S5, many were discarded for various reasons.

In order to implement DR, specific input features must be selected to adequately describe the system. tICA was implemented using each of the characteristics (ϕ/ψ angles, α-carbon distances and angles, or SASA) and the hierarchical agglomeration (5 groups) in the resultant manifold (Figure 9) shows some significant variations between representative spaces. The choice of agglomerative hierarchical clustering over faster but less accurate techniques like K-means or highly parametrized techniques like DBSCAN likely stems from its ability to more accurately capture the underlying structure of the data without being heavily influenced by parameter settings. The hierarchical nature of this method offers a nuanced understanding of data grouping at different levels of granularity, which is crucial to accurately capture the complex biological phenomena represented in the MAP2C/C trajectory and the various ways in which it can be described (Figure 10). Approximately 40% of the MAP2C/C trajectory showed the existence of α-helices, which is clearly observed in the DSSP plot (Figure 10g-i and SI: Figure S6c) of the protein. Assuming that the DR is accurate, this trend and others should also be accurately captured in the latent spaces.

Figure 9 Dimensionally reduced landscapes of MAP2C/C clustered into five labeled clusters (a-c), and the respect place within the trajectory (d-f) using different input features; ϕ/ψ (a), α carbons distances and angles (b), and SASA (c).

Figure 10 Distribution of conformations from the MAP2C/C trajectory on different DR latent spaces; ϕ/ψ (a), α carbons distances and angles (b), and SASA (c). Rolling averages are included to show the overall path of the trajectories (a-c), and several collective variables are plotted, such as radius of gyration (d-f) and total SASA (j-l) as a function of time. The DSSP plots of the secondary structures (g-i) are included for comparison.

In the ϕ/ψ DR (DRϕ/ψ), the backbone conformations are the main contributors to the expression of the conformational landscape. Thus, the state DRϕ/ψα exhibits a high propensity (approximately 7.1 ± 1.9 residues) for α-helices (Table S9), three states (DRϕ/ψβ1, DRϕ/ψβ2, DRϕ/ψβ3) lack instances of α-helices, and a transitional state (DRϕ/ψαβ) that spans a continuous range between these extremes, with partial α-helices characteristics (≈5 residues). In the case of the SASA DR (DRSASA), fingerprints strongly favor solvent intractability instead of only backbone conformations. Because of this, conformations with α-helices are separated into four states (DRSASAα1, DRSASAα2, DRSASAα3, DRSASAα4), and one state (DRSASAβ) with no α-helices. As seen in Figure 9a-c, many of these clusters overlap, as the transition between these states is much more continuous, and there is no significant barrier between the sampled conformations. This can also be seen by the rolling average “path” of the trajectory, shown in the dotted line (Figure 9a-c), which demonstrates that the protein travels within the α and β states with minimal restriction, but is hindered from traversing between them, either indicating poor sampling in the trajectory data or rare phenomena with a high energy barrier. To achieve a robust ensemble from DR/CA, there needs to be some assurance that the centers of each cluster in an ensemble give a good representation of the full trajectory. Similar plots and DR of the other MAP2C trajectories can be found in the Supporting Information (Figures S31–S35). Based on silhouette scores and Davies-Bouldin index (Table S8), SASA produced the best clustering separation of all features, although an investigation of the cluster sizes is warranted. It is important to note that the success of DRSASA with IDP may not produce satisfactory results when applied to ordered proteins, as seen in Figure S38d-i, where GB3 and UBIQ exhibit very different behavior with a preference for DRϕ/ψ.

Several methods were explored to select a representative center from each cluster identified by cluster analysis. An initial approach involved identifying the Euclidean center and choosing the closest conformation on the reduced landscape. However, this method proved to be insufficient due to technical limitations and the potential for outliers to be selected as representatives. As an alternative, the Euclidean center was computed for each cluster on the higher-dimensional landscape, and the structure closest to this point was selected as the center representative. However, this method also had drawbacks and did not yield satisfactory results. To address these limitations, a new approach was used based on the concept of density in the reduced landscape. Specifically, the Gaussian density of the reduced landscape was calculated, and structures residing in the highest-density regions of the latent space were chosen as representative centers for their respective clusters. This method minimized the risk of selecting outliers as representatives, ensuring that the chosen structures accurately represented the characteristics of their clusters.

Using this method to generate CE and compare them with SEs of similar size, we can average the chemical shifts of each to determine which method produces the best ensemble consistently in relation to experimental chemical shift averages. Linear techniques seem to adequately group the trajectories into small ensembles (<10) as seen in Figure 11, although as recorded in Section S3, this ensemble size is insufficient to describe the complex behavior of the protein for the purpose of NMR CS. The results of our study consistently demonstrate the superior performance of CEs over SEs. This is evident in most of the models produced, as shown in the histograms and scatter plots in Figures S48–S50. Although there are some exceptions, such as MAP2C/A ϕ/ψ DR N atoms, MAP2C/B α DR N atoms, and MAP2C/C SASA HN atoms, the trend toward improved performance in smaller ensembles is clear. As the ensemble size approaches 500, the difference between clustered ensembles and sequential ensembles becomes increasingly negligible. This is expected as the sequential ensembles reach convergence in this range, and the addition of new frames has little impact on overall performance. Nonlinear DR techniques are recommended for larger ensemble sizes (>50). Nonlinear DR techniques effectively generated larger ensembles of trajectories due to their ability to capture complex nonlinear relationships, account for higher-order interactions, handle nonuniform and high-dimensional data, and preserve the local and global structure of the data.

Figure 11 Performances of the clustered ensembles, represented as the % of ensembles which outperform sequential ensembles, using different input features and linear dimensionality reduction on different trajectories for HN and N chemical shifts.

To compensate for the additional parametrization required for a nonlinear technique such as tSNE, we devised a protocol to select the best DR based on a metric, integrated silhouette score, SSINT.45 The SSINT combines the clustering quality from both the higher and lower dimensional space, as described in Section S7 in the SI. For each feature tested, the perplexities were scanned as seen in Figures S39 - S44, and the best SSINT DR was selected for clustering at each ensemble size (Figure 12a). The resultant set was compared to ensembles of similar sizes generated sequentially and performed well (Figure 12b). Not all cluster sizes produce better ensembles than traditional sequential ensembles; however, utilizing SSINT, we were able to achieve ensembles with better performance than the averaged CEs derived, and the selected ensembles perform better than SE in ≈85.7% of cluster sizes. The best-selected ensemble (n = 5) is shown in the latent space (Figure 12d) with data point shading representing the number of α-helices, and representative structures from the ensembles can be seen in Figure 12c. It should be noted that the ensembles generated by this methodology are tested primarily for NMR chemical shifts, and further experimental data would better improve the technique and avoid the risk of overfitting.

Figure 12 Representation of the best ensemble selected from (a) the integrated silhouette score, SSINT, (b) the relevant performances derived from CEs compared to SEs, and the performance of that selected by SSINT, (c) a representation of the DR latent space and subsequent clustering, and (d) visual representation of a small cluster with excellent CS agreement.

Similar clustering can be implemented on each of the other phosphorylated trajectories to avoid overfitting of the model (Figure S51–S53), and each time implemented the number of ensembles which performed better than sequential increased, by an average of ≈38%. The amount that the techniques improved the sampling in the subsequent ensembles varied depending on the sampling quality of the original trajectory. Even still, all features and trajectories generate CEs that, on average, performed better than the sequential ensembles, suggesting a robust technique for generating CEs. A compilation of all MAP2C trajectories and the percent of CEs that outperform similar-sized SEs can be seen in Figure 13, highlighting the benefits of both utilizing SASA as a feature for dimensionality reduction and the use of scanning the SSINT to determine the ideal matching of perplexity and cluster size. These ensembles not only match local properties such as chemical shifts, but also exhibit a similar global structure compared to the original trajectories, as indicated by the agreement of properties such as RG and RMSD (Figure S54). This provides further evidence that our method effectively captures the true conformational space of the protein and allows for the identification of representative conformations that accurately reflect the dynamics of the system.

Figure 13 Comparisons of the different dimensionality reductions and clustering, % of ensembles generated through clustering (ENSClust) that outperform those generated from sequential (ENSSeq).

Conclusion

This study illustrates the effectiveness of integrating DR and CA techniques to generate CEs of IDPs that align well with the NMR CS predictions. Within the spectrum of rapid CSPs assessed, Sparta+ emerged as the preeminent choice, albeit with discernible performance disparities dependent on the atom type. The accuracy in predicting the CSs of nitrogen and carbon atoms was notably higher, with a r2 of no less than 0.9 for all trajectories, ordered and disordered. In particular, in the hydrogen atoms, there is a substantial drop in agreement, from 0.95 to 0.72 (Hα) and from 0.84 to 0.49 (HN). This can be attributed to a lower range in ppm for H atoms, coupled with the importance of solvent interactions, which are more important to IDPs than OPs. The former suggests the importance of reporting the error in absolute terms (RMSE) as well as r2, and the latter is given credence by the abnormally large range of performance depending on the MAP2C replicate. Additionally, the methodologies could not effectually capture the influence of phosphorylation, except for ProCS-15, which is explainable as the neural networks were not trained on phosphorylated samples, while the database of ProCS-15 was. Consequently, these findings emphasize the need for advanced methodologies that can more accurately predict CSs within IDPs, particularly for incorporating phosphorylated residues or PTMs.

In the exploration of methodologies to generate CEs that would make such higher-level computation techniques tractable for CS predictions, we utilized GROMOS87 CA to generate CEs, highlighting its limitations for IDPs. With few exceptions in the OPs and small ensembles, these ensembles markedly under-performed in alignment with experimental data. The algorithm’s focus on RMSD as a clustering metric fails to account for the specificity’s of local interactions and conformational dynamics, which are essential for precise NMR CS modeling in IDPs, pointing toward the necessity for a more nuanced approach in selecting clustering criteria. DR served as a preliminary step to clustering to mitigate the computational and analytical challenges posed by the high-dimensional nature of IDP data, enabling a more focused and effective clustering process. Upon applying DR before clustering, several issues arose, such as the selection of input features, the decision to use linear or nonlinear methods, and the calibration of hyperparameters to refine the data analysis process.

This study explored the impact of using different fingerprints for DR on the analysis of IDPs, specifically focusing on the solvent-accessible surface area (SASA), ϕ/ψ dihedral angles, and α-carbon distances and angles. The distinct latent spaces generated from each feature set underscore the importance of feature selection in DR. SASA offers a more comprehensive view of the protein’s external interactions, including transient secondary structures and overall shape. In contrast, ϕ/ψ and α-carbon-based DR focus more on the protein’s backbone conformation and local structure, respectively, which excels particularly for capturing conformational changes in ordered proteins. The study further underscores the insubstantial impact of SASA associated with carbon and nitrogen atoms on capturing the system’s conformational dynamics. The study also indicates that no single feature set universally outperforms the others, suggesting that the choice of DR features should be tailored to the specific analytical goals and characteristics of the protein under study. This performance differentiation underscores the complexity of IDP analysis and the need for a nuanced approach to feature selection in DR techniques.

When choosing between linear and nonlinear DR techniques, the key considerations include the desired ensemble size and the data’s intricacy. For smaller ensembles (<10), linear DR methods such as PCA and TICA are adequate but may not fully capture the complex behaviors necessary for accurate NMR CS predictions. For larger ensembles (>50), nonlinear DR methods, such as tSNE, are preferred due to their ability to handle complex nonlinear relationships and high-dimensional data, effectively generating larger and more representative ensembles. Additionally, the utilization of the techniques outlined in this paper are highly dependent on the input source. Readers are therefore advised that replication might entail the use of verification of trajectories with experimental evidence or enhanced sampling techniques.

The integration of dimensionality reduction (DR) techniques prior to clustering significantly enhanced the clustering process for intrinsically disordered proteins (IDPs), as demonstrated by the creation of conformational ensembles that more precisely mirror the dynamic behaviors of these proteins. This was evidenced by the improved performance of these ensembles over those created through traditional sequential ensemble generation, with ensembles generated using DR and clustering techniques outperforming sequentially generated ensembles in approximately 85.7% of cluster sizes evaluated. This approach not only underscores the effectiveness of DR techniques in refining the clustering process but also highlights the potential of these methods to produce more representative and accurate conformational ensembles of IDPs. While the preliminary results of this investigation are promising, it should be noted that the current method was only trained on one form of experimental data and would benefit from the inclusion of additional sources to avoid overfitting, which is a future endeavor we plan to investigate utilizing more adept neural network approaches.

Data Availability Statement

All molecular dynamics simulation trajectories are available upon request. All input files to run the simulations, as well as our current implementation of the model, can be downloaded from the GitHub repository: https://github.com/Chemistry-Mike/NMR-Clustering-IDP. Additionally, all materials used to generate and run the ProCS from the source can be found available here: https://github.com/martinj80/phaistos.

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.4c00809.(1) Additional trajectory analysis, including RMSD calculated for all trajectories in the nonphosphorylated and phosphorylated MAP2C, hTH1, and ordered trajectories (GB3/UBIQ), with the calculated root mean squared deviation to confirm that the trajectory was not trapped in any local minima or microstates. The variation in RMSD and the combined gyration radius (RG) kernel density estimation is discussed. (2) Analysis and investigation of protein trajectories, DR methods, and further information on the integrity and techniques described in the manuscript. (3) Information on the influence of protein phosphorylation on the chemical shift of 15N by residue, including experimental data and predictions from Sparta+, ShiftX2, and ProCS-15, with specific labeled shifted residues. (4) Performance comparisons of the ensembles generated by the GROMOS algorithm87 and sequentially for different types of atoms in MAP2C/A, MAP2C/C, and all types of atoms in hTH1. (5) DR techniques selected for the study, including principle component analysis (PCA), linear discriminant analysis (LDA), and time-lagged independent component analysis (tICA), with reasons for the selection and implementation details. (6) Details on the challenges in assigning proline 1H N and 15 N atoms in NMR experiments and the performance of various prediction tools, such as Prosecco, Sparta+/ShiftX2, and ProCS-15, for these atoms. (7) Discussion on the importance of selecting the appropriate input fingerprint for DR analysis in molecular dynamics simulations, including dihedral ϕ and ψ angles, α carbon distances and angles, and solvent-accessible surface area. (8) Comparison of different CEs achieved from utilizing different features for the fingerprints of the DR and clustering, with their averaged RG, RMSD, SASA, and EE DIST. (9) Regression plots and performance evaluations for various NMR chemical shift prediction tools, highlighting their deviations from experimental values and the impact of different factors like phosphorylation and protein order (PDF)

Supplementary Material

ci4c00809_si_001.pdf

The authors declare no competing financial interest.

Acknowledgments

The research was financed by the Czech Science Foundation Grant No. 19-14886Y. The project was supported by “The project National Institute of Virology and Bacteriology (Programme EXCELES, ID Project No. LX22NPO5103) - Funded by the European Union - Next Generation EU”. MJ acknowledges support by the research project 2200/04/2024-2026 as part of the “Competition for 2024-2026 Postdoctoral Job Positions at the University of Hradec Králové” at the Faculty of Science, University of Hradec Králové. This work was supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90254). The research was implemented using MetaCentrum and IT4I supercomputer facilities. The work was also supported by the Charles University Grant Agency (GAUK) grant number 344321 and the Charles University project No. SVV 260 666.
==== Refs
References

Toto A. ; Troilo F. ; Visconti L. ; Malagrinò F. ; Bignon C. ; Longhi S. ; Gianni S. Binding Induced Folding: Lessons from the Kinetics of Interaction between NTAIL and XD. Arch. Biochem. Biophys. 2019, 671 , 255–261. 10.1016/j.abb.2019.07.011.31326517
Kulkarni P. ; Uversky V. N. Intrinsically Disordered Proteins: The Dark Horse of the Dark Proteome. Proteomics 2018, 18 , 1800061 10.1002/pmic.201800061.
Perdigão N. ; Rosa A. Dark Proteome Database: Studies on Dark Proteins. High-Throughput 2019, 8 , 8 10.3390/ht8020008.30934744
Uversky V. N. Intrinsically Disordered Proteins and Their “Mysterious” (Meta)Physics. Front. Phys. 2019, 7 , 10 10.3389/fphy.2019.00010.
Lermyte F. Roles, Characteristics, and Analysis of Intrinsically Disordered Proteins: A Minireview. Life (Basel) 2020, 10 , 320 10.3390/life10120320.33266184
Uversky V. N. New Technologies to Analyse Protein Function: An Intrinsic Disorder Perspective. F1000Research 2020, 9 , 101 10.12688/f1000research.20867.1.
Trivedi R. ; Nagarajaram H. A. Intrinsically Disordered Proteins: An Overview. Int. J. Mol. Sci. 2022, 23 , 14050 10.3390/ijms232214050.36430530
Huang J. ; MacKerell A. D. Force Field Development and Simulations of Intrinsically Disordered Proteins. Curr. Opin. Struct. Biol. 2018, 48 , 40–48. 10.1016/j.sbi.2017.10.008.29080468
Bencivenga D. ; Stampone E. ; Roberti D. ; Della Ragione F. ; Borriello A. p27Kip1, an Intrinsically Unstructured Protein with Scaffold Properties. Cells 2021, 10 , 2254 10.3390/cells10092254.34571903
Karlsson E. ; Paissoni C. ; Erkelens A. M. ; Tehranizadeh Z. A. ; Sorgenfrei F. A. ; Andersson E. ; Ye W. ; Camilloni C. ; Jemth P. Mapping the Transition State for a Binding Reaction between Ancient Intrinsically Disordered Proteins. J. Biol. Chem. 2020, 295 , 17698–17712. 10.1074/jbc.RA120.015645.33454008
Dyson H. J. ; Wright P. E. Intrinsically Unstructured Proteins and Their Functions. Nat. Rev. Mol. Cell Biol. 2005, 6 , 197–208. 10.1038/nrm1589.15738986
Brodsky S. ; Jana T. ; Mittelman K. ; Chapal M. ; Kumar D. K. ; Carmi M. ; Barkai N. Intrinsically Disordered Regions Direct Transcription Factor In Vivo Binding Specificity. Mol. Cell 2020, 79 , 459–471.e4. 10.1016/j.molcel.2020.05.032.32553192
Kragelund B. B. ; Schenstrøm S. M. ; Rebula C. A. ; Panse V. G. ; Hartmann-Petersen R. DSS1/SEM1, a Multifunctional and Intrinsically Disordered Protein. Trends Biochem. Sci. 2016, 41 , 446–459. 10.1016/j.tibs.2016.02.004.26944332
Kulkarni P. The Boscombe Valley Mystery: A Lesson in the Perils of Dogmatism in Science. J. Biosci 2021, 46 , 59 10.1007/s12038-021-00179-x.34168102
Babu M. M. ; van der Lee R. ; de Groot N. S. ; Gsponer J. Intrinsically Disordered Proteins: Regulation and Disease. Curr. Opin. Struct. Biol. 2011, 21 , 432–440. 10.1016/j.sbi.2011.03.011.21514144
Bondos S. E. ; Dunker A. K. ; Uversky V. N. Intrinsically Disordered Proteins Play Diverse Roles in Cell Signaling. Cell Commun. Signal. 2022, 20 , 20 10.1186/s12964-022-00821-7.35177069
Dunker A. K. ; Bondos S. E. ; Huang F. ; Oldfield C. J. Intrinsically Disordered Proteins and Multicellular Organisms. Semin. Cell Dev. Biol. 2015, 37 , 44–55. 10.1016/j.semcdb.2014.09.025.25307499
Marín M. ; Ott T. Phosphorylation of Intrinsically Disordered Regions in Remorin Proteins. Front. Plant Sci. 2012, 3 , 86 10.3389/fpls.2012.00086.22639670
Tompa P. ; Schad E. ; Tantos A. ; Kalmar L. Intrinsically Disordered Proteins: Emerging Interaction Specialists. Curr. Opin. Struct. Biol. 2015, 35 , 49–59. 10.1016/j.sbi.2015.08.009.26402567
Korsak M. ; Kozyreva T. Beta Amyloid Hallmarks: From Intrinsically Disordered Proteins to Alzheimer’s Disease. Adv. Exp. Med. Biol. 2015, 870 , 401–421. 10.1007/978-3-319-20164-1_14.26387111
Uversky V. N. Intrinsically Disordered Proteins and Their (Disordered) Proteomes in Neurodegenerative Disorders. Front. Aging Neurosci. 2015, 7 , 18 10.3389/fnagi.2015.00018.25784874
Phillips A. H. ; Kriwacki R. W. Intrinsic Protein Disorder and Protein Modifications in the Processing of Biological Signals. Curr. Opin. Struct. Biol. 2020, 60 , 1–6. 10.1016/j.sbi.2019.09.003.31629249
Ubersax J. A. ; Ferrell Jr J. E. Mechanisms of Specificity in Protein Phosphorylation. Nat. Rev. Mol. Cell Biol. 2007, 8 , 530–541. 10.1038/nrm2203.17585314
Iakoucheva L. M. The Importance of Intrinsic Disorder for Protein Phosphorylation. Nucleic Acids Res. 2004, 32 , 1037–1049. 10.1093/nar/gkh253.14960716
Beveridge R. ; Calabrese A. N. Structural Proteomics Methods to Interrogate the Conformations and Dynamics of Intrinsically Disordered Proteins. Front. Chem. 2021, 9 , 603639 10.3389/fchem.2021.603639.33791275
Lermyte F. Roles, Characteristics, and Analysis of Intrinsically Disordered Proteins: A Minireview. Life 2020, 10 , 320 10.3390/life10120320.33266184
Mishra P. M. ; Verma N. C. ; Rao C. ; Uversky V. N. ; Nandi C. K. Intrinsically Disordered Proteins of Viruses: Involvement in the Mechanism of Cell Regulation and Pathogenesis. Prog. Mol. Biol. Transl. Sci. 2020, 174 , 1–78. 10.1016/bs.pmbts.2020.03.001.32828463
Bakker M. J. ; Sørensen H. V. ; Skepö M. Exploring the Role of Globular Domain Locations on an Intrinsically Disordered Region of p53: A Molecular Dynamics Investigation. J. Chem. Theory Comput. 2024, 20 , 1423–1433. 10.1021/acs.jctc.3c00971.38230670
Pauwels K. ; Lebrun P. ; Tompa P. To Be Disordered or Not To Be Disordered: Is That Still a Question for Proteins in the Cell?. Cell. Mol. Life Sci. 2017, 74 , 3185–3204. 10.1007/s00018-017-2561-6.28612216
Henriques J. ; Arleth L. ; Lindorff-Larsen K. ; Skepö M. On the Calculation of SAXS Profiles of Folded and Intrinsically Disordered Proteins from Computer Simulations. J. Mol. Biol. 2018, 430 , 2521–2539. 10.1016/j.jmb.2018.03.002.29548755
Bakker M. J. ; Mládek A. ; Semrád H. ; Zapletal V. ; Pavlíková Přecechtělová J. Improving IDP Theoretical Chemical Shift Accuracy and Efficiency through a Combined MD/ADMA/DFT and Machine Learning Approach. Phys. Chem. Chem. Phys. 2022, 24 , 27678–27692. 10.1039/D2CP01638A.36373847
Tremblay M.-L. ; Banks A. W. ; Rainey J. K. The Predictive Accuracy of Secondary Chemical Shifts is More Affected by Protein Secondary Structure than Solvent Environment. J. Biomol. NMR 2010, 46 , 257–270. 10.1007/s10858-010-9400-5.20213252
Shen Y. ; Bax A. Sparta+: A Modest Improvement in Empirical NMR Chemical Shift Prediction by Means of an Artificial Neural Network. J. Biomol. NMR 2010, 48 , 13–22. 10.1007/s10858-010-9433-9.20628786
Han B. ; Liu Y. ; Ginzinger S. W. ; Wishart D. S. Shiftx2: Significantly Improved Protein Chemical Shift Prediction. J. Biomol. NMR 2011, 50 , 43–57. 10.1007/s10858-011-9478-4.21448735
Larsen A. S. ; Bratholm L. A. ; Christensen A. S. ; Maher C. ; Jensen J. H. ProCS15: a DFT-Based Chemical Shift Predictor for Backbone and CB Atoms in Proteins. PeerJ. 2015, 3 , e1344 10.7717/peerj.1344.26623185
Pavlíková Přecechtělová J. ; Mládek A. ; Zapletal V. ; Hritz J. Quantum Chemical Calculations of NMR Chemical Shifts in Phosphorylated Intrinsically Disordered Proteins. J. Chem. Theory Comput. 2019, 15 , 5642–5658. 10.1021/acs.jctc.8b00257.31487161
Vícha J. ; Babinský M. ; Demo G. ; Otrusinová O. ; Jansen S. ; Pekárová B. ; Žídek L. ; Munzarová M. L. The Influence of Mg2+ Coordination on 13C and 15N Chemical Shifts in CKI1RD Protein Domain from Experiment and Molecular Dynamics/Density Functional Theory Calculations. Proteins: Struct., Funct., Bioinf. 2016, 84 , 686–699. 10.1002/prot.25019.
Exner T. E. ; Frank A. ; Onila I. ; Möller H. M. Toward the Quantum Chemical Calculation of NMR Chemical Shifts of Proteins. 3. Conformational Sampling and Explicit Solvents Model. J. Chem. Theory Comput. 2012, 8 , 4818–4827. 10.1021/ct300701m.26605634
Fukal J. ; Páv O. ; Buděšínský M. ; Šebera J. ; Sychrovský V. The Benchmark of 31P NMR Parameters in Phosphate: a Case Study on Structurally Constrained and Flexible Phosphate. Phys. Chem. Chem. Phys. 2017, 19 , 31830–31841. 10.1039/C7CP06969C.29171602
Dračínský M. ; Möller H. M. ; Exner T. E. Conformational Sampling by Ab Initio Molecular Dynamics Simulations Improves NMR Chemical Shift Predictions. J. Chem. Theory Comput. 2013, 9 , 3806–3815. 10.1021/ct400282h.26584127
González-Alemán R. ; Hernández-Castillo D. ; Caballero J. ; Montero-Cabrera L. A. Quality Threshold Clustering of Molecular Dynamics: A Word of Caution. J. Chem. Inf. Model. 2020, 60 , 467–472. 10.1021/acs.jcim.9b00558.31532987
Ezerski J. C. ; Cheung M. S. CATS: ATool for Clustering the Ensemble of Intrinsically Disordered Peptides on a Flat Energy Landscape. J. Phys. Chem. B 2018, 122 , 11807–11816. 10.1021/acs.jpcb.8b08852.30362738
Träger S. ; Tamò G. ; Aydin D. ; Fonti G. ; Audagnotto M. ; Peraro M. D. CLoNe: Automated Clustering Based on Local Density Neighborhoods for Application to Biomolecular Structural Ensembles. Bioinformatics 2021, 37 , 921–928. 10.1093/bioinformatics/btaa742.32821900
Ezerski J. C. ; Cheung M. S. CATS: ATool for Clustering the Ensemble of Intrinsically Disordered Peptides on a Flat Energy Landscape. J. Phys. Chem. B 2018, 122 , 11807–11816. 10.1021/acs.jpcb.8b08852.30362738
Appadurai R. ; Koneru J. K. ; Bonomi M. ; Robustelli P. ; Srivastava A. Clustering Heterogeneous Conformational Ensembles of Intrinsically Disordered Proteins with t-Distributed Stochastic Neighbor Embedding. J. Chem. Theory Comput. 2023, 19 , 4711–4727. 10.1021/acs.jctc.3c00224.37338049
Trozzi F. ; Wang X. ; Tao P. UMAP as a Dimensionality Reduction Tool for Molecular Dynamics Simulations of Biomacromolecules: A Comparison Study. J. Phys. Chem. B 2021, 125 , 5022–5034. 10.1021/acs.jpcb.1c02081.33973773
Tribello G. A. ; Gasparotto P. Using Dimensionality Reduction to Analyze Protein Trajectories. Front. Mol. Biosci. 2019, 6 , 46 10.3389/fmolb.2019.00046.31275943
Kukharenko O. ; Sawade K. ; Steuer J. ; Peter C. Using Dimensionality Reduction to Systematically Expand Conformational Sampling of Intrinsically Disordered Peptides. J. Chem. Theory Comput. 2016, 12 , 4726–4734. 10.1021/acs.jctc.6b00503.27588692
McInnes L. ; Healy J. ; Saul N. ; Großberger L. UMAP: Uniform Manifold Approximation and Projection. J. Open Source Softw. 2018, 3 , 861 10.21105/joss.00861.
van der Maaten L. ; Hinton G. Visualizing Data using t-SNE. J. Mach. Learn. Res. 2008, 9 , 2579–2605.
Zapletal V. ; Mládek A. ; Melková K. ; Louša P. ; Nomilner E. ; Jaseňáková Z. ; Kubáň V. ; Makovická M. ; Laníková A. ; Žídek L. ; et al. Choice of Force Field for Proteins Containing Structured and Intrinsically Disordered Regions. Biophys. J. 2020, 118 , 1621–1633. 10.1016/j.bpj.2020.02.019.32367806
Alexa A. ; Schmidt G. ; Tompa P. ; Ogueta S. ; Vázquez J. ; Kulcsár P. ; Kovács J. ; Dombrádi V. ; Friedrich P. The Phosphorylation State of Threonine-220, a Uniquely Phosphatase-Sensitive Protein Kinase A Site in Microtubule-Associated Protein MAP2c, Regulates Microtubule Binding and Stability. Biochemistry 2002, 41 , 12427–12435. 10.1021/bi025916s.12369833
Daubner S. C. ; Le T. ; Wang S. Tyrosine Hydroxylase and Regulation of Dopamine Synthesis. Arch. Biochem. Biophys. 2011, 508 , 1–12. 10.1016/j.abb.2010.12.017.21176768
Nagatsu T. ; Nakashima A. ; Ichinose H. ; Kobayashi K. Human Tyrosine Hydroxylase in Parkinson’s Disease and in Related Disorders. J. Neural Transm. 2019, 126 , 397–409. 10.1007/s00702-018-1903-3.29995172
Louša P. ; Nedozrálová H. ; Župa E. ; Nováček J. ; Hritz J. Phosphorylation of the Regulatory Domain of Human Tyrosine Hydroxylase 1 Monitored using Non-Uniformly Sampled NMR. Biophys. Chem. 2017, 223 , 25–29. 10.1016/j.bpc.2017.01.003.28282625
Dunkley P. R. ; Dickson P. W. Tyrosine Hydroxylase Phosphorylation In Vivo. J. Neurochem. 2019, 149 , 706–728. 10.1111/jnc.14675.30714137
Herrmann J. ; Lerman L. O. ; Lerman A. Ubiquitin and Ubiquitin-Like Proteins in Protein Regulation. Circ. Res. 2007, 100 , 1276–1291. 10.1161/01.RES.0000264500.11888.f0.17495234
Vijay-Kumar S. ; Bugg C. E. ; Cook W. J. Structure of Ubiquitin Refined at 1.8 Å Resolution. J. Mol. Biol. 1987, 194 , 531–544. 10.1016/0022-2836(87)90679-6.3041007
Pratihar S. ; Sabo T. M. ; Ban D. ; Fenwick R. B. ; Becker S. ; Salvatella X. ; Griesinger C. ; Lee D. Kinetics of the Antibody Recognition Site in the Third IgG-Binding Domain of Protein G. Angew. Chem., Int. Ed. 2016, 55 , 9567–9570. 10.1002/anie.201603501.
Ulmer T. S. ; Ramirez B. E. ; Delaglio F. ; Bax A. Evaluation of Backbone Proton Positions and Dynamics in a Small Protein by Liquid Crystal NMR Spectroscopy. J. Am. Chem. Soc. 2003, 125 , 9179–9191. 10.1021/ja0350684.15369375
Berendsen H. ; van der Spoel D. ; van Drunen R. Gromacs: AMessage-Passing Parallel Molecular Dynamics Implementation. Comput. Phys. Commun. 1995, 91 , 43–56. 10.1016/0010-4655(95)00042-E.
Carballo-Pacheco M. ; Strodel B. Comparison of Force Fields for Alzheimer’s A B42: A Case Study for Intrinsically Disordered Proteins. Protein Sci. 2017, 26 , 174–185. 10.1002/pro.3064.27727496
Mu J. ; Liu H. ; Zhang J. ; Luo R. ; Chen H.-F. Recent Force Field Strategies for Intrinsically Disordered Proteins. J. Chem. Inf. Model. 2021, 61 , 1037–1047. 10.1021/acs.jcim.0c01175.33591749
Koder Hamid M. ; Månsson L. K. ; Meklesh V. ; Persson P. ; Skepö M. Molecular Dynamics Simulations of the Adsorption of an Intrinsically Disordered Protein: Force Field and Water Model Evaluation in Comparison with Experiments. Front. Mol. Biosci. 2022, 9 , 958175 10.3389/fmolb.2022.958175.36387274
Henriques J. ; Cragnell C. ; Skepö M. Molecular Dynamics Simulations of Intrinsically Disordered Proteins: Force Field Evaluation and Comparison with Experiment. J. Chem. Theory Comput. 2015, 11 , 3420–3431. 10.1021/ct501178z.26575776
Henriques J. ; Skepö M. Molecular Dynamics Simulations of Intrinsically Disordered Proteins: On the Accuracy of the TIP4P-D Water Model and the Representativeness of Protein Disorder Models. J. Chem. Theory Comput. 2016, 12 , 3407–3415. 10.1021/acs.jctc.6b00429.27243806
Ozenne V. ; Bauer F. ; Salmon L. ; Huang J.-R. ; Jensen M. R. ; Segard S. ; Bernadó P. ; Charavay C. ; Blackledge M. Flexible-Meccano: ATool for the Generation of Explicit Ensemble Descriptions of Intrinsically Disordered Proteins and Their Associated Experimental Observables. Bioinformatics 2012, 28 , 1463–1470. 10.1093/bioinformatics/bts172.22613562
Salmon L. ; Nodet G. ; Ozenne V. ; Yin G. ; Jensen M. R. ; Zweckstetter M. ; Blackledge M. NMR Characterization of Long-Range Order in Intrinsically Disordered Proteins. J. Am. Chem. Soc. 2010, 132 , 8407–8418. 10.1021/ja101645g.20499903
Shen Y. ; Bax A. Protein Backbone Chemical Shifts Predicted from Searching a Database for Torsion Angle and Sequence Homology. J. Biomol. NMR 2007, 38 , 289–302. 10.1007/s10858-007-9166-6.17610132
Li J. ; Bennett K. C. ; Liu Y. ; Martin M. V. ; Head-Gordon T. Accurate Prediction of Chemical Shifts for Aqueous Protein Structure on “Real World” Data. Chem. Sci. 2020, 11 , 3180–3191. 10.1039/C9SC06561J.34122823
Sanz-Hernández M. ; De Simone A. The Prosecco Server for Chemical Shift Predictions in Ordered and Disordered Proteins. J. Biomol. NMR 2017, 69 , 147–156. 10.1007/s10858-017-0145-2.29119515
Parr R. G. ; Yang W. Density-Functional Theory of Atoms and Molecules. Annu. Rev. Phys. Chem. 1995, 46 , 701–728. 10.1146/annurev.pc.46.100195.003413.24341393
Christensen A. S. ; Hamelryck T. ; Jensen J. H. FragBuilder: an Efficient Python Library to Setup Quantum Chemistry Calculations on Peptides Models. PeerJ. 2014, 2 , e277 10.7717/peerj.277.24688855
Stewart J. J. P. Optimization of Parameters for Semiempirical Methods V: Modification of NDDO Approximations and Application to 70 Elements. J. Mol. Model. 2007, 13 , 1173–1213. 10.1007/s00894-007-0233-4.17828561
Handy N. C. ; Cohen A. J. Left-Right Correlation Energy. Mol. Phys. 2001, 99 , 403–412. 10.1080/00268970010018431.
Hoe W.-M. ; Cohen A. J. ; Handy N. C. Assessment of a New Local Exchange Functional OPTX. Chem. Phys. Lett. 2001, 341 , 319 10.1016/S0009-2614(01)00581-4.
Perdew J. P. ; Burke K. ; Ernzerhoff M. Generalized Gradient Approximation Made Simple. Phys. Rev. Lett. 1996, 77 , 3865–3868. 10.1103/PhysRevLett.77.3865.10062328
Perdew J. P. ; Burke K. ; Ernzerhoff M. Generalized Gradient Approximation Made Simple. Phys. Rev. Lett. 1996, 77 , 3865 10.1103/PhysRevLett.77.3865.10062328
Ditchfield R. ; Hehre W. J. ; Pople J. A. Self-Consistent Molecular-Orbital Methods. IX. An Extended Gaussian-Type Basis for Molecular-Orbital Studies of Organic Molecules. J. Chem. Phys. 1971, 54 , 724–728. 10.1063/1.1674902.
Francl M. M. ; Pietro W. J. ; Hehre W. J. ; Binkley J. S. ; Gordon M. S. ; DeFrees D. J. ; Pople J. A. Self-Consistent Molecular-Orbital Methods. XXIII. A Polarization-Type Basis Set for Second-Row Elements. J. Chem. Phys. 1982, 77 , 3654–3665. 10.1063/1.444267.
Gordon M. S. ; Binkley J. S. ; Pople J. A. ; Pietro W. J. ; Hehre W. J. Self-Consistent Molecular-Orbital Methods. 22. Small Split-Valence Basis Sets for Second-Row Elements. J. Am. Chem. Soc. 1982, 104 , 2797–2803. 10.1021/ja00374a017.
Hariharan P. C. ; Pople J. A. The Influence of Polarization Functions on Molecular Orbital Hydrogenation Energies. Theor. Chim. Acta 1973, 28 , 213–222. 10.1007/BF00533485.
Hehre W. J. ; Ditchfield R. ; Pople J. A. Self-Consistent Molecular-Orbital Methods. XII. Further Extensions of Gaussian—Type Basis Sets for Use in Molecular Orbital Studies of Organic Molecules. J. Chem. Phys. 1972, 56 , 2257–2261. 10.1063/1.1677527.
Boomsma W. ; et al. PHAISTOS: A Framework for Markov Chain Monte Carlo Simulation and Inference of Protein Structure. J. Comput. Chem. 2013, 34 , 1697–1705. 10.1002/jcc.23292.23619610
Juhas M. https://github.com/martinj80/phaistos.
McGibbon R. T. ; Beauchamp K. A. ; Harrigan M. P. ; Klein C. ; Swails J. M. ; Hernández C. X. ; Schwantes C. R. ; Wang L.-P. ; Lane T. J. ; Pande V. S. MDTraj: a Modern Open Library for the Analysis of Molecular Dynamics Trajectories. Biophys. J. 2015, 109 , 1528–1532. 10.1016/j.bpj.2015.08.015.26488642
Daura X. ; Gademann K. ; Jaun B. ; Seebach D. ; van Gunsteren W. F. ; Mark A. E. Peptide Folding: When Simulation Meets Experiment. Angew. Chem., Int. Ed. 1999, 38 , 236–240. 10.1002/(SICI)1521-3773(19990115)38:1/2<236::AID-ANIE236>3.0.CO;2-M.
Lindahl E. ; Hess B. ; Van Der Spoel D. GROMACS 3.0: a Package for Molecular Simulation and Trajectory Analysis. Mol. Modeling Annual 2001, 7 , 306–317. 10.1007/s008940100045.
Adzhubei A. A. ; Sternberg M. J. ; Makarov A. A. Polyproline-II Helix in Proteins: Structure and Function. J. Mol. Biol. 2013, 425 , 2100–2132. 10.1016/j.jmb.2013.03.018.23507311
Kjaergaard M. ; Nørholm A.-B. ; Hendus-Altenburger R. ; Pedersen S. F. ; Poulsen F. M. ; Kragelund B. B. Temperature-Dependent Structural Changes in Intrinsically Disordered Proteins: Formation of α–Helices or Loss of Polyproline II?. Protein Sci. 2010, 19 , 1555–1564. 10.1002/pro.435.20556825
Tomasso M. E. ; Tarver M. J. ; Devarajan D. ; Whitten S. T. Hydrodynamic Radii of Intrinsically Disordered Proteins Determined from Experimental Polyproline II Propensities. PLoS Comput. Biol. 2016, 12 , e1004686 10.1371/journal.pcbi.1004686.26727467
Rieloff E. ; Skepö M. The Effect of Multisite Phosphorylation on the Conformational Properties of Intrinsically Disordered Proteins. Int. J. Mol. Sci. 2021, 22 , 11058 10.3390/ijms222011058.34681718
Felli I. C. ; Bermel W. ; Pierattelli R. Exclusively Heteronuclear NMR Experiments for the Investigation of Intrinsically Disordered Proteins: Focusing on Proline Residues. Magn. Reson. 2021, 2 , 511–522. 10.5194/mr-2-511-2021.
Gopal S. M. ; Wingbermühle S. ; Schnatwinkel J. ; Juber S. ; Herrmann C. ; Schäfer L. V. Conformational Preferences of an Intrinsically Disordered Protein Domain: A Case Study for Modern Force Fields. J. Phys. Chem. B 2021, 125 , 24–35. 10.1021/acs.jpcb.0c08702.33382616
Jia M. ; Yuan D. Y. ; Lovelace T. C. ; Hu M. ; Benos P. V. Causal Discovery in High-Dimensional, Multicollinear Datasets. Front. Epidemiol. 2022, 2 , 899655 10.3389/fepid.2022.899655.36778756
