
==== Front
Nat Commun
Nat Commun
Nature Communications
2041-1723
Nature Publishing Group UK London

39251624
51771
10.1038/s41467-024-51771-2
Article
Modelling protein complexes with crosslinking mass spectrometry and deep learning
http://orcid.org/0000-0002-9990-7643
Stahl Kolja 1
Warneke Robert 2
Demann Lorenz 2
Bremenkamp Rica 2
Hormes Björn 2
http://orcid.org/0000-0001-5881-5390
Brock Oliver 34
http://orcid.org/0000-0002-3719-7754
Stülke Jörg jstuelk@gwdg.de

2
http://orcid.org/0000-0001-5999-1310
Rappsilber Juri juri.rappsilber@tu-berlin.de

156
1 https://ror.org/03v4gjf40 grid.6734.6 0000 0001 2292 8254 Technische Universität Berlin, Chair of Bioanalytics, Berlin, Germany
2 https://ror.org/01y9bpm73 grid.7450.6 0000 0001 2364 4210 Georg-August-Universität Göttingen, Department of General Microbiology, Institute for Microbiology & Genetics, GZMB, Göttingen, Germany
3 https://ror.org/03v4gjf40 grid.6734.6 0000 0001 2292 8254 Technische Universität Berlin, Robotics and Biology Laboratory, Berlin, Germany
4 grid.517251.5 Science of Intelligence, Research Cluster of Excellence, Berlin, Germany
5 grid.6363.0 0000 0001 2218 4662 Si-M/“Der Simulierte Mensch”, a Science Framework of Technische Universität Berlin and Charité - Universitätsmedizin Berlin, Berlin, Germany
6 grid.4305.2 0000 0004 1936 7988 Wellcome Centre for Cell Biology, University of Edinburgh, Edinburgh, UK
9 9 2024
9 9 2024
2024
15 78668 9 2023
16 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Scarcity of structural and evolutionary information on protein complexes poses a challenge to deep learning-based structure modelling. We integrate experimental distance restraints obtained by crosslinking mass spectrometry (MS) into AlphaFold-Multimer, by extending AlphaLink to protein complexes. Integrating crosslinking MS data substantially improves modelling performance on challenging targets, by helping to identify interfaces, focusing sampling, and improving model selection. This extends to single crosslinks from whole-cell crosslinking MS, opening the possibility of whole-cell structural investigations driven by experimental data. We demonstrate this by revealing the molecular basis of iron homoeostasis in Bacillus subtilis.

Elucidating the structure of protein complexes is key to understanding life at the molecular level. Here, the authors improve modelling performance on challenging targets by integrating experimental distance restraints from crosslinking mass spectrometry into AlphaFold-Multimer.

Subject terms

Protein structure predictions
Proteomic analysis
Structural biology
https://doi.org/10.13039/501100001659 Deutsche Forschungsgemeinschaft (German Research Foundation) 390540038 390523135 Stülke Jörg Rappsilber Juri https://doi.org/10.13039/100004440 Wellcome Trust (Wellcome) 203149 227166 Stahl Kolja Rappsilber Juri https://doi.org/10.13039/100010269 Wellcome Trust (Wellcome) 227166 Rappsilber Juri issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Solving the structure of protein complexes is key to understanding life at the molecular level. The advent of deep learning-based methods has significantly improved the reliability of single protein structure prediction1. However, the general lack of co-evolutionary information and the lower number of solved structures makes predicting the structure of protein complexes a more difficult task2. Experimental distance restraints, such as proximal residue pairs revealed by crosslinking mass spectrometry (MS), can identify interactions3–5 and their interfaces6. The experimental information can supplement information derived from evolutionary relationships in guiding structure prediction7.

In our previous work, we improved the modelling of challenging single protein targets by leveraging distance restraints in AlphaLink7, a deep learning-based method that integrates crosslinking MS data directly into the pair representation of AlphaFold21 to bias the prediction. Here, we extend AlphaLink to protein complexes. Protein complexes pose a much harder problem because on top of predicting the structure of the individual chains, predicting interactions requires searching a space that grows exponentially with the size of the complex. Crosslinking MS data can help to cut down the search space and reduce the dependency of the prediction on evolutionary information.

In this study, we train with succinimidyl 4,4-azipentanoate (SDA) crosslinks but also successfully test with a different crosslinker. SDA is a soluble crosslinker that provides restraints between residues receptive to its first reaction step, Lys, Ser, Thr, Tyr, and other amino acids on protein surfaces8 and is readily available. SDA crosslinks are of lower resolution (expected < 25 Å Cα-Cα) compared to the photo-AA crosslinks (expected < 15 Å Cα-Cα) we used on single proteins. The lower resolution poses an additional challenge and the weaker evolutionary signal changes the balance of the data from evolutionary information towards crosslink-based distance restraints, which requires more heavy lifting of the network.

We evaluate AlphaLink on challenging heteromeric CASP15 (Critical Assessment of Structure Prediction9) and antibody-antigen targets with simulated crosslinks. Moreover, we validate AlphaLink on in-cell crosslinking MS data5 from Bacillus subtilis that used a different crosslinker, and a virally modified Cullin4-RING ubiquitin ligase (CRL4) complex10,11. We focus the evaluation on heteromeric assemblies since homomers pose a different challenge. The inherent ambiguity of self-links in homo-multimeric assemblies rarely permits distinguishing inter- from intra-chain restraints6.

Results and discussion

Integrating crosslinking MS data substantially improves the prediction quality of challenging CASP15 targets

AlphaLink with distance restraints vastly outperforms AlphaFold-Multimer2 on challenging heteromeric CASP15 targets. It achieves similar or better results than the best performing methods in CASP15, which used up to 120x more sampling to derive predictions. Integrating simulated SDA crosslinks in the modelling of eight challenging heteromeric CASP15 targets (H1129, H1134, H1140, H1141, H1142, H1144, H1166, H1167) substantially improved the DockQ12 score from 0.14 to 0.62 on average, compared to the AlphaFold-Multimer baseline (Fig. 1a) which matches the average DockQ = 0.62 of the best predictions in CASP15. For comparison, AFsample13, one of the top-performing methods in CASP15, averages a DockQ score of 0.56. Except for H1142, including crosslinking MS data always produced at least acceptable solutions according to DockQ score (DockQ ≥ 0.2312). To ensure comparability, we used the same input features as AlphaFold-Multimer (multiple sequence alignments (MSAs) and templates) and fine-tuned AlphaLink on the v2 network weights of AlphaFold-Multimer. Compared to the baseline AlphaFold-Multimer prediction, we increased recycling iterations from 3 to 20 and the number of samples from 25 to 200 (except for H1129 due to computing limitations) to be more in line with methods in CASP15. Increasing sampling and recycling iterations improve the predictions for targets where the crosslinks allow a high degree of flexibility (e.g., H1141) and side chain placement. As a control, Supplementary Fig. 1 shows AlphaLink outperforming AlphaFold-Multimer with the same recycling iterations = 3 and a comparable number of samples (10 samples for AlphaLink vs 25 for AlphaFold-Multimer).Fig. 1 AlphaLink performance on simulated data.

a DockQ comparison on 8 heteromeric CASP15 targets with simulated SDA crosslinks. AlphaLink boxplot shows the highest confidence model out of 200 predictions for N = 10 randomly sampled crosslink sets. The top1 ranked predictions (grey) correspond to submissions that were ranked first by the participants, best prediction is highlighted with an asterisk, AlphaFold-Multimer in blue. The yellow shaded area highlights nanobody, the red shaded area antibody targets. For boxplots, the line shows the median and the whiskers represent the 1.5x interquartile range. b DockQ comparison of AlphaLink and AlphaFold-Multimer on 32 antibody-antigen targets (SAbDab). Error bars represent the 95% confidence interval over N = 10 randomly sampled crosslink sets. Points show the mean. c PAE map comparison of AlphaFold-Multimer and AlphaLink for H1142. d AlphaFold-Multimer prediction for H1142 (cyan) aligned to the crystal structure (green). e AlphaLink prediction for H1142 (cyan) aligned to the crystal structure (green).

AlphaLink underperformed compared to the top groups in CASP15 for three specific targets (H1129, H1134, H1141). In contrast, AFsample excelled on these targets and ranked among the best performers. AFsample increases random exploration in baseline AlphaFold-Multimer by sampling more (up to ~24000 samples for H1166) and turning on dropout during inference which randomly removes information to create more diversity. This strategy can be implemented in AlphaLink. Leveraging crosslink data provides two key benefits: it reduces the search space by concentrating sampling on regions of interest, thereby requiring fewer samples, and it supplements evolutionary information with additional, complementary data. This is particularly valuable for challenging targets, such as antigen-antibody pairs.

For H1129, Yang et al. achieved a DockQ over 0.6 by expanding the MSA search and including the monomer predictions as a template. We also noticed that AlphaFold-Multimer predicts chain B of H1129 separately better than within the complex (0.89 vs 0.79), suggesting a potential problem with the fold-and-dock approach of AlphaFold-Multimer and thus AlphaLink which is partially side-stepped by including the monomer prediction as a template. Including the monomer prediction of chain B of H1129 as a template drastically improved the prediction quality (Supplementary Fig. 2).

H1134 has only a few contacts in the interface and is more flexible (Supplementary Fig. 3a), plausibly explaining the large spread in the DockQ scores in CASP15 (Supplementary Fig. 3b). Such flexible targets especially benefit from increased sampling and recycling. In the case of H1134, the model confidence is also less discriminative (Supplementary Fig. 4). Crosslinks can improve model selection in this case by selecting first by crosslink satisfaction and second by model confidence (see the accordingly modified AlphaLink* in Supplementary Fig. 3b). Training AlphaLink from scratch in the future could improve this by better reflecting crosslinking information in the model confidence.

Originally, in the case of H1141, SDA crosslinks were not restrictive enough (Supplementary Fig. 1). We observe two clusters that differ in the relative orientation of the subunits. Both clusters satisfy all crosslinks (blue in Supplementary Fig. 5). However, the better cluster (model confidence 0.86 versus 0.38) corresponds to the crystal structure (DockQ score 0.72). Increasing sampling and recycling helped to select the correct cluster.

The test set comprised both dimers and multimeric assemblies. The simulated crosslink-derived distance restraints had 10% sequence coverage and 20% false-discovery rate14, resulting in 33 links in the median per protein-protein interaction (including 7 false links, in the median). For comparison purposes, we selected the best predictions based on the highest model confidence2 (0.8 * interface predicted TM-score (ipTM) + 0.2 * predicted TM-score), which allows us to pick close to the best model (Supplementary Fig. 6). Indeed, ipTM and DockQ correlate (ipTM 0.6 equates roughly to DockQ 0.4)2. We compare the performance of AlphaLink with and without crosslinks to exclude the influence of other parameters, such as having trained AlphaLink on larger crops than AlphaFold-Multimer v2.2 and the additional fine-tuning. Indeed, we see that the observed improvements are the result of integrating crosslinks (Supplementary Fig. 7).

Crosslinking MS data improve modelling of challenging antibody-antigen targets

We extended the evaluation by predicting in addition to CASP15 targets also the interactions of 32 recent antibody-antigen targets from the Structural Antibody Database (SAbDab)15 (Fig. 1b). Antibody-antigen targets are challenging because the co-evolutionary signal is lower16. Here, AlphaLink improves the DockQ on average from 0.29 to 0.59 compared to AlphaFold-Multimer. To limit compute, we used three recycling iterations and five samples for both methods. These results confirm what we have seen for CASP15 where the largest improvements were observed on nanobody-antigen (yellow shaded area in Fig. 1a) and antibody-antigen targets (red shaded area in Fig. 1a)16. In these cases, crosslinking MS data drastically aided prediction. AlphaFold-Multimer models are incorrect for 5 out of 6 targets (DockQ < 0.2312) while AlphaLink generates at least medium quality models for 5 out of 6 targets (DockQ > 0.4912). Notably, for H1142, H1166, and H1167, AlphaLink produced better median score predictions than the top-ranked CASP15 submissions (Figs. 1a, c, d). The DockQ score for H1142 improved from 0.01 for AlphaFold-Multimer and 0.1 in CASP15 to 0.855 by leveraging crosslinking MS data. All true links are satisfied, while 100% of the false links are rejected. This demonstrates a high resilience of AlphaLink towards noise when predicting protein complexes, as was already observed for single proteins7. Similarly, the DockQ score for H1166 improved from 0.22 (AlphaFold-Multimer) to 0.8 (AlphaLink), again with crosslink satisfaction and noise rejection being 100%.

A single crosslink obtained in cells can dramatically improve model quality

Encouraged by the success of AlphaLink with simulated data, we modelled 135 dimeric protein-protein interactions (PPIs) with real data5 from Bacillus subtilis cells crosslinked in situ with disuccinimidyl sulfoxide (DSSO)17. This soluble crosslinker provides restraints between primarily Lys but also Ser, Thr, and Tyr groups. The data are sparse (median 1 crosslink per PPI), however even one crosslink can drastically improve the results (Fig. 2a). For example, the model confidence of the CodY-YppF interaction5,18 improves from 0.25 to 0.81 based on a single crosslink (Fig. 2b). For RpoA-RpoC (PDB 6WVK), the DockQ improves from 0.003 to 0.69 based on four in-cell crosslinks (Fig. 2c). Each crosslink would have sufficed to improve model quality substantially (DockQ 0.69-0.7). We note that changed fine-tuning (larger crops) in AlphaLink2 and AlphaFold-Multimer v2.3 alone leads to some improvements, returning correct models in 5 out of 10 cases for RpoA-RpoC, while adding crosslinks improved this to 10 out of 10.Fig. 2 AlphaLink performance on real data from Bacillus subtilis.

a Highest model confidence of AlphaLink v2.2 (N = 5) vs AlphaFold-Multimer v2.2 (M = 5) on the Bacillus subtilis data on K = 135 dimeric protein-protein interactions. Low confidence predictions for both methods (model confidence < 0.5) have higher transparency. b AlphaFold-Multimer (left) and AlphaLink (right) prediction for CodY-YppF with annotated crosslinks (blue - satisfied, red - violated). c Comparison of crystal structure (left), AlphaFold-Multimer prediction (middle) and AlphaLink prediction (right) for RpoA-RpoC with annotated crosslinks (blue - satisfied, red - violated). We omitted one crosslink that is not covered by the crystal structure.

AlphaLink is fine-tuned on model_1, in comparison, AlphaFold-Multimer predicts the targets with 5 different networks. Using different networks improves the model confidence on average by 0.03 points (Supplementary Fig. 8). As seen with the CASP15 targets, crosslinks help to focus sampling on the interesting regions, reducing the amount of sampling required. Overall, integrating crosslinking MS data increases the median model confidence from 0.42 (AlphaFold-Multimer) to 0.6 (AlphaLink). With AlphaLink, 12 additional interactions (total of 46, gain of 35%) reach model confidence > 0.75. Remarkably, the AlphaLink network was not trained on DSSO data, which differ in linked residues, linker length, and density of data from simulated SDA data. Thus, our results suggest the broad applicability of AlphaLink to varying crosslinker chemistry and the typically low data densities achieved in large-scale and whole-cell crosslinking MS experiments.

AlphaLink reveals the long-enigmatic molecular basis of iron homoeostasis in B. subtilis

Our recent in-cell crosslinking investigation of protein interactions in B. subtilis5 revealed that the conserved bacterial global regulator for iron homoeostasis, Fur, interacts with an essential protein of unknown function, YlaN18,19. We confirm here this interaction in a bacterial two-hybrid assay (Fig. 3a). In addition, and in agreement with published structures, both Fur20 and YlaN21 exhibited self-interaction. Analyses of a Fur-regulated promoter (see Fig. 3b) and in vitro binding of Fur to the promoter region (Fig. 3c) demonstrate that YlaN acts as an antagonist of Fur that prevents it from binding to its DNA targets and thus allows expression of the Fur-repressed genes in the absence of iron (see also Supplementary Discussion for detail). Accordingly, we renamed YlaN to Fpa (Fur protein antagonist). This is corroborated by the observation, that the Fpa knockout is viable if iron(III) is added to the medium22. We resorted next to AlphaLink to elucidate the structural basis of Fpa inhibiting Fur. The AlphaLink prediction (model confidence: 0.75) of the B. subtilis Fur-dimer (Fig. 3d) resembles a V-shaped conformation, similar to other Fur proteins which interact with DNA20. AlphaLink predicts a single conformation (model confidence 0.84) consistent with the crosslinking MS data between Lys-74 in the DNA-binding domain of Fur and the Lys residues 23 and 26 of Fpa5,23 (Fig. 3d). This benefited from the integration of crosslinks into the prediction, as AlphaFold-Multimer predicts two conformations for the Fur-Fpa complex, one of which is not supporting an interaction (model confidence 0.41). Fpa directly engages the DNA-binding domains of Fur and disassembles the Fur dimerisation interface, resulting in a breaking of the functional dimer and a complete re-orientation of the DNA-binding domain (see http://www.subtiwiki.uni-goettingen.de/v4/predictedComplex?id=168 and http://www.subtiwiki.uni-goettingen.de/v4/predictedComplex?id=169 for an interactive display of the Fur dimer and the Fur-Fpa complex as well as Supplementary Movie 1). This structure clearly shows why Fur bound by Fpa is incapable of binding DNA. To verify our AlphaLink model of the Fur-Fpa complex, we tested the interaction of a mutant Fur (Fur*, K18A R23E Y60E) version with Fpa. These surface mutations induce a charge reversal and destroy the salt bridges at different contact points of the binding interface between Fur and Fpa. The Fur* protein showed self-interaction as well as an interaction with the wild type Fur protein. In contrast, no interaction was observed between the mutant Fur* protein and Fpa (see Fig. 3a), confirming the predicted AlphaLink model. The proposed rotation and accompanying re-orientation of the DNA-binding domains of Fur explain the loss of the DNA-binding activity of Fur upon interaction with Fpa. Interestingly, the nature of the inducer molecule that causes the release of Fur from DNA in the absence of iron has long been debated. Although direct binding of ferrous iron as a co-repressor has been suggested24, there is no experimental evidence for such a mechanism. On the contrary, Fur alone binds a promotor in the absence of iron in the buffer (Fig. 3b). We show here that it requires the recruitment of the Fur DNA-binding domains by the iron(II)-sensitive Fpa25 to secure the release of Fur from DNA and thereby iron homoeostasis in low iron conditions (Fig. 3e).Fig. 3 The Fpa-Fur interaction controls iron homoeostasis in B. subtilis.

a Bacterial two-hybrid assay to test the interaction between Fur and Fpa. N- and C-terminal fusions of both proteins as well as the mutant Fur* protein to the T18 or T25 domain of the adenylate cyclase (CyaA) were created and the proteins were tested for interaction in E. coli BTH101. Blue colonies indicate an interaction that results in adenylate cyclase activity and subsequent expression of the reporter β-galactosidase. b Gel electrophoretic mobility shift assay of Fur binding to dhbA promoter fragments in the absence of iron. The components added to the assays are shown above the gels. The same results were obtained in three independent experiments. c The impact of Fpa on dhbA promoter activity. Strains carrying a dhbA-lacZ fusion were cultivated in LB with or without added ferric citrate, and promoter activities were determined by quantification of β-galactosidase activities. The values are averages of three independent biological replicates. Standard deviations are shown. d AlphaLink predicts the structure of the Fur-Fpa complex. Left: AlphaLink prediction (model confidence: 0.75) of the Fur dimer. Right: AlphaLink prediction (model confidence: 0.84) of the Fpa-Fur complex with crosslinking MS data. The Fur dimer undergoes a large conformational change. The dimerization region is shown in transparent. All crosslinks (shown as blue lines) are satisfied in the prediction. e Model for the control of Fur activity by Fpa. At high iron concentrations, the Fur dimer binds its target DNA sequences in the promoter regions of genes involved in iron homoeostasis. The Fpa protein binds ferrous iron and is unable to interact with Fur. If iron gets limiting, apo-Fpa forms a complex with Fur, resulting in the release of Fur from its DNA targets, and thus in expression of genes for iron uptake.

Crosslinking MS-driven modelling of a multi-protein complex

Finally, we challenged AlphaLink to model, with the help of real sulfo-SDA data, the 6-subunit CRL4DCAF1-CtD/Vprmus/SAMHD1 assembly11 (360 kDa, 3118 AA) (Fig. 4), using a crystal structure and cryo-EM density11 as ground truth. The accessory protein Vpr from certain simian immunodeficiency viruses targets the DCAF1 substrate receptor of host Cullin4-RING ubiquitin ligases (CRL4), to recruit and mark the restriction factor SAMHD1 for proteasomal degradation and thus to stimulate virus replication11,26. The v2.2 network weights of AlphaFold-Multimer fail to return meaningful structures of this assembly. This is overcome by moving to v2.3 network weights, which have been trained on larger complexes. Leveraging crosslinks improves the model confidence from 0.56 (AlphaFold-Multimer v2.3) to 0.64 (AlphaLink v2.3) (Fig. 4a). The CRL4 subunit CUL4A shows substantial movements, resulting in three conformational states according to cryo-EM data (Fig. 4b). AlphaFold-Multimer predicts a contact between CUL4A and DCAF1 that is not supported by the cryo-EM density (Fig. 4c). By contrast, AlphaLink predicts no contact (Fig. 4d). The density doesn’t contain SAMHD1 (shown in light red in Figs. 4c, d). SAMHD1 has both crosslinks to CUL4A (grey) and DDB1 (purple) which influences the arrangement. In addition, crosslinks allow AlphaLink to position the viral Vpr protein inside the experimental density in agreement with the crystal structure (PDB 6ZX9) (Figs. 4e, green, f), while AlphaFold-Multimer places Vpr incorrectly. Our predicted Vpr-DCAF1 interaction model achieved a DockQ score of 0.56 compared to 0.04 by AlphaFold-Multimer (Fig. 4f).Fig. 4 AlphaLink performance on real data from CRL4 with v2.3 weights.

a AlphaFold-Multimer (v2.3) (cyan) and AlphaLink (v2.3) (green) prediction for a virally modified CRL4 assembly. SAMHD1 is not visualised for clarity. b Schematic view of the three states of Cullin4 observed in the EM density (EMD-10612 to EMD-10614). SAMHD1 is not part of the density. c Schematic view of the AlphaFold-Multimer prediction, showing a contact not present in the density (b). d Schematic view of the AlphaLink prediction. SAMHD1 is shown in light red. e Placement of the viral Vpr protein for AlphaFold-Multimer (cyan) and AlphaLink (green) in the EM density. f Predictions of DCAF1-Vpr compared to the crystal structure (PDB 6ZX9).

Discussion

Our findings demonstrate the successful extension of AlphaLink, an experiment-assisted AI approach, to predict protein complex structures. We can now start visualising at pseudo-atomic resolution protein-protein interactions inside cells, by leveraging the very scarce data of whole-cell crosslinking MS. With this breakthrough, protein-protein interaction studies move from delivering links (who) to structures (how), inside cells and at scale. Importantly, in-cell crosslinking reveals protein interactions that are often lost upon cell lysis5. While lysis is frequently required for large-scale study of protein complexes, doing so without prior crosslinking challenges, especially transient and fragile features of complexes. AlphaLink with whole-cell crosslinking will therefore accelerate discovering hitherto hidden aspects of biology, to advance our understanding of life and widen our therapeutic options in cases of disease.

Methods

Crosslink simulation

We simulate SDA crosslinks with XWalk27 with a 25 Å Cα-Cα cutoff and trypsin digestion. XWalk identifies cross-linkable residue pairs based on solvent accessibility, crosslinker, and peptide digestion. We add FDR = 20% of noise to match the expected FDR. At least one crosslink is always incorrect, the actual FDR can therefore be much higher. False positive crosslinks can be any residue pair > 25 Å Cα-Cα, where one residue is a Lys, Ser, Thr, and Tyr. In real data, there is often an affinity towards Lys and clustering of crosslinks. These biases are not reflected in the simulation since we uniformly subsample the crosslink candidates. This doesn’t play a big role in training, and we can see that it translates to real data, but it may give better coverage during testing than what we can normally expect. The link-level FDR is simulated by shuffling the crosslinks and counting the number of incorrect links observed so far. The coverage is set to 10% and corresponds to the sequence coverage based on the longer sequence. We sample inter- and intra-protein crosslinks independently.

Integration of crosslinks

We integrate crosslinks in AlphaLink the same way we did in the monomer version7. We add a crosslink embedding layer to the neural network that projects the soft label contact map into the 128-d z-space. The projection is added to the pair representation (z). This way, the crosslinking information influences retrieval of the co-evolutionary information and the coupled updates with the MSA representation enables noise rejection.

Fine-tuning of alphafold-multimer

We switched to Uni-Fold28 since OpenFold29 didn’t support multimers. To avoid training Uni-Fold from scratch, we fine-tune the weights provided by Deepmind. For v2 we use AlphaFold-Multimer 2.2.4 weights as the starting point (https://github.com/deepmind/alphafold/releases/tag/v2.2.4). AlphaFold-Multimer was trained on PDBs deposited before 2018-04-30, predating CASP13. For v3, we use the AlphaFold-Multimer 2.3.0 weights (https://github.com/deepmind/alphafold/releases/tag/v2.3.0). AlphaFold-Multimer was trained on PDBs deposited before 2021-09-30, predating CASP15. The networks are refined on 11424 protein complexes with a total of 34054 chains from the DIPS-Plus30 training set with simulated SDA crosslinking data. DIPS-Plus contains PDBs deposited before May 2021, predating CASP15. We use Uni-Fold v2.1.0 (https://github.com/dptech-corp/Uni-Fold/releases/tag/v2.1.0). MSAs were generated with the reduced database setting. We train and test with model_1. Since we focus on heteromers, we sample homomeric crosslinks like heteromeric crosslinks during training to have more samples.

For training, we follow the refinement training regime outlined in the AlphaFold-Multimer paper but expand the crop size to 640AA to increase interface exposure and the number of crosslinks we see during training. We train on 4 A100 GPUs for 10 days. We used early stopping on the validation set which consists of proteins from CAMEO31 released after 2022.

Evaluation set up

For the comparison, we use the official predictions from CASP15. The CASP15 targets were classified as TBM/FM (template-based modelling / free modelling), meaning that there is a partial template for a subunit and/or the assembly. Our main comparison point is NBIS-AF2-multimer which corresponds to standard AlphaFold-Multimer v2. We use the same MSAs as NBIS-AF2-multimer provided by Arne Elofsson (http://duffman.it.liu.se/casp15/). We increased the recycling iterations to 20, the original comparison can be found in Supplementary Fig. 1. We randomly sample 10 crosslink sets with 10% coverage and 20% FDR for each target and predict each with a different seed for a total of 200 seeds (10 in the original comparison). We only relax the best sample (chosen by model confidence) per crosslink set. The maximum MSA cluster size is 512 sequences.

For the SAbDab comparison, we compile a new data set consisting of 33 recent antibody-antigen targets between 01-01-2022 and 11-10-2023 which represent challenging protein complexes due to the lower evolutionary signal. We only include targets which have a crystal structure with a resolution of 3 Å or better and a single assembly, to simplify the evaluation. We do not remove targets that may have homologues in the training set.

The primary evaluation metric is the DockQ score. We compute the DockQ scores for AlphaLink with the official CASP15 evaluation scripts (https://git.scicore.unibas.ch/schwede/casp15_ema).

The DockQ score is the average of the fraction of native contacts (Fnat), the interface RMSD (iRMS)32, and the ligand RMSD (LRMS)32. The interface includes all contacting residues. Residues are in contact if they are from different chains and at least one heavy atom is within 5 Å. The iRMS increases the cutoff to 10 Å. The Fnat then corresponds to the recall of the native contacts. The iRMS is the RMSD of the backbone atoms in the interface. The LRMS is the RMSD after superimposing the larger of the two structures onto the smaller one. The final DockQ score is the average DockQ score over all interfaces for protein complexes with more than two chains.

We only relax the best prediction per crosslink subset to save compute time. Relaxing only slightly changes the final scores.

For the Bacillus subtilis predictions, we use the same MSAs as O’Reilly et al5. We predict the targets again with AlphaFold-Multimer v2.2 to be comparable. The results with the original AlphaFold-Multimer v2.1 predictions from the study are shown in Supplementary Fig. 9. The model confidence might not be comparable. Except for a few targets, there are no crystal structures available for Bacillus subtilis which is why we have to resort to model confidence as an indicator of improvement. Supplementary Fig. 10a shows the correlation between the model confidence and the DockQ score and further the relationship between the model confidences of AlphaLink and AlphaFold-Multimer with respect to the DockQ score (Supplementary Fig. 10b). Although, on these hard targets, the model confidence is an overestimation, a better model confidence generally translates into a better DockQ score.

For the Cullin4 complex, we use the v3 weights to predict the structures. We always predict the full complex and use the same MSAs for both AlphaFold-Multimer and AlphaLink. We compare the prediction of DCAF1-Vpr to the crystal structure (PDB 6ZX9) with the T4 tag removed. The EM densities correspond to the EMDB accession codes: EMD-10611 (core), EMD-10612 (conformational state-1), EMD-10613 (state-2) and EMD-10614 (state-3). There are on average 21 crosslinks per interface.

The structures are visualised with PyMol v2.5.0 and ChimeraX 1.7.1.

Strains, media and growth conditions

E. coli DH5α and Rosetta DE3 (28) were used for cloning and for the expression of recombinant proteins, respectively. All B. subtilis strains used in this study are derivatives of the laboratory strain 168. They are listed in Supplementary Table 1. B. subtilis and E. coli were grown in Luria-Bertani (LB) or in sporulation (SP) medium33,34. For growth assays and the in vivo interaction experiments, B. subtilis was cultivated in LB, SP, or CSE-Glc minimal medium34,35. CSE-Glc is a chemically defined medium that contains sodium succinate (6 g/l), potassium glutamate (8 g/l), and glucose (1 g/l) as the carbon sources35. Iron sources were added as indicated. The media were supplemented with ampicillin (100 µg/ml), kanamycin (50 µg/ml), chloramphenicol (5 µg/ml), or erythromycin and lincomycin (2 and 25 µg/ml, respectively) if required. LB and SP plates were prepared by the addition of Bacto Agar (Difco) (17 g/l) to the medium. All oligonucleotides used in this study are listed in Supplementary Table 2.

DNA manipulation

Transformation of E. coli and plasmid DNA extraction were performed using standard procedures33. All commercially available plasmids, restriction enzymes, T4 DNA ligase and DNA polymerases were used as recommended by the manufacturers. B. subtilis was transformed with plasmids, genomic DNA or PCR products according to the two-step protocol34. Transformants were selected on SP plates containing erythromycin (2 µg/ml) plus lincomycin (25 µg/ml), chloramphenicol (5 µg/ml), kanamycin (10 µg/ml), or spectinomycin (250 µg/ml). DNA fragments were purified using the QIAquick PCR Purification Kit (Qiagen, Hilden, Germany). DNA sequences were determined by the dideoxy chain termination method33.

Construction of mutant strains by allelic replacement

Deletion of the fur and fpa genes was achieved by transformation of B. subtilis 168 or GP879 with a PCR product constructed using oligonucleotides to amplify DNA fragments flanking the target genes and an appropriate intervening resistance cassette36. The integrity of the regions flanking the integrated resistance cassette was verified by sequencing PCR products of about 1100 bp amplified from chromosomal DNA of the resulting mutant strains.

Phenotypic analysis

In B. subtilis, amylase activity was detected after growth on plates containing nutrient broth (7.5 g/l), 17 g Bacto agar/l (Difco) and 5 g hydrolysed starch/l (Connaught). Starch degradation was detected by sublimating iodine onto the plates.

Quantitative studies of lacZ expression in B. subtilis were performed as follows: cells were grown in CSE-Glc or LB medium supplemented with iron sources as indicated. Cells were harvested at OD600 of 0.5 to 0.8. b-Galactosidase specific activities were determined with cell extracts obtained by lysozyme treatment34. One unit of β-galactosidase is defined as the amount of enzyme which produces 1 nmol of o-nitrophenol per min at 28 °C.

Plasmid constructions

To express the Fur and Fpa proteins carrying a N-terminal His-tag in E. coli, the fur and fpa genes were amplified using chromosomal DNA of B. subtilis 168 as the template and appropriate oligonucleotides that attached specific restriction sites to the fragment. Those were: BamHI and XhoI for cloning fur in pET-SUMO (Invitrogen, Germany), and BamHI and SalI for cloning fpa in pWH84437. The resulting plasmids were pGP3589 and pGP2583 for Fur and Fpa, respectively.

For overexpression of fpa in B. subtilis, we constructed plasmid pGP3897. For this purpose, the fpa gene was amplified and cloned between the BamHI and SalI site of the expression vector pBQ20038.

Plasmid pAC739 was used to construct a translational fusion of the dhbA promoter region to the promoterless lacZ gene. For this purpose, the promoter region was amplified using oligonucleotides that attached EcoRI and BamHI restriction to the ends of the products. The fragments were cloned between the EcoRI and BamHI sites of pAC7. The resulting plasmid was pGP3594.

Protein expression and purification

E. coli Rosetta(DE3) was transformed with the plasmid pGP37140, pGP2583, and pGP3589 encoding His-tagged versions of PtsH, Fpa, and Fur, respectively. For overexpression, cells were grown in 2x LB and expression was induced by the addition of isopropyl 1-thio-β-D-galactopyranoside (final concentration, 1 mM) to exponentially growing cultures (OD600 of 0.8). The His-tagged proteins were purified in 1x ZAP buffer (50 mM Tris-HCl, 200 mM NaCl, pH 7.5). Cells were lysed by four passes (18,000 p.s.i.) through an HTU DIGI-F press (G. Heinemann, Germany). After lysis, the crude extract was centrifuged at 46,400 × g for 60 min and then passed over a Ni2+nitrilotriacetic acid column (IBA, Göttingen, Germany). The proteins were eluted with an imidazole gradient. After elution, the fractions were tested for the desired protein using SDS-PAGE. The purified proteins were concentrated in a Vivaspin turbo 15 (Sartorius) centrifugal filter device (cut-off 5 or 50 kDa). The protein samples were stored at −80 °C until further use. The protein concentration was determined according to the method of Bradford41 using the Bio-Rad dye binding assay and bovine serum albumin as the standard.

Electromobility shift assay (EMSA) with DNA

To analyse the binding of Fur to the dhbA promoter region, we performed EMSA assays with a 284 bp dhbA promoter fragment that carries the Fur binding site and purified Fur, Fpa, and PtsH proteins. 200 ng of DNA and 80 pmoI of the proteins were used. The samples were first prepared without the proteins only with DNA, buffer and water and heated for 2 minutes at 95 °C. Then the proteins were added in different combinations and the samples were incubated for 30 minutes at 37 °C. Meanwhile, the EMSA gels were applied to a pre run at 90 V for 30 minutes immersed in TBE buffer (28). Afterwards, 2 µl of the loading dye were added and the samples were loaded into the gel pockets. The gel was run for 3 hours at 110 V. Then, the gels were immersed in TBE containing HDGreen® fluoreszence dye (Intas, Germany). After 2 minutes the gels were photographed under UV light.

Bacterial two-hybrid assay

Primary protein-protein interactions were identified by bacterial two-hybrid (BACTH) analysis42. The BACTH system is based on the interaction-mediated reconstruction of Bordetella pertussis adenylate cyclase (CyaA) activity in E. coli BTH101. Functional complementation between two fragments (T18 and T25) of CyaA as a consequence of the interaction between bait and prey molecules results in the synthesis of cAMP, which is monitored by measuring the β-galactosidase activity of the cAMP-CAP-dependent promoter of the E. coli lac operon. Plasmids pUT18C and p25N allow the expression of proteins fused to the T18 and T25 fragments of CyaA, respectively. For these experiments, we used the plasmids pGP3868-pGP3875, which encode N-and C-terminal fusions of T18 or T25 to fur and fpa. The plasmids were obtained by cloning the fur and fpa between the KpnI and BamHI sites of pUT18C and p25N42. The mutant fur* allele was purchased from Eurofins Genomics (Germany) and then amplified and cloned as the wild type fur gene. The resulting plasmids were then used for co-transformation of E. coli BTH101 and the protein-protein interactions were then analysed by plating the cells on LB plates containing 100 µg/ml ampicillin, 50 µg/ml kanamycin, 40 µg/ml X-Gal (5-bromo-4-chloro-3-indolyl-ß-D-galactopyranoside), and 0.5 mM IPTG (isopropyl-ß-D-thiogalactopyranoside). The plates were incubated for a maximum of 36 h at 28 °C.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Supplementary Information

Reporting Summary

Description of Additional Supplementary Files

Supplementary Movie 1

Source data

Transparent Peer Review file

Source data

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-024-51771-2.

Acknowledgements

We thank Andrea Graziadei and David Schwefel for their feedback on the manuscript and David Schwefel for helping to identify potential mutations. Christina Herzberg is acknowledged for the help with some reporter assays. We are grateful to the Uni-Fold team for providing a fully open-source and trainable reimplementation of AlphaFold-Multimer (https://github.com/dptech-corp/Uni-Fold). Figure 3/panel e, created with BioRender.com, released under a Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International license. This research was supported by the Wellcome Trust [Grant number 227166] (JR, KS) and the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany´s Excellence Strategy – EXC 2008 – project number 390540038 (JR) – UniSysCat (JR) and under Germany’s Excellence Strategy – EXC 2002/1 “Science of Intelligence” – project number 390523135 (OB), as well as CRC1565 [Grant number 469281184] (JS) and SPP1879 [project number STU214/ 16-2] (JS). The Wellcome Centre for Cell Biology is supported by core funding from the Wellcome Trust [203149].

Author contributions

Software development: K.S. Data Analysis: K.S., L.D., R.B., R.W., B.H. Supervision: J.R., O.B., J.S. Supplement information: L.D., R.B., R.W., B.H., J.S. Writing- initial draft: K.S., J.R., O.B. Writing- editing and revision: K.S., J.R., O.B., L.D., R.B., R.W., B.H., J.S.

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.

Funding

Open Access funding enabled and organized by Projekt DEAL.

Data availability

The AlphaLink models generated in this study based on experimental crosslinks have been deposited as integrative/hybrid models in PDB-Dev43 under accession codes PDBDEV_00000221 and PDBDEV_G_1000003. The previously published data we used in this study are available in the PRIDE and EM database under accession code PXD020453 (Cullin4 crosslinks), EMD-10611 (core), EMD-10612 (state-1), EMD-10613 (state-2) and EMD-10614 (state-3); JPST001796 and PXD035508 JPST001797 and PXD035519 JPST001791 and PXD035362 (Bacillus subtilis). Source data are provided in this paper.

Code availability

The code for AlphaLink is deposited at https://github.com/Rappsilber-Laboratory/AlphaLink2. Model weights were deposited at Zenodo: 10.5281/zenodo.8007238.

Competing interests

The other authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Jumper J Highly accurate protein structure prediction with AlphaFold Nature 2021 596 583 589 10.1038/s41586-021-03819-2 34265844
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).34265844 10.1038/s41586-021-03819-2
2. Evans, R. et al. Protein complex prediction with alphafold-multimer. Preprint at bioRxiv 2021.10.04.463034 (2022).
3. Tang, X. & Bruce, J. E. Chemical cross-linking for protein–protein interaction studies. in Mass Spectrometry of Proteins and Peptides: Methods and Protocols (eds. Lipton, M. S. & Paša-Tolic, L.) 283–293 (Humana Press, 2009).
4. Lenz S Reliable identification of protein-protein interactions by crosslinking mass spectrometry Nat. Commun. 2021 12 3564 10.1038/s41467-021-23666-z 34117231
Lenz, S. et al. Reliable identification of protein-protein interactions by crosslinking mass spectrometry. Nat. Commun. 12, 3564 (2021).34117231 10.1038/s41467-021-23666-z
5. O’Reilly FJ Protein complexes in cells by AI-assisted structural proteomics Mol. Syst. Biol. 2023 19 e11544 10.15252/msb.202311544 36815589
O’Reilly, F. J. et al. Protein complexes in cells by AI-assisted structural proteomics. Mol. Syst. Biol. 19, e11544 (2023).36815589 10.15252/msb.202311544
6. Maiolica A Structural analysis of multiprotein complexes by cross-linking, mass spectrometry, and database searching Mol. Cell. Proteom. 2007 6 2200 2211 10.1074/mcp.M700274-MCP200
Maiolica, A. et al. Structural analysis of multiprotein complexes by cross-linking, mass spectrometry, and database searching. Mol. Cell. Proteom. 6, 2200–2211 (2007).10.1074/mcp.M700274-MCP200
7. Stahl, K., Graziadei, A., Dau, T., Brock, O. & Rappsilber, J. Protein structure prediction with in-cell photo-crosslinking mass spectrometry and deep learning. Nat. Biotechnol. 41,1810–1819 (2023).
8. Belsom A Schneider M Fischer L Brock O Rappsilber J Serum albumin domain structures in human blood serum by mass spectrometry and computational biology Mol. Cell. Proteom. 2016 15 1105 1116 10.1074/mcp.M115.048504
Belsom, A., Schneider, M., Fischer, L., Brock, O. & Rappsilber, J. Serum albumin domain structures in human blood serum by mass spectrometry and computational biology. Mol. Cell. Proteom. 15, 1105–1116 (2016).10.1074/mcp.M115.048504
9. Moult J Pedersen JT Judson R Fidelis K A large-scale experiment to assess protein structure prediction methods Proteins 1995 23 ii v 10.1002/prot.340230303 8710822
Moult, J., Pedersen, J. T., Judson, R. & Fidelis, K. A large-scale experiment to assess protein structure prediction methods. Proteins 23, ii–v (1995).8710822 10.1002/prot.340230303
10. Mahon C Krogan NJ Craik CS Pick E Cullin E3 ligases and their rewiring by viral factors Biomolecules 2014 4 897 930 10.3390/biom4040897 25314029
Mahon, C., Krogan, N. J., Craik, C. S. & Pick, E. Cullin E3 ligases and their rewiring by viral factors. Biomolecules 4, 897–930 (2014).25314029 10.3390/biom4040897
11. Banchenko S Structural insights into Cullin4-RING ubiquitin ligase remodelling by Vpr from simian immunodeficiency viruses PLoS Pathog. 2021 17 e1009775 10.1371/journal.ppat.1009775 34339457
Banchenko, S. et al. Structural insights into Cullin4-RING ubiquitin ligase remodelling by Vpr from simian immunodeficiency viruses. PLoS Pathog. 17, e1009775 (2021).34339457 10.1371/journal.ppat.1009775
12. Basu S Wallner B DockQ: a quality measure for protein-protein docking models PLoS One 2016 11 e0161879 10.1371/journal.pone.0161879 27560519
Basu, S. & Wallner, B. DockQ: a quality measure for protein-protein docking models. PLoS One 11, e0161879 (2016).27560519 10.1371/journal.pone.0161879
13. Wallner, B. AFsample: improving multimer prediction with AlphaFold using massive sampling. Bioinformatics 39, btad573 (2023).
14. Fischer L Rappsilber J Quirks of error estimation in cross-linking/mass spectrometry Anal. Chem. 2017 89 3829 3833 10.1021/acs.analchem.6b03745 28267312
Fischer, L. & Rappsilber, J. Quirks of error estimation in cross-linking/mass spectrometry. Anal. Chem. 89, 3829–3833 (2017).28267312 10.1021/acs.analchem.6b03745
15. Dunbar J SAbDab: the structural antibody database Nucleic Acids Res. 2014 42 D1140 D1146 10.1093/nar/gkt1043 24214988
Dunbar, J. et al. SAbDab: the structural antibody database. Nucleic Acids Res. 42, D1140–D1146 (2014).24214988 10.1093/nar/gkt1043
16. Yin R Feng BY Varshney A Pierce BG Benchmarking alphafold for protein complex modeling reveals accuracy determinants Protein Sci. 2022 31 e4379 10.1002/pro.4379 35900023
Yin, R., Feng, B. Y., Varshney, A. & Pierce, B. G. Benchmarking alphafold for protein complex modeling reveals accuracy determinants. Protein Sci. 31, e4379 (2022).35900023 10.1002/pro.4379
17. Kao, A. et al. Development of a novel cross-linking strategy for fast and accurate identification of cross-linked peptides of protein complexes. Mol. Cell. Proteomics 10, M110.002212 (2011).
18. Pedreira T Elfmann C Stülke J The current state of SubtiWiki, the database for the model organism Bacillus subtilis Nucleic Acids Res. 2022 50 D875 D882 10.1093/nar/gkab943 34664671
Pedreira, T., Elfmann, C. & Stülke, J. The current state of SubtiWiki, the database for the model organism Bacillus subtilis. Nucleic Acids Res. 50, D875–D882 (2022).34664671 10.1093/nar/gkab943
19. Wicke D Meißner J Warneke R Elfmann C Stülke J Understudied proteins and understudied functions in the model bacterium bacillus subtilis—a major challenge in current research Mol. Microbiol. 2023 120 19 18 10.1111/mmi.15053
Wicke, D., Meißner, J., Warneke, R., Elfmann, C. & Stülke, J. Understudied proteins and understudied functions in the model bacterium bacillus subtilis—a major challenge in current research. Mol. Microbiol. 120, 19–18 (2023).10.1111/mmi.15053
20. Butcher J Sarvan S Brunzelle JS Couture J-F Stintzi A Structure and regulon of campylobacter jejuni ferric uptake regulator fur define apo-fur regulation Proc. Natl Acad. Sci. USA. 2012 109 10047 10052 10.1073/pnas.1118321109 22665794
Butcher, J., Sarvan, S., Brunzelle, J. S., Couture, J.-F. & Stintzi, A. Structure and regulon of campylobacter jejuni ferric uptake regulator fur define apo-fur regulation. Proc. Natl Acad. Sci. USA. 109, 10047–10052 (2012).22665794 10.1073/pnas.1118321109
21. Xu L Crystal structure of S. aureus YlaN, an essential leucine rich protein involved in the control of cell shape Proteins 2007 68 438 445 10.1002/prot.21377 17469204
Xu, L. et al. Crystal structure of S. aureus YlaN, an essential leucine rich protein involved in the control of cell shape. Proteins 68, 438–445 (2007).17469204 10.1002/prot.21377
22. Peters JM A Comprehensive, CRISPR-based functional analysis of essential genes in bacteria Cell 2016 165 1493 1506 10.1016/j.cell.2016.05.003 27238023
Peters, J. M. et al. A Comprehensive, CRISPR-based functional analysis of essential genes in bacteria. Cell 165, 1493–1506 (2016).27238023 10.1016/j.cell.2016.05.003
23. Elfmann C Stülke J PAE viewer: a webserver for the interactive visualization of the predicted aligned error for multimer structure predictions and crosslinks Nucleic Acids Res. 2023 51 W404 W410 10.1093/nar/gkad350 37140053
Elfmann, C. & Stülke, J. PAE viewer: a webserver for the interactive visualization of the predicted aligned error for multimer structure predictions and crosslinks. Nucleic Acids Res. 51, W404–W410 (2023).37140053 10.1093/nar/gkad350
24. Lee J-W Helmann JD Functional specialization within the fur family of metalloregulators Biometals 2007 20 485 499 10.1007/s10534-006-9070-7 17216355
Lee, J.-W. & Helmann, J. D. Functional specialization within the fur family of metalloregulators. Biometals 20, 485–499 (2007).17216355 10.1007/s10534-006-9070-7
25. Boyd, J. M. et al. YlaN is an iron(II) binding protein that functions to relieve Fur-mediated repression of gene expression in Staphylococcus aureus. bioRxiv 2023.10.03.560778 (2023).
26. Fregoso OI Evolutionary toggling of Vpx/Vpr specificity results in divergent recognition of the restriction factor SAMHD1 PLoS Pathog. 2013 9 e1003496 10.1371/journal.ppat.1003496 23874202
Fregoso, O. I. et al. Evolutionary toggling of Vpx/Vpr specificity results in divergent recognition of the restriction factor SAMHD1. PLoS Pathog. 9, e1003496 (2013).23874202 10.1371/journal.ppat.1003496
27. Kahraman A Malmström L Aebersold R Xwalk: computing and visualizing distances in cross-linking experiments Bioinformatics 2011 27 2163 2164 10.1093/bioinformatics/btr348 21666267
Kahraman, A., Malmström, L. & Aebersold, R. Xwalk: computing and visualizing distances in cross-linking experiments. Bioinformatics 27, 2163–2164 (2011).21666267 10.1093/bioinformatics/btr348
28. Li, Z. et al. Uni-Fold: An open-source platform for developing protein folding models beyond alphafold. Preprint at bioRxiv 2022.08.04.502811 (2022).
29. Ahdritz, G. et al. OpenFold: Retraining AlphaFold2 yields new insights into its learning mechanisms and capacity for generalization. Nat Methods. 21, 1–11 (2024).
30. Townshend, R., Bedi, R., Suriana, P. & Dror, R. End-to-end learning on 3d protein structure for interface prediction. Adv. Neural Inf. Process. Syst. 32, (2019).
31. Leemann M Automated benchmarking of combined protein structure and ligand conformation prediction Proteins 2023 91 1912 1924 10.1002/prot.26605 37885318
Leemann, M. et al. Automated benchmarking of combined protein structure and ligand conformation prediction. Proteins 91, 1912–1924 (2023).37885318 10.1002/prot.26605
32. Méndez R Leplae R De Maria L Wodak SJ Assessment of blind predictions of protein-protein interactions: current status of docking methods Proteins 2003 52 51 67 10.1002/prot.10393 12784368
Méndez, R., Leplae, R., De Maria, L. & Wodak, S. J. Assessment of blind predictions of protein-protein interactions: current status of docking methods. Proteins 52, 51–67 (2003).12784368 10.1002/prot.10393
33. Sambrook, J., Fritsch, E. F., Maniatis, T. & Others. Molecular cloning: a laboratory manual. (Cold spring harbor laboratory press, 1989).
34. Kunst F Rapoport G Salt stress is an environmental signal affecting degradative enzyme synthesis in Bacillus subtilis J. Bacteriol. 1995 177 2403 2407 10.1128/jb.177.9.2403-2407.1995 7730271
Kunst, F. & Rapoport, G. Salt stress is an environmental signal affecting degradative enzyme synthesis in Bacillus subtilis. J. Bacteriol. 177, 2403–2407 (1995).7730271 10.1128/jb.177.9.2403-2407.1995
35. Schmalisch MH Bachem S Stülke J Control of the bacillus subtilis antiterminator protein GlcT by phosphorylation. elucidation of the phosphorylation chain leading to inactivation of GlcT J. Biol. Chem. 2003 278 51108 51115 10.1074/jbc.M309972200 14527945
Schmalisch, M. H., Bachem, S. & Stülke, J. Control of the bacillus subtilis antiterminator protein GlcT by phosphorylation. elucidation of the phosphorylation chain leading to inactivation of GlcT. J. Biol. Chem. 278, 51108–51115 (2003).14527945 10.1074/jbc.M309972200
36. Diethmaier C A novel factor controlling bistability in bacillus subtilis: the YmdB protein affects flagellin expression and biofilm formation J. Bacteriol. 2011 193 5997 6007 10.1128/JB.05360-11 21856853
Diethmaier, C. et al. A novel factor controlling bistability in bacillus subtilis: the YmdB protein affects flagellin expression and biofilm formation. J. Bacteriol. 193, 5997–6007 (2011).21856853 10.1128/JB.05360-11
37. Schirmer F Ehrt S Hillen W Expression, inducer spectrum, domain structure, and function of MopR, the regulator of phenol degradation in Acinetobacter calcoaceticus NCIB8250 J. Bacteriol. 1997 179 1329 1336 10.1128/jb.179.4.1329-1336.1997 9023219
Schirmer, F., Ehrt, S. & Hillen, W. Expression, inducer spectrum, domain structure, and function of MopR, the regulator of phenol degradation in Acinetobacter calcoaceticus NCIB8250. J. Bacteriol. 179, 1329–1336 (1997).9023219 10.1128/jb.179.4.1329-1336.1997
38. Martin-Verstraete I Débarbouillé M Klier A Rapoport G Interactions of wild-type and truncated LevR of Bacillus subtilis with the upstream activating sequence of the levanase operon J. Mol. Biol. 1994 241 178 192 10.1006/jmbi.1994.1487 8057358
Martin-Verstraete, I., Débarbouillé, M., Klier, A. & Rapoport, G. Interactions of wild-type and truncated LevR of Bacillus subtilis with the upstream activating sequence of the levanase operon. J. Mol. Biol. 241, 178–192 (1994).8057358 10.1006/jmbi.1994.1487
39. Weinrauch Y Msadek T Kunst F Dubnau D Sequence and properties of comQ, a new competence regulatory gene of Bacillus subtilis J. Bacteriol. 1991 173 5685 5693 10.1128/jb.173.18.5685-5693.1991 1715859
Weinrauch, Y., Msadek, T., Kunst, F. & Dubnau, D. Sequence and properties of comQ, a new competence regulatory gene of Bacillus subtilis. J. Bacteriol. 173, 5685–5693 (1991).1715859 10.1128/jb.173.18.5685-5693.1991
40. Pietack N In vitro phosphorylation of key metabolic enzymes from Bacillus subtilis: PrkC phosphorylates enzymes from different branches of basic metabolism J. Mol. Microbiol. Biotechnol. 2010 18 129 140 20389117
Pietack, N. et al. In vitro phosphorylation of key metabolic enzymes from Bacillus subtilis: PrkC phosphorylates enzymes from different branches of basic metabolism. J. Mol. Microbiol. Biotechnol. 18, 129–140 (2010).20389117
41. Bradford MM A rapid and sensitive method for the quantitation of microgram quantities of protein utilizing the principle of protein-dye binding Anal. Biochem. 1976 72 248 254 10.1016/0003-2697(76)90527-3 942051
Bradford, M. M. A rapid and sensitive method for the quantitation of microgram quantities of protein utilizing the principle of protein-dye binding. Anal. Biochem. 72, 248–254 (1976).942051 10.1016/0003-2697(76)90527-3
42. Karimova G Pidoux J Ullmann A Ladant D A bacterial two-hybrid system based on a reconstituted signal transduction pathway Proc. Natl Acad. Sci. Usa. 1998 95 5752 5756 10.1073/pnas.95.10.5752 9576956
Karimova, G., Pidoux, J., Ullmann, A. & Ladant, D. A bacterial two-hybrid system based on a reconstituted signal transduction pathway. Proc. Natl Acad. Sci. Usa. 95, 5752–5756 (1998).9576956 10.1073/pnas.95.10.5752
43. Vallat B Webb B Westbrook JD Sali A Berman HM Development of a prototype system for archiving integrative/hybrid structure models of biological macromolecules Structure 2018 26 894 904.e2 10.1016/j.str.2018.03.011 29657133
Vallat, B., Webb, B., Westbrook, J. D., Sali, A. & Berman, H. M. Development of a prototype system for archiving integrative/hybrid structure models of biological macromolecules. Structure 26, 894–904.e2 (2018).29657133 10.1016/j.str.2018.03.011
