
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae128
bbae128
Problem Solving Protocol
AcademicSubjects/SCI01060
IMPRINTS.CETSA and IMPRINTS.CETSA.app: an R package and a Shiny application for the analysis and interpretation of IMPRINTS-CETSA data
https://orcid.org/0009-0000-7430-3467
Gerault Marc-Antoine Department of Oncology and Pathology, Karolinska Institutet, 171 77 Stockholm, Sweden
Aix-Marseille Univ, INSERM, CNRS, Institut Paoli-Calmettes, CRCM, Marseille Protéomique, F-13009 Marseille, France

Granjeaud Samuel Aix-Marseille Univ, INSERM, CNRS, Institut Paoli-Calmettes, CRCM, Marseille Protéomique, F-13009 Marseille, France

Camoin Luc Aix-Marseille Univ, INSERM, CNRS, Institut Paoli-Calmettes, CRCM, Marseille Protéomique, F-13009 Marseille, France

Nordlund Pär Department of Oncology and Pathology, Karolinska Institutet, 171 77 Stockholm, Sweden
Institute of Molecular and Cell Biology, A*STAR, 138673, Singapore

https://orcid.org/0000-0001-9386-5967
Dai Lingyun Institute of Molecular and Cell Biology, A*STAR, 138673, Singapore
Department of Geriatrics, and Shenzhen Clinical Research Centre for Geriatrics, The First Affiliated Hospital, School of Medicine, Southern University of Science and Technology, Shenzhen 518020, China

Corresponding authors. Lingyun Dai, Department of Geriatrics, and Shenzhen Clinical Research Centre for Geriatrics, The First Affiliated Hospital, School of Medicine, Southern University of Science and Technology, Shenzhen 518020, China. E-mail: lingyun.dai@outlook.com; Pär Nordlund, Department of Oncology and Pathology, Karolinska Institutet, 171 77 Stockholm, Sweden. E-mail: par.nordlund@ki.se; Luc Camoin, Aix-Marseille Univ, INSERM, CNRS, Institut Paoli-Calmettes, CRCM, Marseille Protéomique, F-13009 Marseille, France. E-mail: luc.camoin@inserm.fr
5 2024
31 3 2024
31 3 2024
25 3 bbae12815 10 2023
10 2 2024
11 3 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (https://creativecommons.org/licenses/by-nc/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact journals.permissions@oup.com

Abstract

IMPRINTS-CETSA (Integrated Modulation of Protein Interaction States—Cellular Thermal Shift Assay) provides a highly resolved means to systematically study the interactions of proteins with other cellular components, including metabolites, nucleic acids and other proteins, at the proteome level, but no freely available and user-friendly data analysis software has been reported. Here, we report IMPRINTS.CETSA, an R package that provides the basic data processing framework for robust analysis of the IMPRINTS-CETSA data format, from preprocessing and normalization to visualization. We also report an accompanying R package, IMPRINTS.CETSA.app, which offers a user-friendly Shiny interface for analysis and interpretation of IMPRINTS-CETSA results, with seamless features such as functional enrichment and mapping to other databases at a single site. For the hit generation part, the diverse behaviors of protein modulations have been typically segregated with a two-measure scoring method, i.e. the abundance and thermal stability changes. We present a new algorithm to classify modulated proteins in IMPRINTS-CETSA experiments by a robust single-measure scoring. In this way, both the numerical changes and the statistical significances of the IMPRINTS information can be visualized on a single plot. The IMPRINTS.CETSA and IMPRINTS.CETSA.app R packages are freely available on GitHub at https://github.com/nkdailingyun/IMPRINTS.CETSA and https://github.com/mgerault/IMPRINTS.CETSA.app, respectively. IMPRINTS.CETSA.app is also available as an executable program at https://zenodo.org/records/10636134.

cellular thermal shift assay
integrated modulation of protein interaction states
mass spectrometry-based proteomics
statistical analysis
algorithm
Swedish Research Council 10.13039/501100004359 Swedish Cancer Society 10.13039/501100002794 Knut and Alice Wallenberg Foundation 10.13039/501100004063 National Natural Science Foundation of China 10.13039/501100001809 32070748 82003814 Excellent Scientific and Technological Innovation Training Program of Shenzhen RCYX20210706092040048 Singapore’s National Research Foundation NRF-CRP22–2019-0003 IBISA Plateforme Technologique Aix-Marseille Cancéropôle PACA Provence-Alpes-Côte d’Azur Région Institut Paoli-Calmettes and the Centre de Recherche en Cancérologie de Marseille Fonds Européen de Developpement Régional and Plan Cancer
==== Body
pmcINTRODUCTION

Proteins and their interactions with other biomolecules (such as metabolites, nucleic acids and other proteins) underlie the biochemical operations of almost all cellular events [1, 2]. Several decades of research have been devoted to the development of technologies and methods for measuring various aspects of proteins including amino acid sequences, expression levels, post-translational modifications and three-dimensional structures. Typically, the measurement of protein–biomolecule interactions involves working with purified proteins and/or cell lysate, making the direct detection of protein interaction states (PRINTS) under real physiological conditions a technical challenge [3–5].

The Cellular Thermal Shift Assay (CETSA), which was originally reported by our group, is an in vivo biophysical assay for measuring protein–ligand interactions in intact cells [6]. The principle of operation is based on ligand-induced thermal stabilization of target proteins. In 2014, a proof-of-concept paper showing the coupling of quantitative mass spectrometry (MS) with CETSA has demonstrated the power of MS-CETSA (also known as thermal proteome profiling, TPP) to systematically measure drug–target engagement and identify markers of drug efficacy and toxicity at the proteomic level [7]. Since then, several MS-CETSA variants have been developed, including isothermal dose response (ITDR)–MS-CETSA [8], thermal proximity co-aggregation (TPCA) [9, 10], two-dimensional (2D)-TPP, proteome integral solubility alteration (PISA) [11] and isothermal shift assay (iTSA) [12]. Several statistical software or packages have been developed to analyze these proteome-wide thermal stability data, including the TPP and TPP2D package [13, 14], the mineCETSA package [15], the NPARC package [16], the Rtpca package [17], the Inflect software [18] and the ProSAP software [19].

We note that most of these MS-CETSA applications have focused on the identification and validation of the direct interactions of small molecules or drugs with proteins [20–23]. TPCA, however, is an extension of classical MS-CETSA or TPP into the proteome-wide study of protein complex (including protein–protein interaction, PPI) dynamics in situ [24, 25]. To extract PRINTS modulations in different cellular states, we have implemented a new multi-dimensional CETSA format, named IMPRINTS (Integrated Modulation of Protein Interaction States)-CETSA [26]. The newly developed multi-dimensional IMPRINTS-CETSA method provides a highly resolved means to study the interactions of proteins with other cellular components (including metabolites, nucleic acids and other proteins) in intact cells and tissues at the proteome level [5]. In a typical IMPRINTS-CETSA experiment, a time course or concentration gradient of drug action or a cellular process [5] is studied. The IMPRINTS-CETSA analysis has been successfully applied to address important biological questions such as dissecting protein dynamics between different cell cycle phases [26], monitoring the time dependence of drug action on cellular states and comparing drug-resistant versus drug-sensitive cells [27].

The IMPRINTS-CETSA experiment involves the allocation of multiple (typically three) biological replicates of samples to be compared after thermal challenge at an isothermal temperature using a tandem mass tag (TMT) multiplex set and the arrangement of several TMT sets from numerous isotherms (typically six) along the melting range (typically 37–64°C). By packaging multiple biological replicates as well as multiple cell conditions into the same TMT multiplex set, the experimental errors introduced during the sample preparation and MS experiment are minimized [5]. Although the proof of concept has been published in Dai et al. [26], numerical analysis and intuitive interpretation are still a challenging task for many biologists who are interested in generating their own IMPRINTS-CETSA datasets.

Here, we not only present the IMPRINTS.CETSA R package for data analysis but also provide the first interactive Shiny-based graphical user interface, IMPRINTS.CETSA.app, for analyzing and interpreting IMPRINTS-CETSA data in a much more user-friendly manner.

Functionality overview

IMPRINTS-CETSA data preprocessing, normalization and visualization

The workflow starts with the pre-processing of the raw protein abundances from each heating temperature, which are typically exported (in a txt format file) from quantitative proteomics software after TMT-based quantification using mass spectrometry (Figure 1A). Currently, the most widely used software supporting TMT-based quantification are Proteome Discoverer (Thermo Scientific) and MaxQuant [28]. Both are supported by our two R packages. After joining the ProteinGroup files from each isotherm, all contaminants or reverse proteins and proteins without quantitative information are removed. Then, the protein isoform ambiguity is resolved using the principle of parsimony. Once the data have been filtered by a minimal measurement completeness and rearranged into a format of a single protein per row, the data are normalized based on the principle of minimal variance [29, 30]. Finally, the batch effect, if any, is removed by the removeBatchEffect function from the limma R package [31]. The data are now log2-transformed, so that the ratios in the original data correspond to the differences in the transformed data (Figure S1).

Figure 1 Data preprocessing, normalization and visualization for IMPRINTS-CETSA dataset. (A) IMPRINTS-CETSA dataset generation and processing workflow. To examine the integrated modulation of protein interaction states (IMPRINTS) between treatment and control samples, the samples are aliquoted and subjected to CETSA-style heat treatment typically at a minimum of six temperatures. The soluble protein fractions in the samples are separated from the denatured proteins by centrifugation and collected for tryptic digestion. Samples from each temperature are then multiplexed using TMT-based quantitative proteomics and analyzed with a mass spectrometer. The raw data are then processed by the Proteome Discoverer or MaxQuant software to obtain the raw protein abundance information for each sample. After data normalization and calculation, the relative log2 fold changes between treatment and control samples for each temperature/replicate combination can now be obtained and visualized as bar plots. The hit list can be generated using either the 2D-score method or the I-score method in the IMPRINTS.CETSA package. The IMPRINTS.CETSA.app package is a Shiny-based user-friendly interface allows the user to easily perform the data processing described above as well as to easily visualize and interpret the IMPRINTS-CETSA dataset. (B) IMPRINTS profiles examples classified into four categories: no significant changes (the NN group), only an abundance change (the CN group), only a stability change (the NC group) or changes in both (the CC group). This categorization can also be visualized in the scatter plot obtained with the 2D-score method.

The relative fold changes between treatments for each temperature/replicate combination can now be obtained and visualized as bar plots (Figure 1A). From these data, the user can map proteins to known protein complexes, calculate the correlation of the IMPRINTS profile among protein complex subunits and more importantly, segregate proteins into a few categories. For example, the modulations of PRINTS at the proteome scale can be assessed by two measures: the relative changes in the protein abundance and thermal stability. The protein abundance change is the relative fold change of protein levels at the lowest measurement temperature (typically 37°C), while the thermal stability change is the mean of the relative fold changes of protein levels at other measurement temperatures minus the protein abundance change. Such categorization allows us to distinguish proteins that have no significant changes (the NN group), only a stability change (the NC group), only an abundance change (the CN group) or changes in both (the CC group) (Figure 1B).

The layout of IMPRINTS.CETSA.app

The layout of the IMPRINTS.CETSA.app, an R package containing a Shiny application for processing, analyzing and interpreting IMPRINTS-CETSA data, has seven main tabs (Figure S2). The Home tab provides an introduction to the history and development of CETSA. Under the Analysis tab, all the data preprocessing and normalization steps for IMPRINTS-CETSA datasets can be performed in a user-friendly manner. The user can freely choose to perform the scoring using either the classic 2D-score or the new intercept score (I-score). Once the hit list is obtained, the user can plot the IMPRINTS profiles of the selected proteins and perform several different visualizations (e.g. heatmap) under the Visualization tab. To help the user to better interpret the results we also provide gene ontology enrichment modules. For example, under the Barplots network tab within the Gene Ontology tab, the user can generate an interactive protein–protein interaction (PPI) network with the corresponding IMPRINTS profile of each protein within each node. Under the Cell tab, the user can also map the proteins of interest to their subcellular localization and then plot an interactive image of a cell where each protein is represented as a dot and mapped to its corresponding cellular localization (Figure 1A). Finally, the user not only can upload their own data in each tab but can also keep the data permanently on their local machine using the Database tab to facilitate their analysis and perform dataset comparisons (Figure S2). A tutorial video is provided within the application to help the user to learn how to use it.

METHODOLOGY

Scoring by the 2D-score method

The two-measures scoring method, named as 2D-score, was developed based on the evaluation of the changes on both the protein abundance and thermal stability axes. Typically, we always include a set of measurements of soluble proteins heated at 37°C. The mean of the log2 fold changes of protein levels between different treatments (such as different cell cycle phases) at 37°C represent the abundance score.

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $abundance. score=\hat{FC_{37{}^{\circ}C}}$\end{document} , where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\hat{FC_{37{}^{\circ}C}}$\end{document} is the mean of the log2 fold changes at 37°C.

The thermal stability score is calculated by taking the mean difference of the mean of relative fold changes of protein levels at other measurement temperatures with the protein abundance score.

\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $stability. score=\frac{1}{\mathrm{N}}\sum_{\mathrm{k}=1}^{\mathrm{N}}\left[\hat{{\mathrm{FC}}_{{\mathrm{T}}_{\mathrm{k}}}} - \mathrm{abundance}.\mathrm{score}\right]$\end{document} , where \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $N$\end{document} is the number of temperatures other than 37°C and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} $\hat{FC_{T_k}}$\end{document} is the mean of the log2 fold changes at the temperature \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${T}_k$\end{document}, other than 37°C.

The proteins with a significant abundance.score are considered abundance hits and proteins with a significant stability.score are considered stability hits. Such scoring allows us to classify proteins into four categories: no significant changes (the NN group), only a stability change (the NC group), only an abundance change (the CN group) or changes in both (the CC group) which can be visualized by plotting the stability.score against the abundance.score (Figure 1B).

Scoring by the I-score method

We have also worked out a robust single-measure scoring method. To determine a statistically robust score to rank the best hits, we first assessed the mean of the differences as performed in Martinez-Val et al. [32]. However, the mean metric can reduce the overall score in the case where significant fold changes only occur at a single temperature. Therefore, we chose to compute a weighted least-squared regression on the absolute value of the fold changes, ranked in descending order, with greater weight given to the two largest fold changes. From this regression, we can obtain the intercept value, which is then multiplied by the sign of the mean fold changes to keep track of whether the protein is stabilized or destabilized (Figure 2A). We then applied a z-score normalization to these values for scoring, so the final score is called the I-score. In this way, we were able to obtain a single-value score to rank each protein. We then computed a moderated t-test for each temperature fold change and obtained what we called the combined P-value, which is the Fisher’s test P-value of the two P-values from the two largest fold changes (Figure 2B). Finally, for each protein, the –log10(combined P-value) relative to the I-score can be visualized in a volcano plot. The more the protein is in the upper-left/right corner of the plot, the more confident we can be that the PRINTS of that protein is modulated (Figure 2C). In order to obtain the threshold curve, a false discovery rate (FDR) correction is applied to the combined P-value.

Figure 2 The I-score calculation for an IMPRINTS-CETSA dataset. (A) Scheme of the method used to obtain the Intercept-score (I-score). (B) Scheme depicting the method used to obtain the combined P-value. (C) Volcano plot showing I-score versus \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{upgreek} \usepackage{mathrsfs} \setlength{\oddsidemargin}{-69pt} \begin{document} ${-\mathit{\log}}_{10}\left( Fisher\ P- value\right).$\end{document} The results indicates that protein A has a greater thermal shift than protein B.

Build-in datasets for demo analysis and data structure

The typical data structure is illustrated in Figure S1. Basically, the original data acquired at each isotherm contains the raw protein abundances for each protein group per row under different combinations of replicates and treatments (e.g. B1_Ctrl represents the biological replicate 1 of the Ctrl sample, B3_Treated represents the biological replicate 1 of the Treated sample). After joining the datasets from different isotherms, the data structure is rearranged to the format of one unique protein group per row. In the rearranged dataset, the columns containing protein abundance information are typically named with a combination of heating temperature, replicate and treatment (e.g. 52C_B1_Ctrl represents the biological replicate 1 of the Ctrl sample from the 52°C isotherm dataset).

To facilitate the testing of the functionalities of IMPRINTS.CETSA.app, we have included the two published IMPRINTS-CETSA datasets from Dai et al. [26]. Both datasets measure the IMPRINTS between different cell cycle phases of K562 cells. In the first one named ‘elutriation’, the cells were synchronized by elutriation method to allow the collection of cell fractions enriched at G1, S and G2 phases based on the cell size difference. In the second one named ‘chemarrest’, cell cycle arrest was induced by a double thymidine block method at the G1/S border (G1S), S and prometaphase (PM) phases.

Moreover, to facilitate the demonstration of the analysis of a typical IMPRINTS-CETSA dataset, we have subset the original elutriation dataset (which contains >5000 proteins in each set) and included only ~500 proteins as a small example so that the testing of the analysis part of the Shiny application can be performed much faster.

Infrastructures for development

Both packages are written primarily in the R language and depend on version 4.0 or greater. Many analytical functions and visualizations depend on different R packages, such as the ggplot2 package for non-interactive visualizations as well as the plotly and visNetwork packages for all interactive visualizations. For gene ontology analysis, the STRINGdb [33] and clusterProfiler [34] packages from Bioconductor are used. The cellular localization information of proteins is obtained from the Protein Atlas database at https://www.proteinatlas.org [35]. All package dependencies and infrastructure of IMPRINTS.CETSA.app are summarized in Figure S2. In addition, to facilitate the access and use of IMPRINTS.CETSA.app, we have compiled it as an executable program for all Windows users as a standalone application, which is available at https://zenodo.org/records/10636134.

RESULTS AND DISCUSSION

Case study 1: IMPRINTS during the progression of human cell cycle

To demonstrate the functionality of our packages, we first used a published elutriation-based cell cycle synchronization dataset [26]. The cell cycle mediates cell growth, genome duplication and division, with many proteins playing important regulatory roles [36]. Using the 2D-score method with an FDR of 1% and a median absolute deviation (MAD) of 2.5, a total of 727 proteins were found to be modulated along either the protein abundance and/or thermal stability axes (Figure S3A). Using the I-score method with an FDR of 1% and an I-score threshold of 1.5, a total of 1019 proteins were found as hits. The cyclin-B1 (CCNB1) protein and the cyclin-dependent kinase 1 (CDK1) protein, which form a central protein complex regulator that drives cells through G2 phase and mitosis, were ranked as the top hits when comparing G2 to G1 phase (Figure 3A). In both cases, we found approximately 10 times fewer hits in the G2-versus-G1 phase compared to the S-versus-G1 phase (Figure 3B), which is consistent with the previous observation that there are only small differences between these two gap phases in CETSA and their similar underlying biochemical programming [26]. Compared to the hits from the 2D-score method, 577 proteins were identified in common and 442 proteins were uniquely found as hits by the I-score method (Figure S3A). Among these 442 proteins, at least six proteins are known to be involved in cell cycle regulation based on the annotations from the Uniprot database, including cell division cycle and apoptosis regulator protein 1 (CCAR1), G1/S-specific cyclin-D3 (CCND3), anaphase-promoting complex subunit 6 (CDC16), cell division cycle 5-like protein (CDC5L), cell division cycle protein 27 homolog (CDC27) and cyclin-dependent kinase 2 (CDK2). Moreover, an over-representation analysis of the WikiPathway and KEGG databases on the hits found by the two methods indicated that the GO term cell cycle was significant with the hits found by the I-score method with a P-value lower than 0.05; suggesting the usefulness and sensitivity of the newly developed I-score method (Figure S3B).

Figure 3 Hit analysis by the two scoring methods on the elutriation-based cell cycle synchronization dataset. (A) Scatter plot from the elutriation-based cell cycle synchronization dataset obtained by the 2D-score method. (B) Volcano plot from the elutriation-based cell cycle synchronization dataset obtained by the I-score method. The results indicated that in both methods, there were approximately 10 times more hits in the phase S versus G1 than in the G2 versus G1. The top ten hits obtained by each method in each comparison group are labeled.

Case study 2: IMPRINTS after the treatment with fluorouracil-based compounds

Next, we performed a similar analysis on a published IMPRINTS-CETSA dataset that studies the interaction of fluorouracil-based compounds with proteins in living breast cancer cells [27]. 5-Fluorouracil (5-FU), a classic antimetabolite chemotherapeutic agent, is widely used as a first-line treatment for many cancers [37]. Once inside the cell, 5-FU is first metabolized to the nucleosides 5-fluorouridine (FUR) and 5-fluorodeoxyuridine (FUDR) and then metabolized further to exert diverse and complex effects on multiple pathways, including nucleotide metabolism, as well as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA) synthesis/turnover [38]. Although 5-FU is widely used and studied, the complete understanding of the key proteins and pathways at different stages of MOA is still lacking. Liang et al. have used the IMPRINTS-CETSA methods to systematically investigate the cellular metabolism and downstream effects of fluorouracil-based compounds including 5-FU, FUR and FUDR in living cells [27].

In this case study, we focused on the dataset investigating the time-dependent effects of FUR and FUDR on MCF7 cells after 2 and 12 h of treatment, respectively, resulting in four different conditions: FUR-2 h, FUDR-2 h, FUR-12 h and FUDR-12 h. At the early time point, using the 2D-score method with the same thresholds as above, FDR of 1% and a MAD of 2.5, the main modulated targets include thymidylate synthase (TYMS) (Figure 4A and B), the well-known major protein target of 5-FU [39]. In addition, the abasic site processing protein HMCES (HMCES), which is involved in DNA repair, appears as one of the most prominent hits only with FUDR treatment, whereas pseudouridylate synthase 1 homolog (PUS1) and pseudouridylate synthase 7 homolog (PUS7), which are related to the RNA modification process, appear as two of the most prominent hits only with FUR treatment, highlighting the subtle difference between the two closely related 5-FU metabolites in triggering downstream effects [27] (Figure 4A). The same top targets were also found when using the I-score method with the same thresholds as above, FDR of 1% and an I-score threshold of 1.5, although the total number of hits is much smaller (Figure 4B). The I-score method also identified two other key enzymes related to pyrimidine metabolism in the top 10 list: thymidine kinase, cytosolic (TK1) and thymidylate kinase (DTYMK) [8, 40]. In both methods, the hits and the over-representation analysis indicated a time dependency of protein interaction dynamics (Figure S4B). For example, the terms DNA damage response and miRNA regulation of DNA damage response became significant only after 12 h treatment with FUR in both two scoring methods (Figure S4B). Similarly, the terms biomarkers for pyrimidine metabolism disorders and fluoropyrimidine activity were found to be significant after 12 h treatment with FUDR (Figure S4B).

Figure 4 Hit analysis by the two scoring methods on the 5-FU dataset after 2 and 12 h of treatment. (A) Scatter plot from the 5-FU dataset after 2 and 12 h of treatment obtained by the 2D-score method. The top 10 hits obtained by each method in each comparison group are labeled. (B) Volcano plot from the 5-FU dataset after 2 and 12 h of treatment obtained by the I-score method.

CONCLUSION

Here, we present two packages for the analysis of IMPRINTS-CETSA datasets. The IMPRINTS.CETSA package provides the basic framework for robust data analysis. The IMPRINTS.CETSA.app offers a user-friendly interface for interactive analysis. Two scoring methods, 2D-score and I-score, are provided to efficiently rank targets. Two case studies have been conducted to benchmark our two scoring methods. The first case study shows that the PRINTS of many regulatory or cell cycle proteins are modulated throughout the cell cycle. The second case study shows that IMPRINTS-CETSA can follow the time course of different drug responses. Our results indicate the different sensitivities of the two scoring methods. In addition to the classical 2D-score, the parallel use of the I-score could provide a different perspective on hits as well as a more complete picture of modulated protein interaction states, i.e. reflecting the drug effects and implicated cellular processes.

We believe that our new packages will make the analysis of IMPRINTS-CETSA data more accessible, especially to those with limited knowledge of the R language programming. In addition, to further facilitate the access and use of IMPRINTS.CETSA.app, an executable program has been compiled as a standalone application for all Windows users. This will make the installation and use more user-friendly. As future work, we plan to provide a compiled version of IMPRINTS.CETSA.app that is compatible with MacOS and Linux. We also plan to reduce the number of dependencies in order to reduce the size of our packages and to avoid potential problems that could arise from R package updates. Finally, the current version of IMPRINT-CETSA data analysis focuses only on the protein level. In the next major release, peptide-level analysis should be worked out to distinguish specific functional proteoforms [41]. We will follow the issue tracking on GitHub to fix problems that we didn’t anticipate and to continue to improve based on the feedback from the community.

We believe that once the analytical barrier has been overcome, the widespread adoption of the IMPRINTS-CETSA methodology will have a major impact on the understanding of biological and disease processes and drug actions in the future. It also provides the opportunity to discover novel mechanistic protein targets and biomarkers for intervention and therapy.

Key Points

We provide a basic framework to robustly analyze the IMPRINTS-CETSA (Integrated Modulation of Protein Interaction States–Cellular Thermal Shift Assay) data format.

A user-friendly and interactive Shiny interface is provided to facilitate the analysis and interpretation of the IMPRINTS-CETSA dataset for those with limited knowledge of the R language programming.

We also present a new algorithm to classify modulated proteins in IMPRINTS-CETSA experiments through a robust single-measure scoring.

Supplementary Material

Figure_S1_IMPRINTSCETSAapp_detailed_workflow_bbae128

Figure_S2_IMPRINTSCETSAapp_structure_and_dependencies_bbae128

Figure_S3_elutriation_case_study_results_bbae128

Figure_S4_5FU_case_study_results_bbae128

Table_S1_bbae128

ACKNOWLEDGEMENTS

We would like to thank all our co-workers who helped us test the IMPRINTS.CETSA and IMRPINTS.CETSA.app packages and gave us important feedback, with special thanks to Nayana Prabhu and Ying Yu Liang from the Institute of Molecular and Cell Biology (IMCB) at A*STAR in Singapore and Anderson Ramos and Tamoghna Mandal from the Department of Oncology and Pathology at Karolinska Institutet in Stockholm.

FUNDING

This work was funded by the Swedish Research Council, the Swedish Cancer Society, the Knut and Alice Wallenberg Foundation. It was also supported by the National Natural Science Foundation of China [32070748, 82003814]; Excellent Scientific and Technological Innovation Training Program of Shenzhen [RCYX20210706092040048]; Singapore’s National Research Foundation [NRF-CRP22–2019-0003]. Marseille Proteomics (marseille-proteomique.univ-amu.fr) is supported by IBISA (Infrastructures Biologie Santé et Agronomie), Plateforme Technologique Aix-Marseille, the Cancéropôle PACA, the Provence-Alpes-Côte d’Azur Région, the Institut Paoli-Calmettes and the Centre de Recherche en Cancérologie de Marseille, Fonds Européen de Developpement Régional and Plan Cancer.

DATA AVAILABILITY

All datasets analyzed in this study are publicly available. The IMPRINTS.CETSA and IMPRINTS.CETSA.app R packages are freely available on GitHub at https://github.com/nkdailingyun/IMPRINTS.CETSA and https://github.com/mgerault/IMPRINTS.CETSA.app, respectively. The compiled version of the Shiny application of IMPRINTS.CETSA.app is available as an executable program at https://zenodo.org/records/10636134. The code used to obtain results from our case studies is available on request to the corresponding authors.

Author Biographies

Marc-Antoine Gerault is now a PhD student at the Karolinska Institutet, Sweden. His research interests focus on developing bioinformatics tools in the field of proteomics involving statistics, machine and deep learning.

Samuel Granjeaud is a research engineer at Inserm, working on the CRCM’s bioinformatics platform in France. His work focuses on proteomics and cytometry data analysis.

Luc Camoin, PhD, is the operational manager of the Marseille Proteomics platform at the Cancer Research Center of Marseille in France. His work is dedicated to the proteomics field.

Pär Nordlund is a Professor at the Karolinska Institutet, Sweden. He is originally a structural biologist and the inventor of the cellular thermal shift assay (CETSA). His group now works mainly on understanding cancer processes using CETSA.

Lingyun Dai, PhD, is a Principal Investigator and Clinical Associate Professor at the First Affiliated Hospital and School of Medicine of Southern University of Science and Technology. He works mainly on functional proteomics, systems pharmacology and CETSA technique.
==== Refs
References

1. Zuckerkandl E , PaulingL. Molecules as documents of evolutionary history. J Theor Biol 1965;8 (2 ):357–66.5876245
2. Vidal M , CusickME, BarabásiAL. Interactome networks and human disease. Cell 2011;144 (6 ):986–98.21414488
3. Budayeva HG , CristeaIM. A mass spectrometry view of stable and transient protein interactions. Adv Exp Med Biol 2014;806 :263–82.24952186
4. Rogawski R , SharonM. Characterizing endogenous protein complexes with biological mass spectrometry. Chem Rev 2022;122 (8 ):7386–414.34406752
5. Dai L , PrabhuN, LiangYY, et al. Horizontal cell biology: monitoring global changes of protein interaction states with the proteome-wide cellular thermal shift assay (CETSA). Annu Rev Biochem 2019;88 :383–408.30939043
6. Martinez Molina D , JafariR, IgnatushchenkoM, et al. Monitoring drug target engagement in cells and tissues using the cellular thermal shift assay. Science 2013;341 (6141 ):84–7.23828940
7. Savitski MM , ReinhardFBM, FrankenH, et al. Tracking cancer drugs in living cells by thermal profiling of the proteome. Science 2014;346 (6205 ):1255784.25278616
8. Lim YT , PrabhuN, DaiL, et al. An efficient proteome-wide strategy for discovery and characterization of cellular nucleotide-protein interactions. PloS One 2018;13 (12 ):e0208273.30521565
9. Tan CSH , GoKD, BisteauX, et al. Thermal proximity coaggregation for system-wide profiling of protein complex dynamics in cells. Science 2018;359 (6380 ):1170–7.29439025
10. Teitz J , SanderJ, SarkerH, Fernandez-PatronC. Potential of dissimilarity measure-based computation of protein thermal stability data for determining protein interactions. Brief Bioinform 2023;24 (3):bbad143.
11. Gaetani M , SabatierP, SaeiAA, et al. Proteome integral solubility alteration: a high-throughput proteomics assay for target deconvolution. J Proteome Res 2019;18 (11 ):4027–37.31545609
12. Ball KA , WebbKJ, ColemanSJ, et al. An isothermal shift assay for proteome scale drug-target identification. Commun Biol 2020;3 (1 ):75.32060372
13. Childs D , KurzawaN, FrankenH, et al. TPP: analyze thermal proteome profiling (TPP) experiments. R package version 3.28.0. Bioconductor 2023. 10.18129/B9.bioc.TPP.
14. Kurzawa N , BecherI, SridharanS, et al. A computational method for detection of ligand-binding proteins from dose range thermal proteome profiles. Nat Commun 2020;11 (1 ):5783.33188197
15. Dziekan JM , YuH, ChenD, et al. Identifying purine nucleoside phosphorylase as the target of quinine using cellular thermal shift assay. Sci Transl Med 2019;11 (473):eaau3174.
16. Childs D , BachK, FrankenH, et al. Nonparametric analysis of thermal proteome profiles reveals novel drug-binding proteins. Mol Cell Proteomics 2019;18 (12 ):2506–15.31582558
17. Kurzawa N , MateusA, SavitskiMM. Rtpca: an R package for differential thermal proximity coaggregation analysis. Bioinformatics 2021;37 (3 ):431–3.32717044
18. McCracken NA , Peck JusticeSA, WijeratneAB, MosleyAL. Inflect: optimizing computational workflows for thermal proteome profiling data analysis. J Proteome Res 2021;20 (4 ):1874–88.33660510
19. Ji HC , LuX, ZhengZX, et al. ProSAP: a GUI software tool for statistical analysis and assessment of thermal stability data. Brief Bioinform 2022;23 (3):bbac057.
20. Martinez Molina D , NordlundP. The cellular thermal shift assay: a novel biophysical assay for in situ drug target engagement and mechanistic biomarker studies. Annu Rev Pharmacol Toxicol 2016;56 :141–61.26566155
21. Dai L , LiZ, ChenD, et al. Target identification and validation of natural products with label-free methodology: a critical review from 2005 to 2020. Pharmacol Ther 2020;216 :107690.32980441
22. Sun J , PrabhuN, TangJ, et al. Recent advances in proteome-wide label-free target deconvolution for bioactive small molecules. Med Res Rev 2021;41 (6 ):2893–926.33533067
23. Gao P , LiuY-Q, XiaoW, et al. Identification of antimalarial targets of chloroquine by a combined deconvolution strategy of ABPP and MS-CETSA. Mil Med Res 2022;9 (1 ):30.35698214
24. Lyu H-N , FuC, ChaiX, et al. Systematic thermal analysis of the Arabidopsis proteome: thermal tolerance, organization, and evolution. Cell Syst 2023;14 (10 ):883–894.e4.37734376
25. Sun S , ZhengZ, WangJ, et al. Improved in situ characterization of protein complex dynamics at scale with thermal proximity co-aggregation. Nat Commun 2023;14 (1 ):7697.38001062
26. Dai L , ZhaoT, BisteauX, et al. Modulation of protein-interaction states through the cell cycle. Cell 2018;173 (6 ):1481–1494.e13.29706543
27. Liang YY , BacanuS, SreekumarL, et al. CETSA interaction proteomics define specific RNA-modification pathways as key components of fluorouracil-based cancer drug cytotoxicity. Cell Chem Biol 2022;29 (4 ):572–585.e8.34265272
28. Cox J , HeinMY, LuberCA, et al. Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ. Mol Cell Proteomics 2014;13 (9 ):2513–26.24942700
29. Huber W , vonHeydebreckA, SültmannH, et al. Variance stabilization applied to microarray data calibration and to the quantification of differential expression. Bioinformatics 2002;18 :S96–104.12169536
30. Liu X , LiN, LiuS, et al. Normalization methods for the analysis of unbalanced transcriptome data: a review. Front Bioeng Biotechnol 2019;7 :358.32039167
31. Ritchie ME , PhipsonB, WuD, et al. Limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 2015;43 (7 ):e47–7.25605792
32. Martinez-Val A , Bekker-JensenDB, SteigerwaldS, et al. Spatial-proteomics reveals phospho-signaling dynamics at subcellular resolution. Nat Commun 2021;12 (1 ):7113.34876567
33. Szklarczyk D , GableAL, NastouKC, et al. The STRING database in 2021: customizable protein-protein networks, and functional characterization of user-uploaded gene/measurement sets. Nucleic Acids Res 2021;49 (D1 ):D605–d612.33237311
34. Wu T , HuE, XuS, et al. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. Innovation (Camb) 2021;2 (3 ):100141.34557778
35. Uhlen M , OksvoldP, FagerbergL, et al. Towards a knowledge-based human protein atlas. Nat Biotechnol 2010;28 (12 ):1248–50.21139605
36. Ly T , AhmadY, ShlienA, et al. A proteomic chronology of gene expression through the cell cycle in human myeloid leukemia cells. Elife 2014;3 :e01630.24596151
37. Grem JL . 5-fluorouracil: forty-plus and still ticking. A review of its preclinical and clinical development. Invest New Drugs 2000;18 (4 ):299–313.11081567
38. Vodenkova S , BuchlerT, CervenaK, et al. 5-fluorouracil and other fluoropyrimidines in colorectal cancer: past, present and future. Pharmacol Ther 2020;206 :107447.31756363
39. Zhang N , YinY, XuSJ, ChenWS. 5-fluorouracil: mechanisms of resistance and reversal strategies. Molecules 2008;13 (8 ):1551–69.18794772
40. Mathews CK . Deoxyribonucleotide metabolism, mutagenesis and cancer. Nat Rev Cancer 2015;15 (9 ):528–39.26299592
41. Aebersold R , AgarJN, AmsterIJ, et al. How many human proteoforms are there? Nat Chem Biol 2018;14 (3 ):206–14.29443976
