
==== Front
Anal Chem
Anal Chem
ac
ancham
Analytical Chemistry
0003-2700
1520-6882
American Chemical Society

39213479
10.1021/acs.analchem.4c02650
Technical Note
pyMS-Vis, an Open-Source Python Application for Visualizing and Investigating Deconvoluted Top-Down Mass Spectrometric Experiments: A Histone Proteoform Case Study
https://orcid.org/0000-0001-6107-3666
Pesavento James J. *†
Bindra Megan S. †
Das Udayan †
Rommelfanger Sarah R. ‡§
https://orcid.org/0000-0003-3575-3224
Zhou Mowei ∥⊥
Paša-Tolić Ljiljana ∥
Umen James G. ‡§
† Saint Mary’s College of California, Moraga, California 94575, United States
‡ Donald Danforth Plant Science Center, St. Louis, Missouri 63132, United States
§ Washington University in St. Louis, St. Louis, Missouri 63130, United States
∥ Environmental Molecular Science Laboratory, Pacific Northwest National Laboratory, Richland, Washington 99354, United States
* Tel: 925-631-4430; Email: jjp6@stmarys-ca.edu.
30 08 2024
17 09 2024
96 37 1472714733
21 05 2024
19 08 2024
16 08 2024
© 2024 The Authors. Published by American Chemical Society
2024
The Authors
https://creativecommons.org/licenses/by-nc-nd/4.0/ Permits non-commercial access and re-use, provided that author attribution and integrity are maintained; but does not permit creation of adaptations or other derivative works (https://creativecommons.org/licenses/by-nc-nd/4.0/).

We report the development of an open-source Python application that provides quantitative and qualitative information from deconvoluted liquid-chromatography top-down mass spectrometry (LC-TDMS) data sets. This simple-to-use program allows users to search masses-of-interest across multiple LC-TDMS runs and provides visualization of their ion intensities and elution characteristics while quantifying their abundances relative to one another. Focusing on proteoform-rich histone proteins from the green microalga Chlamydomonas reinhardtii, we were able to quantify proteoform abundances across different growth conditions and replicates in minutes instead of hours typically needed for manual spreadsheet-based analysis. This resulted in extending previously published qualitive observations on Chlamydomonas histone proteoforms into quantitative ones, leading to an exciting new discovery on alpha-amino termini processing exclusive to histone H2A family members. Lastly, the script was intentionally developed with readability and customizability in mind so that fellow mass spectrometrists can modify the code to suit their lab-specific needs.

National Institutes of Health 10.13039/100000002 R01 GM126557 Environmental Molecular Sciences Laboratory 10.13039/100019109 10.46936/expl.proj.2019.51056/60006669 Division of Molecular and Cellular Biosciences 10.13039/100000152 MCB 1943493 Division of Molecular and Cellular Biosciences 10.13039/100000152 MCB 1515220 document-id-old-9ac4c02650
document-id-new-14ac4c02650
ccc-price
==== Body
pmcIntroduction

A typical proteomics study focusing on intact proteins by mass spectrometry (MS), referred to as top-down mass spectrometry (TDMS), often incorporates liquid chromatography (LC) to facilitate the separation and identification of proteoforms1 followed by intact mass (MS1) and tandem mass (MS2) measurements. The resulting data sets can be highly complex with tens of thousands of molecules detected in a single experiment. The raw LC-MS data is then deconvoluted to reduce spectral complexity. This occurs by employing decharging and deisotoping2 steps that effectively transform ion intensity for all charge states into a single, summed intensity of its monoisotopic mass, which is then aggregated across all identified charge states. Since deconvolution occurs on the MS1 and MS2 (or more) level, the resulting lists of mass features are information rich and can be used for further proteomic investigations. Often, this processed data is used to first identify, and often simultaneously to quantify, proteoforms using such software as ProsightPD, MASH Explorer,3 TopPIC,4 and MSPathFinderT.5 When multiple condition/experimental data sets are simultaneously analyzed together, the collective output is a comparative proteome-wide landscape that may identify important condition-specific proteoforms and generate new hypotheses. These software suites provide many tools and options to satisfy the needs of a broad user base; however, as research lines become more specific and directed, specialized analytical approaches are sometimes needed and are not always available from commercial software. Analytical bottlenecks that are not addressed by existing software can be solved through labor intensive manual analyses and/or by creation of custom-developed software or scripts.

Over the past ten years, the Python scripting language has been steadily gaining popularity for its utility in processing mass spectral data. The use of Python in data analysis has become widespread due to its multiplatform compatibility, being an easy-to-learn coding language, and the availability of third-party packages for customization and extensibility. Additionally, it is an open-source language that offers promising computational solutions to questions raised by mass spectrometrists across many diverse fields. Many currently available mass spectrometry-based Python scripts are packages that contain useful methods/modules that can be integrated into more application-specific scripts. For example, pyOpenMS6 and Pyteomics7 offer an extensive collection of specialized modules such as reading/writing different mass spectral file formats, peak-picking algorithms, and retention time prediction. Other scripts have been developed for more specialized investigations, such as PTM detection in bottom-up MS data8 or processing and visualization of tandem MS (MS2) spectra.9 These packages contain useful analytical methods and visualization for small molecule or bottom-up data sets; however, Python scripts that offer downstream analysis and visualization of TDMS data sets have yet to be developed.

For targeted TDMS studies where the proteoforms-of-interest have been characterized with high confidence, it is informative to compare intraexperimental ion abundances of related proteoforms, such as post-translational modifications (PTMs) or splice variants, to estimate their respective abundances.10,11 This approach is different than the interexperimental abundance quantitation offered in many free and commercially available proteomics analysis software (e.g., spectral counting, isobaric mass labeling, etc.). An example of an intraexperimental targeted TDMS analysis is representing acetylated proteoform abundances as a fraction of all proteoforms for a given protein. To do this manually requires deconvolution of MS1 spectra, determination of the masses-of-interest and their elution profiles, conversion of mass intensities to relative percentages, then entry into a spreadsheet to facilitate multifile or multiproteoform calculations (see Rommelfanger12 for this type of analysis on acetylated histone H4). Lastly, the processed spreadsheet data are often graphed as a histogram of protein intensity relative ratios (PIRRs10) versus proteoform mass. Extending this approach across many experimental conditions adds additional time and effort which can severely limit scalability. In this technical note, we describe pyMS-Vis, a stand-alone Python application for MS1 data analysis and visualization that performs many of the aforementioned manual calculations and graphing. pyMS-Vis takes deconvoluted LC-MS files as input, along with user-specified parameters (e.g., masses-of-interest, scan range, etc.), and outputs interactive graphs displaying an intensity and elution window for each specified mass. Furthermore, the program can process multiple data sets together, report relative abundances with standard deviations, and quickly produce a histogram of intensities for each mass. No knowledge of Python programming is needed to run the program which is freely available on GitHub in both Python script and Windows/Mac OS executable formats. For users who are either beginner or advanced Python users, we provide extensively commented code to assist in further customization.

Experimental Section

Data Preparation

The LC-TDMS histone data sets (.raw) were obtained from MassIVE (# MSV000088458),12 processed using default settings into *.mzML format by ProteoWizard’s MSConvert (ver. 3.0.21335–519be8f) and deconvoluted using the default settings in either TopFD (ver. 1.6.4) or FLASHDECONV (ver. 3.0.0-pre-HEAD-2022–08–16). If centroided *.mzML data sets were needed, the “peak Picking” option was selected in MSConvert prior to deconvolution. The output files were *.msAlign or *.tsv (deconvoluted mass and intensity lists) and .ms1 ft (feature map), which were used as input in our Python script.

Python

The script was generated using Python version 3.10.11 with the following packages installed: NumPy (ver. 1.23.0), matplotlib (ver. 3.5.2), Bokeh (ver. 3.1.1), and Pandas (ver. 1.4.3). The script was run successfully in Windows 10 (x64), Windows 11 (ARM), and MacOS Sonoma (ARM). The code is posted on Github with extensive comments that provide more information on additional features of the script and information on the classes and methods operating therein (https://github.com/pesavent/pyMS-Vis). Tutorial videos are available on YouTube (https://tinyurl.com/pyMS-Vis).

Cell Cycle Synchronization, Histone Extraction, and LC-MS Analysis

Chlamydomonas reinhardtii strain 21gr (CC-1690) was grown in 500 mL flasks containing 300 mL of HSM media, immersed in 25 C water baths and sparged with 1% CO2. Cells were synchronized in diurnal conditions by alternating 12h light, 12h dark cycles with equal fluences of red (630 nm) and blue (450 nm) light (150 μE m–2 s–1 each). Cultures were diluted daily to maintain cell density between 1 × 105 – 2 × 106 cells/mL. Cells were harvested at ZT4 (precommitment), ZT8 (postcommitment), ZT11 (S/M) and ZT22 (G0) by centrifugation. Nuclei from these cells were isolated and then treated with salt and acid to extract histone proteins as previously described.12 Purified histones were then analyzed by LC-MS by first separating proteoforms on a C18 column coupled with an Orbitrap Fusion Lumos or Eclipse as previously described.12 The resulting LC-TDMS data sets were processed as described above.

Results and Discussion

Raw LC-MS Data Preparation Prior to Script Analysis

The workflow to convert raw LC-TDMS data into a format compatible with our Python application is shown in Figure 1. First, raw LC-MS files are converted to mzML13 and then deconvoluted with software such as FLASHDeconv14 or TopFD.15 These programs generate *.tsv or *.msAlign files, respectively, which contain a list of deconvoluted masses, their relative intensities, and most abundant charge state for each scan. Additional outputs include a feature map (*.feature or *.ms1 ft file) that, when plotted, is a heatmap representing ion intensity of each mass bounded by its LC elution window. The deconvoluted spectra are then passed through a database search engine which uses precursor masses (MS1) and corresponding fragment ions (MS2) to identify proteoforms. Once proteoforms-of-interest are identified and validated, the same deconvoluted files can be used to search and quantify respective proteoform masses and their ion abundances from the MS1 data (see Figure 1 and our previous work12). Manual quantitative analysis of specific proteoform abundances requires filtering based on mass error tolerance and elution windows, followed by spreadsheet input for summing and averaging intensities for all detected proteoforms. pyMS-Vis expedites this analysis by taking deconvoluted spectra and user-provided search parameters to perform these calculations and output graphs for data visualization and inspection. We previously developed a Python script (SMC) that uses deconvoluted spectral files to search for MS2 spectra from a specific precursor mass across an LC-MS experiment and combine that information for a better fragment ion signal-to-noise ratio resulting in a higher confidence PTM identification and localization.12 The combined use of these two Python applications, pyMS-Vis and SMC, should accelerate proteoform analysis and quantitation of TDMS data sets. If positional isomers are of interest, a newly developed R package IsoForma(16) could be combined with pyMS-Vis for deeper proteoform characterization.

Figure 1 A typical workflow needed to visualize proteoforms in the pyMS-Vis application has four major steps. After multiple LC-MS data sets are acquired, the *.raw files must be converted to the universal, community standard mzML format. Deconvolution by FLASHDeconv or TopFD provides the MS1 and MS2 information needed for proteoform identification through either TopPIC or ProsightPD. After proteoforms are manually validated for multiple LC-MS experiments, they can be quantified and visualized by reanalyzing the deconvoluted information using the pyMS-Vis application.

Data Entry and Search Parameter Features

While pyMS-Vis can be used to analyze any deconvoluted LC-MS proteomics data, we first chose to test the script using LC-MS data from histone proteins which exhibit a high degree of proteoform complexity due to gene duplicates and paralog diversity combined with an extraordinarily high degree of post-translational modification. Since histone PTMs are often dynamic, questions arise about how histone proteoforms change abundance across different conditions. We previously characterized 86 histone proteoforms in the green microalga Chlamydomonas reinhardtii, and for many of these we were able to provide relative quantification.15 We decided to use five high-resolution LC-MS data sets (MassIVE # MSV000088458) generated in that study to test pyMS-Vis since the histone proteoforms identified were extensively validated through manual analysis. Within the histone family members, we focused on the most common histone H2A proteoforms, which range from ∼13,400 to ∼13 800 Da, and include four closely related protein isoforms which all can be multiply acetylated. Unless otherwise noted, analysis was performed on LC-MS data sets deconvoluted by FLASHDeconv.

The basic inputs and parameters for targeting proteoform masses within a deconvoluted MS1 file are shown in Figure 2 and Figure 3. As files are loaded, they can be binned in different experimental groups allowing for simple statistics to be performed on technical or biological replicates (Figure 2). After file selection and group assignment, the next parameter selection in the ’Static’ mode defaults to a list of masses, the minimum and maximum retention time, and a standard mass error tolerance (Figure 3Ai). In this mode, a quick survey of the data quality and determination of future search conditions can be performed. For instance, when a hypothetical series of masses are searched for, more than one elution window of the same mass may appear at significantly different retention times. Such differences often suggest the presence of unrelated proteins that coincidentally have the same queried mass, which can be verified by analyzing the corresponding MS2 spectra. In this case, it is important to filter out these mases by limiting retention time or scan number before quantification. If this is the case, the user can select “Dynamic” mode and limit the analytical window on a per-file basis (Figure 3Bi and 3Ci). In situations where more flexibility and open-ended mass searching are desired, instead of entering discrete mass values, the user can set a mass range along with a mass step size and a mass tolerance (Figure 3Ci). Two advanced parameters include a mass calibration adjustment on the deconvoluted masses and an off-by-one-dalton search option (not shown). Explained in greater detail below, once the user processes the files, we use the Python package Bokeh for visualization. Bokeh generates interactive graphs that allow the user to perform operations such as pan and zoom, while also offering the ability to export each graph as a portable network graphics (png) file. The pyMS-Vis application automatically saves the interactive graphs as an *.html file and the quantitative information as a *.csv file in the same directory as the script.

Figure 2 pyMS-Vis can analyze multiple files and provide adjustable search options such as monoisotopic mass, mass range and LC elution range. Deconvoluted files (as *.msAlign or *.tsv) are loaded into the application based on experimental condition. For example, these data sets are from asynchronously grown Chlamydomonas and grouped as “Asynch”.

Figure 3 Graphical outputs after running the pyMS-Vis application provide visualization of mass intensities for histone H2A and their standard deviations, elution profiles and mass feature maps. (A-C, boxed blue) These three panels represent three separate graphical outputs corresponding to the “Static” (Ai), “Dynamic” (Bi), and “Dynamic” with mass range (Ci) search parameters. Each output represents averaged relative abundances of searched masses, along with their corresponding standard deviation (A-C, (ii)), and the elution profiles of each searched mass (A-C, (iii)). (D) This interactive heatmap shows the abundance and elution window of masses within the selected range of 13 450 Da and 13 800 Da (top) with the isotopic distribution of the selected mass 13 586 Da (bottom). (E) The previously published and annotated H2A MS1 ion intensities match closely with the quantification of the corresponding deconvoluted data (adapted from Figure 4A in Rommelfanger et al.12 and color coded to match the corresponding queried masses in Biii).

Interactive Graphical Outputs Provide Insights into Proteoform Abundances and Quality of the Deconvoluted MS1 Spectra

Upon entering the “Static” mode parameters for histone H2A, the script calculates the total intensity of all masses found and displays them as a bar graph where each mass intensity is represented as a percentage of the total. If replicates are included, the program also displays the standard deviation from the mean (Figure 3A-C, (ii)). Each file can be inspected by clicking on the corresponding tab, which then displays the corresponding elution profile as a histogram of ion intensities for each identified mass (Figure 3A-C, (iii)). Additionally, each mass is color coded and if multiple masses are queried, they can be hidden/displayed by clicking on the respective mass in the legend to reduce visual complexity. In the case of histone H2A, many of the identified masses had elution windows that correlated with their previously reported retention times.15 However, some masses consistent with H2A appeared outside of the typical elution window and were found to be unrelated proteins (e.g., masses eluting ∼10,200 s in Figure 3A middle panel). These masses could be excluded in “Dynamic” mode by narrowing the analytical window, and the resulting data for H2A were a more accurate representation of relative abundances with reduced variances (Figure 3B). Searching the entire monoisotopic mass range of canonical histone H2A proteoforms (∼13,400 to ∼13 800 Da) with a step size of 3 Da and a mass tolerance of 0.5 Da incorporated more masses into the quantification (Figure 3C). The relative abundances and errors matched closely to the previous manual searches where masses were individually entered (compare middle panel (ii) of Figure 3C to Figure 3A,B). The total computational time for the analysis of 5 asynchronous Chlamydomonas LC-MS data sets by pyMS-Vis was approximately 45 s.

If a feature map (*.ms1 ft or *.feature file) is placed in the same directory as the deconvoluted file, the user can select whether to plot its data as an interactive heatmap (Figure 3D, top). A mass can be individually selected in the heatmap, which will automatically display its isotopic distribution, if present, below (Figure 3D, bottom). Furthermore, the feature file’s heatmap serves as a quality control for the histograms generated by pyMS-Vis (Figure 3A-C, (iii)). The feature map may also be used reiteratively to quantify other masses-of-interest present in the deconvoluted LC-MS data. For example, the user can return to the Dynamic or Static search mode and adjust the parameters to search for new masses and/or alter other search parameters.

We compared the relative abundance histograms generated from our script for histone H2A proteoforms to the corresponding summed MS1 data that was manually quantified previously15 and found a high level of agreement (compare Figures 3A-C, (ii), to Figure 3E). This finding validates use of FLASHDeconv as a robust and accurate deconvolution algorithm for most of the H2A proteoforms. We decided to test whether deconvolution by TopFD generated similar proteoform abundances by pyMS-Vis, since a recent study reported better proteoform identification rates when a multisoftware approach was used to reanalyze the same TDMS data sets.17 We found pyMS-Vis analysis of both FLASHDeconv- and TopFD-deconvoluted data produced strikingly similar (and accurate) lists of histone H2A monoisotopic masses and their corresponding ion intensities, despite employing different deconvolution algorithms (Supplemental Figure 1).

Rapid Visualization of Histone Proteoform Abundances Across Chlamydomonas’ Cell Cycle Reveals Unexpected PTM Dynamics

After confirming that our script reliably quantifies MS1 intensities of H2A from asynchronous Chlamydomonas, we tested its ability to quantify H2A proteoforms from 4 separate cell cycle phases, with each phase having three biological replicates. The cells were diurnally synchronized and samples collected during precommitment (PreC), postcommitment (PostC), S/M, and G0 phase. Chlamydomonas has a noncanonical cell cycle program called multiple fission that deviates from the standard binary fission cell cycle: cells may grow more than 2-fold during a prolonged G1 phase and under optimal conditions my enlarge by over 10-fold in size during G1.18 At the end of G1 phase they undergo a series of n rapid alternating S phases and mitoses (S/M) and produce 2n daughters. Commitment occurs in mid-G1 phase and defines a point of no return: When growth is stopped by withdrawing light or nutrients, preC cells arrest in G0 phase, whereas postC cells will complete at least one cycle of S/M even if no further growth has occurred. Replication dependent histones genes, like most cell cycle genes, are expressed only during S/M phase when they are required for packaging newly replicated DNA into chromatin. We targeted the same H2A masses entered in the asynchronous Chlamydomonas analysis (Figure 3) and ensured the H2A proteoform elution windows were similar and controlled for across all data sets. We were able to quickly obtain profiles of H2A proteoforms (Figure 4), along with standard deviations in abundance, across the 12 data sets from four cell cycle stages (pyMS-Vis computation time ∼1 min). Upon checking data quality, it was determined that the third biological replicate of the PostC cell cycle stage contained very little signal and displayed H2A intensities and elution characteristics that differed significantly from the other PostC replicates (as well as the other cell cycle stages). After a second LC-TDMS analysis of this sample, it was determined that significant protein loss occurred during histone preparation (data not shown). The addition of this data led to larger error bars in the relative abundance measurements compared to the other data (data not shown). When we reanalyzed the deconvoluted files without this outlier data set, it was immediately obvious that masses 13488, 13545, and 13700 were specifically enriched in S/M phase relative to all others (Figure 4A). After analyzing the corresponding MS2 data for these masses, we identified them as histone H2A.1, H2A.2 and H2A.3, respectively, each lacking the alpha-amino acetylation (ΔNα-ac) normally present on all H2A proteoforms. Interestingly, after revisiting the retention times in the saved pyMS-Vis output files, each of the three Nα-ac H2A proteoforms eluted later from the C18 resin relative to their respective ΔNα-ac form (data not shown). These LC elution characteristics are consistent with the greater hydrophobicity of Nα-ac proteoforms versus nonacetylated versions.

Figure 4 Application of the Python script in the analysis of Chlamydomonas histones during its cell cycle reveals significant H2A proteoform changes. (A) Relative intensities of canonical H2A proteoforms were quantified during the PreC, PostC, S/M and G0 of its cell cycle (in replicate). (B) Enlargement of H2A.3 masses in (A) shows clear abundance increases of mass 13700 (B–I) and decrease of mass 13742 (B–II) during S/M phase. (C) A model describing the possible pathways for the addition or removal of H2A alpha-amino acetylation (Nα-ac) during mitosis.

We have previously identified these three ΔNα-ac H2A proteoforms (13488, 13545, and 13700 Da) in asynchronous Chlamydomonas at small, yet significant abundances (∼1% of all canonical H2A).12 In that study, we found these forms to be highly variable after qualitative MS inspection, which is corroborated quantitatively by the larger standard deviations relative to their mean abundance shown here (see Figure 3A-C, left panels). Because the cells in those previous samples were asynchronous, only a small fraction of each sample was in S/M phase, though we suspect some residual synchrony may have been present and caused variability in the relative abundance of ΔNα-ac H2A proteoform abundance between biological replicates.

The observation of extremely high levels of ΔNα-ac (∼45% of all canonical H2A) during S/M phase may have significant implications in protein processing and histone maturation19 during the algal cell cycle. A model illustrating possible removal of Nα-ac from parental H2A or cotranslational addition of Nα-ac to newly synthesized H2A proteins in Chlamydomonas is shown in Figure 4C. If Nα-ac occurs on a protein, it is thought to happen cotranslationally (i.e., at the time of protein synthesis during translation) by evolutionarily conserved Nα acetyltransferases (NATs) and is considered permanent since no Nα-ac deacetylases have been discovered. While delayed H2A Nα-acetylation is the most likely mechanism, the data do not rule out potential removal of Nα-ac on parental H2A by a yet-to-be-discovered histone Nt-deacetylase. It should be noted that both Chlamydomonas H2A.Z and H4 are nearly all Nα-acetylated, have a respective amino-terminal sequence of S1-G-K-G and S1-G-R-G (vs A1-G-R-G for canonical H2A), and do not show a delay in alpha-amino terminal acetylation during the cell cycle or have significant levels of the corresponding ΔNα-ac proteoform in asynchronous cultures (less than 1%; data not shown). This novel, temporal Nα acetylation unique to Chlamydomonas canonical H2A histones may play an important role in algal chromatin maturation during its multiple fission cycles.

Conclusions

pyMS-Vis has been successfully used to quickly compare multiple deconvoluted TDMS data sets, resulting in informative MS1 abundance profiles for multiple histone H2A proteoforms. We demonstrated how this application could be used to determine asynchronous or cell-cycle synchronized Chlamydomonas histone proteoform abundances from replicate LC-MS data sets. The quantified MS1 data revealed the degree of variability in elution profiles and relative intensities of histone H2A proteoforms. Additionally, a dramatic change in H2A masses was observed during S/M phase of the cell cycle. These results informed our MS2 analytical workflow in TDValidator20 (Proteinaceous Inc.) leading to the exciting discovery of potentially novel alpha-amino processing of H2A. This script can be easily applied to many other nonhistone proteins for which there is deconvoluted LC-MS data (i.e., not just top-down, but also middle-down and bottom-up deconvoluted data). Our hope is that this tool will be valuable for other mass spectrometry applications and serve as a backbone for further customization.

Supporting Information Available

The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.analchem.4c02650.A comparison between FLASHDeconv and TopFD deconvoluted file inputs and pyMSVis outputs (Supplemental Figure S1), Discussion of the Python code, packages used, and future developments of the script, Integration of pyMS-Vis with other software/scripts/workflows (PDF)

Supplementary Material

ac4c02650_si_001.pdf

Author Present Address

⊥ Zhejiang University, Hangzhou, Zhejiang 310058, China

The authors declare no competing financial interest.

Acknowledgments

This material is based partially upon work supported by the National Science Foundation under Grant No. MCB 1943493 (JJP). A portion of this research was performed on a project award (10.46936/expl.proj.2019.51056/60006669) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. Work in the laboratory of JGU was supported by National Institutes of Health grant R01 GM126557 and National Science Foundation Grant MCB 1515220.
==== Refs
References

Aebersold R. ; Agar J. N. ; Amster I. J. J. ; Baker M. S. ; Bertozzi C. R. ; Boja E. S. ; Costello C. E. ; Cravatt B. F. ; Fenselau C. ; Garcia B. A. ; Ge Y. ; Gunawardena J. ; Hendrickson R. C. ; Hergenrother P. J. ; Huber C. G. ; Ivanov A. R. ; Jensen O. N. ; Jewett M. C. ; Kelleher N. L. ; Kiessling L. L. ; Krogan N. J. ; Larsen M. R. ; Loo J. A. ; Ogorzalek Loo R. R. ; Lundberg E. ; MacCoss M. J. ; Mallick P. ; Mootha V. K. ; Mrksich M. ; Muir T. W. ; Patrie S. M. ; Pesavento J. J. J. ; Pitteri S. J. ; Rodriguez H. ; Saghatelian A. ; Sandoval W. ; Schlüter H. ; Sechi S. ; Slavoff S. A. ; Smith L. M. ; Snyder M. P. ; Thomas P. M. ; Uhlén M. ; Van Eyk J. E. ; Vidal M. ; Walt D. R. ; White F. M. ; Williams E. R. ; Wohlschlager T. ; Wysocki V. H. ; Yates N. A. ; Young N. L. ; Zhang B. How Many Human Proteoforms Are There?. Nat. Chem. Biol. 2018, 14 (3 ), 206–214. 10.1038/nchembio.2576.29443976
Senko M. W. ; Beu S. C. ; McLafferty F. W. Determination of Monoisotopic Masses and Ion Populations for Large Biomolecules from Resolved Isotopic Distributions. J. Am. Soc. Mass Spectrom. 1995, 6 (4 ), 229–233. 10.1016/1044-0305(95)00017-8.24214167
Wu Z. ; Roberts D. S. ; Melby J. A. ; Wenger K. ; Wetzel M. ; Gu Y. ; Ramanathan S. G. ; Bayne E. F. ; Liu X. ; Sun R. ; Ong I. M. ; McIlwain S. J. ; Ge Y. MASH Explorer: A Universal Software Environment for Top-Down Proteomics. J. Proteome Res. 2020, 19 (9 ), 3867–3876. 10.1021/acs.jproteome.0c00469.32786689
Kou Q. ; Xun L. ; Liu X. TopPIC: A Software Tool for Top-down Mass Spectrometry-Based Proteoform Identification and Characterization. Bioinformatics 2016, 32 (22 ), 3495–3497. 10.1093/bioinformatics/btw398.27423895
Park J. ; Piehowski P. D. ; Wilkins C. ; Zhou M. ; Mendoza J. ; Fujimoto G. M. ; Gibbons B. C. ; Shaw J. B. ; Shen Y. ; Shukla A. K. ; Moore R. J. ; Liu T. ; Petyuk V. A. ; Tolić N. ; Paša-Tolić L. ; Smith R. D. ; Payne S. H. ; Kim S. Informed-Proteomics: Open-Source Software Package for Top-down Proteomics. Nat. Methods 2017, 14 (9 ), 909–914. 10.1038/nmeth.4388.28783154
Röst H. L. ; Schmitt U. ; Aebersold R. ; Malmström L. pyOpenMS: A Python-Based Interface to the OpenMS Mass-Spectrometry Algorithm Library. Proteomics 2014, 14 (1 ), 74–77. 10.1002/pmic.201300246.24420968
Levitsky L. I. ; Klein J. A. ; Ivanov M. V. ; Gorshkov M. V. Pyteomics 4.0: Five Years of Development of a Python Proteomics Framework. J. Proteome Res. 2019, 18 (2 ), 709–714. 10.1021/acs.jproteome.8b00717.30576148
Barente A. S. ; Villén J. A Python Package for the Localization of Protein Modifications in Mass Spectrometry Data. J. Proteome Res. 2023, 22 (2 ), 501–507. 10.1021/acs.jproteome.2c00194.36315500
Bittremieux W. Spectrum_utils : A Python Package for Mass Spectrometry Data Processing and Visualization. Anal. Chem. 2020, 92 (1 ), 659–661. 10.1021/acs.analchem.9b04884.31809021
Pesavento J. J. ; Mizzen C. A. ; Kelleher N. L. Quantitative Analysis of Modified Proteins and Their Positional Isomers by Tandem Mass Spectrometry: Human Histone H4. Anal. Chem. 2006, 78 (13 ), 4271–4280. 10.1021/ac0600050.16808433
Wu Z. ; Tiambeng T. N. ; Cai W. ; Chen B. ; Lin Z. ; Gregorich Z. R. ; Ge Y. Impact of Phosphorylation on the Mass Spectrometry Quantification of Intact Phosphoproteins. Anal. Chem. 2018, 90 (8 ), 4935–4939. 10.1021/acs.analchem.7b05246.29565561
Rommelfanger S. R. ; Zhou M. ; Shaghasi H. ; Tzeng S. ; Evans B. S. ; Umen J. G. ; Pesavento J. J. An Improved Top-Down Mass Spectrometry Characterization of Chlamydomonas Reinhardtii Histones and Their Post-Translational Modifications. J. Am. Soc. Mass Spectrom. 2021, 32 , 1671 10.1021/jasms.1c00029.34165968
Martens L. ; Chambers M. ; Sturm M. ; Kessner D. ; Levander F. ; Shofstahl J. ; Tang W. H. ; Römpp A. ; Neumann S. ; Pizarro A. D. ; Montecchi-Palazzi L. ; Tasman N. ; Coleman M. ; Reisinger F. ; Souda P. ; Hermjakob H. ; Binz P.-A. ; Deutsch E. W. mzML—a Community Standard for Mass Spectrometry Data. Molecular & Cellular Proteomics 2011, 10 (1 ), R110.000133 10.1074/mcp.R110.000133.
Jeong K. ; Kim J. ; Gaikwad M. ; Hidayah S. N. ; Heikaus L. ; Schlüter H. ; Kohlbacher O. FLASHDeconv: Ultrafast, High-Quality Feature Deconvolution for Top-Down Proteomics. Cell Systems 2020, 10 (2 ), 213–218.e6. 10.1016/j.cels.2020.01.003.32078799
Basharat A. R. ; Zang Y. ; Sun L. ; Liu X. TopFD: A Proteoform Feature Detection Tool for Top-Down Proteomics. Anal. Chem. 2023, 95 (21 ), 8189–8196. 10.1021/acs.analchem.2c05244.37196155
Degnan D. J. ; Lewis L. A. ; Bramer L. M. ; McCue L. A. ; Pesavento J. J. ; Zhou M. ; Bilbao A. IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data. J. Proteome Res. 2024, 23 , 3318 10.1021/acs.jproteome.3c00681.38421884
Tabb D. L. ; Jeong K. ; Druart K. ; Gant M. S. ; Brown K. A. ; Nicora C. ; Zhou M. ; Couvillion S. ; Nakayasu E. ; Williams J. E. ; Peterson H. K. ; McGuire M. K. ; McGuire M. A. ; Metz T. O. ; Chamot-Rooke J. Comparing Top-Down Proteoform Identification: Deconvolution, PrSM Overlap, and PTM Detection. J. Proteome Res. 2023, 22 , 2199 10.1021/acs.jproteome.2c00673.37235544
Cross F. R. ; Umen J. G. The Chlamydomonas Cell Cycle. Plant Journal 2015, 82 (3 ), 370–392. 10.1111/tpj.12795.
Demetriadou C. ; Koufaris C. ; Kirmizis A. Histone N-Alpha Terminal Modifications: Genome Regulation at the Tip of the Tail. Epigenetics and Chromatin 2020, 13 (1 ), 1–13. 10.1186/s13072-020-00352-w.31918747
Fornelli L. ; Srzentić K. ; Huguet R. ; Mullen C. ; Sharma S. ; Zabrouskov V. ; Fellers R. T. ; Durbin K. R. ; Compton P. D. ; Kelleher N. L. Accurate Sequence Analysis of a Monoclonal Antibody by Top-Down and Middle-Down Orbitrap Mass Spectrometry Applying Multiple Ion Activation Techniques. Anal. Chem. 2018, 90 (14 ), 8421–8429. 10.1021/acs.analchem.8b00984.29894161
