
==== Front
Protein Sci
Protein Sci
10.1002/(ISSN)1469-896X
PRO
Protein Science : A Publication of the Protein Society
0961-8368
1469-896X
John Wiley & Sons, Inc. Hoboken, USA

10.1002/pro.5174
PRO5174
Tools for Protein Science
Tools for Protein Science
BracketMaker: Visualization and optimization of chemical protein synthesis
Evangelista and Kay
Evangelista Judah L. https://orcid.org/0000-0002-9458-7441
1
Kay Michael S. 1 kay@biochem.utah.edu

1 Department of Biochemistry University of Utah Salt Lake City Utah USA
* Correspondence
Michael S. Kay, Department of Biochemistry, University of Utah, Salt Lake City, UT, USA.
Email: kay@biochem.utah.edu

14 9 2024
10 2024
14 9 2024
33 10 10.1002/pro.v33.10 e517426 8 2024
28 5 2024
28 8 2024
© 2024 The Author(s). Protein Science published by Wiley Periodicals LLC on behalf of The Protein Society.
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the terms of the http://creativecommons.org/licenses/by-nc-nd/4.0/ License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non‐commercial and no modifications or adaptations are made.

Abstract

Chemical protein synthesis (CPS), in which custom peptide segments of ~20–60 aa are produced by solid‐phase peptide synthesis and then stitched together through sequential ligation reactions, is an increasingly popular technique. The workflow of CPS is often depicted with a “bracket” style diagram detailing the starting segments and the order of all ligation, desulfurization, and/or deprotection steps to obtain the product protein. Brackets are invaluable tools for comparing multiple possible synthetic approaches and serve as blueprints throughout a synthesis. Drawing CPS brackets by hand or in standard graphics software, however, is a painstaking and error‐prone process. Furthermore, the CPS field lacks a standard bracket format, making side‐by‐side comparisons difficult. To address these problems, we developed BracketMaker, an open‐source Python program with built‐in graphic user interface (GUI) for the rapid creation and analysis of CPS brackets. BracketMaker contains a custom graphics engine which converts a text string (a protein sequence annotated with reaction steps, introduced herein as a standardized format for brackets) into a high‐quality vector or PNG image. To aid with new syntheses, BracketMaker's “AutoBracket” tool automatically performs retrosynthetic analysis on a set of segments to draft and rank all possible ligation orders using standard native chemical ligation, protection, and desulfurization techniques. AutoBracket, in conjunction with an improved version of our previously reported Automated Ligator (Aligator) program, provides a pipeline to rapidly develop synthesis plans for a given protein sequence. We demonstrate the application of both programs to develop a blueprint for 65 proteins of the minimal Escherichia coli ribosome.

chemical protein synthesis
E. coli ribosome
graphic design tools
native chemical ligation
Python tools
National Institutes of Health 10.13039/100000002 source-schema-version-number2.0
cover-dateOctober 2024
details-of-publishers-convertorConverter:WILEY_ML3GV2_TO_JATSPMC version:6.4.8 mode:remove_FC converted:14.09.2024
Evangelista JL , Kay MS . BracketMaker: Visualization and optimization of chemical protein synthesis. Protein Science. 2024;33 (10 ):e5174. 10.1002/pro.5174

Review Editor: Nir Ben‐Tal
==== Body
pmc1 INTRODUCTION

Chemical protein synthesis (CPS) is a powerful technique to create proteins for biochemical studies or drug discovery. Proteins that are difficult or impossible to obtain by recombinant means, such as those containing non‐canonical aa, post‐translational modifications, site‐specific labels, or even fully mirror‐image D‐proteins, can be produced by CPS with high homogeneity and precise atomic control (Agouridas et al., 2020; Kent, 2017; Kent, 2019; Kulkarni et al., 2018; Tan et al., 2020). The general principle involves producing peptides of up to ~50 aa by solid‐phase peptide synthesis (SPPS), joining peptide segments together through chemoselective ligation reactions (most commonly native chemical ligation, NCL (Conibear et al., 2018)), and folding into a functional protein. Since the inception of NCL three decades ago (Dawson et al., 1994; Muir & Kent, 1993), CPS has grown from a specialty technique into a field of its own, with dedicated international conferences (Bello et al., 2019) as well as a publicly available protein chemical synthesis database (PCS‐db) with information on hundreds of synthesized proteins reported in literature (Agouridas et al., 2017). As the community grows and aims for larger and more diverse protein targets, the field needs data science tools to manage the complexity of information in a typical CPS project.

The scope of CPS is continually widening thanks to many important chemical advances over the years (Agouridas et al., 2019). Peptide thioesters, the required starting materials for NCL, can be readily produced by common Fmoc‐SPPS as masked thioester surrogates (most commonly hydrazides (Fang et al., 2011; Flood et al., 2018; Zheng et al., 2013)). Segments can be assembled into proteins convergently using a combination of masked thioesters and removable Cys protecting groups (e.g., Acm (Maity et al., 2016), Thz (Bang & Kent, 2004), and Tfa‐Thz (Huang et al., 2016)). Additional ligation sites beyond Cys have been unlocked by desulfurization reactions (originally converting Cys to Ala (Wan & Danishefsky, 2007; Yan & Dawson, 2001) but expanded to many thiolated aa (Jin & Li, 2018; Kulkarni et al., 2018; Rohde & Seitz, 2010) including the commercially available β‐thio‐valine [penicillamine] (Chen et al., 2008)), as well as alternative ligation chemistries (e.g., serine–threonine ligation [STL] (Liu & Li, 2018), α‐ketoacid‐hydroxylamine ligation [KAHA] (Rohrbacher et al., 2015), diselenide‐selenoester ligation (Kulkarni et al., 2019; Mousa et al., 2017), and others (Asahina et al., 2017; Blanco‐Canosa et al., 2015; Terrier et al., 2016)). With a growing list of options in the modern protein chemist's toolbox, a wide variety of proteins are now accessible; on the other hand, every synthetic approach requires careful planning and consideration of available tools.

To plan a CPS strategy, one typically sketches a “bracket” diagram detailing the synthetic route from starting segments to product protein (Figure 1). The CPS bracket is a convenient shorthand and visualization tool, a blueprint referenced throughout a CPS project, and an unofficially required accompaniment to published syntheses. Despite their widespread use, there is no standard way to make these figures. Published syntheses vary both in visual style and in portrayed information, complicating the side‐by‐side comparison of different strategies. Drafting several possible brackets is extremely helpful for large proteins with multiple possible ligation sites, but this process is arduous (and prone to typos) in standard graphics software or by hand. As CPS increases in popularity, the field can greatly benefit from custom tools to plan and visualize syntheses.

FIGURE 1 Example CPS bracket depicting our lab's previous synthesis of DapA (Weinstock et al., 2014). The color‐coded reference sequence of the final protein is located above the bracket with special aa bolded and underlined. The bracket depicts the order of all reactions and the identity of all starting segments and isolated products. This image was produced entirely in BracketMaker (SVG format) from text input with default program settings. See description of the makeBracketFigure() function (Supplementary Information) for details on the text‐to‐image conversion process.

Our Aligator program (Jacobsen et al., 2017) finds and ranks all possible combinations of ligation sites within an input protein sequence, reducing the overwhelming number of possible strategies to a short list of preferred options based on the properties of starting segments. However, the program's output (a list of segment sets with scores) only provides partial information about the difficulty of synthesis, and segments must be assembled into brackets to get a complete picture. Bracket‐building requires even more decisions about optimal ligation in order to minimize yield loss, and these decisions are hampered by the tedious nature of visualizing and assessing possible routes. Further computational assistance is necessary to develop complete synthesis plans and accurately assess their difficulty. To facilitate CPS decision‐making and bracket drawing, we have developed BracketMaker, a standalone Python‐language program with a custom interactive interface for creation and analysis of CPS brackets.

In this work, we define the essential elements of the bracket figure and encode these elements in a novel plain‐text format, which serves as the input for BracketMaker's graphics engine. We then describe the logic behind BracketMaker's primary function, translating this “bracket notation” to a publication‐quality vector image within an interactive workspace. Next, we introduce our Auto‐Bracket tool, which performs retrosynthetic analysis using standard NCL chemistry to generate and rank possible ligation orders for a set of segments. This tool can be used in conjunction with Aligator to rapidly identify the most feasible routes toward a given protein. We conclude with our perspective on the current scope of CPS and assess the path toward a major goal in the field, the synthesis of all 65 subunits of a mirror‐image D‐ribosome. With our user‐friendly computational tools, we aim to promote cohesiveness within the CPS field, lower the barrier to entry for newcomers, and help extend the reach of CPS to larger and more complex protein targets.

2 RESULTS AND DISCUSSION

2.1 Bracket notation: A condensed format for CPS brackets

To develop our first‐of‐its‐kind graphics engine, we first had to define a set of instructions that could be fed to the program and transformed into a complete bracket figure, as well as a format in which these instructions could be saved and retrieved for future editing. For this tool to be useful, providing the instructions must be faster and simpler than drawing a bracket by traditional means. After assessing the variety of brackets in literature, we determined that the essential elements of a bracket figure are as follows: The full protein sequence: This sequence is included as a key, allowing segment labels in the bracket itself to be abbreviated. All amino acids are shown as they are in the final protein and differences are marked on segments. Segments are typically abbreviated to only the first and last aa, as these residues are most relevant for evaluating strategies.

Ligation reactions: Segments are joined two at a time, giving the bracket its characteristic shape as ligation intermediates gradually combine to form the final protein. Each reaction shows the initial starting segments, the product isolated after all one‐pot transformations, and a brief description of reaction conditions. Sequential reactions are stacked so one may trace the full pathway from any single segment to the final product.

Side‐chain modifications: Deprotection, desulfurization, and other side‐chain modifications are essential to many strategies, and any differences from the final protein should be clearly labeled on relevant segments. If a reaction is performed on a purified ligation intermediate, the segment should be shown both before and after modification; if performed one‐pot with ligation, only the final product must be shown.

C‐terminal modifications: Unless otherwise specified, segments are assumed to contain an amine at their N‐terminus and carboxylic acid at their C‐terminus. In most CPS strategies, however, all segments except the last contain a reactive group at their C‐terminus (e.g., a thioester or hydrazide for NCL) that should be clearly indicated. The level of detail for this label varies, but simple text is usually sufficient when standard techniques are used.

Assembling this information by drawing shapes and manually editing text labels is tedious. Due to the repetitive nature of the figure, however, we reasoned that most of the pieces of the bracket, including all intermediates and the final protein, could be extrapolated if we knew only the starting segments and order of reaction steps. Using this principle, we developed a simple plain‐text code that uniquely identifies each CPS strategy, called “bracket notation” (Figure 2), and an algorithm to translate this code into a figure.

FIGURE 2 Example of “bracket notation,” the text language of BracketMaker encoding all steps of CPS. (a) Text format describing our lab's previous synthesis of DapA (Weinstock et al., 2014) with (b) key of each sequence feature. Text has been colored for clarity to match highlighting in BracketMaker interface (see Figure 3); black = single‐letter aa, gray = C‐terminus label, red = special aa, orange = reaction step. The user manual on our BracketMaker GitHub page contains additional syntax details. This input text is sufficient to generate the output vector image in Figure 1.

Bracket notation is similar to FASTA format, except the protein sequence is marked with several annotations using specific characters. Any special aa that cannot be described with a single‐letter code is named within parentheses, and side‐chain labels or protecting groups are indicated with a hyphen (e.g., (C‐Acm)). Changes in side‐chain identity during a synthesis (e.g., C to A via desulfurization) are defined with an arrow from starting to product aa (e.g., (C > A)). Ligation junctions are inserted between segments with numbers in square brackets (e.g., [1],[2],[3]), and any side‐chain modification steps are described with keywords—either one‐pot with NCL (e.g., [1,Acm]) or in an explicit reaction afterward (e.g., [1][Acm]). The number of each ligation specifies its order relative to the final protein; for consistency across different strategies, especially in imbalanced brackets, the final ligation is always [1]. Finally, C‐terminal modifications can be added to each segment with a hyphen, or if many segments have the same C‐terminus, can be left out of the bracket notation and described elsewhere with a global setting. Altogether, this code is sufficient to describe most CPS strategies using standard NCL methodology and can easily be adapted to other types of chemistry. From a bracket‐notation text string, BracketMaker's graphics engine extrapolates all information necessary to produce a complete bracket figure; for instance, converting text in Figure 2 to the full image in Figure 1. Additional informational elements, such as chemical structures of reactive groups or reaction details, are often specific to each individual project; for simplicity, these are left out of bracket notation, and optional text placeholders are included instead in the bracket image template.

3 THE BracketMaker PROGRAM

Once we built a function capable of replicating most published brackets, we wanted to expand this program into an interactive bracket‐building tool. When planning a new CPS project, one must carefully consider ligation sites to divide a protein into SPPS‐compatible segments, then choose an efficient ligation order to maximize final product yield. We thus sought to develop a tool that could be used for this retrosynthetic analysis—beginning with a protein sequence, gradually making decisions about ligation strategy, and ending up with a detailed plan of all reaction steps and requisite aa modifications (and multiple backup plans, in case of unforeseen obstacles encountered during synthesis). Importantly, we wanted this tool to be easily accessible and not require programming ability. To this end, we developed BracketMaker, a free standalone executable containing an interactive workspace for designing brackets, with custom functions to score brackets and rapidly draft multiple strategies.

BracketMaker (Figure 3) consists of a Python‐based GUI in which text (bracket notation) is interpreted and converted to a bracket image. The image dynamically updates in response to text edits, allowing users to gradually build brackets by adding ligation reactions and other sequence annotations. From this interface, the user may also control other image settings, such as individual segment colors and default text placeholders. Brackets and figure settings can easily be saved and re‐opened in the GUI, or brackets may be output as a standard vector graphics image for further editing in other graphics software such as Adobe Illustrator (output .svg files can also be viewed in most web browsers, making this a convenient format for sharing plans). To help familiarize users with bracket notation, the GUI color highlights different parts of the syntax and provides tools to build brackets with a single click (described in the next section). Multiple brackets can be loaded into the open workspace and cycled through with arrow buttons, allowing easy visual comparison of different synthesis plans. This feature also simplifies retrosynthetic analysis, as a partially finished bracket can be “cloned” and revisited later to explore alternative ways of filling out the remainder. Overall, BracketMaker replaces the tedious manual process of designing, comparing, and sharing brackets with an intuitive and user‐friendly workspace.

FIGURE 3 The BracketMaker interface with an example bracket. Text input box contains bracket notation text, color‐coded by the program. A bracket preview image is displayed below and automatically updates as text is changed. The figure settings menu on the right shows all default settings. Note that the reference sequence at the top of the output vector image (Figure 1) is absent from this preview image.

3.1 Bracket assembly and optimization tools

In the early stages of a CPS project, it is helpful to work through many different routes of synthesis and compare their difficulty, a practice we refer to as “bracketology.” The process involves two challenges that can be greatly aided by data science and visualization tools: dividing the protein into segments that can be produced by SPPS in good quality (usually up to ~50 aa), and finding the most efficient order of joining segments together to minimize accumulating yield losses. Our Aligator (Automated Ligator) program (Jacobsen et al., 2017) assists in choosing ligation junctions by ranking possible ligation strategies (i.e., sets of segments) by factors such as segment length, predicted segment solubility, and the amino acids at ligation junctions (applying penalties for desulfurizations and poor‐yielding thioesters). Judging the true difficulty of a strategy, however, requires full consideration of the path from every starting segment to the final product in addition to the properties of the segments themselves. Segments are joined sequentially two at a time into increasingly larger intermediates. With large proteins consisting of four or more segments, many ligation orders are possible, and choosing the optimal order can be a complex task of its own. Furthermore, it is vital to have backup plans in case an impassible obstacle, such as an insoluble intermediate segment, is encountered. Thus, we sought to provide tools within BracketMaker to help users assemble segments into efficient brackets and judge the feasibility of different plans.

3.1.1 Automated bracket assembly (Auto‐Bracket)

First, we designed a bracket assembly algorithm that converts an initial protein sequence to a feasible bracket by making gradual changes to the bracket notation text. This process is essentially a retrosynthetic analysis from the final ligation backward (i.e., upward). When a ligation junction is chosen, the two starting segments are assessed and sequence modifications are applied when necessary. With standard NCL chemistry, these include (a) protection of N‐terminal Cys in left‐half segments to avoid cyclization and oligomerization, (b) thiolation of any aa other than Cys at ligation sites (e.g., substitution of Ala with Cys or of Val with penicillamine [Pen]), and (c) protection of native Cys during desulfurization. Appropriate desulfurization and/or deprotection steps are added after NCL to convert side chains back to their native form in the final protein. After repeating these steps for each junction, building a bracket stepwise from the bottom up, one arrives at initial segments produced by SPPS with all required modifications.

With a language to describe CPS steps, replicating these retrosynthetic analysis steps in our Auto‐Bracket tool was relatively straightforward. The main difficulty was deciding what ligation order should be chosen by default, as the most logical order (and the orders that are even possible) depends on starting segments, types of ligation junctions used, and the chemistry available to individual users. To address this problem, we next sought to develop a general computational solution to find an optimal order for each unique set of segments.

3.1.2 Bracket evaluation and optimization

Equipped with a rapid method of generating logically sound brackets, we expanded Auto‐Bracket to work through every possible permutation of steps and produce a list of all feasible brackets for any given set of segments. To discriminate between different ligation orders, we developed an initial set of bracket scores representing the efficiency of synthesis routes: Max Path, Avg Path, and Steps.

In general, an efficient bracket maximizes final protein yield by minimizing the path length (i.e., number of sequential reactions) of any given segment. In our bracket scoring system, only reactions that require purification of products are counted toward a segment's path length. In hydrazide NCL, the steps of thioester activation and ligation are carried out in a single reaction tube, and thus these sequential reactions count only as a single path (Figure 4a). Similarly, deprotection methods that can be performed immediately after NCL without purification, such as the conversion of Tfa‐Thz to Cys (Figure 4a), do not count toward path length. Desulfurization reactions and some deprotection reactions, such as that of Cys‐Acm, usually cannot be performed one‐pot with NCL (Figure 4b), and thus these reactions are counted in each segment's path length. NCL reactions with poor predicted yield are given additional penalties: +1 path for poor‐yielding thioesters (generally I, L, K, T, and V (Hackeng et al., 1999)) or +2 paths for sterically hindered thiols (e.g., V [as Pen]). The simplest way of choosing between two brackets is to choose the one with shorter yield‐limiting path length (i.e., smaller Max Path), and if this score is tied, shorter path lengths among all segments on average (i.e., smaller Avg Path).

FIGURE 4 Auto‐Bracket predicts an optimal ligation order for a given set of segments. (a), (b) Reactions that are considered in Auto‐Bracket's retrosynthetic analysis process and in bracket scoring. The step of thioester activation (a) is left out of brackets, as this generally occurs in the same reaction as NCL. Solid arrows represent reactions that contribute toward a segment's path length, and hollow arrows represent one‐pot reactions that do not count toward path length. (c)–(h) Auto‐Bracket's top result for model six‐segment proteins containing various types of ligation junctions, including all ideal thioesters paired with Cys sites (c); one poor thioester at the central ligation site (d); one central Ala ligation site (e); multiple adjacent Ala ligation sites (f); all ideal thioesters and Cys sites but without one‐pot Cys deprotection (g); and a realistic example with multiple problematic ligation junctions (h). Each segment's path length score is listed above the bracket. Winning brackets are those that minimize path length scores for all segments, and in many situations, a lopsided bracket will score better than a fully convergent one.

Thioesters incompatible with hydrazide‐NCL (D, E, N, P, and Q) are usually absent from synthetic strategies due to exceptionally poor yield. These thioesters are forbidden by default in Aligator (Jacobsen et al., 2017) and do not appear in any of our analyzed brackets. Whether or not these thioesters are given path length penalties can be customized in user settings.

Finally, it is also helpful to count total steps in the synthesis, as this score indicates the length of time required to complete a project. The Steps score correlates with Avg Path (more steps = more path lengths) and is a secondary measure of project complexity, less important than the yield‐limiting lengths of individual paths. Using this simple reaction‐counting metric to sort brackets, Auto‐Bracket considers all possible synthesis routes and returns a single winning strategy (or optionally, a sorted list) with most efficient bracket scores (lowest Max Path, Avg Path, and Steps, in order of sort priority).

After testing all possible options, Auto‐Bracket ultimately returns sensible brackets for any given set of segments using standard NCL. Auto‐Bracket's top choices for several example sets of model segments are shown in Figure 4. As expected, a convergent strategy (Figure 4c) generally scores better than a linear one (i.e., entirely N‐to‐C or C‐to‐N), but in many cases one of the many possible intermediate (or “semi‐convergent”) strategies is preferred over fully convergent. For example, a poor ligation site in the center of a protein should not be used for the final reaction between two multi‐segment intermediates, but instead a lopsided bracket should be used so this problem reaction occurs earlier between less precious segments (Figure 4d); in terms of scores, the winning bracket applies path length penalties to as few segments as possible. In strategies with a mix of Ala and Cys sites, the requirement of all Ala ligations and desulfurizations to occur before native Cys ligations may outright forbid a convergent ligation order if Ala is centrally located (Figure 4e). Multiple adjacent Ala ligations can be performed and desulfurized simultaneously after the final one (Figure 4f), and thus the specific location of Ala sites is more important for bracket difficulty than the number of Ala sites. Finally, if Cys protecting groups must be removed in an explicit separate reaction step after NCL (which can be toggled in Auto‐Bracket settings), a linear (N‐to‐C) or semi‐convergent bracket may score better than fully convergent (Figure 4g). A more realistic set of segments, with multiple Ala sites, poor thioesters, and internal Cys that must be protected, will have a particular optimal ligation order that minimizes all path length penalties (Figure 4h).

Our bracket scores provide a mathematical basis for the intuitive logic of bracketology. When given the seven segments previously used to assemble the largest protein published by our group, the 312‐aa DapA (Weinstock et al., 2014), as well as the chemical limitations we had at the time (i.e., no one‐pot desulfurization or deprotection steps), Auto‐Bracket provides the strategy that was actually used (as shown in Figure 1). Similarly, Auto‐Bracket predictions for some of the largest proteins produced by hydrazide‐NCL chemistry in recent years match well with actual strategies reported by authors (discussed further below). Overall, Auto‐Bracket provides a rapid alternative to manually drafting multiple brackets for comparison and mimics the process of an experienced protein chemist to assemble and compare multiple brackets.

3.1.3 Aligator‐to‐BracketMaker computational pipeline

The combination of Aligator and Auto‐Bracket allows us to quickly build and explore CPS strategies starting only from a protein sequence. The sequence is first processed in Aligator to obtain possible segment sets, and each set (i.e., each line of Aligator's output) is converted via Auto‐Bracket to a feasible ligation strategy. The final output from the process is a list of possible brackets, each representing a unique possible set of segments. We have included an “Aligator Import” tool in BracketMaker to provide a direct pipeline.

When choosing a synthesis strategy for a protein, one should consider the properties of initial segments (Aligator scores) and the efficiency of the bracket (BracketMaker scores). It is up to an individual user to decide, for example, if a strategy with difficult starting segments but an efficient ligation plan is preferable to one with easy starting segments but a long reaction path. For more nuanced analyses, the BracketMaker interface enables unique methods of visually exploring data, such as modifying color settings to highlight a known problem segment from Aligator, and then cycling through possible ligation orders to minimize this segment's path length. With our computational CPS tools, the information needed to make strategic decisions is readily accessible.

3.2 The D‐ribosome: A test case for computational CPS

A powerful advantage of CPS is the ability to produce mirror‐image D‐proteins (made of D‐amino acids), which cannot be obtained through other means. D‐proteins have enabled techniques such as racemic crystallography (Agouridas et al., 2020; Kent, 2019) for facile refinement of protein structures and mirror‐image phage display (Schumacher et al., 1996) for the discovery of proteolytically stable D‐peptide drugs (Liu et al., 2016; Welch et al., 2010). Fewer than 30 D‐protein syntheses have been reported (Agouridas et al., 2020; Rohden et al., 2021), and many of these are among the largest proteins synthesized to date (Fan et al., 2021; Pech et al., 2017; Wang et al., 2016; Weidmann et al., 2019; Weinstock et al., 2014; Xu et al., 2017; Xu & Zhu, 2022), each a remarkable achievement. To enable more routine access to larger D‐proteins, a long‐term goal in the CPS field is the synthesis of a complete mirror‐image ribosome, which could produce virtually any size of D‐protein via in vitro translation.

Only a few ribosomal subunits have been produced so far. The Liu group synthesized a eukaryotic ribosomal protein S25 (RpS25, 6 segments) as an early demonstration of convergent assembly enabled by peptide hydrazides (Fang et al., 2012). Our group synthesized L31 (3 segments) to demonstrate the use of temporary solubilizing tags in the ligation of difficult segments (Jacobsen et al., 2016). Finally, the Zhu group synthesized three subunits, L5 (4 segments), L18 (3 segments), and L25 (2 segments), and demonstrated that these proteins can co‐assemble with RNA into functional 5S ribonucleoprotein complexes in both L and D chirality (Ling et al., 2020). Beyond ribosomal subunits themselves, great progress has been made toward other foundational tools for mirror‐image biology (Rohden et al., 2021), including D versions of DNA polymerases (Jiang et al., 2017; Pech et al., 2017; Wang et al., 2016; Xu et al., 2017), DNA ligase (Weidmann et al., 2019), and RNA polymerase (Xu & Zhu, 2022) capable of binding mirror‐image L‐DNA or L‐RNA and producing functional nucleic acid polymers. The largest proteins among these, Pfu DNA polymerase (Fan et al., 2021) and T7 RNA polymerase (Xu & Zhu, 2022), were produced by the Zhu group using a “split fragment” approach in which each protein was divided into complementing fragments of an accessible length for CPS—two for Pfu (6 and 9 segments) and three for T7 (5, 6, and 8 segments)—that can co‐assemble into a functional enzyme. Altogether, this set represents a range of brackets of varying segment length and complexity up to the largest reported so far, providing several examples to compare Auto‐Bracket predictions with successful strategies employed by leaders in the CPS field.

We first reproduced the published brackets for ribosomal subunits (L5, L18, L25, L31, and RpS25) and split polymerase fragments (Pfu‐N, Pfu‐C, T7‐N, T7‐M, and T7‐C) and counted their bracket scores: the Max Path among all segments and Avg Path among all segments (applying path length penalties to our default list of poor thioesters), and total Rxn Steps. We then used Auto‐Bracket to assemble the same set of segments used by authors into every possible ligation order, with default program settings reflecting widely used hydrazide‐NCL chemistry: Tfa‐Thz as the N‐terminal Cys protecting group (with one‐pot deprotection), Acm as the internal Cys protecting group (without one‐pot deprotection), and default choices for penalized poor thioesters. All possible ligation orders were scored and ranked, and the winning bracket with minimal path length scores (i.e., the first displayed result) was compared to author strategies (Supplementary Table S1 and Supplementary Figure S1).

In cases where the bracket is either very simple (2‐segment L25, 3‐segment L18) or constrained by the relative placement of Cys and Ala ligation sites (4‐segment L5, 5‐segment T7‐C), only one or two ligation orders are possible. In these cases, the Auto‐Bracket winner either matches exactly to the ligation order used by authors, or both orders are tied. For 3‐segment L31, which contains a single poor thioester, the author strategy (N‐to‐C, conducting the penalized ligation second) differs from Auto‐Bracket's suggested result (C‐to‐N, penalized ligation first). At the time of this protein's publication, however, one‐pot deprotection of N‐terminal Cys was not widely used and may have been unavailable to the authors; re‐running the same input without one‐pot deprotection adds an extra step to the C‐to‐N bracket, and in this case, the Auto‐Bracket winner matches the author strategy (N‐to‐C). This comparison demonstrates how the optimal ligation order for a set of segments depends heavily on the type of chemistry used, and individual users should adopt settings matching their available chemistry.

For 6‐segment RpS25, the winning ligation order predicted by Auto‐Bracket is the most convergent possible assembly, matching the authors' choice exactly (#1 of 10). For other large proteins, author brackets can be found early in the output result list, including for 6‐segment Pfu‐C (#4 of 42), 6‐segment T7‐M (#4 of 42), and 8‐segment T7‐N (#4 of 28). In these cases, the authors and Auto‐Bracket both use convergent strategies overall, with subtle differences due to poor thioester penalties. The Auto‐Bracket winner always places poor thioester ligations as early as possible, applying the +1 path length penalty to fewer segments and reducing the average path length by a small amount (∆Avg Path = 0.17–0.5) compared to the authors' chosen order. If these penalties are ignored, the author brackets all tie for first place. It should be noted that the poor thioester list used for our analysis (which may be customized before each Auto‐Bracket run) represents BracketMaker default settings and does not necessarily reflect actual yields from author brackets. In reality, the success of a given ligation involves several hard‐to‐predict factors beyond the choice of thioester—such as solubility of segments and product, or contributions of secondary structure to reaction kinetics—which are unaccounted in bracket scores. This comparison demonstrates that our method of scoring brackets by counting path lengths is not a complete measurement of synthetic difficulty, but is a useful framework for separating the best from the worst possible ligation orders.

For 9‐segment Pfu‐N, the exact strategy used by authors is not found among Auto‐Bracket's 1430 possible brackets. In the author's strategy, desulfurization was conducted in two steps: on the left half of the protein after assembling the first 5 segments, and again on the final full‐length protein. Our assembly algorithm, however, assumes that any number of Cys can be desulfurized simultaneously, and thus will always use a single combined desulfurization when possible to avoid redundant steps. The closest match among Auto‐Bracket results, using the same order as authors but with a combined desulfurization, ranks #188 of 1430, with significant score differences between this and the Auto‐Bracket winner (∆Max Path = 2, ∆Avg Path = 0.89). The authors' final ligation uses a Leu thioester, applying a poor thioester penalty to all segments; however, this reaction had good yield in the reported synthesis and does not realistically deserve this penalty. When Auto‐Bracket is re‐run without including Leu as a poor thioester, the author bracket for Pfu‐N ties for #1 of 1430 (along with 7 other orders), and the original Auto‐Bracket winner ties for #9 (along with 59 others). This comparison demonstrates how small changes in user input settings can have great impact on the overall rank order of Auto‐Bracket predictions, and users should choose settings reflecting their personal preferences, experience, and available chemistry options. Overall, Auto‐Bracket predicts reasonable ligation orders for a given set of segments using logic similar to experienced protein chemists, quickly providing a starting list of options for further nuanced analysis.

3.3 Processing the Escherichia coli ribosome with Aligator and Auto‐Bracket

The minimal E. coli ribosome (including 11 accessory factors) is an ambitious synthetic target comprised of 65 proteins ranging from 38 to 890 aa (see our previous Aligator report (Jacobsen et al., 2017) for more information on these proteins). To simplify this challenge, we sought to apply our bracket‐building tools to lay out an initial blueprint and identify the most daunting obstacles.

We first processed the set of 65 protein sequences (21 30S subunits, 33 50S subunits, and 11 translation accessory factors) in Aligator to obtain up to 1000 top‐ranked strategies (i.e., sets of possible segments divided at valid ligation junctions) for each protein. Using our original default program settings for valid ligation junctions (restricted to junctions between an acceptable thioester aa [A, C, F, G, H, I, K, L, M, R, S, T, V, W, or Y] and thiol aa [C or A]) (Jacobsen et al., 2017) and a reasonable segment length limit (60 aa, less generous than in our previous report), we were unable to find valid strategies for 5 of the 65 proteins, as the gap between possible ligation junctions in these sequences was too large. Thus, we modified the program to accept additional thiol sites, which required substantial improvements to Aligator's processing capability to handle the increased number of possible segment combinations (see Supplementary Methods). For this analysis, we chose to include V as an acceptable third thiol site, as penicillamine (Pen, or β‐thio‐valine) is commercially available as an Fmoc‐SPPS building block (in L or D) and can be desulfurized simultaneously with Ala (Kulkarni et al., 2018). Pen ligations typically suffer from extremely hampered ligation kinetics (Haase et al., 2008) and thus are penalized more so than poor thioesters in both Aligator and BracketMaker scoring algorithms (poor thioesters paired with Pen receive combined penalties).

With an expanded list of possible ligation sites (allowing C, A, and V thiols) and a 60‐aa segment limit, we obtained valid strategies for all 65 ribosomal proteins ranging from 1 to 22 segments. Using the Aligator Import tool in BracketMaker, we applied Auto‐Bracket to assemble each set (up to 1000 per protein) into its most efficient calculated bracket with default settings for hydrazide‐NCL as described above. From the resulting list of up to 1000 brackets per protein, we identified the top overall bracket (smallest Max and Avg path lengths) for each (Figure 5, Supplementary Table S2, and Supplementary Figure S2).

FIGURE 5 Assessment of synthetic difficulty of Escherichia coli ribosome proteins. Each point represents one protein, and the height (y‐axis) of each point represents the lowest possible Max Path score among the optimized brackets (produced by Auto‐Bracket) of up to 1000 Aligator strategies (i.e., possible segment sets) per protein. Points are grouped into difficulty categories based on Max Path, and each category (Easy, Medium, Hard, and Danger Zone) is shown with its protein count. The shape and color of each point represent the types of thiol aa used at ligation junctions in each synthesis (in the specific best‐scoring bracket whose height is shown), as indicated in the figure legend. Five single‐segment proteins (<60 aa) are omitted from the graph; these have a Max Path of 0.

3.3.1 Initial assessment of ribosome brackets

Due to differences in the way Aligator and BracketMaker rank strategies (e.g., Aligator does not account for ligation order, and BracketMaker does not account for segment length or solubility), the top Aligator strategy differs from the best bracket in approximately half of cases within this ribosomal set, especially in larger strategies. In the protein L5, for example, the winning Aligator strategy uses four segments of ideal length (30–50 aa each, 179 aa total) joined at C, A, and V ligation sites with a Max Path of 6, with the V ligation contributing a +2 path penalty. In contrast, the winning L5 bracket in BracketMaker uses entirely A ligation sites for a significantly improved Max Path of 4, but the segments (three >50 aa and one 12 aa) score poorly in Aligator because of their lengths.

Other differences between Aligator and BracketMaker rankings are related to the number of A sites used and their relative placement to C sites. Every A ligation is given a −2 score penalty in Aligator, but in BracketMaker, multiple adjacent A sites result in only +1 additional path length from the combined desulfurization step regardless of their number (as in Figure 4f). The differences between Aligator and BracketMaker scores highlight the need for a nuanced analysis of synthesis plans and human interpretation of the output scores.

To assess the difficulty of the D‐ribosome project as a whole, we have assigned each protein a predicted difficulty rating—Easy, Medium, Hard, or in the “Danger Zone”—based on our own experience and precedents from literature. The Max Path score is a good gauge of synthesis complexity, as increases in this score correlate with protein size as well as diversity of ligation junctions (Figure 5). Thus, we can use Max Path as a simple rule of thumb (easily counted by eye) to initially judge synthesis difficulty.

3.3.2 “Easy” proteins (max path 0–2)

In the ribosome, 5 proteins (L30 and L32‐36) can potentially be made as a single segment up to 60 aa, 3 proteins (S21, L27, and L31) can be accessed through a single NCL at an ideal site (i.e., good thioester with Cys), and 6 proteins (L19, L22, L33, L12, S15, and S20) require a single NCL followed by desulfurization at an Ala site. These proteins have few possible division points, and the best choice (prioritizing assembly of soluble, medium‐length segments ~40 aa) is often obvious without computational assistance. Unexpected problems encountered during synthesis—segment insolubility, impurities from SPPS, or especially poor ligation efficiency—can often be overcome by brute force (i.e., starting with more material and accepting low yield) or by introducing additional side‐chain modifications such as temporary solubilizing tags (Fulcher et al., 2019; Jacobsen et al., 2016) or chemical auxiliaries for templated (proximity‐aided) ligation (Giesler et al., 2020). We therefore define this category as “Easy,” as these proteins should be generally accessible with standard CPS techniques.

Using Max Path as the primary score for grouping as we do here, the Easy category could also include brackets of 3–4 segments if they used only ideal Cys ligation junctions in a convergent order (in our E. coli ribosome set, Cys is sparse (Jacobsen et al., 2017) and no brackets of this description are found). In our judgment, it would be fair to group these ideal all‐Cys brackets in the Easy category. As more segments are introduced, however, special attention must be paid to initial quality of starting segments to avoid accumulating impurities in the product. Even “easy” brackets, therefore, require close assessment of initial segments, the full ligation path, and an individual lab's capabilities.

3.3.3 “Medium” proteins (max path 3–4)

About half of the ribosome (31 proteins) can be characterized by brackets of 2–4 segments with slightly sub‐optimal ligation paths. In these proteins, even the preferred ligation strategy faces at least one challenge, such as multiple A ligations requiring simultaneous desulfurization (L3, L4, L7, L9, L18, L20, L35, S5, S10, S16), unavoidable poor thioesters (IF‐1, L15, L16, L21, L25, S6, S7, S9, S19), a combination of A and C sites constraining possible ligation order (L10, L11, L14, L17, S11, S13, S14), or all A ligations with internal C in some segments necessitating Acm protection (L5, L6, L10, L28, L35, S12, S18).

Although the number of segments in these proteins may be small, choosing an optimal ligation strategy involves weighing one or more of these yield‐diminishing factors along with properties of initial segments. Therefore, these brackets may be considered of “Medium” difficulty. Proteins of this size should be generally within reach of standard techniques (up to ~200 aa in this ribosomal set) but can be significant research undertakings if problems are encountered.

Again, using only Max Path as our difficulty gauge, ideal convergent brackets containing 8 or even 16 segments could qualify as “Medium” (see hypothetical bracket scores in Figure 4). Intuitively this classification would not make sense, as the current record among published brackets is 10 segments (Xu et al., 2017). However, it is evident from this analysis that as more segments are introduced, the possibility of an ideal ligation strategy diminishes rapidly, and thus large brackets with low Max Path are extremely rare.

3.3.4 “Hard” proteins (max path 5–6)

As proteins increase in size and number of segments, challenges compound. The “Hard” category (12 ribosomal proteins) spans the largest range in terms of protein size (~100–400 aa) and the number of starting segments (2–9). The reason for each protein's difficulty rating varies, but in each case, challenges are reflected in the Max Path score. Any single V ligation plus desulfurization contributes a path length of +4, and thus two 2‐segment brackets using a single V site make it into this category, either because the V is paired with a poor thioester (L24) or requires internal C‐Acm protection (S17). Other brackets contain poor thioesters at especially difficult locations, either at the only available C site that must be used for the final ligation (IF‐3), or adjacent to other poor‐thioester sites leading to sequential poor‐yielding ligations (L1, L13, S3). Additional brackets contain similar challenges as described for the Medium category, but in higher frequency (L2, S4, S8, EF‐Ts, RF1, RRF).

Proteins with brackets of this size and difficulty have been achieved by leading labs in the field (Fan et al., 2021; Jiang et al., 2017; Premdjee et al., 2021; Sun & Brik, 2019; Weinstock et al., 2014; Xu et al., 2017; Xu & Zhu, 2022), but currently require intensive time and resources to produce. Yield losses at far‐downstream steps can be devastating, so proper planning is vital. Thus, syntheses in this category could benefit greatly from computational assistance as well as ligation‐ and solubility‐enhancing tools described previously (although these chemical tools have yet to be demonstrated in context of a large protein ≥300 aa).

3.3.5 The “danger zone” (max path 7+)

Finally, some ribosome brackets contain unprecedented obstacles, extending into the “Danger Zone” (8 in this set). Four of these proteins contain more segments than any published bracket on record (RF3 = 11 segments, S1 = 13 segments, EF‐G = 17 segments, and IF2 = 22 segments), and these present the same common challenges as described for Hard and Medium brackets but in higher frequency. Two proteins (EF‐Tu 1 and EF‐Tu 2, of nearly identical sequence) use three separate V ligations in their top brackets (all performed at top‐level between starting segments), and one (RF2) contains a V ligation followed by a poor thioester in its limiting path; these V ligations are likely to be significant yield bottlenecks. The most surprising member of this category is an 8‐segment protein of modest length (S2, 241 aa) with one large segment bordered by two V ligations in every possible Aligator strategy, requiring these ligations to be performed consecutively and leading to the largest Max Path (11) in the ribosomal set.

None of these proteins, save for S2, are of unprecedented Max Path (to our knowledge, the longest path length in any published bracket is 9, from a large protein assembled in near‐linear fashion (Xu et al., 2017)). Obtaining acceptable yields of product protein over this many sequential yield losses, however, requires starting segments to be produced at large scale. In addition to the tools we have already described, the synthesis of these proteins may rely on additional protein engineering, such as introducing acceptable mutations to provide additional ligation sites or splitting into independent subunits that can associate non‐covalently to form a functional complex. The combination of both approaches has been demonstrated in some of the most complex syntheses to date (Fan et al., 2021; Xu & Zhu, 2022), and these methods provide further opportunities for computational assistance to predict optimal mutations and/or stable sub‐domains. Furthermore, this analysis highlights the need for continued development of new synthesis tools, both to enhance poor‐yielding ligations and unlock alternative ligation sites.

Overall, thanks to many recent advances in CPS methodology, most of the ribosome appears to be within current or near‐future reach. The D‐ribosome project, however, will be a massive undertaking likely requiring combining diverse approaches and contributions from many groups. We, therefore, urge labs developing new CPS methodology to consider ribosomal proteins as model systems to showcase tools, as each synthesis and the data it provides will be a valuable stepping stone toward this shared goal.

4 CONCLUSION AND FUTURE DIRECTIONS

CPS is primed to become a data‐science driven field. As protein targets increase in size, synthetic challenges rapidly accumulate; however, an increasingly diverse chemical toolkit provides many solutions. As all protein targets and peptide segments are unique, every synthetic strategy requires deliberate planning and consideration of options, and our computational tools greatly assist with decision‐making and visualization. The enhanced processing speed of our updated Aligator enables consideration of additional ligation junctions and provides overhead for more sophisticated segment evaluation in the future. Additional factors to consider in future segment‐ and bracket‐scoring algorithms include the predicted formation of unintended byproducts (such as racemized side chains and aspartimides (Subirós‐Funosas et al., 2011)), placement of quality‐enhancing pseudoprolines (Wöhr et al., 1996) and isoacyl dipeptides (Yoshiya et al., 2007), nuanced prediction of solubility and structure (as well as suggestions for temporary solubilizing tags (Fulcher et al., 2019; Jacobsen et al., 2016)), and feedback from additional real‐world CPS data to improve the accuracy of scores. With BracketMaker's framework to perform mathematical calculations with brackets, we now have a starting point for holistic consideration of an entire synthesis, and future efforts will be aimed at fully integrating the two programs (e.g., weighing bracket scores based on starting and intermediate segment properties). In future updates, the retrosynthetic analysis process of our bracket assembly tool could easily be adapted to other types of ligation chemistry, protecting groups, and side‐chain modifications given a defined order of steps. For example, selenocysteine (Sec) and selective Sec/Cys dechalcogenation could be added as an option for Ala ligation sites in segments with internal Cys, avoiding the need for Acm protection/deprotection (Kulkarni et al., 2019; Mousa et al., 2017). Furthermore, ligation sites beyond Cys and Ala may be considered with techniques such as KAHA (Rohrbacher et al., 2015) and STL (Liu & Li, 2018), and a more advanced bracket‐building algorithm could include rules for combining these techniques and traditional NCL in a single synthesis with proper timing for chemical compatibility. Efforts to compile the varied methodology in existing CPS literature are already underway (Agouridas et al., 2017), and our standardized bracket format can assist with future database projects.

In addition, Aligator and BracketMaker are designed to bring CPS to a wider user base. As CPS finds increasingly varied applications in bioscience and drug discovery (Agouridas et al., 2020; Kent, 2019; Tan et al., 2020), it is on the cusp of becoming a commonplace technique. The modern chemical methodology of CPS is well established, robust, and based on commercially available reagents, making the technology theoretically accessible to any lab with SPPS capability, but the complexity of information is a major barrier to entry. With the tools described here, however, anyone can begin with a target protein sequence and obtain a step‐by‐step synthesis plan in a matter of minutes. Both Aligator and BracketMaker are open‐source and freely available to download, and we look forward to feedback from the community as these programs see more widespread use. Ultimately, we aim to increase accessibility of protein synthesis, facilitate communication across the CPS field, and equip researchers with the information needed to take on increasingly ambitious proteins.

5 METHODS

5.1 Instructions for installation and use

BracketMaker is available as a standalone executable for Windows and MacOS. The BracketMaker executable, its open‐source Python 3 code, and a user manual with detailed syntax and program instructions are freely available to download from the program's GitHub page (https://github.com/kay-lab/BracketMaker). See Supplementary Materials for a detailed program description.

Our updated Aligator v2 (now capable of high‐throughput protein processing) is available on its GitHub page (https://github.com/kay-lab/Aligator). See our original Aligator report (Jacobsen et al., 2017) for a general program description and Supplementary Materials for v2 updates.

5.2 Selection and processing of test data

5.2.1 DapA bracket

All references to the DapA sequence and bracket refer to the exact ligation strategy used for L‐DapA in our 2014 report (Weinstock et al., 2014), reproduced here in BracketMaker format. This bracket represents the largest and most complex CPS reported by our lab.

5.2.2 E. coli ribosome

The minimal E. coli ribosome consists of 65 total proteins, including all protein subunits of the 30S and 50S complexes as well as 11 key accessory factors. Definition and rationale of this test set are described in our initial Aligator report (Jacobsen et al., 2017). FASTA sequences for these proteins were obtained from UniProtKB.

Ribosome protein sequences were processed in Aligator with default thioesters (I, L, K, T, V accepted; A, C, F, H, G, M, R, S, W, Y preferred), an expanded set of thiols (C, A, and V), and a 60‐aa segment limit to obtain a list of up to 1000 top strategies each. The resulting “<Protein Name> All Strategies.txt” files were processed in BracketMaker via our Aligator Import tool to obtain the best‐scoring ligation order for each set of segments, giving a list of up to 1000 top brackets per protein in bracket notation (the strategies for IF2 were too large for Auto‐Bracket to process [22 segments] and resulted in memory errors, so we divided these strategies into smaller groups of segments separated by two available C sites [which must be used last in either order] and ran Auto‐Bracket separately, then manually combined the top brackets from each to obtain the top full IF2 bracket). The following Auto‐Bracket settings were used, representing available hydrazide‐NCL chemistry commonly used in our lab: One‐Pot Desulfurization = FALSE, One‐Pot Cys Deprotection = FALSE, Cys Protecting Group = “Acm,” One‐Pot N‐terminal Cys Deprotection = TRUE, N‐terminal Cys Protecting Group = “Tfa‐Thz.” The resulting list of brackets, each representing a unique combination of segments, were sorted by Max Path and then by Average Path to obtain the top‐ranked bracket (lowest path lengths) for each protein.

AUTHOR CONTRIBUTIONS

Judah L. Evangelista: Writing – original draft; conceptualization; investigation; methodology; software; data curation; visualization; funding acquisition. Michael S. Kay: Conceptualization; writing – review and editing; supervision; funding acquisition.

CONFLICT OF INTEREST STATEMENT

The author declares no conflicts of interest.

Supporting information

Data S1 Supporting Information

ACKNOWLEDGMENTS

The authors wish to thank Patrick Erickson and Paul Spaltenstein for their coding expertise and helpful discussions about Aligator, Michael Jacobsen for training in CPS and insightful manuscript feedback, and Shradda Nayak for graphic design guidance. This work was supported by NIH grants (T32‐GM122740 to J.L.E. and U54‐AI170856 to M.S.K).
==== Refs
REFERENCES

Agouridas V , El Mahdi O , Cargoet M , Melnyk O . A statistical view of protein chemical synthesis using NCL and extended methodologies. Bioorg Med Chem. 2017;25 :4938–4945.28578993
Agouridas V , El Mahdi O , Diemer V , Cargoët M , Monbaliu J‐CM , Melnyk O . Native chemical ligation and extended methods: mechanisms, catalysis, scope, and limitations. Chem Rev. 2019;119 :7328–7443.31050890
Agouridas V , El Mahdi O , Melnyk O . Chemical protein synthesis in medicinal chemistry. J Med Chem. 2020;63 :15140–15152.33236900
Asahina Y , Kawakami T , Hojo H . One‐pot native chemical ligation by combination of two orthogonal thioester precursors. Chem Commun. 2017;53 :2114–2117.
Bang D , Kent SB . A one‐pot total synthesis of crambin. Angew Chem Int Ed Engl. 2004;43 :2534–2538.15127445
Bello C , Hartrampf N , Walport LJ , Conibear AC . Protein chemistry looking ahead: 8(th) chemical protein synthesis meeting 16‐19 June 2019, Berlin, Germany. Cell Chem Biol. 2019;26 :1349–1354.31626782
Blanco‐Canosa JB , Nardone B , Albericio F , Dawson PE . Chemical protein synthesis using a second‐generation N‐acylurea linker for the preparation of peptide‐thioester precursors. J Am Chem Soc. 2015;137 :7197–7209.25978693
Chen J , Wan Q , Yuan Y , Zhu J , Danishefsky SJ . Native chemical ligation at valine: a contribution to peptide and glycopeptide synthesis. Angew Chem Int Ed Engl. 2008;47 :8521–8524.18833563
Conibear AC , Watson EE , Payne RJ , Becker CFW . Native chemical ligation in protein synthesis and semi‐synthesis. Chem Soc Rev. 2018;47 :9046–9068.30418441
Dawson PE , Muir TW , Clark‐Lewis I , Kent SB . Synthesis of proteins by native chemical ligation. Science. 1994;266 :776–779.7973629
Fan C , Deng Q , Zhu TF . Bioorthogonal information storage in L‐DNA with a high‐fidelity mirror‐image Pfu DNA polymerase. Nat Biotechnol. 2021;39 :1548–1555.34326549
Fang GM , Li YM , Shen F , Huang YC , Li JB , Lin Y , et al. Protein chemical synthesis by ligation of peptide hydrazides. Angew Chem. 2011;50 :7645–7649.21648030
Fang G‐M , Wang J‐X , Liu L . Convergent chemical synthesis of proteins by ligation of peptide hydrazides. Angew Chem Int Ed Engl. 2012;51 :10347–10350.22968928
Flood DT , Hintzen JCJ , Bird MJ , Cistrone PA , Chen JS , Dawson PE . Leveraging the Knorr pyrazole synthesis for the facile generation of thioester surrogates for use in native chemical ligation. Angew Chem Int Ed Engl. 2018;57 :11634–11639.29908104
Fulcher JM , Petersen ME , Giesler RJ , Cruz ZS , Eckert DM , Francis JN , et al. Chemical synthesis of Shiga toxin subunit B using a next‐generation traceless "helping hand" solubilizing tag. Org Biomol Chem. 2019;17 :10237–10244.31793605
Giesler RJ , Erickson PW , Kay MS . Enhancing native chemical ligation for challenging chemical protein syntheses. Curr Opin Chem Biol. 2020;58 :37–44.32745915
Haase C , Rohde H , Seitz O . Native chemical ligation at valine. Angew Chem Int Ed Engl. 2008;47 :6807–6810.18626881
Hackeng TM , Griffin JH , Dawson PE . Protein synthesis by native chemical ligation: expanded scope by using straightforward methodology. Proc Natl Acad Sci USA. 1999;96 :10068–10073.10468563
Huang YC , Chen CC , Gao S , Wang YH , Xiao H , Wang F , et al. Synthesis of L‐ and D‐ubiquitin by one‐pot ligation and metal‐free desulfurization. Chem A Eur J. 2016;22 :7623–7628.
Jacobsen MT , Erickson PW , Kay MS . Aligator: a computational tool for optimizing total chemical synthesis of large proteins. Bioorg Med Chem. 2017;25 :4946–4952.28651912
Jacobsen MT , Petersen ME , Ye X , Galibert M , Lorimer GH , Aucagne V , et al. A helping hand to overcome solubility challenges in chemical protein synthesis. J Am Chem Soc. 2016;138 :11775–11782.27532670
Jiang W , Zhang B , Fan C , Wang M , Wang J , Deng Q , et al. Mirror‐image polymerase chain reaction. Cell Discov. 2017;3 :17037.29051832
Jin K , Li X . Advances in native chemical ligation–desulfurization: a powerful strategy for peptide and protein synthesis. Chem A Eur J. 2018;24 :17397–17404.
Kent S . Chemical protein synthesis: inventing synthetic methods to decipher how proteins work. Bioorg Med Chem. 2017;25 :4926–4937.28687227
Kent SBH . Novel protein science enabled by total chemical synthesis. Protein Sci. 2019;28 :313–328.30345579
Kulkarni SS , Sayers J , Premdjee B , Payne RJ . Rapid and efficient protein synthesis through expansion of the native chemical ligation concept. Nat Rev Chem. 2018;2 :0122.
Kulkarni SS , Watson EE , Premdjee B , Conde‐Frieboes KW , Payne RJ . Diselenide‐selenoester ligation for chemical protein synthesis. Nat Protoc. 2019;14 :2229–2257.31227822
Ling JJ , Fan C , Qin H , Wang M , Chen J , Wittung‐Stafshede P , et al. Mirror‐image 5S ribonucleoprotein complexes. Angew Chem Int Ed Engl. 2020;59 :3724–3731.31841243
Liu H , Li X . Serine/threonine ligation: origin, mechanistic aspects, and applications. Acc Chem Res. 2018;51 :1643–1655.29979577
Liu M , Li X , Xie Z , Xie C , Zhan C , Hu X , et al. D‐peptides as recognition molecules and therapeutic agents. Chem Rec. 2016;16 :1772–1786.27255896
Maity SK , Jbara M , Laps S , Brik A . Efficient palladium‐assisted one‐pot deprotection of (acetamidomethyl)cysteine following native chemical ligation and/or desulfurization to expedite chemical protein synthesis. Angew Chem Int Ed Engl. 2016;55 :8108–8112.27126503
Mousa R , Notis Dardashti R , Metanis N . Selenium and selenocysteine in protein chemistry. Angew Chem Int Ed Engl. 2017;56 :15818–15827.28857389
Muir TW , Kent SB . The chemical synthesis of proteins. Curr Opin Biotechnol. 1993;4 :420–427.7763972
Pech A , Achenbach J , Jahnz M , Schulzchen S , Jarosch F , Bordusa F , et al. A thermostable D‐polymerase for mirror‐image PCR. Nucleic Acids Res. 2017;45 :3997–4005.28158820
Premdjee B , Andersen AS , Larance M , Conde‐Frieboes KW , Payne RJ . Chemical synthesis of phosphorylated insulin‐like growth factor binding protein 2. J Am Chem Soc. 2021;143 :5336–5342.33797881
Rohde H , Seitz O . Ligation‐desulfurization: a powerful combination in the synthesis of peptides and glycopeptides. Biopolymers. 2010;94 :551–559.20593472
Rohden F , Hoheisel JD , Wieden HJ . Through the looking glass: milestones on the road towards mirroring life. Trends Biochem Sci. 2021;46 :931–943.34294544
Rohrbacher F , Wucherpfennig TG , Bode JW . Chemical protein synthesis with the KAHA ligation. Top Curr Chem. 2015;363 :1–31.25761549
Schumacher TN , Mayr LM , Minor DL Jr , Milhollen MA , Burgess MW , Kim PS . Identification of D‐peptide ligands through mirror‐image phage display. Science. 1996;271 :1854–1857.8596952
Subirós‐Funosas R , El‐Faham A , Albericio F . Aspartimide formation in peptide chemistry: occurrence, prevention strategies and the role of N‐hydroxylamines. Tetrahedron. 2011;67 :8595–8606.
Sun H , Brik A . The journey for the total chemical synthesis of a 53 kDa protein. Acc Chem Res. 2019;52 :3361–3371.31536331
Tan Y , Wu H , Wei T , Li X . Chemical protein synthesis: advances, challenges, and outlooks. J Am Chem Soc. 2020;142 :20288–20298.
Terrier VP , Adihou H , Arnould M , Delmas AF , Aucagne V . A straightforward method for automated Fmoc‐based synthesis of bio‐inspired peptide crypto‐thioesters. Chem Sci. 2016;7 :339–345.29861986
Wan Q , Danishefsky SJ . Free‐radical‐based, specific desulfurization of cysteine: a powerful advance in the synthesis of polypeptides and glycopolypeptides. Angew Chem Int Ed Engl. 2007;46 :9248–9252.18046687
Wang Z , Xu W , Liu L , Zhu TF . A synthetic molecular system capable of mirror‐image genetic replication and transcription. Nat Chem. 2016;8 :698–704.27325097
Weidmann J , Schnolzer M , Dawson PE , Hoheisel JD . Copying life: synthesis of an enzymatically active mirror‐image DNA‐ligase made of D‐amino acids. Cell Chem Biol. 2019;26 :645–651.e3.30880154
Weinstock MT , Jacobsen MT , Kay MS . Synthesis and folding of a mirror‐image enzyme reveals ambidextrous chaperone activity. Proc Natl Acad Sci USA. 2014;111 :11679–11684.25071217
Welch BD , Francis JN , Redman JS , Paul S , Weinstock MT , Reeves JD , et al. Design of a potent D‐peptide HIV‐1 entry inhibitor with a strong barrier to resistance. J Virol. 2010;84 :11235–11244.20719956
Wöhr T , Wahl F , Nefzi A , Rohwedder B , Sato T , Sun X , et al. Pseudo‐prolines as a solubilizing, structure‐disrupting protection technique in peptide synthesis. J Am Chem Soc. 1996;118 :9218–9227.
Xu W , Jiang W , Wang J , Yu L , Chen J , Liu X , et al. Total chemical synthesis of a thermostable enzyme capable of polymerase chain reaction. Cell Discov. 2017;3 :17008.28265464
Xu Y , Zhu TF . Mirror‐image T7 transcription of chirally inverted ribosomal and functional RNAs. Science. 2022;378 :405–412.36302022
Yan LZ , Dawson PE . Synthesis of peptides and proteins without cysteine residues by native chemical ligation combined with desulfurization. J Am Chem Soc. 2001;123 :526–533.11456564
Yoshiya T , Taniguchi A , Sohma Y , Fukao F , Nakamura S , Abe N , et al. "O‐acyl isopeptide method" for peptide synthesis: synthesis of forty kinds of "O‐acyl isodipeptide unit" Boc‐Ser/Thr(Fmoc‐Xaa)‐OH. Org Biomol Chem. 2007;5 :1720–1730.17520140
Zheng JS , Tang S , Qi YK , Wang ZP , Liu L . Chemical synthesis of proteins using peptide hydrazides as thioester surrogates. Nat Protoc. 2013;8 :2483–2495.24232250
