
==== Front
Nat Commun
Nat Commun
Nature Communications
2041-1723
Nature Publishing Group UK London

39251606
52048
10.1038/s41467-024-52048-4
Article
Site-specific template generative approach for retrosynthetic planning
http://orcid.org/0000-0002-3728-0021
Shee Yu 1
Li Haote 1
Zhang Pengpeng 1
http://orcid.org/0000-0003-1907-1565
Nikolic Andrea M. 1
http://orcid.org/0000-0001-6388-7743
Lu Wenxin 1
http://orcid.org/0000-0003-3811-0662
Kelly H. Ray 2
Manee Vidhyadhar 2
Sreekumar Sanil 2
Buono Frederic G. 2
Song Jinhua J. 2
http://orcid.org/0000-0001-8741-7236
Newhouse Timothy R. timothy.newhouse@yale.edu

1
http://orcid.org/0000-0002-3262-1237
Batista Victor S. victor.batista@yale.edu

1
1 https://ror.org/03v76x132 grid.47100.32 0000 0004 1936 8710 Department of Chemistry, Yale University, New Haven, CT USA
2 grid.418412.a 0000 0001 1312 9717 Chemical Development, Boehringer Ingelheim Pharmaceuticals Inc, Ridgefield, CT USA
6 9 2024
6 9 2024
2024
15 781820 3 2024
26 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Retrosynthesis, the strategy of devising laboratory pathways by working backwards from the target compound, is crucial yet challenging. Enhancing retrosynthetic efficiency requires overcoming the vast complexity of chemical space, the limited known interconversions between molecules, and the challenges posed by limited experimental datasets. This study introduces generative machine learning methods for retrosynthetic planning. The approach features three innovations: generating reaction templates instead of reactants or synthons to create novel chemical transformations, allowing user selection of specific bonds to change for human-influenced synthesis, and employing a conditional kernel-elastic autoencoder (CKAE) to measure the similarity between generated and known reactions for chemical viability insights. These features form a coherent retrosynthetic framework, validated experimentally by designing a 3-step synthetic pathway for a challenging small molecule, demonstrating a significant improvement over previous 5-9 step approaches. This work highlights the utility and robustness of generative machine learning in addressing complex challenges in chemical synthesis.

Enhancing retrosynthetic efficiency requires overcoming the vast complexity of chemical space, the limited known interconversions between molecules, and the challenges posed by limited experimental datasets. Here, the authors introduce generative machine learning methods for retrosynthetic planning that generate reaction templates.

Subject terms

Method development
Cheminformatics
Synthetic chemistry methodology
National Energy Research Scientific Computing Center (NERSC), a U.S. Department of Energy Office of Science User Facility located at Lawrence Berkeley National Laboratory, operated under Contract No. DE-AC0205CH11231 using NERSC award BES-ERCAP0024372issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Retrosynthesis is the design of deconstructing complex molecules into simpler building blocks, a concept originally developed by Corey as a means to educate students to conduct multistep synthesis1. This intellectual framework laid the foundation for the development of ComputerAided Synthesis Planning (CASP), a field that emerged to assist chemists in navigating various paths of synthesis, playing a pivotal role in augmenting human capabilities for refining a synthetic approach2,3. In the earlier stages, systems based on expert rules provided valuable insights for chemists4–7. As organic chemistry advanced, encompassing broader chemical space and synthetic methodologies, recent advancements in CASP have shifted from rule-based to precedent-based approaches8. This shift was facilitated by large-scale extraction of reaction rules9. The process progressed from manual creation to automated extraction from extensive chemical datasets. Several extraordinary software packages have emerged due to this transition which empowered CASP tools to tap into repositories of historical reaction data8,10,11. Grzybowski and others12,13 further introduced user-purpose-driven tools for route optimization, demonstrating remarkable success through experimental validations14–18. Furthermore, the integration of machine learning (ML) methods has marked the latest chapter in the ongoing evolution of CASP19,20. ML models offer promising alternatives and can be broadly categorized as selection-based, semi-template, or generation-based methods21 (see Fig. 1a).Fig. 1 Common machine learning methods for retrosynthesis and our approach.

a Reactants and templates can be selected or generated based on a target compound using different machine learning models. Template generation is used in this work. b A structured latent space is incorporated in one of the models in this work. Sampling in the latent space can give different reaction templates for given products. c Reduction of synthetic steps for a key intermediate for active pharmaceutical ingredients (API). OPRD 2024 refers to ref. 62.

Selection-based methods, such as reactant selection and template selection methods, aim to choose appropriate molecules or reaction rules from the given sets. Reactant selection methods22,23 involve ranking molecules from a collection of candidates based on the target compounds. While reactant selection methods have the advantage of ensuring the chosen molecules are valid, their effectiveness relies on the availability of reactants in the candidate sets. Template selection methods24–30 rank the reaction templates in terms of their applicability to the target molecules. These templates capture subgraph patterns representing the change in atoms and bonds during a reaction. Notably, the RDChiral repository by Coley et al.31 offers template extraction methods and a collection of reaction templates in the form of SMARTS strings. Template selection methods simplify the reaction representation to a single template instead of multiple reactants. In addition, the same template can be applied to different target compounds instead of having multiple sets of reactants for the target compounds, thereby providing a higher coverage of reaction space. However, like reactant selection methods, template selection methods are dependent on the coverage and diversity of available templates within predefined reaction rules.

Semi-template methods32–37 involve the identification of reaction centers, synthons, or leaving groups, followed by the prediction of corresponding reactants based on these rules. Some semi-template methods34,36,37 are akin to selection-based methods, where reactants are obtained by predicting reaction centers and selecting from a collection of leaving groups, or by selecting necessary edits on molecular graphs. Other semi-template methods adopt generation components, in which reactants are generated from products and identified synthons or rules.

Generation-based methods are not bound by the sets of available reactants or templates and hold promise to map wider areas of chemical space. These include template-free methods38–53 that treat reactant generation as a translation task, aiming to predict the reactants directly from the given products without having in-dataset reaction rules. They therefore bear the potential to explore a wider range of possible reactions.

In this study, we introduce template generation which represents a new distinct category of generation-based methods for retrosynthetic planning. Template generation models employ the Sequence-to-Sequence (S2S) architecture trained to translate product information into reaction templates, as opposed to generating reactants. The capability of template generation thus extends beyond the available templates or predefined reaction rules of template selection-based approaches, enabling the discovery of novel reaction templates that expand the scope of retrosynthetic planning. The combination of generated reaction templates and the “RunReactants” function from RDKit, offer an efficient means to swiftly identify templates that yield grammatically coherent reactants from given products. This facilitates the exploration of previously uncharted chemical reactions and pathways.

One of the major benefits of using template generation is the ease of checking the reaction validity. During the transformation of a reaction template, the product is guaranteed to be converted to the reactant with exact matching of atoms indices and relevant functional groups from the description of the template. In comparison to reactant generative models, this feature greatly reduces the uncertainty in the produced reactants which might not correspond to any known reactions or have key atom mismatches due to problems during decoding.

Our template generation method introduces a design where site-specific templates (SST) are generated along with target compounds with labeled reaction centers (i.e., center-labeled products, CLP) that specify the reaction centers. This results in the generation of concise and informative sets of templates that are different from the templates available in the RDChiral repository31. Through benchmarking with a public dataset, the performance of our approach is demonstrated.

The second design is a sampling generative model (sampling model) for template generation conditioned by target compounds. S2S models, such as those employed in the template-free methods, predict pathways deterministically and do not have a sampling process or definition of latent space. In contrast, our sampling model has a latent space, enabling the generation, interpolation, extrapolation, and distance measurement of various templates (Fig. 1b). Deterministic models that take target compounds as inputs and generate templates are also developed in this work. Importantly, the encoder of the model can incorporate positional embedding for reaction centers, enabling users to specify specific reacting sites during prediction. Results are benchmarked on the USPTO-FULL dataset.

Our sampling model, based on the conditional kernel elastic autoencoder (CKAE)54, is the first of its kind in the field of retrosynthesis. This model conditions on corresponding products during training, allowing interpolation and extrapolation of reaction templates in latent space to generate new reaction templates during the sampling process. The latent space also provides a measure of distances between reaction templates, allowing us to identify the closest reaction reference within the dataset, or determine the similarity between two chemical reactions. Previous works on assessing reaction similarity and reaction classification use physicochemical properties55–57, molecular fingerprints58,59, or reaction SMILES strings as input60. In this work, SSTs and CLPs are used to evaluate the similarity of reactions. Schwaller et al.60 include reaction conditions such as catalysts and solvents in the reaction SMILES strings, while the fingerprint methods from Schneider et al.58 and Ghiandoni et al.59 require that reactants are separated from reagents. Our method is similar to Schwaller et al.60 in terms of using strings as input, but SSTs and CLPs provide a more concise way to represent reactions (without reagents) and carry additional information about atom mapping, like the method from ref. 59.

With SSTs and generation methods in place, our approach is validated through the practical application of synthesis. A library of potent anti-cancer agents was recently reported by Boehringer Ingelheim61. One of the key intermediates for the synthesis of these anti-cancer compounds is compound 1, a cyclohexanone with a quaternary stereogenic center in the α-position containing an alkyne moiety (Fig. 1c)62,63. Our objective was to develop a more step-efficient route to synthesize compound 1. The route proceeds over 3 steps, compared to prior approaches that required 5–9 steps, including a recent process involving Grignard-mediated epoxide opening as a key step in a 5-step route starting from commercial starting materials62,63. Reducing the number of steps in a synthetic process is enabling to develop scalable and more sustainable approaches, while also reducing the amount of time necessary for each batch64,65. Our experimental validation demonstrates the practicality and reliability of the retrosynthetic predictions, suggesting their underlying promise to address a wide spectrum of synthetic challenges.

Results

Site-specific templates and center-labeled products

Reaction templates that only apply to reaction centers within the target compounds are referred to as site-specific templates (SST). These are different from RDChiral templates which involve a broader structural context31 since SSTs do not differentiate neighboring atoms or special functional groups when matching substructures within the target compounds. The presence of center-labeled products (CLP) is a pre-requisite for the effective use of SSTs. Such labeling is essential to avoid ambiguity when a SST can be applied to multiple sites within a target compound. Examples of SST and CLP are shown in Fig. 2a where the “*” symbol represents the reaction centers. To prepare SSTs, the radius parameter in RDChiral is set to 0 (while RDChiral normally sets radius to 1 which captures 1 bond away from the reaction centers) and special functional groups are removed. Therefore, neighboring atoms and distal functional groups are not included in SSTs. Also, explicit degrees and explicit numbers of hydrogens are not included in the SSTs. To prepare CLPs, RDChiral also has implementations to capture the changed atoms, so the centers can be labeled for target compounds. Further explanations and examples are provided in the Supplementary Information (Sec. 1 and Sec. 2).Fig. 2 Schematic retrosynthetic workflow for Models A and B.

a Workflow of Model A. b Workflow of Model B. Model B has reaction center embeddings and does not have center-labeled products in the output. Detailed descriptions of the models are provided in the Supplementary Information (Sec. 3 and Sec. 4).

Deterministic model performance

Figure 2 shows a schematic representation of the deterministic models, Model A and Model B. Model A takes a target compound as an input and translates it into SSTs and CLPs. CLPs specify how the SSTs should be applied to the target compound. Model B takes the target and the specific reaction centers and generates templates corresponding to those specific sites (see Supplementary Information Sec. 3 for more information and Supplementary Fig. 3 for a comparison of the models).

Figure 3 shows the comparative analysis of the performance for Models A and B (highlighted in red), in terms of Top-K accuracy, as compared to state-of-the-art methods.

Top-K accuracy measures the percentage of top-K predictions containing reactants that precisely match the ground truth reactants in the testing set. The Top-K results are derived from the beam search method, where the product of the next token probabilities is used to rank the output templates and their corresponding precursors. This ranking is referred to as the beam score in this work. Figure 3 includes results for the original USPTO-Full testing dataset as well as for a cleaned testing set to address errors related to atom mappings such as solvent and reagent atoms erroneously considered as part of the reactions. The cleaned testing set is prepared by removing reactions containing reactants that are the 50 most frequently observed spectators in the USPTO-FULL dataset. The size of the cleaned testing set is 90.7% of the original set of 95k reactions.Fig. 3 USPTO-Full Top-K accuracy of retrosynthesis models.

GLN28, LocalRetro29, and Neuralsym24 in black are template-based selection methods. GraphRetro34, RetroPrime35, and RetroExplainer36 in yellow are semi-template methods. GTA45, Tied-Transformer50, MEGAN47, Transformer44, and R-SMILES52 in green are template-free generation methods. This work (in red) uses a template generation method. Reactant-based selection methods are not included due to out-of-memory for the USPTO-FULL dataset21. 1Indicates that if the correct reactants contain one of the 50 most commonly seen spectators in the USPTO-Full dataset, the reaction is removed from the test set. 2Indicates that reaction centers are provided. 3Indicates that the maximum number of reaction centers is two.

Model A, which does not use reaction centers, performs comparably well to other methods. The cleaned set allows for higher accuracy although it may inadvertently exclude some reactions where the common spectators actually participate as reactants. Model B leverages reaction center information. On the cleaned set, Model B reaches a performance milestone, achieving an accuracy rate as high as 80% for Top-10 predictions (see Supplementary Information Sec. 9 for details).

RetroExplainer36, with semi-template components, demonstrates remarkable prediction accuracy owing to its data modeling approach and the utilization of a set of leaving groups. However, this approach may experience variations in performance when handling uncommon scenarios or leaving groups not explicitly represented in the dataset. R-SMILES52, a template-free generation-based method, introduced the root-aligned SMILES representation to ensure minimal edit distances between product and reactant SMILES. Through this customized string representation and data augmentation, they achieved the highest accuracy among template-free methods. Nonetheless, data augmentation is not utilized in this work, leaving room for potential improvements in accuracy for future endeavors.

Top-K accuracy is not the sole criterion for evaluating retrosynthetic methods; explainability and inference time are equally important factors. The explainability of the template generation approach is facilitated by atom mappings from templates. Neuralsym24 is widely used for benchmarking multistep tree search methods due to its fast inference time. Our template generative approach has a similar order of magnitude of inference time as Neuralsym (101 s), while other methods operate at an order of 102 s or above, as shown in ref. 21. In this reference21, batch size optimization or multi-process multi-GPU acceleration are not implemented, so a batch size of 1 is used for comparison. R-SMILES, for instance, has an inference time at the order of 103 s. Although both R-SMILES and our approach are generation methods, the template generation approach using SSTs has a much shorter string representation than reactants and does not require augmented SMILES inputs to reach high accuracy, resulting in significantly shorter inference times.

In addition, an analysis of the Top-K accuracy considering different numbers of reaction centers for Model B is shown. Over half of the test reactions possess one or two reaction centers, following the same distribution of reaction center counts of the training set. Consequently, for test reactions with a maximum of two reaction centers, Model B achieved the highest Top-K accuracy compared to other center counts, with the Top-10 accuracy reached 90% (see last row of Fig. 3 and Supplementary Information Sec. 9), showcasing exceptional predictive capabilities in scenarios characterized by a limited number of reaction centers. This also aligns with the precursor selection process illustrated in “Experimental validation”, where two reaction centers are consistently utilized for Model B. The high Top-K accuracy achieved by Model B for reactions with few reaction centers is particularly significant, as it corresponds to real-world applications where a majority of reactions feature a low number of reaction centers. For instance, 90% of the dataset comprises reactions with no more than four reaction centers (see Supplementary Information Sec. 9).

Sampling model with latent space

A sampling generative model, which exploits a sampling process with a latent space, is different from the deterministic approach. To the best of our knowledge, the application of a sampling model for retrosynthetic planning has not been explored. Model C is built upon the architecture of Conditional Kernel-Elastic Autoencoder (CKAE)54. In Model C, both the input and output consist of combinations of SSTs and CLPs. The goal of Model C, akin to a variational autoencoder, is to reconstruct the input with latent space compression. Comparing to previous CKAE molecular generation models where conditions are represented by specific values or molecular properties, the CKAE model as applied to Model C utilizes the SMILES representation of target molecules as conditions. During the sampling process, a target compound is provided as the condition and latent vectors are sampled, different SSTs and CLPs, which correspond to the same target compound condition, can then be generated for different latent vectors.

In addition to generative sampling, the encoder of Model C offers a valuable referencing feature. It maps the input into a latent space with a distance regularized by a modified maximum mean discrepancy loss (m-MMD)54. This distance furnishes a quantifiable metric for assessing the similarity between reactions, aiding in evaluating and understanding the differences between chemical transformations. Such capability enables the identification of similar reactions within the dataset.

The distance between various chemical transformations in latent space can be used to interpolate between chemical reactions. This can be useful when searching for a reaction that could be an intermediate between two known chemical reactions. In Fig. 4a, an interpolation process is visualized. Initially, two reaction templates are selected, represented by the top and bottom templates and the latent vectors in the latent space. These templates serve as the starting points to explore the intermediates. This interpolation allowed the discovery of the templates corresponding to each of the latent vectors along the path between the two originals. It can be observed that the middle templates and reactants form a blending of the starting templates and reactants. This observation provides evidence that the latent space captures chemical information, showing the distance measure between various chemical transformations.Fig. 4 Interpolation of templates in the latent space of Model C and reactants from Model A and Model C outputs.

a The intermediates of the top and bottom latent representations are decoded. b Selected reactants for 2-, 3-, 4-substituted cyclohexanone derivatives as target compounds.

To illustrate the differences between Model A (deterministic) and Model C (generative sampling), the single-step predictions of 2-, 3-, and 4-substituted cyclohexanone derivatives are examined. Model A and Model C are compared because, unlike Model B, they both do not take reaction center information as input, and this comparison highlights the effects of latent sampling on model outputs. Based on the acquired results, representative precursors are selected for all three target molecules. As shown in Fig. 4b, Model A suggestions are primarily based on functional group interconversions and protection reactions. While Model C also proposes these transformations, diverse precursors and reactions are also proposed. These examples complement the intuitive bias of many synthetic chemists and point to areas of opportunity for the creative development of novel chemical transformations. Please refer to Supplementary Information Sec. 7 to see experiments and examples of how the models can generate reactions that extend beyond the available templates.

Regarding the usage of each model: Model A should be used when high-accuracy predictions are needed without explicit reaction center information. Model B is suitable when there are specific insights or constraints regarding reaction centers, requiring precise control over disconnections. Model C is ideal for seeking greater diversity and potential for unconventional transformations, as it leverages generative sampling to explore a broader chemical space without predefined reaction centers.

Experimental validation

Developing inexpensive, rapid, and robust methods for the synthesis of bioactive molecules is one of the key goals in pharmaceutical chemistry66. Herein, we utilized our Model B, chosen because of its high accuracy and reaction center embedding, for establishing the shortest route for the synthesis of a target compound. Figure 5 highlights how Model B can be used to navigate multiple options for retrosynthesis (see Supplementary Information Sec. 8 for more pharmaceutical examples). The top five ranked precursors, based on beam scores or synthetic accessibility (SA) scores67 as implemented in RDKit68, are shown. Each level corresponds to a new prediction by the single-step model to reach an intermediate. This tool highlights the interactive nature of the model with a human expert who selects intermediates for further analysis. Both ketone and aldehyde precursors for the target compound are highly ranked. Several intermediates and subsequent retrosynthetic steps were examined. The aldehyde was chosen due to subsequent retrosynthetic evaluation showing it to be a highly enabling retron. The next retrosynthetic step from the aldehyde included the alpha-allyl cyclohexanone, which facilitates the application of the highly robust Tsuji–Trost allylation.Fig. 5 Retrosynthetic planning for compound 1 by Model B.

The top five precursors, ranked by beam scores or SA scores, are displayed for each compound in left-to-right order. The red boxes highlight the selected precursors.

Figure 6a serves as a reference point derived from Model C. The left-hand side illustrates the allylation step that we employed in our synthesis. On the right-hand side, the reference is obtained by encoding the allylation template and the product labeled with the reaction center into Model C’s latent space. This process allows us to identify the closest latent vectors from the training dataset, and that closest reference corresponds to the reaction shown on the right-hand side of Fig. 6a. Interestingly, the exact chemical transformation that was suggested had previously been conducted, but is not in the USPTO-FULL dataset. This highlights how our approach compliments other synthetic planning tools, such as Reaxys and SciFinder.Fig. 6 Synthesis of compound 1.

a Reference found with Model C for the allylation step. b Experimental procedure for the selected route: (i) 2-methylcyclohexanone (1 equiv.), allyl methyl carbonate (3 equiv.), Pd2(dba)3 (5 mol % Pd), t-BuXPhos (11 mol %), R-TRIP (10 mol %), 3 Å MS, CyH, 45 °C, 5 days, 49%; (ii) O3, CH2Cl2, − 78 °C, then PPh3 (2 equiv.), −78 °C → rt, 16 h, 86%; iii) NfF (1.05 equiv.), BTTP (6 equiv.), DMF, −30 °C → rt, 19 h, 78%. The enantiomeric ratio is reported by Pupo et al.69.

In order to synthesize the enantiomerically enriched target molecule, we applied the enantioselective Pd-catalyzed Tsuji–Trost allylation of a ketone and applied conditions recently reported by Pupo et al.69. The prior literature protocol for this substrate reported an enantiomeric ratio of 95.5:4.5. The allylated intermediate was treated with ozone in order to obtain the ketoaldehyde derivative in good yield (Fig. 6b). For the final step, a modified procedure by Boltukhina et al70. was applied to form the alkyne in 78% yield. The overall yield of our 3-step route is 33%, despite our route not having undergone process optimization. It should be stated that further process optimization is expected to improve the efficiency of this approach, although this proof-of-concept demonstrates the ability to develop step-efficient routes. This experimental procedure serves as evidence that the newly developed ML models can facilitate the development of synthetic routes for pharmaceutically significant molecules and enhance existing routes.

An alternative to the route presented in Fig. 6b, an even shorter route to compound 1, could be one entailing direct α-alkynylation of 2-methylcyclohexanone. Methods for direct introduction of an alkyne moiety next to a ketone are scarce and rely on substitution with electrophilic alkyne species (selected examples71–76). Most commonly used in modern organic chemistry are hypervalent iodine reagents such as Waser’s or Ochiai’s reagent77. While this method would furnish the target molecule in fewer synthetic steps, it would have to be followed by the separation of two enantiomers since enantioselective α-alkynylation of ketones has not yet been reported.

Discussion

In this work, a string-based approach for retrosynthesis planning is introduced, utilizing generative models to address the challenges posed by the vast chemical space and synthesis complexity. Specifically, this work introduces template generation as a new category in machine learning methods for computer-aided synthesis planning. Two types of generative models are developed, including deterministic generative models (Model A and Model B) and a sampling generative model that utilizes CKAE (Model C).

Model A and Model B are benchmarked on the USPTO-FULL dataset. Particularly, Model B can incorporate reaction center information, enabling the generation of templates that apply to the specified reacting sites. On the other hand, Model C represents a pioneering application of sampling method from latent space, capable of generating diverse reactions. The design of Model C defines distances between reactions, which allows Model C to identify the closest reference from the dataset for newly generated templates, making it a suitable tool for generating and validating a wide range of potential reactions.

This work presents two approaches for single-step synthetic planning, high-accuracy deterministic models and high-diversity sampling models. The capability of specifying reacting sites, the availability of relevant reaction references, and the successful results of experimental validations on an important pharmaceutically relevant intermediate make the models valuable tools in guiding retrosynthetic analysis.

Methods

Training details

In total, 10% dropout was applied to all attention matrices and embedding vectors. ADAM optimizer78 was used with a learning rate of 5 × 10−5. Gradient normalization79 was set to 1.0. During training, each token in the input to the encoders is replaced by a mask token for Model A and Model B with the probability of 0.15.

Model architecture

Models A, B, and C each has 6 layers of transformer encoders and decoders as implemented in ref. 80. For Models A and B, 8 attention heads and an embedding size of 256 are used. For Model C, 16 attention heads and an embedding size of 512 are used.

The reaction center embeddings for Model B are achieved by adding the embedding of the reaction center token “*” at the specific position of the atoms similar to the concept of positional embeddings.

Model C is constructed based on the conditional kernel elastic autoencoder model54, with a 5120-dimensional latent space. The conditions are embeddings of target compounds and are also achieved by 6 layers of transformer encoders and 16 heads with an embedding size of 51280. These embeddings are then compressed into 10 embedding vectors by a linear layer and concatenated with the input embedding and the latent space. See Supplementary Information Sec. 4 and Supplementary Fig. 5 for more details and visualization of the architecture.

Beam search

To derive multiple possible predictions, beam search44 is used across all models. During decoding, the transformer decoder attends to the encoder output and the sequence that had been generated. The decoder outputs probabilities of all possible tokens for the next position in the sequence. Beam search maintains a fixed-size set of candidate sequences, the number that the method keeps is called the beam size B. The top B most probable sequences at each decoding step are selected to proceed to the next step of decoding until the stopping criteria of maximum allowed length are reached or an End Of Sequence (<EOS>) token is output.

For the Top-K accuracy test, beam search with a beam size of 50 is used during all decoding processes. At each decoding step, the model outputs the 50 most probable candidate tokens and continues the sequence until the stopping criteria are met.

The diversity of deterministic models is solely derived from the beam search process, as this type of model lacks a latent space for sampling. Consequently, generating novel reactions using a deterministic model through beam search can be challenging. In contrast, the sampling model, equipped with a latent space, can generate diverse and novel reactions more effectively.

Synthesis

Details of the synthesis, such as reaction conditions, purification, and NMR spectra, are provided in the Supplementary Information.

Supplementary information

Supplementary Information

Peer Review File

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-024-52048-4.

Acknowledgements

V.S.B. acknowledges a generous allocation of high-performance computing time from the National Energy Research Scientific Computing Center (NERSC), a U.S. Department of Energy Office of Science User Facility located at Lawrence Berkeley National Laboratory, operated under Contract No. DE-AC0205CH11231 using NERSC award BES-ERCAP0024372. The development of the methodology was supported by the NSF CCI grant (VSB, Award Number 2124511). The applications and experiments were supported by Boehringer Ingelheim.

Author contributions

The machine learning methods are developed by Y.S. and H.L., with equal contributions, under the guidance of V.S.B. The experimental validations are conducted by A.M.N. and P.Z., with equal contribution, and W.L., under the guidance of T.R.N. The experimental design and execution were advised and supervised by H.R.K., V.M., S.S., F.B., J.J.S., and T.R.N. The initial draft of the manuscript was primarily written by Y.S., with contributions from all authors during the final draft preparation.

Peer review

Peer review information

Nature Communications thanks the anonymous reviewers for their contribution to the peer review of this work. A peer review file is available.

Data availability

The 50 most commonly seen spectators are obtained from the USPTO-Full reaction file in RDChiral GitHub Repository31. While the train-validation-test split of the USPTO-Full dataset is obtained from the GitHub repository of ref. 44. Experimental data, such as the NMR spectra, are provided in the Supplementary Information. All data are available from the corresponding authors upon request.

Code availability

A user-friendly interface was developed, and all pre-trained models from this work can be accessed on models.batistalab.com.

Competing interests

V.S.B., H.L., and Y.S. have filed a patent application related to the work described in this manuscript. The patent is assigned to Yale University. The authors confirm that the patent filing does not affect the integrity or objectivity of the research presented. The remaining authors declare no competing interests.

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yu Shee, Haote Li.
==== Refs
References

1. James Corey E Todd Wipke W Computer-assisted design of complex organic syntheses: pathways for molecular synthesis can be devised with a computer and equipment for graphical communication Science 1969 166 178 192 10.1126/science.166.3902.178 17731475
James Corey, E. & Todd Wipke, W. Computer-assisted design of complex organic syntheses: pathways for molecular synthesis can be devised with a computer and equipment for graphical communication. Science 166, 178–192 (1969).17731475 10.1126/science.166.3902.178
2. Thakkar, A. J. The coming of the computer age to organic chemistry: recent approaches to systematic synthesis analysis. Computers Chem. 3–18 (2006).
3. Alan R Quantitative correlation of physical and chemical properties with chemical structure: utility for prediction Chem. Rev. 2010 110 5714 5789 10.1021/cr900238d 20731377
Alan, R. et al. Quantitative correlation of physical and chemical properties with chemical structure: utility for prediction. Chem. Rev. 110, 5714–5789 (2010).20731377 10.1021/cr900238d
4. Wipke, W. T. & Howe, W. J. Computerassisted Organic Synthesis (ACS Publications, 1977).
5. Gelernter HL Empirical explorations of SYNCHEM: the methods of artificial intelligence are applied to the problem of organic synthesis route discovery Science 1977 197 1041 1049 10.1126/science.197.4308.1041 17836062
Gelernter, H. L. et al. Empirical explorations of SYNCHEM: the methods of artificial intelligence are applied to the problem of organic synthesis route discovery. Science 197, 1041–1049 (1977).17836062 10.1126/science.197.4308.1041
6. Bauer J Fontain E Forstmeyer D Ugi I Interactive generation of organic reactions by igor 2 and the PC-assisted discovery of a new reaction Tetrahedron Comput. Methodol. 1988 1 129 132 10.1016/0898-5529(88)90017-6
Bauer, J., Fontain, E., Forstmeyer, D. & Ugi, I. Interactive generation of organic reactions by igor 2 and the PC-assisted discovery of a new reaction. Tetrahedron Comput. Methodol. 1, 129–132 (1988).10.1016/0898-5529(88)90017-6
7. Hanessian S Franco J Larouche B The psychobiological basis of heuristic synthesis planning-man, machine and the Chiron approach Pure Appl. Chem. 1990 62 1887 1910 10.1351/pac199062101887
Hanessian, S., Franco, J. & Larouche, B. The psychobiological basis of heuristic synthesis planning-man, machine and the Chiron approach. Pure Appl. Chem. 62, 1887–1910 (1990).10.1351/pac199062101887
8. Ravitz O Data-driven computer aided synthesis design Drug Discov. Today.: Technol. 2013 10 e443 e449 10.1016/j.ddtec.2013.01.005 24050141
Ravitz, O. Data-driven computer aided synthesis design. Drug Discov. Today.: Technol. 10, e443–e449 (2013).24050141 10.1016/j.ddtec.2013.01.005
9. Cook A Computer-aided synthesis design: 40 years on Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2012 2 79 107
Cook, A. et al. Computer-aided synthesis design: 40 years on. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2, 79–107 (2012).
10. Bøgevig A Route design in the 21st century: the ic synth software tool as an idea generator for synthesis prediction Org. Process Res. Dev. 2015 19 357 368 10.1021/op500373e
Bøgevig, A. et al. Route design in the 21st century: the ic synth software tool as an idea generator for synthesis prediction. Org. Process Res. Dev. 19, 357–368 (2015).10.1021/op500373e
11. Davies IW The digitization of organic synthesis Nature 2019 570 175 181 10.1038/s41586-019-1288-y 31190012
Davies, I. W. The digitization of organic synthesis. Nature 570, 175–181 (2019).31190012 10.1038/s41586-019-1288-y
12. Gothard CM Rewiring chemistry: algorithmic discovery and experimental validation of one-pot reactions in the network of organic chemistry Angew. Chem. Int. Ed. 2012 51 7922 7927 10.1002/anie.201202155
Gothard, C. M. et al. Rewiring chemistry: algorithmic discovery and experimental validation of one-pot reactions in the network of organic chemistry. Angew. Chem. Int. Ed. 51, 7922–7927 (2012).10.1002/anie.201202155
13. Kowalik M Parallel optimization of synthetic pathways within the network of organic chemistry Angew. Chem. Int. Ed. 2012 51 7928 7932 10.1002/anie.201202209
Kowalik, M. et al. Parallel optimization of synthetic pathways within the network of organic chemistry. Angew. Chem. Int. Ed. 51, 7928–7932 (2012).10.1002/anie.201202209
14. Klucznik T Efficient syntheses of diverse, medicinally relevant targets planned by computer and executed in the laboratory Chem 2018 4 522 532 10.1016/j.chempr.2018.02.002
Klucznik, T. et al. Efficient syntheses of diverse, medicinally relevant targets planned by computer and executed in the laboratory. Chem 4, 522–532 (2018).10.1016/j.chempr.2018.02.002
15. Mikulak-Klucznik B Computational planning of the synthesis of complex natural products Nature 2020 588 83 88 10.1038/s41586-020-2855-y 33049755
Mikulak-Klucznik, B. et al. Computational planning of the synthesis of complex natural products. Nature 588, 83–88 (2020).33049755 10.1038/s41586-020-2855-y
16. Lin Y Reinforcing the supply chain of umifenovir and other antiviral drugs with retrosynthetic software Nat. Commun. 2021 12 7327 10.1038/s41467-021-27547-3 34916512
Lin, Y. et al. Reinforcing the supply chain of umifenovir and other antiviral drugs with retrosynthetic software. Nat. Commun. 12, 7327 (2021).34916512 10.1038/s41467-021-27547-3
17. Hardy MA Nan B Wiest Olaf Sarpong R Strategic elements in computer-assisted retrosynthesis: a case study of the pupukeanane natural products Tetrahedron 2022 104 132584 10.1016/j.tet.2021.132584 36743342
Hardy, M. A., Nan, B., Wiest, Olaf & Sarpong, R. Strategic elements in computer-assisted retrosynthesis: a case study of the pupukeanane natural products. Tetrahedron 104, 132584 (2022).36743342 10.1016/j.tet.2021.132584
18. Lin Y Zhang R Wang D Cernak T Computer-aided key step generation in alkaloid total synthesis Science 2023 379 453 457 10.1126/science.ade8459 36730413
Lin, Y., Zhang, R., Wang, D. & Cernak, T. Computer-aided key step generation in alkaloid total synthesis. Science 379, 453–457 (2023).36730413 10.1126/science.ade8459
19. Filipa de Almeida A Moreira R Rodrigues T Synthetic organic chemistry driven by artificial intelligence Nat. Rev. Chem. 2019 3 589 604 10.1038/s41570-019-0124-0
Filipa de Almeida, A., Moreira, R. & Rodrigues, T. Synthetic organic chemistry driven by artificial intelligence. Nat. Rev. Chem. 3, 589–604 (2019).10.1038/s41570-019-0124-0
20. Struble TJ Current and future roles of artificial intelligence in medicinal chemistry synthesis J. Med. Chem. 2020 63 8667 8682 10.1021/acs.jmedchem.9b02120 32243158
Struble, T. J. et al. Current and future roles of artificial intelligence in medicinal chemistry synthesis. J. Med. Chem. 63, 8667–8682 (2020).32243158 10.1021/acs.jmedchem.9b02120
21. Zhong Z Recent advances in deep learning for retrosynthesis Wiley Interdiscip. Rev.: Comput. Mol. Sci. 2024 14 e1694
Zhong, Z. et al. Recent advances in deep learning for retrosynthesis. Wiley Interdiscip. Rev.: Comput. Mol. Sci. 14, e1694 (2024).
22. Guo Z Wu S Ohno M Yoshida R Bayesian algorithm for retrosynthesis J. Chem. Inf. Model. 2020 60 4474 4486 10.1021/acs.jcim.0c00320 32975943
Guo, Z., Wu, S., Ohno, M. & Yoshida, R. Bayesian algorithm for retrosynthesis. J. Chem. Inf. Model. 60, 4474–4486 (2020).32975943 10.1021/acs.jcim.0c00320
23. Lee, H. et al. Retcl: a selection-based approach for retrosynthesis via contrastive learning. Preprint at. https://arxiv.org/abs/2105.00795 (2021).
24. Segler MH Waller MP Neuralsymbolic machine learning for retrosynthesis and reaction prediction Chem. A Eur. J. 2017 23 5966 5971 10.1002/chem.201605499
Segler, M. H. & Waller, M. P. Neuralsymbolic machine learning for retrosynthesis and reaction prediction. Chem. A Eur. J. 23, 5966–5971 (2017).10.1002/chem.201605499
25. Coley CW Rogers L Green WH Jensen KF Computer-assisted retrosynthesis based on molecular similarity ACS Cent. Sci. 2017 3 1237 1245 10.1021/acscentsci.7b00355 29296663
Coley, C. W., Rogers, L., Green, W. H. & Jensen, K. F. Computer-assisted retrosynthesis based on molecular similarity. ACS Cent. Sci. 3, 1237–1245 (2017).29296663 10.1021/acscentsci.7b00355
26. Ishida S Terayama K Kojima R Takasu K Okuno Y Prediction and interpretable visualization of retrosynthetic reactions using graph convolutional networks J. Chem. Inf. Model. 2019 59 5026 5033 10.1021/acs.jcim.9b00538 31769668
Ishida, S., Terayama, K., Kojima, R., Takasu, K. & Okuno, Y. Prediction and interpretable visualization of retrosynthetic reactions using graph convolutional networks. J. Chem. Inf. Model. 59, 5026–5033 (2019).31769668 10.1021/acs.jcim.9b00538
27. Fortunato ME Coley CW Barnes BC Jensen KF Data augmentation and pretraining for template-based retrosynthetic prediction in computer-aided synthesis planning J. Chem. Inf. Model. 2020 60 3398 3407 10.1021/acs.jcim.0c00403 32568548
Fortunato, M. E., Coley, C. W., Barnes, B. C. & Jensen, K. F. Data augmentation and pretraining for template-based retrosynthetic prediction in computer-aided synthesis planning. J. Chem. Inf. Model. 60, 3398–3407 (2020).32568548 10.1021/acs.jcim.0c00403
28. Dai, H., Li, C., Coley, C., Dai, B. & Song, L Retrosynthesis prediction with conditional graph logic network. Adv. Neural Inf. Process. Syst. 32 (2019).
29. Chen S Jung Y Deep retrosynthetic reaction prediction using local reactivity and global attention JACS Au 2021 1 1612 1620 10.1021/jacsau.1c00246 34723264
Chen, S. & Jung, Y. Deep retrosynthetic reaction prediction using local reactivity and global attention. JACS Au 1, 1612–1620 (2021).34723264 10.1021/jacsau.1c00246
30. Seidl P Improving few-and zero-shot reaction template prediction using modern Hopfield networks J. Chem. Inf. Model. 2022 62 2111 2120 10.1021/acs.jcim.1c01065 35034452
Seidl, P. et al. Improving few-and zero-shot reaction template prediction using modern Hopfield networks. J. Chem. Inf. Model. 62, 2111–2120 (2022).35034452 10.1021/acs.jcim.1c01065
31. Coley CW Green WH Jensen KF Rdchiral: an RDKit wrapper for handling stereochemistry in retrosynthetic template extraction and application J. Chem. Inf. Model. 2019 59 2529 2537 10.1021/acs.jcim.9b00286 31190540
Coley, C. W., Green, W. H. & Jensen, K. F. Rdchiral: an RDKit wrapper for handling stereochemistry in retrosynthetic template extraction and application. J. Chem. Inf. Model. 59, 2529–2537 (2019).31190540 10.1021/acs.jcim.9b00286
32. Yan C Retroxpert: decompose retrosynthesis prediction like a chemist Adv. Neural Inf. Process. Syst. 2020 33 11248 11258
Yan, C. et al. Retroxpert: decompose retrosynthesis prediction like a chemist. Adv. Neural Inf. Process. Syst. 33, 11248–11258 (2020).
33. Shi, C., Xu, M., Guo, H., Zhang, M. & Tang, J. A graph to graphs framework for retrosynthesis prediction. In International Conference on Machine Learning 8818–8827 (PMLR, 2020).
34. Somnath VR Bunne C Coley C Krause A Barzilay R Learning graph models for retrosynthesis prediction Adv. Neural Inf. Process. Syst. 2021 34 9405 9415
Somnath, V. R., Bunne, C., Coley, C., Krause, A. & Barzilay, R. Learning graph models for retrosynthesis prediction. Adv. Neural Inf. Process. Syst. 34, 9405–9415 (2021).
35. Wang X A diverse, plausible and transformer-based method for single-step retrosynthesis predictions Chem. Eng. J. 2021 420 129845 10.1016/j.cej.2021.129845
Wang, X. et al. A diverse, plausible and transformer-based method for single-step retrosynthesis predictions. Chem. Eng. J. 420, 129845 (2021).10.1016/j.cej.2021.129845
36. Wang Y Retrosynthesis prediction with an interpretable deep-learning framework based on molecular assembly tasks Nat. Commun. 2023 14 6155 10.1038/s41467-023-41698-5 37788995
Wang, Y. et al. Retrosynthesis prediction with an interpretable deep-learning framework based on molecular assembly tasks. Nat. Commun. 14, 6155 (2023).37788995 10.1038/s41467-023-41698-5
37. Zhong W Yang Z Chen CY Retrosynthesis prediction using an end-to-end graph generative architecture for molecular graph editing Nat. Commun. 2023 14 3009 10.1038/s41467-023-38851-5 37230985
Zhong, W., Yang, Z. & Chen, C. Y. Retrosynthesis prediction using an end-to-end graph generative architecture for molecular graph editing. Nat. Commun. 14, 3009 (2023).37230985 10.1038/s41467-023-38851-5
38. Liu B reaction prediction using neural sequence-to-sequence models ACS Cent. Sci. 2017 3 1103 1113 10.1021/acscentsci.7b00303 29104927
Liu, B. et al. reaction prediction using neural sequence-to-sequence models. ACS Cent. Sci. 3, 1103–1113 (2017).29104927 10.1021/acscentsci.7b00303
39. Karpov, P., Godin, G. & Tetko, I. V. A transformer model for retrosynthesis. In International Conference on Artificial Neural Networks, 817–830 (Springer, 2019).
40. Chen, B., Shen, T., Jaakkola, T. S. & Barzilay, R. Learning to make generalizable and diverse predictions for retrosynthesis. Preprint at. https://arxiv.org/abs/1910.09688 (2019).
41. Lee AA Molecular transformer unifies reaction prediction and retrosynthesis across pharma chemical space Chem. Commun. 2019 55 12152 12155 10.1039/C9CC05122H
Lee, A. A. et al. Molecular transformer unifies reaction prediction and retrosynthesis across pharma chemical space. Chem. Commun. 55, 12152–12155 (2019).10.1039/C9CC05122H
42. Lin K Xu Y Pei J Lai L Automatic retrosynthetic route planning using template-free models Chem. Sci. 2020 11 3355 3364 10.1039/C9SC03666K 34122843
Lin, K., Xu, Y., Pei, J. & Lai, L. Automatic retrosynthetic route planning using template-free models. Chem. Sci. 11, 3355–3364 (2020).34122843 10.1039/C9SC03666K
43. Zheng S Rao J Zhang Z Xu J Yang Y Predicting retrosynthetic reactions using self-corrected transformer neural networks J. Chem. Inf. Model. 2019 60 47 55 10.1021/acs.jcim.9b00949 31825611
Zheng, S., Rao, J., Zhang, Z., Xu, J. & Yang, Y. Predicting retrosynthetic reactions using self-corrected transformer neural networks. J. Chem. Inf. Model. 60, 47–55 (2019).31825611 10.1021/acs.jcim.9b00949
44. Tetko IV Karpov P Van Deursen R Godin G State-of-the-art augmented nlp transformer models for direct and single-step retrosynthesis Nat. Commun. 2020 11 5575 10.1038/s41467-020-19266-y 33149154
Tetko, I. V., Karpov, P., Van Deursen, R. & Godin, G. State-of-the-art augmented nlp transformer models for direct and single-step retrosynthesis. Nat. Commun. 11, 5575 (2020).33149154 10.1038/s41467-020-19266-y
45. Seo, S. W. et al. Gta: graph truncated attention for retrosynthesis. In Proceedings of the AAAI Conference on Artificial Intelligence Vol. 35, 531–539 (Association for the Advancement of Artificial Intelligence (AAAI), 2021).
46. Mao K Xiao X Xu T Rong Y Huang J Zhao P Molecular graph enhanced transformer for retrosynthesis prediction Neurocomputing 2021 457 193 202 10.1016/j.neucom.2021.06.037
Mao, K., Xiao, X., Xu, T., Rong, Y., Huang, J. & Zhao, P. Molecular graph enhanced transformer for retrosynthesis prediction. Neurocomputing 457, 193–202 (2021).10.1016/j.neucom.2021.06.037
47. Sacha M edit graph attention network: modeling chemical reactions as sequences of graph edits J. Chem. Inf. Model. 2021 61 3273 3284 10.1021/acs.jcim.1c00537 34251814
Sacha, M. et al. edit graph attention network: modeling chemical reactions as sequences of graph edits. J. Chem. Inf. Model. 61, 3273–3284 (2021).34251814 10.1021/acs.jcim.1c00537
48. Mann V Venkatasubramanian V Retrosynthesis prediction using grammar-based neural machine translation: an information-theoretic approach Computers Chem. Eng. 2021 155 107533 10.1016/j.compchemeng.2021.107533
Mann, V. & Venkatasubramanian, V. Retrosynthesis prediction using grammar-based neural machine translation: an information-theoretic approach. Computers Chem. Eng. 155, 107533 (2021).10.1016/j.compchemeng.2021.107533
49. Ucak UV Kang T Ko J Lee J Substructure-based neural machine translation for retrosynthetic prediction J. Cheminformatics 2021 13 4 10.1186/s13321-020-00482-z
Ucak, U. V., Kang, T., Ko, J. & Lee, J. Substructure-based neural machine translation for retrosynthetic prediction. J. Cheminformatics 13, 4 (2021).10.1186/s13321-020-00482-z
50. Kim E Lee D Kwon Y Park MS Choi YS Valid, plausible, and diverse retrosynthesis using tied two-way transformers with latent variables J. Chem. Inf. Model. 2021 61 123 133 10.1021/acs.jcim.0c01074 33410697
Kim, E., Lee, D., Kwon, Y., Park, M. S. & Choi, Y. S. Valid, plausible, and diverse retrosynthesis using tied two-way transformers with latent variables. J. Chem. Inf. Model. 61, 123–133 (2021).33410697 10.1021/acs.jcim.0c01074
51. Irwin R Dimitriadis S He J Bjerrum EJ Chemformer: a pre-trained transformer for computational chemistry Mach. Learn.: Sci. Technol. 2022 3 015022
Irwin, R., Dimitriadis, S., He, J. & Bjerrum, E. J. Chemformer: a pre-trained transformer for computational chemistry. Mach. Learn.: Sci. Technol. 3, 015022 (2022).
52. Zhong Z Root-aligned smiles: a tight representation for chemical reaction prediction Chem. Sci. 2022 13 9023 9034 10.1039/D2SC02763A 36091202
Zhong, Z. et al. Root-aligned smiles: a tight representation for chemical reaction prediction. Chem. Sci. 13, 9023–9034 (2022).36091202 10.1039/D2SC02763A
53. Ucak UV Ashyrmamatov I Ko J Lee J Retrosynthetic reaction pathway prediction through neural machine translation of atomic environments Nat. Commun. 2022 13 1186 10.1038/s41467-022-28857-w 35246540
Ucak, U. V., Ashyrmamatov, I., Ko, J. & Lee, J. Retrosynthetic reaction pathway prediction through neural machine translation of atomic environments. Nat. Commun. 13, 1186 (2022).35246540 10.1038/s41467-022-28857-w
54. Li H Kernel-elastic autoencoder for molecular design PNAS Nexus 2024 3 168 10.1093/pnasnexus/pgae168
Li, H. et al. Kernel-elastic autoencoder for molecular design. PNAS Nexus 3, 168 (2024).10.1093/pnasnexus/pgae168
55. Chen L Gasteiger J Organic reactions classified by neural networks: Michael additions, Friedel–Crafts alkylations by alkenes, and related reactions Angew. Chem. Int. Ed. Engl. 1996 35 763 765 10.1002/anie.199607631
Chen, L. & Gasteiger, J. Organic reactions classified by neural networks: Michael additions, Friedel–Crafts alkylations by alkenes, and related reactions. Angew. Chem. Int. Ed. Engl. 35, 763–765 (1996).10.1002/anie.199607631
56. Chen L Gasteiger J Knowledge discovery in reaction databases: landscaping organic reactions by a self-organizing neural network J. Am. Chem. Soc. 1997 119 4033 4042 10.1021/ja960027b
Chen, L. & Gasteiger, J. Knowledge discovery in reaction databases: landscaping organic reactions by a self-organizing neural network. J. Am. Chem. Soc. 119, 4033–4042 (1997).10.1021/ja960027b
57. Satoh H Classification of organic reactions: similarity of reactions based on changes in the electronic features of oxygen atoms at the reaction sites J. Chem. Inf. Computer Sci. 1998 38 210 219 10.1021/ci9701190
Satoh, H. et al. Classification of organic reactions: similarity of reactions based on changes in the electronic features of oxygen atoms at the reaction sites. J. Chem. Inf. Computer Sci. 38, 210–219 (1998).10.1021/ci9701190
58. Schneider N Lowe DM Sayle RA Landrum GA Development of a novel fingerprint for chemical reactions and its application to large-scale reaction classification and similarity J. Chem. Inf. Model. 2015 55 39 53 10.1021/ci5006614 25541888
Schneider, N., Lowe, D. M., Sayle, R. A. & Landrum, G. A. Development of a novel fingerprint for chemical reactions and its application to large-scale reaction classification and similarity. J. Chem. Inf. Model. 55, 39–53 (2015).25541888 10.1021/ci5006614
59. Ghiandoni GM Development and application of a data-driven reaction classification model: comparison of an electronic lab notebook and medicinal chemistry literature J. Chem. Inf. Model. 2019 59 4167 4187 10.1021/acs.jcim.9b00537 31529948
Ghiandoni, G. M. et al. Development and application of a data-driven reaction classification model: comparison of an electronic lab notebook and medicinal chemistry literature. J. Chem. Inf. Model. 59, 4167–4187 (2019).31529948 10.1021/acs.jcim.9b00537
60. Schwaller P Mapping the space of chemical reactions using attention-based neural networks Nat. Mach. Intell. 2021 3 144 152 10.1038/s42256-020-00284-w
Schwaller, P. et al. Mapping the space of chemical reactions using attention-based neural networks. Nat. Mach. Intell. 3, 144–152 (2021).10.1038/s42256-020-00284-w
61. Abbott, J. et al. Annulated 2-amino-3cyano thiophenes and derivatives for the treatment of cancer. US Patent 11,945,812 (2024).
62. Tan Z Development of a scalable synthesis toward a kras g12c inhibitor building block bearing an all-carbon quaternary stereocenter, part 2: asymmetric synthesis via shi epoxidation Org. Process Res. Dev. 2024 28 78 91 10.1021/acs.oprd.3c00363
Tan, Z. et al. Development of a scalable synthesis toward a kras g12c inhibitor building block bearing an all-carbon quaternary stereocenter, part 2: asymmetric synthesis via shi epoxidation. Org. Process Res. Dev. 28, 78–91 (2024).10.1021/acs.oprd.3c00363
63. Leung JC Development of a scalable synthesis toward a kras g12c inhibitor building block bearing an all-carbon quaternary stereocenter, part 1: from discovery route to kilogram-scale production Org. Process Res. Dev. 2024 28 67 77 10.1021/acs.oprd.3c00362
Leung, J. C. et al. Development of a scalable synthesis toward a kras g12c inhibitor building block bearing an all-carbon quaternary stereocenter, part 1: from discovery route to kilogram-scale production. Org. Process Res. Dev. 28, 67–77 (2024).10.1021/acs.oprd.3c00362
64. Newhouse T Baran PS Hoffmann RW The economies of synthesis Chem. Soc. Rev. 2009 38 3010 3021 10.1039/b821200g 19847337
Newhouse, T., Baran, P. S. & Hoffmann, R. W. The economies of synthesis. Chem. Soc. Rev. 38, 3010–3021 (2009).19847337 10.1039/b821200g
65. Colberg J K(Mimi) Hii K Koenig SG Importance of green and sustainable chemistry in the chemical industry: a joint virtual issue between acs sustainable chemistry & engineering and organic process research & development Org. Process Res. Dev. 2022 26 2176 2178 10.1021/acs.oprd.2c00171
Colberg, J., K(Mimi) Hii, K. & Koenig, S. G. Importance of green and sustainable chemistry in the chemical industry: a joint virtual issue between acs sustainable chemistry & engineering and organic process research & development. Org. Process Res. Dev. 26, 2176–2178 (2022).10.1021/acs.oprd.2c00171
66. Eastgate MD Schmidt MA Fandrick KR On the design of complex drug candidate syntheses in the pharmaceutical industry Nat. Rev. Chem. 2017 1 0016 10.1038/s41570-017-0016
Eastgate, M. D., Schmidt, M. A. & Fandrick, K. R. On the design of complex drug candidate syntheses in the pharmaceutical industry. Nat. Rev. Chem. 1, 0016 (2017).10.1038/s41570-017-0016
67. Ertl P Schuffenhauer A Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions J. Cheminformatics 2009 1 1 11 10.1186/1758-2946-1-8
Ertl, P. & Schuffenhauer, A. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. J. Cheminformatics 1, 1–11 (2009).10.1186/1758-2946-1-8
68. Landrum, G. et al. Rdkit: Open-source Cheminformatics https://scholar.google.com/citations?view_op=view_citation&hl=zh-TW&user=xr9paY0AAAAJ&citation_for_view=xr9paY0AAAAJ:J_g5lzvAfSwC (2006).
69. Pupo G Properzi R List B Asymmetric catalysis with CO2: the direct α-allylation of ketones Angew. Chem. Int. Ed. 2016 55 6099 6102 10.1002/anie.201601545
Pupo, G., Properzi, R. & List, B. Asymmetric catalysis with CO2: the direct α-allylation of ketones. Angew. Chem. Int. Ed. 55, 6099–6102 (2016).10.1002/anie.201601545
70. Boltukhina EV Sheshenev AE Lyapkalo IM Convenient synthesis of nonconjugated alkynyl ketones from keto aldehydes by a chemoselective one-pot nonaflation—base catalyzed elimination sequence Tetrahedron 2011 67 5382 5388 10.1016/j.tet.2011.05.095
Boltukhina, E. V., Sheshenev, A. E. & Lyapkalo, I. M. Convenient synthesis of nonconjugated alkynyl ketones from keto aldehydes by a chemoselective one-pot nonaflation—base catalyzed elimination sequence. Tetrahedron 67, 5382–5388 (2011).10.1016/j.tet.2011.05.095
71. Kende AS Fludzinski P Chloroacetylenes as Michael acceptors. ii. direct ethynylation and vinylation of tertiary enolates Tetrahedron Lett. 1982 23 2373 2376 10.1016/S0040-4039(00)87345-1
Kende, A. S. & Fludzinski, P. Chloroacetylenes as Michael acceptors. ii. direct ethynylation and vinylation of tertiary enolates. Tetrahedron Lett. 23, 2373–2376 (1982).10.1016/S0040-4039(00)87345-1
72. Nishimura Y Amemiya R Yamaguchi M α-ethynylation reaction of ketones using catalytic amounts of trialkylgallium base Tetrahedron Lett. 2006 47 1839 1843 10.1016/j.tetlet.2005.12.133
Nishimura, Y., Amemiya, R. & Yamaguchi, M. α-ethynylation reaction of ketones using catalytic amounts of trialkylgallium base. Tetrahedron Lett. 47, 1839–1843 (2006).10.1016/j.tetlet.2005.12.133
73. Utaka A Cavalcanti LN Silva LF Electrophilic alkynylation of ketones using hypervalent iodine Chem. Commun. 2014 50 3810 3813 10.1039/C4CC00608A
Utaka, A., Cavalcanti, L. N. & Silva, L. F. Electrophilic alkynylation of ketones using hypervalent iodine. Chem. Commun. 50, 3810–3813 (2014).10.1039/C4CC00608A
74. Wegener M Kirsch SF The reactivity of 4-hydroxy-and 4-silyloxy-1, 5-allenynes with homogeneous gold (i) catalysts Org. Lett. 2015 17 1465 1468 10.1021/acs.orglett.5b00348 25739001
Wegener, M. & Kirsch, S. F. The reactivity of 4-hydroxy-and 4-silyloxy-1, 5-allenynes with homogeneous gold (i) catalysts. Org. Lett. 17, 1465–1468 (2015).25739001 10.1021/acs.orglett.5b00348
75. Wang J Protecting-group-free syntheses of ent-kaurane diterpenoids:[3+ 2+ 1] cycloaddition/cycloalkenylation approach J. Am. Chem. Soc. 2020 142 2238 2243 10.1021/jacs.9b13722 31968171
Wang, J. et al. Protecting-group-free syntheses of ent-kaurane diterpenoids:[3+ 2+ 1] cycloaddition/cycloalkenylation approach. J. Am. Chem. Soc. 142, 2238–2243 (2020).31968171 10.1021/jacs.9b13722
76. Jang D Choi M Chen J Lee C Enantioselective total synthesis of (+)garsubellin A Angew. Chem. 2021 133 22917 22921 10.1002/ange.202109193
Jang, D., Choi, M., Chen, J. & Lee, C. Enantioselective total synthesis of (+)garsubellin A. Angew. Chem. 133, 22917–22921 (2021).10.1002/ange.202109193
77. Hari DP Caramenti P Waser J Cyclic hypervalent iodine reagents: enabling tools for bond disconnection via reactivity umpolung Acc. Chem. Res. 2018 51 3212 3225 10.1021/acs.accounts.8b00468 30485071
Hari, D. P., Caramenti, P. & Waser, J. Cyclic hypervalent iodine reagents: enabling tools for bond disconnection via reactivity umpolung. Acc. Chem. Res. 51, 3212–3225 (2018).30485071 10.1021/acs.accounts.8b00468
78. Kingma, D. P. & Ba, J. Adam: a method stochastic optimization. Preprint at. https://arxiv.org/abs/1412.6980 (2014).
79. Chen, Z., Badrinarayanan, V., Lee, C. Y. & Rabinovich, A. Gradnorm: gradient normalization for adaptive loss balancing in deep multitask networks. In International Conference on Machine Learning, 794–803 (PMLR, 2018).
80. Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30 (2017).
