
==== Front
bioRxiv
BIORXIV
bioRxiv
2692-8205
Cold Spring Harbor Laboratory

39282447
10.1101/2024.09.05.611466
preprint
1
Article
CryptKeeper: a negative design tool for reducing unintentional gene expression in bacteria
http://orcid.org/0000-0001-6219-3168
Roots Cameron T. Conceptualization Methodology Software Visualization Writing – Original Draft Preparation Writing – Review & Editing
http://orcid.org/0000-0003-0888-7358
Barrick Jeffrey E. Conceptualization Funding Acquisition Methodology Writing – Original Draft Preparation Writing – Review & Editing *
Department of Molecular Biosciences, Center for Systems and Synthetic Biology, The University of Texas at Austin, Austin, Texas 78712, U.S.A.
* Correspondence: jbarrick@cm.utexas.edu
05 9 2024
2024.09.05.611466https://creativecommons.org/licenses/by/4.0/ This work is licensed under a Creative Commons Attribution 4.0 International License, which allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use.
nihpp-2024.09.05.611466.pdf
Foundational techniques in molecular biology—such as cloning genes, tagging biomolecules for purification or identification, and overexpressing recombinant proteins—rely on introducing non-native or synthetic DNA sequences into organisms. These sequences may be recognized by the transcription and translation machinery in their new context in unintended ways. The cryptic gene expression that sometimes results has been shown to produce genetic instability and mask experimental signals. Computational tools have been developed to predict individual types of gene expression elements, but it can be difficult for researchers to contextualize their collective output. Here, we introduce CryptKeeper, a software pipeline that visualizes predictions of bacterial gene expression signals and estimates the translational burden possible from a DNA sequence. We investigate several published examples where cryptic gene expression in E. coli interfered with experiments. CryptKeeper accurately postdicts unwanted gene expression from both eukaryotic virus infectious clones and individual proteins that led to genetic instability. It also identifies off-target gene expression elements that resulted in truncations that confounded protein purification. Incorporating negative design using CryptKeeper into reverse genetics and synthetic biology workflows can help to mitigate cloning challenges and avoid unexplained failures and complications that arise from unintentional gene expression.

plasmid instability
reliability and reproducibility
computational DNA sequence design
design-build-test cycle
recombinant protein overexpression
Defense Advanced Research Projects AgencyHR0011-17-2-0052 National Institutes of HealthR01GM088344 Army Research OfficeW911NF-20-1-0195 National Science FoundationIOS-2103208 MCB-2123996
==== Body
pmcIntroduction

RNAs and proteins may be unexpectedly transcribed and translated from a DNA sequence. This type of cryptic gene expression can complicate studying and engineering biological systems. Cryptic gene expression can occur when natural DNA sequences contain promoters and ribosome-binding sites that are not annotated because they are redundant, antisense, or internal to genes. It can also emerge when sequences are moved into a new cellular context (e.g., cloning eukaryotic sequences in E. coli) (1–4) or as a consequence of engineered changes to sequences (e.g., combining genetic parts, optimizing codon usage, or introducing artificial watermarks) (5,6). Cryptic gene expression products that interfere with the intended function of an engineered DNA construct may cause a design to be deemed a failure. Worse yet, cryptic gene expression may occur unbeknownst to researchers, causing them to misinterpret experimental results (7,8). Unintentional expression of genes, truncated pieces of genes, or out-of-frame products is often burdensome or even toxic to a host organism, creating a strong selection pressure favoring cells with mutations in the engineered DNA sequence (1–4,9). Rapid evolution of escape mutants that eliminate these or other sources of burden can be one reason that certain sequences are unreliable or even unconstructable (10,11).

Both the process and products of gene expression can be burdensome to a cell. Studies of recombinant protein overexpression have shown that growth rates of bacterial cells decrease in proportion to how much of their translational capacity, usually determined by the number of ribosomes, is redirected to expressing exogenous proteins (12,13). Expression of some proteins is also directly deleterious due to their activities, whether they are enzymes that rewire metabolism in ways that redirect limiting resources away from cellular replication or disrupt cellular homeostasis in other ways (11,12). It is rarer for transcription of RNA alone to cause an appreciable burden on a bacterial cell, but it has been documented in yeast protein overexpression systems (14,15).

Negative design is the process of eliminating undesirable qualities to engineer a safer or more effective system (16). The relative rates at which ribosomes initiate translation from different start codons in E. coli and other bacteria can be accurately predicted from characteristics of their ribosome-binding sites and surrounding sequences (17–19). Therefore, one negative design strategy for solving problems stemming from cryptic protein expression is to concentrate on redesigning a DNA sequence to eliminate the potential for unwanted translation, whether or not there is any evidence that a relevant mRNA is transcribed. Because translation and transcription are coupled in bacteria, disrupting translation is also expected to reduce RNA levels by short-circuiting transcription and promoting mRNA degradation (20,21).

Another negative design approach would be to eliminate cryptic transcription so mRNAs are not produced in the first place. Unfortunately, tools for predicting bacterial promoters and terminators currently have limited accuracy, only make qualitative predictions, and/or are not openly accessible (22–40). Furthermore, many promoter prediction tools use coding sequences as their non-promoter group during development, an assumption which could lead to systematically misclassifying cryptic promoters found within ORFs (30,37,40). Even so, predictions of promoters and terminators could provide additional context for interpreting predictions of translated reading frames and warn of other potential problems with a sequence design, such as the accidental production of inhibitory transcripts that are antisense to known genes.

Here we describe CryptKeeper, an open-source software tool that integrates and displays predictions of Escherichia coli gene expression elements in engineered DNA constructs, such as plasmids. CryptKeeper is designed to allow users to evaluate the potential for cryptic gene expression that may interfere with the construction or function of a DNA sequence. We demonstrate the utility of CryptKeeper by using it to analyze the results of several prior studies in which researchers identified cryptic gene expression that was problematic and then redesigned their sequences to avoid it.

Methods

Software overview

CryptKeeper integrates the output of several tools that predict bacterial gene expression elements from DNA sequences (Fig. 1A). It accepts input sequences in GenBank or FASTA format. It displays predictions of translation initiation sites, promoters, Rho-dependent terminators, and intrinsic (Rho-independent) terminators. These predictions are summarized with a translational burden score and displayed in an interactive visualization so that a user can evaluate whether there is the potential for cryptic gene expression that may interfere with the function or stability of their DNA sequence. CryptKeeper is a Python package. It and all of its dependencies can be installed as Bioconda packages.

Translational Burden Prediction

Thermodynamic models can accurately predict the relative rates at which ribosomes initiate translation from different start codons (17–19). CryptKeeper calculates the translation initiation rates at all start codons in the input sequence using OSTIR version 1.1.2 (Roots et al., 2021). The translational burden of a DNA construct on an E. coli host cell is expected to be proportional to the number of ribosomes bound to new mRNAs transcribed from it (12,13). If rates of translation elongation and termination are fast relative to initiation and uniform, then ribosome occupancy of a given open-reading frame (ORF) will be directly proportional to the rate of initiation at its start codon and its length. Therefore, we summarize these results as a translational burden score for each ORF that is the product of its predicted translation initiation rate and its length in base pairs. Very short ORFs (<45 bases) are unlikely to contribute much to the overall burden of a construct. They are predicted by Cryptkeeper but are not shown in its graphical output (to avoid unnecessary visual clutter) when using the default settings.

Promoter and Terminator Prediction

CryptKeeper displays predictions of E. coli σ70 promoters from a fork of the Promoter Calculator version 1.2.2 (29) that we created to add multithreading, reduce memory usage, and make it installable as a Bioconda package. Intrinsic (Rho-independent) terminators are predicted using TransTermHP version 2.09 (28). Rho-dependent terminators are predicted using a fork of RhoTermPredict version 3.4.0 (23) that we created to make it installable as a Bioconda package. CryptKeeper does not attempt to integrate transcription predictions into an overall score because these they are less accurate and complete than translation predictions. For example, the Promoter Calculator only predicts σ70 promoter initiation strength with a coefficient of determination of 0.45 for a test set of plasmid-encoded promoters in E. coli (29). While some tools exist that predict promoters that use alternative sigma factors, they are classifiers that do not quantitatively predict strength or are not open-source tools that can be run at the command line (30,37,40). By contrast, prediction of intrinsic terminators reaches >90% accuracy and specificity (28) and should generalize to many other bacterial species. However, there is almost always some read-through of these terminators (42,43) and this characteristic is not predicted by current tools. Predictions of Rho-dependent terminators may also generalize across bacteria, but they tend to have indistinct boundaries and even less is known about how well their presence and efficiencies are predicted by current algorithms (Di Salvo et al., 2019). Despite these current shortcomings, predictions of transcription initiation and termination elements may provide additional context to the user and may be sufficient for spotting problems in certain cases.

Output

For an input DNA construct (Fig. 1B), CryptKeeper uses the Bokeh Python library (44) to output an HTML document that includes an interactive plot (Fig. 1C). This plot displays stacked boxes associated with different ORFs. The height of each box is proportional to the predicted rate of translation initiation at the start codon of the ORF, which makes its area proportional to the translational burden score. Boxes are also colored according to their burden scores on a linear scale. Promoter and terminator predictions are shown on two inner tracks, one for each DNA strand. The number of these predictions shown can be adjusted by the user. The default is to show the three strongest promoters, Rho-dependent terminators, and intrinsic terminators per kilobase of the input sequence. If a GenBank file was used as the input, features annotated in this file are shown in the central track. The CryptKeeper plot can be zoomed and rescaled, and it displays information about each predicted feature and annotation on mouseover. To facilitate further analysis by users, a table describing each predicted element is provided below the plot and in a separate comma-separated values (CSV) output file.

Test Datasets

We tested CryptKeeper on DNA sequences from six published studies that encountered and characterized cryptic gene expression from plasmids in E. coli (Table 1). In each case, the relevant sequences were recreated in silico. When sufficient information was available, the entire plasmid sequences were reconstructed and analyzed. All of these studies report how researchers introduced mutations that resolved their issues, which allowed us to further examine how well CryptKeeper output tracks with the experimentally validated outcomes of redesigning DNA sequences. Plasmid annotations were based on GenBank records (45), pLannotate predictions (41), and descriptions in the relevant studies.

Results

Virus Infectious Clone Case Studies

Cloning a plant or animal virus into an E. coli vector makes it possible to use standard molecular biology workflows to modify its sequence. These infectious clone plasmids are a widespread reverse genetics tool for studying viruses and their applications in biotechnology. Because the virus DNA in an infectious clone is replicating in a cell that is evolutionarily distant from its eukaryotic host, it should not produce toxic virus proteins or infectious particles. However, sometimes there are cryptic gene expression elements in virus sequences that direct transcription and translation in E. coli cells. These products can cause significant translational burden or toxicity, leading to plasmid instability.

In a study that cloned Zika virus into a bacterial plasmid to create an infectious clone, researchers identified two putative E. coli promoters they designated ECP1 and ECP2 within the nucleotide sequence of the E envelope protein that they suspected were responsible for the expression of toxic products (1). CryptKeeper predicts ECP1 as the strongest promoter within the Zika genome. ECP2 is also identified by CryptKeeper, but it is among the weaker promoter predictions in the construct, so it is not shown in the plot with the default settings. The researchers found that introducing point mutations in both ECP1 and ECP2, as well as ECP1 alone, allowed them to stabilize the infectious clone. Their sequence changes reduced the transcription initiation rate predicted by CryptKeeper for ECP1 by 67% and did not change the predicted rate of ECP2 (Fig. 2A). Both promoters are located approximately 1800 bases upstream of an ORF with a predicted translational burden that is similar to what we found explained instability in other case studies. However, these researchers did not test for cryptic translation, so we cannot determine whether the instability they observed was due to translational burden or a toxic effect of a protein product.

In another study, researchers found that a Dengue virus infectious clone plasmid was unstable in E. coli, which they attributed to cryptic bacterial promoters within its 5ʹ untranslated region (UTR) (46). Subsequent studies produced stabilized Dengue virus infectious clones through a variety of approaches. Several copies of the TetR binding site tetO were sufficient for its binding to the 5ʹ UTR to prevent transcription initiation from the upstream promoters (47). Using a low-copy bacterial artificial chromosome instead of a high-copy plasmid also stabilized an infectious Dengue virus clone, presumably because this reduced all cryptic expression (47,48). Recently, researchers constructed a stable infectious clone in a high-copy plasmid by introducing a synthetic intron within the NS1 coding region (3). This addition interrupts translation of the long Dengue polyprotein in E. coli cells where it is not spliced, but the intron is removed and the virus RNA becomes infectious when transfected into mammalian cells. CryptKeeper predicts several promoters in the Dengue 5ʹ UTR, including a weak promoter overlapping with previously predicted ones. It also predicts a strong E. coli ribosome-binding site that initiates translation beginning at M126 of the Dengue polyprotein (Fig 2B). Adding the synthetic intron introduces a stop codon that interrupts this very long reading frame (Fig 2B), which reduces the predicted translational burden from this ORF by 88.4% and reduces the total translational burden of the complete infectious clone plasmid by 48.6% (Fig. 2C).

Eukaryotic Transgene Case Studies

Gene expression costs are not restricted to plasmids that encode viruses. Individual proteins or protein complexes are often cloned into bacteria to study their functions or for biomanufacturing. Long or toxic ORFs in these constructs can present cloning issues similar to those of the infectious clone plasmids.

When cloned in E. coli, the mouse mdr1a cDNA was found to contain a bacterial promoter and ribosome binding site near the 5ʹ end of its ORF (4). These elements contributed to instability. Mutating the start codon associated with the ribosome binding site (RBS) resulted in a stable plasmid. CryptKeeper predicts both the cryptic promoter and the cryptic RBS reported in the study. The researchers eliminated cryptic translation in E. coli by changing the M107 ATG start codon of mdr1a to CTG. CryptKeeper predicts that this edit should completely abolish expression of the highly burdensome ORF (Fig. 3A, B), thereby reducing the burden of the cDNA portion of the plasmid by 60.9% (Fig 3D).

Another similar study found that the SCN1A cDNA encoding the human sodium channel Nav1.1 contains a cryptic promoter and translation initiation site that results in strong expression of a truncated product in E. coli (2). Introduction of a β-globin/IgG chimeric intron containing an in-frame stop codon was used to disrupt E. coli translation and establish plasmid stability in this case. CryptKeeper detects both the promoter and translation initiation site suspected of causing cryptic gene expression (Fig. 3C). Interruption of the cryptic ORF by the introduced intron reduces the predicted translational burden associated with its initiation site by 78.4% (Fig. 3B) and the score for the complete plasmid by 44.3% (Fig. 3D).

Protein Truncation Case Studies

Experimental complications from cryptic translation are not limited to instability from burden. Truncated proteins produced by internal translation initiation sites can disrupt fusions to purification tags, antibody epitopes, or fluorescent reporters, decoupling these sequences from the protein of interest. It has been demonstrated that mutating predicted translation initiation sites can be used eliminate the unintentional production of truncated proteins (49). CryptKeeper is able to detect and visualize these internal ribosome binding sites, which can help researchers diagnose and redesign their sequences to prevent the unintentional translation of truncated proteins.

Researchers who cloned the yeast GCN5 gene in E. coli to express and purify the yeast SAGA histone acetyltransferase observed a smaller product that was suspected to be a proteolytic degradation product of full-length SAGA on SDS-PAGE gels (7). Editing the construct to disrupt an internal RBS or start codon reduced or eliminated this band, revealing that the smaller product was a truncated protein resulting from cryptic translation. CryptKeeper predicts that the rate of translation initiation at this internal start codon is ~350% than that of the upstream start codon for the full-length protein (Fig. 4A). The mutated RBS sequence used in the study reduced the truncated protein’s predicted expression to just 8% of the unmodified truncation and 29% of the full-length protein. As shown in the study, mutating this start codon from GTG to GTT entirely eliminated the truncated product.

A NISTmAb human antibody gene sequence for E. coli expression produced a shorter protein that copurified with the complete antibody (6). Later, it was demonstrated that this product was a truncated heavy chain produced by an RBS and GTG start codon that were unintentionally introduced during the codon optimization process (8). In a construct in which the researchers used this putative RBS and GTG codon to drive sfGFP expression to investigate the source of this unwanted product, mutating the start codon to GTC fully eliminated fluorescence. CryptKeeper identifies the unintended RBS in the sfGFP expression construct and correctly predicts that no full-length protein will be produced after the GTG to GTC start codon mutation (Fig. 4B).

Discussion

DNA synthesis, assembly, and cloning workflows are critical for a wide variety of bioengineering tasks, including vector construction, protein purification, enzyme engineering, genetic circuit design, and more. Cryptic gene expression can disrupt these workflows and obfuscate experimental results, leading to abandoning constructs, time-consuming troubleshooting, or incorrect conclusions. Since many cloning failures go unexplained and unpublished, problems with cryptic gene expression are undoubtedly underreported. Currently, there is no freely accessible, open-source solution for integrating output from the ecosystem of tools for predicting gene expression elements into a visual dashboard that makes potential design issues immediately evident to a researcher. We show that CryptKeeper can effectively diagnose these issues, as described in the troubleshooting case studies.

Ultimately, the utility of CryptKeeper is limited by the computational tools that are available for predicting gene expression. For E. coli, these challenges currently include high rates of false-positives/false-negatives and poor quantitative accuracy when predicting promoters, even with state-of-the-art algorithms trained on large sets of experimental data (22,29,30,33,37,40). Tools for predicting transcription initiation rates driven by alternative sigma factors, transcription driven by T7 RNA polymerase, and quantitative predictions of terminator read-through are also needed to complete the picture of transcription in this host. We expect tools for predicting gene expression to improve as laboratory automation makes more extensive training sets available and as new machine learning approaches are adopted (e.g., large language models) (50,51). Ideally, it would be possible to tailor these tools for different bacterial hosts used as alternative chassis for cloning (e.g. Vibrio natriegens) (52) or for specific bioengineering applications (e.g., Pseudomonas putida) (53). Another goal for the field should be to extend these approaches to widely used eukaryotic chassis, such as budding yeast (Saccharomyces cerevisiae), where cryptic gene expression also poses challenges (54).

The main summary output from CryptKeeper is a translational burden score for each ORF in the input sequence. This score reflects, in relative terms, how much of a cell’s capacity for translation is expected to be redirected to this ORF, as this has been shown to be the major cause of burden for many constructs (10,12,13). The translational burden score is currently calculated simply as the translation initiation rate multiplied by the ORF length. More detailed models could account for how rare codons that slow translation or mRNA structures that act as pause sites exacerbate this burden by leading to more ribosomes than expected from the simple model becoming sequestered on certain mRNAs. This effect has been experimentally demonstrated by comparing constructs with rare codons early versus late in a reading frame (12). Incorporating these refinements into CryptKeeper’s score could be especially important for evaluating burden from cloning eukaryotic sequences with very different codon usage into E. coli plasmids, as is the case when constructing virus infectious clones.

CryptKeeper is most useful as a tool for negative design. In this paradigm, one takes care to avoid issues that could arise from off-target interactions when engineering a system. Other examples of negative design in synthetic biology include adding genetic insulators between modules (55), avoiding crosstalk between metabolic pathways (56), and editing DNA sequences to remove mutational hotspots (57,58). In the context of negative design, it is fine to overpredict problems, as long as alternatives without these potential issues exist in the space of possible designs. Biological sequences have so many degrees of freedom that this is often the case. For example, one can eliminate potential ribosome binding sites by making synonymous changes in one or a few codons or eliminate a start codon with single amino acid substitution. Researchers can—and we argue, should—take precautionary steps to avoid off-target translation and unintentional translational burden even if it is unclear whether any RNA containing an ORF will be transcribed. Future versions of CryptKeeper could automate the redesign of sequences for this objective. Natural gene sequences have experienced selection against off-target gene expression elements in their original biological contexts (59,60). CryptKeeper can help researchers follow suit and apply negative design to sequences they create de novo or transplant into new contexts to improve the reliability and reproducibility of synthetic biology.

Acknowledgements

We thank Sean Leonard and Peng Geng for helpful discussions and software testing. We also thank Davy Wybiral for helping to resolve a Python Package Index naming conflict.

Funding

This work was supported by the Defense Advanced Research Projects Agency (HR0011-17-2-0052), National Institutes of Health (R01GM088344), Army Research Office (W911NF-20-1-0195), and National Science Foundation (IOS-2103208 and MCB-2123996).

Data Availability Statement

Data used for testing CryptKeeper is available at https://github.com/barricklab/CryptKeeper. The current version of the repository has been archived on Zenodo (DOI:10.5281/zenodo.13308762).

Fig. 1. (A) CryptKeeper overview. Predictions of gene expression signals in an input DNA sequence are integrated into an interactive plot and used to calculate an overall translational burden score. (B) pSB1C3-K3174006, an example of a plasmid engineered to express a transgene. It contains a pUC origin of replication, chloramphenicol resistance cassette, and BioBrick K3174006, which expresses mTagBFP (K592100) under control of a medium-strength constitutive promoter (J23110) and a medium-strength ribosome binding site (B0032) (10). Figure adapted from pLannotate output (41). (C) CryptKeeper output for pSB1C3-K3174006. The outermost tracks display protein-coding sequences on the forward and reverse strands as stacked boxes with heights proportional to their predicted translation initiation rates and colors and areas proportional to their individual translational burden scores. The next inner two tracks display predictions of RNA expression signals on each strand: promoters (green), Rho-dependent terminators (red), and intrinsic terminators (purple). The central track displays annotations from the input GenBank file (matching panel B).

Fig. 2. Case studies of virus infectious clone redesign. (A) Predicted strengths of promoters in a Zika virus infectious clone plasmid before and after redesign, compared to the promoter in BioBrick K3174006. (B) CryptKeeper burden plots for the Dengue virus sequence in an infectious clone plasmid before and after adding an intron that disrupts the ORF that makes the largest contribution to burden. (C) Predicted translational burden of a Dengue virus infectious clone before and after redesign compared to the predicted burden of the mTagBFP ORF from BioBrick K3174006.

Figure 3. Case studies of eliminating cryptic translation from eukaryotic transgenes. (A) CryptKeeper translation predictions for a cloned mouse mdr1a cDNA sequence before and after mutating the start codon associated with a cryptic ORF. (B) CryptKeeper burden predictions for the mTagBFP ORF from BioBrick K3174006, the cryptic ORF of mdr1a before and after mutating its start codon, and the cryptic ORF of SCN1A before and after introducing an engineered intron. (C) CryptKeeper translation predictions for a human SCN1A cDNA sequence before and after redesigning it to include an engineered intron. (D) Predicted total burden of plasmid pSB1C-K3174006, the complete mdr1a cDNA before and after mutating the cryptic ORF start codon, and the full plasmid encoding SCN1A before and after introducing the engineered intron.

Figure 4. Case studies of redesign to eliminate truncated protein expression. (A) CryptKeeper translation predictions for a GCN5 expression cassette for producing yeast SAGA histone acetyltransferase. Predictions are shown before and after mutating an internal start codon. (B) CryptKeeper translation predictions for a construct consisting of a fragment of the codon-optimized antibody NISTmAB ORF (nucleotides 555 to 732) placed upstream of a sfGFP reporter. Predictions are shown before and after mutating an internal start codon in the NISTmAB ORF.

Table 1. Test Datasets

Cloned Sequence	Complication	Solution	Citation	
Zika virus	Plasmid Instability	Eliminate two promoters	Chen et al. 2018	
Dengue virus	Plasmid Instability	Insert artificial intron	Holliday et al. 2023	
Mouse mdr1a cDNA	Plasmid Instability	Eliminate ribosome binding site	Pluchino et al. 2015	
Human SCN1A cDNA	Plasmid Instability	Eliminate promoter and insert artificial intron	DeKeyser et al. 2021	
Yeast GCN5	Truncated Protein	Eliminate internal ribosome binding site and/or start codon	Jennings et al. 2016	
Human NISTmAb	Truncated Protein	Eliminate internal start codon	Leith et al. 2019	

Material Availability Statement

CryptKeeper is open-source software released under a GPL-3.0 license. Source code, instructions, and example data are available at https://github.com/barricklab/CryptKeeper. The current version of the repository has been archived on Zenodo (DOI:10.5281/zenodo.13308762). Additionally, CryptKeeper can be installed as Bioconda package.

Conflict of Interest Disclosure

The authors declare no conflicts of interest.
==== Refs
References

1. Chen Y. , Liu T. , Zhang Z. , Chen M. , Rong L. , Ma L. , Yu B. , Wu D. , Zhang P. , Zhu X. , Huang X. , Zhang H. , & Li Y.-P . (2018). Novel genetically stable infectious clone for a Zika virus clinical isolate and identification of RNA elements essential for virus production. Virus Research, 257 , 14–24. 10.1016/j.virusres.2018.08.016 30144463
2. DeKeyser J.-M. , Thompson C. H. , & George A. L . (2021). Cryptic prokaryotic promoters explain instability of recombinant neuronal sodium channels in bacteria. Journal of Biological Chemistry, 296 , 100298. 10.1016/j.jbc.2021.100298 33460646
3. Holliday M. , Corliss L. , & Lennemann N. J . (2023). Construction and rescue of a DNA-launched DENV2 infectious clone. Viruses, 15 (2 ), Article 2. 10.3390/v15020275
4. Pluchino K. M. , Esposito D. , Moen J. K. , Hall M. D. , Madigan J. P. , Shukla S. , Procter L. V. , Wall V. E. , Schneider T. D. , Pringle I. , Ambudkar S. V. , Gill D. R. , Hyde S. C. , & Gottesman M. M . (2015). Identification of a cryptic bacterial promoter in mouse (mdr1a) Pglycoprotein cDNA. PLOS ONE, 10 (8 ), e0136396. 10.1371/journal.pone.0136396 26309032
5. Espah Borujeni A. , Zhang J. , Doosthosseini H. , Nielsen A. A. K. , & Voigt C. A . (2020). Genetic circuit characterization by inferring RNA polymerase movement and ribosome usage. Nature Communications, 11 (1 ), 5001. 10.1038/s41467-020-18630-2
6. Reddy P. T. , Brinson R. G. , Hoopes J. T. , McClung C. , Ke N. , Kashi L. , Berkmen M. , & Kelman Z . (2018). Platform development for expression and purification of stable isotope labeled monoclonal antibodies in Escherichia coli. mAbs, 10 (7 ), 992–1002. 10.1080/19420862.2018.1496879 30060704
7. Jennings M. J. , Barrios A. F. , & Tan S . (2016). Elimination of truncated recombinant protein expressed in Escherichia coli by removing cryptic translation initiation site. Protein Expression and Purification, 121 , 17–21. 10.1016/j.pep.2015.12.001 26739786
8. Leith E. M. , O’Dell W. B. , Ke N. , McClung C. , Berkmen M. , Bergonzo C. , Brinson R. G. , & Kelman Z . (2019). Characterization of the internal translation initiation region in monoclonal antibodies expressed in Escherichia coli. Journal of Biological Chemistry, 294 (48 ), 18046–18056. 10.1074/jbc.RA119.011008 31604819
9. Umenhoffer K. , Fehér T. , Balikó G. , Ayaydin F. , Pósfai J. , Blattner F. R. , & Pósfai G . (2010). Reduced evolvability of Escherichia coli MDS42, an IS-less cellular chassis for molecular and synthetic biology applications. Microbial Cell Factories, 9 (1 ), 38. 10.1186/1475-2859-9-38 20492662
10. Radde N. , Mortensen G. A. , Bhat D. , Shah S. , Clements J. J. , Leonard S. P. , McGuffie M. J. , Mishler D. M. , & Barrick J. E . (2024). Measuring the burden of hundreds of BioBricks defines an evolutionary limit on constructability in synthetic biology. Nature Communications, 15 (1 ), 6242. 10.1038/s41467-024-50639-9
11. Rugbjerg P. , Myling-Petersen N. , Porse A. , Sarup-Lytzen K. , & Sommer M. O. A . (2018). Diverse genetic error modes constrain large-scale bio-based production. Nature Communications, 9 (1 ), 787. 10.1038/s41467-018-03232-w
12. Ceroni F. , Algar R. , Stan G.-B. , & Ellis T . (2015). Quantifying cellular capacity identifies gene expression designs with reduced burden. Nature Methods, 12 (5 ), 415–418. 10.1038/nmeth.3339 25849635
13. Scott M. , Gunderson C. W. , Mateescu E. M. , Zhang Z. , & Hwa T . (2010). Interdependence of cell growth and gene expression: origins and consequences. Science, 330 (6007 ), 1099–1102. 10.1126/science.1192588 21097934
14. Kafri M. , Metzl-Raz E. , Jona G. , & Barkai N . (2016). The cost of protein production. Cell Reports, 14 (1 ), 22–31. 10.1016/j.celrep.2015.12.015 26725116
15. Segall-Shapiro T. H. , Meyer A. J. , Ellington A. D. , Sontag E. D. , & Voigt C. A . (2014). A ‘resource allocator’ for transcription based on a highly fragmented T7 RNA polymerase. Molecular Systems Biology, 10 (7 ), 742. 10.15252/msb.20145299 25080493
16. Richardson J. S. , & Richardson D. C . (2002). Natural β-sheet proteins use negative design to avoid edge-to-edge aggregation. Proceedings of the National Academy of Sciences, 99 (5 ), 2754–2759. 10.1073/pnas.052706099
17. Reis A. C. , & Salis H. M . (2020). An automated model test system for systematic development and improvement of gene expression models. ACS Synthetic Biology, 9 (11 ), 3145–3156. 10.1021/acssynbio.0c00394 33054181
18. Salis H. M. , Mirsky E. A. , & Voigt C. A . (2009). Automated design of synthetic ribosome binding sites to control protein expression. Nature Biotechnology, 27 (10 ), 946–950. 10.1038/nbt.1568
19. Seo S. W. , Yang J.-S. , Kim I. , Yang J. , Min B. E. , Kim S. , & Jung G. Y . (2013). Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metabolic Engineering, 15 , 67–74. 10.1016/j.ymben.2012.10.006 23164579
20. Deana A. , & Belasco J. G . (2005). Lost in translation: the influence of ribosomes on bacterial mRNA decay. Genes & Development, 19 (21 ), 2526–2533. 10.1101/gad.1348805 16264189
21. Kim S. , Wang Y.-H. , Hassan A. , & Kim S . (2024). Re-defining how mRNA degradation is coordinated with transcription and translation in bacteria. bioRxiv, 2024.04.18.588412. 10.1101/2024.04.18.588412
22. de Avila e Silva S. , Echeverrigaray S. , & Gerhardt G. J. L . (2011). BacPP: Bacterial promoter prediction—A tool for accurate sigma-factor specific assignment in enterobacteria. Journal of Theoretical Biology, 287 , 92–99. 10.1016/j.jtbi.2011.07.017 21827769
23. Di Salvo M. , Puccio S. , Peano C. , Lacour S. , & Alifano P . (2019). RhoTermPredict: An algorithm for predicting Rho-dependent transcription terminators based on Escherichia coli, Bacillus subtilis and Salmonella enterica databases. BMC Bioinformatics, 20 (1 ), 117. 10.1186/s12859-019-2704-x 30845912
24. Feng C.-Q. , Zhang Z.-Y. , Zhu X.-J. , Lin Y. , Chen W. , Tang H. , & Lin H . (2019). iTerm-PseKNC: A sequence-based tool for predicting bacterial transcriptional terminators. Bioinformatics, 35 (9 ), 1469–1477. 10.1093/bioinformatics/bty827 30247625
25. Gardner P. P. , Barquist L. , Bateman A. , Nawrocki E. P. , & Weinberg Z . (2011). RNIE: Genome-wide prediction of bacterial intrinsic terminators. Nucleic Acids Research, 39 (14 ), 5845–5852. 10.1093/nar/gkr168 21478170
26. Huang Y.-K. , Yu C.-H. , & Ng I.-S . (2024). Precise strength prediction of endogenous promoters from Escherichia coli and J-series promoters by artificial intelligence. Journal of the Taiwan Institute of Chemical Engineers, 160 , 105211. 10.1016/j.jtice.2023.105211
27. Jin Y. , Ma H. , Xu Z. Z. , & Lu Z. J . (2023). BATTER: Accurate prediction of rho-dependent and rho-independent transcription terminators in metagenomes bioRxiv, 2023.10.02.560326. 10.1101/2023.10.02.560326
28. Kingsford C. L. , Ayanbule K. , & Salzberg S. L . (2007). Rapid, accurate, computational discovery of Rho-independent transcription terminators illuminates their relationship to DNA uptake. Genome Biology, 8 (2 ), R22. 10.1186/gb-2007-8-2-r22 17313685
29. LaFleur T. L. , Hossain A. , & Salis H. M . (2022). Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria. Nature Communications, 13 (1 ), 5159. 10.1038/s41467-022-32829-5
30. Lai H.-Y. , Zhang Z.-Y. , Su Z.-D. , Su W. , Ding H. , Chen W. , & Lin H . (2019). iProEP: A computational predictor for predicting promoter. Molecular Therapy - Nucleic Acids, 17 , 337–346. 10.1016/j.omtn.2019.05.028 31299595
31. Lesnik E. A. , Sampath R. , Levene H. B. , Henderson T. J. , McNeil J. A. , & Ecker D. J . (2001). Prediction of rho-independent transcriptional terminators in Escherichia coli. Nucleic Acids Research, 29 (17 ), 3583–3594. 10.1093/nar/29.17.3583 11522828
32. Lin H. , Deng E.-Z. , Ding H. , Chen W. , & Chou K.-C . (2014). iPro54-PseKNC: A sequence-based predictor for identifying sigma-54 promoters in prokaryote with pseudo k-tuple nucleotide composition. Nucleic Acids Research, 42 (21 ), 12961–12972. 10.1093/nar/gku1019 25361964
33. Liu B. , & Li K . (2019). iPromoter-2L2.0: Identifying promoters and their types by combining smoothing cutting window algorithm and sequence-based features. Molecular Therapy - Nucleic Acids, 18 , 80–87. 10.1016/j.omtn.2019.08.008 31536883
34. Nadiras C. , Eveno E. , Schwartz A. , Figueroa-Bossi N. , & Boudvillain M . (2018). A multivariate prediction model for Rho-dependent termination of transcription. Nucleic Acids Research, 46 (16 ), 8245–8260. 10.1093/nar/gky563 29931073
35. Naville M. , Ghuillot-Gaudeffroy A. , Marchais A. , & Gautheret D . (2011). ARNold: A web tool for the prediction of Rho-independent transcription terminators. RNA Biology, 8 (1 ), 11–13. 10.4161/rna.8.1.13346 21282983
36. Solovyev V. & Salamov A . (2011). Automatic annotation of microbial genomes and metagenomic sequences. In Metagenomics and its Applications in Agriculture, Biomedicine and Environmental Studies (Ed. Li R.W. ), Nova Science Publishers, p.61–78. (pp. 61–78).
37. Xiao X. , Hu Z. , Luo Z. , & Xu Z . (2023). iPSI(2L)-EDL: A two-layer predictor for identifying promoters and their types based on ensemble deep learning. Current Bioinformatics, 19 (4 ), 327–340. 10.2174/0115748936264316230926073231
38. Zhai W. , Duan Y. , Zhang X. , Xu G. , Li H. , Shi J. , Xu Z. , & Zhang X . (2022). Sequence and thermodynamic characteristics of terminators revealed by FlowSeq and the discrimination of terminators strength. Synthetic and Systems Biotechnology, 7 (4 ), 1046–1055. 10.1016/j.synbio.2022.06.003 35845313
39. Zhang H. , Li J. , Hu F. , Lin H. , & Ma J . (2024). AMter: An end-to-end model for transcriptional terminators prediction by extracting semantic feature automatically based on attention mechanism. Concurrency and Computation: Practice and Experience, 36 (13 ), e8056. 10.1002/cpe.8056
40. Zhang M. , Jia C. , Li F. , Li C. , Zhu Y. , Akutsu T. , Webb G. I. , Zou Q. , Coin L. J. M. , & Song J . (2022). Critical assessment of computational tools for prokaryotic and eukaryotic promoter prediction. Briefings in Bioinformatics, 23 (2 ), bbab551. 10.1093/bib/bbab551 35021193
41. McGuffie M. J. , & Barrick J. E . (2021). pLannotate: Engineered plasmid annotation. Nucleic Acids Research, 49 (W1 ), W516–W522. 10.1093/nar/gkab374 34019636
42. Chen Y.-J. , Liu P. , Nielsen A. A. K. , Brophy J. A. N. , Clancy K. , Peterson T. , & Voigt C. A . (2013). Characterization of 582 natural and synthetic terminators and quantification of their design constraints. Nature Methods, 10 (7 ), 659–664. 10.1038/nmeth.2515 23727987
43. Tarnowski M. J. , & Gorochowski T. E . (2022). Massively parallel characterization of engineered transcript isoforms using direct RNA sequencing. Nature Communications, 13 (1 ), 434. 10.1038/s41467-022-28074-5
44. Bokeh Development Team. (2018). Bokeh: Python library for interactive visualization (Version 3.4.1) [Computer software]. http://www.bokeh.pydata.org
45. Sayers E. W. , Bolton E. E. , Brister J. R. , Canese K. , Chan J. , Comeau D. C. , Connor R. , Funk K. , Kelly C. , Kim S. , Madej T. , Marchler-Bauer A. , Lanczycki C. , Lathrop S. , Lu Z. , Thibaud-Nissen F. , Murphy T. , Phan L. , Skripchenko Y. , … Sherry S. T . (2021). Database resources of the National Center for Biotechnology Information. Nucleic Acids Research, 50 (D1 ), D20–D26. 10.1093/nar/gkab1112
46. Li D. , Aaskov J. , & Lott W. B . (2011). Identification of a cryptic prokaryotic promoter within the cDNA encoding the 5′ end of dengue virus RNA genome. PLOS ONE, 6 (3 ), e18197. 10.1371/journal.pone.0018197 21483867
47. Pu S.-Y. , Wu R.-H. , Tsai M.-H. , Yang C.-C. , Chang C.-M. , & Yueh A . (2014). A novel approach to propagate flavivirus infectious cDNA clones in bacteria by introducing tandem repeat sequences upstream of virus genome. Journal of General Virology, 95 (7 ), 1493–1503. 10.1099/vir.0.064915-0 24728712
48. Usme-Ciro J. A. , Lopera J. A. , Enjuanes L. , Almazán F. , & Gallego-Gomez J. C . (2014). Development of a novel DNA-launched dengue virus type 2 infectious clone assembled in a bacterial artificial chromosome. Virus Research, 180 , 12–22. 10.1016/j.virusres.2013.12.001 24342140
49. Whitaker W. R. , Lee H. , Arkin A. P. , & Dueber J. E . (2015). Avoidance of truncated proteins from unintended ribosome binding sites within heterologous protein coding sequences. ACS Synthetic Biology, 4 (3 ), 249–257. 10.1021/sb500003x 24931615
50. Stephenson A. , Lastra L. , Nguyen B. , Chen Y.-J. , Nivala J. , Ceze L. , & Strauss K . (2023). Physical laboratory automation in synthetic biology. ACS Synthetic Biology, 12 (11 ), 3156–3169. 10.1021/acssynbio.3c00345 37935025
51. Zhang S. , Fan R. , Liu Y. , Chen S. , Liu Q. , & Zeng W . (2023). Applications of transformer-based language models in bioinformatics: A survey. Bioinformatics Advances, 3 (1 ), vbad001. 10.1093/bioadv/vbad001 36845200
52. Weinstock M. T. , Hesek E. D. , Wilson C. M. , & Gibson D. G . (2016). Vibrio natriegens as a fast-growing host for molecular biology. Nature Methods, 13 (10 ), 849–851. 10.1038/nmeth.3970 27571549
53. Martínez-García E. , & de Lorenzo V . (2024). Pseudomonas putida as a synthetic biology chassis and a metabolic engineering platform. Current Opinion in Biotechnology, 85 , 103025. 10.1016/j.copbio.2023.103025 38061264
54. Wei W. , Hennig B. P. , Wang J. , Zhang Y. , Piazza I. , Sanchez Y. P. , Chabbert C. D. , Adjalley S. H. , Steinmetz L. M. , & Pelechano V . (2019). Chromatin-sensitive cryptic promoters putatively drive expression of alternative protein isoforms in yeast. Genome Research, 29 (12 ), 1974–1984. 10.1101/gr.243378.118 31740578
55. Lou C. , Stanton B. , Chen Y.-J. , Munsky B. , & Voigt C. A . (2012). Ribozyme-based insulator parts buffer synthetic circuits from genetic context. Nature Biotechnology, 30 (11 ), 1137–1142. 10.1038/nbt.2401
56. Agapakis C. M. , Ducat D. C. , Boyle P. M. , Wintermute E. H. , Way J. C. , & Silver P. A . (2010). Insulation of a synthetic hydrogen metabolism circuit in bacteria. Journal of Biological Engineering, 4 (1 ), 3. 10.1186/1754-1611-4-3 20184755
57. Jack B. R. , Leonard S. P. , Mishler D. M. , Renda B. A. , Leon D. , Suárez G. A. , & Barrick J. E . (2015). Predicting the genetic stability of engineered DNA sequences with the EFM Calculator. ACS Synthetic Biology, 4 (8 ), 939–943. 10.1021/acssynbio.5b00068 26096262
58. Menuhin-Gruman I. , Arbel M. , Amitay N. , Sionov K. , Naki D. , Katzir I. , Edgar O. , Bergman S. , & Tuller T . (2022). Evolutionary Stability Optimizer (ESO): A novel approach to identify and avoid mutational hotspots in DNA sequences while maintaining high expression levels. ACS Synthetic Biology, 11 (3 ), 1142–1151. 10.1021/acssynbio.1c00426 34928133
59. Itzkovitz S. , Hodis E. , & Segal E . (2010). Overlapping codes within protein-coding sequences. Genome Research, 20 (11 ), 1582–1589. 10.1101/gr.105072.110 20841429
60. Yang C. , Hockenberry A. J. , Jewett M. C. , & Amaral L. A. N . (2016). Depletion of Shine-Dalgarno sequences within bacterial coding regions is expression dependent. G3 (Bethesda, Md.), 6 (11 ), 3467–3474. 10.1534/g3.116.032227 27605518
