
==== Front
Brief Bioinform
Brief Bioinform
bib
Briefings in Bioinformatics
1467-5463
1477-4054
Oxford University Press

10.1093/bib/bbae229
bbae229
Problem Solving Protocol
AcademicSubjects/SCI01060
AnnoView enables large-scale analysis, comparison, and visualization of microbial gene neighborhoods
Wei Xin Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Tan Huagang Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Lobb Briallen Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Zhen William Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Wu Zijing Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Parks Donovan H Australian Centre for Ecogenomics, School of Chemistry and Molecular Biosciences, University of Queensland, St Lucia, QLD 4072, Brisbane, Australia

Neufeld Josh D Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Moreno-Hagelsieb Gabriel Department of Biology, Wilfrid Laurier University, 75 University Avenue West, Waterloo, ON, Canada

Doxey Andrew C Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada

Corresponding author. Department of Biology and Waterloo Centre for Microbial Research, University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada. E-mail: acdoxey@uwaterloo.ca
Xin Wei and Huagang Tan co-first authors.

5 2024
14 5 2024
14 5 2024
25 3 bbae22922 1 2024
02 4 2024
26 4 2024
© The Author(s) 2024. Published by Oxford University Press.
2024
https://creativecommons.org/licenses/by/4.0/ This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.

Abstract

The analysis and comparison of gene neighborhoods is a powerful approach for exploring microbial genome structure, function, and evolution. Although numerous tools exist for genome visualization and comparison, genome exploration across large genomic databases or user-generated datasets remains a challenge. Here, we introduce AnnoView, a web server designed for interactive exploration of gene neighborhoods across the bacterial and archaeal tree of life. Our server offers users the ability to identify, compare, and visualize gene neighborhoods of interest from 30 238 bacterial genomes and 1672 archaeal genomes, through integration with the comprehensive Genome Taxonomy Database and AnnoTree databases. Identified gene neighborhoods can be visualized using pre-computed functional annotations from different sources such as KEGG, Pfam and TIGRFAM, or clustered based on similarity. Alternatively, users can upload and explore their own custom genomic datasets in GBK, GFF or CSV format, or use AnnoView as a genome browser for relatively small genomes (e.g. viruses and plasmids). Ultimately, we anticipate that AnnoView will catalyze biological discovery by enabling user-friendly search, comparison, and visualization of genomic data. AnnoView is available at http://annoview.uwaterloo.ca

bioinformatics
microbial genomics
gene neighborhoods
genomic context
genome visualization
functional annotation
Natural Sciences and Engineering Research Council of Canada 10.13039/501100000038 RGPIN-2019-04266
==== Body
pmcBACKGROUND

Gene neighborhoods are regions of genomes containing sets of co-localized genes. For bacteria and archaea, the organization of genes into clusters or operons can provide insight into gene function, regulation and evolutionary relationships [1–3]. The identification of specialized gene clusters within genomes can help identify diverse physiological and functional traits including biosynthetic gene clusters [4, 5], transport and uptake systems [6] and antibiotic resistance and virulence traits [7–10]. Comparison of orthologous gene neighborhoods can reveal functional diversification [7, 11, 12], identify cases of horizontal transfer and genomic rearrangements [13–15], predict functional associations between genes [2, 16, 17] and infer operons and pathways [18, 19].

There are multiple strategies for analyzing sets of neighboring genes and their patterns of co-localization across genomes. Tools such as Artemis/Artemis Comparison Tool [20], IGV [21], Mauve [22] or Jbrowse [23] can be used for pairwise or small-scale gene neighborhood comparison involving one or a few species. These tools are valuable for visualizing and comparing genomic regions of interest and their associated functional annotations. Synteny mapping tools can also aid in the comparison of gene neighborhoods [24–26]. Additionally, public databases and comparative genomics platforms can provide gene neighborhood visualization features, which include the NCBI [27], UCSC [28], STRING [29] and Ensembl databases [30]. For more advanced gene neighborhood analysis, including comparative analysis across larger sets of genomes, several tools and portals have been developed [27, 31–41] (see Supplementary Table 1 for a comparison, see Supplementary Data available online at http://bib.oxfordjournals.org/). These can provide a more detailed and focused view on gene neighborhoods than genome browsers, facilitate comparison across species in pre-computed databases, and exploration of predicted gene functions from different annotation methods.

Despite existing tools for gene neighborhood comparison, there is still a need for a platform that: (i) is intuitive and easy to use; (ii) contains a wide diversity of genomes across the full bacterial and archaeal tree of life; (iii) is taxonomically and functional annotated in a comprehensive and consistent manner; (iv) is not restricted to a single reference database and allows users to upload their own custom datasets; (v) is capable of comparing relatively large numbers of genomic regions simultaneously; (vi) is capable of exploring ‘unannotated’ genes and proteins of unknown function [42]. In order to address the need for a platform with such functionality, here we introduce AnnoView, a versatile and easy-to-use web platform for microbial gene neighborhood exploration. AnnoView enables large-scale gene neighborhood search and visualization across 30 238 bacterial genomes and 1672 archaeal genomes. Through integration with AnnoTree [43] and the Genome Taxonomy Database (GTDB) [44], AnnoView allows users to identify and explore gene neighborhoods of interest across large sets of genomes, and cluster and visualize them based on gene function composition, with a choice of different annotation methods including Kyoto Encyclopedia of Genes and Genomes (KEGG) [45, 46], Pfam and TIGRFAM [47]. Additionally, AnnoView allows users to upload custom genome datasets, such as sets of gene neighborhoods from different species, metagenomic contigs, small bacterial genomes, viral genomes and plasmids. Users can easily share their AnnoView sessions using session URL links or download their data for offline editing.

METHODS

Construction of the AnnoView database

The AnnoView database was constructed using information from the AnnoTree R95 database and GTDB (R95) [43, 44]. Gene and functional annotations are derived from the gene prediction and annotation pipeline described previously [43]. In brief, Prodigal [48] with default parameters was used for gene calling from whole genomic FASTA files, resulting in GBK/GFF files and predicted protein-coding sequences for all organisms. Both Pfam and TIGRFAM annotation were performed on all protein sequences using HMMER [49] as well as pfam_scan.pl (ftp://ftp.ebi.ac.uk/pub/databases/Pfam/Tools/) for Pfam domains specifically. For KEGG annotation, DIAMOND [50] was used to identify homologous sequence clusters within the UniRef100 database [51] and existing annotations were transferred. DIAMOND was run with the following parameters: E-value ≤1e-5, identity ≥30%, and query-to-subject and subject-to-query alignment coverage ≥70%. This pipeline was used to functionally annotate the genes from 30 238 bacterial genomes and 1672 archaeal genomes.

For the AnnoView database, short names of and descriptions of functional annotations were retrieved from the above KEGG, Pfam and TIGRFAM annotation data. Gene coordinates and orientations were extracted from the Prodigal-generated GFF files. An AnnoView SQL database was constructed to store predicted protein sequences, gene functional annotation data (i.e. KEGG, Pfam and TIGRFAM), gene coordinates and strand information (i.e., + or -), CDS length, as well as GTDB taxonomy information. All protein sequences in the AnnoView database were then used to build a DIAMOND (2.1.8) database [50] that can be searched to find similar proteins to a query. Once DIAMOND searches are performed by the user, the backend server will run MySQL queries on all sequence hits to identify proteins within a ± 10 kb window size. Protein-coding genes that span the edges of this window are excluded. This process is highly efficient due to the presence of pre-computed data and indexes in the MySQL database, allowing for quick execution of MySQL queries.

Web server implementation

The AnnoView web server adopts a client–server website design (Supplementary Figure 1, see Supplementary Data available online at http://bib.oxfordjournals.org/). For the frontend client, we used react.js to build the user interface (https://reactjs.org/) and d3.js (https://d3js.org) for visualizing gene neighborhoods. The Python-based Flask web framework (https://flask.palletsprojects.com) was used for backend server development. Pre-computed gene neighborhood data are stored in a MySQL (https://www.mysql.com/) database. Gene neighborhood data resulting from user queries and user-uploaded data are stored in a MongoDB (https://www.mongodb.com/) database. The pseudocode for the algorithm that clusters gene neighborhoods based on a gene of interest (Supplementary Figure 2, see Supplementary Data available online at http://bib.oxfordjournals.org/) is a modified version of the Smith–Waterman algorithm for local sequence alignment [52] applied to vectors of gene lists where each gene is labeled by its annotation.

Upload workflow for NCBI data retrieval and gene annotations

In addition to retrieval of gene neighborhood information from the AnnoTree [43] and GTDB [44] database, a bioinformatic workflow was outlined for users to retrieve gene neighborhood information directly from the NCBI database, which includes additional genomic information not available in the GTDB. The workflow involves: (i) retrieval of gene neighborhood datasets in GBK format from NCBI, (ii) annotation of gene neighborhoods with taxonomic information, (iii) KEGG and Pfam functional annotation of gene neighborhoods, (iv) sorting of gene neighborhoods by selection of a gene/protein of interest as the ‘center gene’ and (v) visualization of gene neighborhoods in AnnoView. The detailed NCBI-based workflow and associated code is accessible at the following GitHub repository: https://github.com/satellite830/AnnoView

RESULTS

The AnnoView platform

AnnoView is a gene neighborhood browser for microbial genomes available at http://annoview.uwaterloo.ca/. There are two modules in AnnoView: (i) the ‘upload’ module allows users to upload genomic neighborhoods in GBK, GFF and CSV format and (ii) the ‘GTDB search’ module allows users to query the GTDB/AnnoView genome database for matches to an input protein sequence. After a protein query is searched against the AnnoView database, the server returns genomic neighborhoods surrounding identified matches (genes corresponding to protein homologs). Pre-computed functional annotations in KEGG, Pfam and TIGRFAM are then available to the user for further exploration and visualization. The protein homology search criteria include E-value, coverage cut-off and a choice of database to search (bacteria/archaea).

After the homology search, the user can further select a subset of protein hits for genomic neighborhood display by taxonomy or by E-value on an ‘intermediate’ results page. This intermediary page solves potential issues of overwhelming number of sequence matches, which need further filtering. The ‘Select by taxonomy’ option is intended for a user interested in exploring genomic neighborhoods in a specific lineage, whereas ‘Select by E-value’ is intended for users to examine only top-scoring genomic neighborhoods based on sequence similarity to a query protein.

The interactive browser-component of AnnoView (Figure 1) allows users to explore multiple gene neighborhoods (up to 500) in a single window. A toolbar panel at the top of the browser allows users to perform different functions. These include a zoom-in and zoom-out function, and a download button for retrieving gene neighborhood information in CSV format. Downloaded CSV data can be customized by the user to edit existing gene neighborhood information or add taxonomic and functional annotations. The gene neighborhood plot can also be exported in vector-quality SVG format suitable for publication or offline editing by image-editing software (e.g. Adobe Illustrator, Inkscape and Affinity Designer).

Figure 1 The AnnoView gene neighborhood browser and its key features. (1) Stretch; (2) Compress; (3) Zoom in; (4) Zoom out; (5) Reset the page; (6) Recenter image; (7) Download the gene neighborhood in CSV format; (8) Export image in SVG format; (9) Search box; (10) Display taxonomy labels; (11) Color genes by metadata; the default is colored by KEGG annotations; (12) Taxonomy labels; (13) Delete selected tracks; (14) Gene neighborhood; (15) Gene information when hovering on a gene; (16) More options when right-clicking on a gene; (17) Legend.

To navigate to a specific gene, users can search for gene names in the search box. The search box is also useful for navigating genes that share a general term such as ‘flagellar’ to highlight multiple matches (e.g. ‘flagellar M-ring protein FliF’ and ‘flagellar motor switch protein FliG’). Taxonomy labels allow users to choose the taxonomy level to display in the visualization panel. Users can change the displayed metadata to color gene neighborhoods based on different functional annotations.

The visualization panel is an interactive interface that allows exploration of basic gene information, clustering of gene neighborhoods, editing of the plot and data retrievals, such as protein sequences and gene metadata (Figure 1). For each gene neighborhood displayed in the visualization panel, its taxonomy information is displayed on the left side of the panel. The taxonomic level can be changed under the taxonomy labels in the toolbar panel. To view gene information, the user can either hover over a gene or right-click on a gene for more information (e.g. protein sequence and annotation details). More options are available when right-clicking on a gene for the user to edit the gene neighborhood plot in AnnoView. After choosing a center gene, gene neighborhoods can be clustered using the clustering algorithm implemented in AnnoView, which is based on the similarity of neighboring gene content (see Methods, Supplementary Figure 2 and Supplementary Figure 3, see Supplementary Data available online at http://bib.oxfordjournals.org/). Furthermore, users can flip gene neighborhood orientations, change gene colors and add genes to the legend. The vertical order of gene neighborhoods displayed can also be changed based on a user’s preference. For example, the user may wish to manually re-order neighborhoods based on phylogenetic relationships. A unique token is generated for each online session and can be used for access for up to a month. Because of the unique web token, the URL for the user’s current session can also be shared with other users over the web to enable data sharing.

Comparison with existing tools: An advantage of AnnoView over previously developed tools for gene neighborhood visualization include an ability to handle many (500) tracks simultaneously to enable larger-scale visualization, a unique integration with the GTDB database, and an ability to not only explore pre-computed annotations but also perform custom sequence searches against the GTDB as well as custom data uploads. This flexibility handles multiple use cases that cannot be addressed by many other tools (Supplementary Table 1, see Supplementary Data available online at http://bib.oxfordjournals.org/). Additionally, AnnoView has been designed to be interactive: e.g. users can dynamically re-configure (re-arrange, re-align, re-color, etc.) gene neighborhood tracks for visualization and to generate publication-quality figures. These interactive capabilities further distinguish AnnoView from comparable tools.

EXAMPLE WORKFLOWS

Example 1: Using AnnoTree-DB search to explore the McrA operon

We envision that a common use case for AnnoView will be for users to search our pre-computed genome database for matches to a gene/protein of interest. For instance, using AnnoView, a user may explore gene neighborhoods surrounding methyl-coenzyme M reductase subunit alpha (McrA) in archaea (Figure 2). Methyl-coenzyme M reductase (MCR) is an enzyme that catalyzes the final step of methanogenesis in archaea [53]. Searching a representative McrA protein sequence from Methanosarcina barkeri (https://uniprot.org/uniprotkb/P07962/) in the AnnoView database returns 379 detected homologs. As mentioned in the previous section, there are two ways of filtering the database matches: either select hits by taxonomy or by the top N (where N is defined by the user) number of hits ranked by E-value. Here, we select McrA homologous sequences in Methanosarcina, a genus of archaea known for methanogenesis. The gene neighborhood visualization shows that mcrA is part of a highly conserved mcrBDCGA gene cluster, which forms the known operon structure of methanogenic MCRs [54]. The MCR operon consists of the essential genes mcrA, mcrB and mcrG, as well as two accessory genes mcrC and mcrD.

Figure 2 An example illustrating an AnnoTree database search using AnnoView. (a) AnnoTree search interface. (b) The intermediate page for user to select gene neighborhood by taxonomy or by E-value. (c) The genomic context of selected McrA homologs in the Methanosarcina genus.

Example 2: Visualizing a custom gene neighborhood dataset (the SLR4 gene locus)

Another envisioned use case for AnnoView is to upload a custom gene neighborhood dataset and explore it further using AnnoView’s gene neighborhood browser (Figure 3A and B). The gene neighborhood dataset can be generated from public genomic databases such as NCBI or user-generated genomic datasets in GBK or GFF format. To visualize the gene neighborhoods labeled by different functional annotations than those provided by GenBank (e.g. protein product name), such as Pfam or KEGG annotation, a bioinformatic workflow is provided at https://github.com/satellite830/AnnoView. This workflow outlines how users can download a CSV table from AnnoView, add customized annotations and taxonomy information to the table, and upload the data back to AnnoView for further exploration (See Methods for a detailed workflow).

Figure 3 AnnoView exploration of a custom uploaded gene neighborhood dataset. (a) Workflow of generating the gene neighborhood dataset. The gene neighborhood dataset in GBK format was downloaded from NCBI, then the neighboring genes were functionally annotated by Pfam. The gene neighborhood dataset along with its protein functional annotations was uploaded to AnnoView for visualization. (b) AnnoView upload interface. (c) The gene neighborhoods of Slr4 and its homologs. The species-specific insertion in C. mytili is surrounded with a box.

To demonstrate this workflow, we focus on an example from a previous publication [55], where AnnoView was used to visualize gene neighborhoods surrounding a newly identified protein called Slr4 from Pseudoalteromonas tunicata and its homologs in related species (Figure 3C). Gene neighborhoods were annotated and visualized based on Pfam annotations and all annotations occurring more than twice have been highlighted (Figure 3C). In this study, Slr4 was functionally characterized as a novel surface layer (S-layer) protein that forms a proteinaceous protective layer surrounding cells [55]. The genomic context provides additional supportive evidence of this function, as it reveals a conserved pattern of Type II secretion (T2SS) pathway genes immediately upstream of Slr4, which have been previously implicated in S-layer protein secretion (Figure 3C). Because dedicated T2SS systems are a common feature of S-layer genes [56, 57], it is likely that this conserved operon upstream of Slr4 plays a similar function and serves as a dedicated T2SS pathway for secreting this novel S-layer protein family.

In addition to visualizing conserved neighborhood features, another important function of AnnoView visualization is to reveal differences in genomic context. An example of this is a species-specific genomic insertion in the Slr4 neighborhood of Cognaticolwellia mytili (Figure 3C). The insertion includes two additional ORFs annotated as PF18765 (polymerase beta, nucleotidyltransferase) and PF01934 (ribonuclease HepT-like), which are two genes in the recently discovered type II toxin/antitoxin (HEPN/MNT) system [58]. The insertion of this toxin/antitoxin system into the T2SS gene cluster is consistent with the highly mobile nature of type II toxin/antitoxin systems [59].

Example 3: Visual comparison of Betacoronavirus genomes

While primarily intended as a tool to explore microbial gene neighborhoods around genes of interest, AnnoView is also capable of whole genome visualization and comparative analysis for relatively smaller sequences, such as viral genomes, plasmids, and in some cases, smaller bacterial genomes. An AnnoView visualization of several Betacoronavirus genomes reveals several interesting and biologically relevant differences in genome structure (Figure 4). The overall genome structure (gene order and composition) of different Betacoronavirus species is highly conserved. Important genes such as ORF1ab, N (nucleocapsid protein), S (spike protein), M (membrane protein) and E (envelope protein) are conserved in all Betacoronavirus genomes. However, HE (hemagglutinin esterase) is only present in the Embecovirus subgroup. The hemagglutinin esterase is an enzyme that facilitates attachment and destruction of sialic acid receptors present on the host cell surface, which aids in molecular recognition of specific target cells [60, 61]. The unique insertion of the HE gene into genomes within the Embecovirus lineage likely resulted from a horizontal transfer event originating from influenza C/D [62]. This proposed insertion event likely occurred in ancestral members of the Embecovirus subgenus before the divergence of other Betacoronavirus subgroups, resulting in the unique presence of the HE protein in the Embecovirus subgroup.

Figure 4 Visualization of Betacoronavirus genomes using AnnoView. Genome tracks are sorted by the genome phylogeny from NCBI Virus. The small box surrounds a unique split gene in the rabbit coronavirus HKU14 genome.

A second genomic difference revealed by AnnoView visualization is the presence of four ORFs between the ORF1ab and downstream HE gene within the rabbit coronavirus HKU14, in comparison to other genomes from the Embecovirus subgroup, which have zero or one ORF (NS2a) between ORF1ab and HE (Figure 4). Previous work has shown that this unique pattern in the rabbit coronavirus genome have resulted from the splitting of an ancestral NS2a gene into four smaller ORFs [63].

CONCLUSION

AnnoView is a versatile gene neighborhood visualization tool that enables rapid and user-friendly exploration of microbial genomes. The AnnoView database is integrated with AnnoTree, which provides comprehensive and consistent taxonomic classification for all genomes within the Genome Taxonomy Database. In addition, users can explore their own user-uploaded genomic datasets, including orthologous gene neighborhoods across sets of species or small size genomes. AnnoView has already enabled genomic analysis and visualization for several studies [7, 55, 64–66]. In the future, we hope to add more features to AnnoView, such as expanding the set of functional annotation methods available to users (e.g. including the InterPro suite of annotations [67]), updating AnnoView on a regular basis with the most up-to-date GTDB data, allowing users to customize their genome neighborhood criteria (e.g. window size), and displaying other genomic features, such as tRNAs, rRNAs, CRISPR loci and pseudogenes. We envision AnnoView being a useful tool for gene neighborhood exploration, visualization, and comparison of microbial genomes with applications for genetics and genomics research more broadly.

Key Points

Bioinformatic analysis of gene neighborhoods is a key aspect of microbial genomics

Despite a range of available tools, gene neighborhood comparison remains to be a challenge

We introduce AnnoView, a web platform for large-scale gene neighborhood exploration across the microbial tree of life

AnnoView allows researchers to explore genomes and pre-computed functional annotations from the Genome Taxonomy Database (GTDB) and AnnoTree database

Users can also upload their own custom genomic datasets for exploration, comparison and visualization

Supplementary Material

SuppFigures_BIB_bbae229

SupTable1_bbae229

FUNDING

Natural Sciences and Engineering Research Council of Canada (Grant / Award Number: RGPIN-2019-04266).

DATA AVAILABILITY

The data underlying this article is available in Zenodo, at https://dx.doi.org/10.5281/zenodo.3998874 for the bacterial database and https://dx.doi.org/10.5281/zenodo.4001160 for the archaeal database.
==== Refs
References

1. Huynen M , SnelB, LatheW, BorkP. Predicting protein function by genomic context: quantitative evaluation and qualitative inferences. Genome Res 2000;10 :1204–10.10958638
2. Korbel JO , JensenLJ, Von MeringC, BorkP. Analysis of genomic context: prediction of functional associations from conserved bidirectionally transcribed gene pairs. Nat Biotechnol 2004;22 :911–7.15229555
3. Galperin MY , KooninEV. Who’s your neighbor? New computational approaches for functional genomics. Nat Biotechnol 2000;18 :609–13.10835597
4. Cimermancic P , MedemaMH, ClaesenJ, et al. Insights into secondary metabolism from a global analysis of prokaryotic biosynthetic gene clusters. Cell 2014;158 :412–21.25036635
5. Kautsar SA , BlinK, ShawS, et al. MIBiG 2.0: a repository for biosynthetic gene clusters of known function. Nucleic Acids Res 2020;48 :D454–8.31612915
6. Crits-Christoph A , BhattacharyaN, OlmMR, et al. Transporter genes in biosynthetic gene clusters predict metabolite characteristics and siderophore activity. Genome Res 2021;31 :239–50.33361114
7. Wei X , Moreno-HagelsiebG, GlickBR, DoxeyAC. Comparative analysis of adenylate isopentenyl transferase genes in plant growth-promoting bacteria and plant pathogenic bacteria. Heliyon 2023;9 :e13955.36938451
8. Lobb B , DoxeyAC. Novel function discovery through sequence and structural data mining. Curr Opin Struct Biol 2016;38 :53–61.27289211
9. Doxey AC , MansfieldMJ, MontecuccoC. Discovery of novel bacterial toxins by genomics and computational biology. Toxicon 2018;147 :2–12.29438679
10. Li J , TaiC, DengZ, et al. VRprofile: gene-cluster-detection-based profiling of virulence and antibiotic resistance traits encoded within genome sequences of pathogenic bacteria. Brief Bioinform 2018;19 :566–74.28077405
11. Cascales E . The type VI secretion toolkit. EMBO Rep 2008;9 :735–41.18617888
12. Liu Z , CheemaJ, VigourouxM, et al. Formation and diversification of a paradigm biosynthetic gene cluster in plants. Nat Commun 2020;11 :5354.33097700
13. Zhang S , LebretonF, MansfieldMJ, et al. Identification of a botulinum neurotoxin-like toxin in a commensal strain of Enterococcus faecium. Cell Host Microbe 2018;23 :169–176.e6.29396040
14. Mansfield MJ , AdamsJB, DoxeyAC. Botulinum neurotoxin homologs in non-clostridium species. FEBS Lett 2015;589 :342–8.25541486
15. Mansfield MJ , Sugiman-MarangosSN, MelnykRA, DoxeyAC. Identification of a diphtheria toxin-like gene family beyond the Corynebacterium genus. FEBS Lett 2018;592 :2693–705.30058084
16. Overbeek R , FonsteinM, D’SouzaM, et al. The use of gene clusters to infer functional coupling. Proc Natl Acad Sci U S A 1999;96 :2896–901.10077608
17. Dandekar T , SnelB, HuynenM, BorkP. Conservation of gene order: a fingerprint of proteins that physically interact. Trends Biochem Sci 1998;23 :324–8.9787636
18. Zhao S , KumarR, SakaiA, et al. Discovery of new enzymes and metabolic pathways by using structure and genome context. Nature 2013;502 :698–702.24056934
19. Salgado H , Moreno-HagelsiebG, SmithTF, Collado-VidesJ. Operons in Escherichia coli: genomic analyses and predictions. Proc Natl Acad Sci U S A 2000;97 :6652–7.10823905
20. Carver TJ , RutherfordKM, BerrimanM, et al. ACT: the Artemis comparison tool. Bioinformatics 2005;21 :3422–3.15976072
21. Thorvaldsdóttir H , RobinsonJT, MesirovJP. Integrative genomics viewer (IGV): high-performance genomics data visualization and exploration. Brief Bioinform 2013;14 :178–92.22517427
22. Darling ACE , MauB, BlattnerFR, PernaNT. Mauve: multiple alignment of conserved genomic sequence with rearrangements. Genome Res 2004;14 :1394–403.15231754
23. Diesh C , StevensGJ, XieP, et al. JBrowse 2: a modular genome browser with views of synteny and structural variation. Genome Biol 2023;24 :74.37069644
24. Wang Y , TangH, DebarryJD, et al. MCScanX: a toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Res 2012;40 :e49.22217600
25. Tang H , BomhoffMD, BrionesE, et al. SynFind: compiling syntenic regions across any set of genomes on demand. Genome Biol Evol 2015;7 :3286–98.26560340
26. Chen IMA , ChuK, PalaniappanK, et al. The IMG/M data management and analysis system v.6.0: new tools and advanced capabilities. Nucleic Acids Res 2021;49 :D751–63.33119741
27. Cleary AM , FarmerAD. Genome context viewer (GCV) version 2: enhanced visual exploration of multiple annotated genomes. Nucleic Acids Res 2023;51 :W225–31.37207325
28. Raney BJ , BarberGP, Benet-PagèsA, et al. The UCSC genome browser database: 2024 update. Nucleic Acids Res 2024;52 :D1082–8.37953330
29. Szklarczyk D , FranceschiniA, WyderS, et al. STRING v10: protein-protein interaction networks, integrated over the tree of life. Nucleic Acids Res 2015;43 (D1 ):D447–52.25352553
30. Cunningham F , AllenJE, AllenJ, et al. Ensembl 2022. Nucleic Acids Res 2022;50 :D988–95.34791404
31. Price M , ArkinA. A fast comparative genome browser for diverse bacteria and archaea. PLoS One. 2024;19 :e0301871.
32. Pereira J . GCsnap: interactive snapshots for the comparison of protein-coding genomic contexts. J Mol Biol 2021;433 :166943.33737026
33. Li X , ChenF, ChenY. Gcluster: a simple-to-use tool for visualizing and comparing genome contexts for numerous genomes. Bioinformatics 2020;36 :3871–3.32221617
34. Harrison KJ , deCrécy-LagardV, ZallotR. Gene graphics: a genomic neighborhood data visualization web application. Bioinformatics 2018;34 :1406–8.29228171
35. Gumerov VM , ZhulinIB. TREND: a platform for exploring protein function in prokaryotes based on phylogenetic, domain architecture and gene neighborhood analyses. Nucleic Acids Res 2020;48 :W72–6.32282909
36. Garcia PS , JauffritF, GrangeasseC, Brochier-ArmanetC. GeneSpy, a user-friendly and flexible genomic context visualizer. Bioinformatics 2019;35 :329–31.29912383
37. Botas J , Rodríguez Del RíoÁ, Giner-LamiaJ, Huerta-CepasJ. GeCoViz: genomic context visualisation of prokaryotic genes from a functional and evolutionary perspective. Nucleic Acids Res 2022;50 :W352–7.35639770
38. Vallenet D , CalteauA, DuboisM, et al. MicroScope: an integrated platform for the annotation and exploration of microbial gene functions through genomic, pangenomic and metabolic comparative analysis. Nucleic Acids Res 2020;48 :D579–89.31647104
39. Dieckmann MA , BeyversS, Nkouamedjo-FankepRC, et al. EDGAR3.0: comparative genomics and phylogenomics on a scalable infrastructure. Nucleic Acids Res 2021;49 :W185–92.33988716
40. Dehal PS , JoachimiakMP, PriceMN, et al. MicrobesOnline: an integrated portal for comparative and functional genomics. Nucleic Acids Res 2010;38 :D396–400.19906701
41. Saha CK , Sanches PiresR, BrolinH, et al. FlaGs and webFlaGs: discovering novel biology through the analysis of gene neighbourhood conservation. Bioinformatics 2021;37 :1312–4.32956448
42. Lobb B , TremblayBJ-M, Moreno-HagelsiebG, DoxeyAC. An assessment of genome annotation coverage across the bacterial tree of life. Microb Genom 2020;6 :e000341. 10.1099/mgen.0.000341.
43. Mendler K , ChenH, ParksDH, et al. AnnoTree: visualization and exploration of a functionally annotated microbial tree of life. Nucleic Acids Res 2019;47 :4442–8.31081040
44. Parks DH , ChuvochinaM, WaiteDW, et al. A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life. Nat Biotechnol 2018;36 :996–1004.30148503
45. Kanehisa M , FurumichiM, TanabeM, et al. KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res 2017;45 :D353–61.27899662
46. Finn RD , BatemanA, ClementsJ, et al. Pfam: the protein families database. Nucleic Acids Res 2014;42 :D222–30.24288371
47. Haft DH , SelengutJD, RichterRA, et al. TIGRFAMs and genome properties in 2013. Nucleic Acids Res 2012;41 :D387–95.23197656
48. Hyatt D , ChenG-L, LocascioPF, et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 2010;11 :119.20211023
49. Finn RD , ClementsJ, EddySR. HMMER web server: interactive sequence similarity searching. Nucleic Acids Res 2011;39 (suppl ):W29–37.21593126
50. Buchfink B , XieC, HusonDH. Fast and sensitive protein alignment using DIAMOND. Nat Methods 2015;12 :59–60.25402007
51. Suzek BE , WangY, HuangH, et al. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics 2015;31 :926–32.25398609
52. Smith TF , WatermanMS. Identification of common molecular subsequences. J Mol Biol 1981;147 :195–7.7265238
53. Ermler U , GrabarseW, ShimaS, et al. Crystal structure of methyl-coenzyme M reductase: the key enzyme of biological methane formation. Science 1997;278 :1457–62.9367957
54. Gendron A , AllenKD. Overview of diverse methyl/alkyl-coenzyme M reductases and considerations for their potential heterologous expression. Front Microbiol 2022;13 :867342.35547147
55. Ali S , JenkinsB, ChengJ, et al. Slr4, a newly identified S-layer protein from marine Gammaproteobacteria, is a major biofilm matrix component. Mol Microbiol 2020;114 :979–90.32804439
56. Tomás JM . The main Aeromonas pathogenic factors. ISRN Microbiol 2012;2012 :256261.23724321
57. Boot HJ , PouwelsPH. Expression, secretion and antigenic variation of bacterial S-layer proteins. Mol Microbiol 1996;21 :1117–23.8898381
58. Yao J , ZhenX, TangK, et al. Novel polyadenylylation-dependent neutralization mechanism of the HEPN/MNT toxin/antitoxin system. Nucleic Acids Res 2020;48 :11054–67.33045733
59. Fraikin N , GoormaghtighF, Van MelderenL. Type II toxin-antitoxin systems: evolution and revolutions. J Bacteriol 2020;202 :e00763-19.
60. de Groot RJ . Structure, function and evolution of the hemagglutinin-esterase proteins of corona- and toroviruses. Glycoconj J 2006;23 :59–72.16575523
61. Zeng Q , LangereisMA, vanVlietALW, et al. Structure of coronavirus hemagglutinin-esterase offers insight into corona and influenza virus evolution. Proc Natl Acad Sci U S A 2008;105 :9065–9.18550812
62. Woo PCY , HuangY, LauSKP, YuenK-Y. Coronavirus genomics and bioinformatics analysis. Viruses 2010;2 :1804–20.21994708
63. Lau SKP , WooPCY, YipCCY, et al. Isolation and characterization of a novel Betacoronavirus subgroup a coronavirus, rabbit coronavirus HKU14, from domestic rabbits. J Virol 2012;86 :5481–96.22398294
64. Wei X , WentzT, LobbB, et al. Identification of divergent botulinum neurotoxin homologs in Paeniclostridium ghonii. bioRxiv 2022; 2022.08.17.504336.
65. Hodgins HP , ChenP, LobbB, et al. Ancient clostridium DNA and variants of tetanus neurotoxins associated with human archaeological remains. Nat Commun 2023;14 :5475.37673908
66. Wei X , LobbB, WangK, et al. Identification of a botulinum neurotoxin-like gene cluster in Bacillus toyonensis. bioRxiv 2023; 2023.07.21.550100.
67. Hunter S , ApweilerR, AttwoodTK, et al. InterPro: the integrative protein signature database. Nucleic Acids Res 2009;37 :D211–5.18940856
