
==== Front
Proc Natl Acad Sci U S A
Proc Natl Acad Sci U S A
PNAS
Proceedings of the National Academy of Sciences of the United States of America
0027-8424
1091-6490
National Academy of Sciences

39186661
202414223
10.1073/pnas.2414223121
commCommentaryevolutionEvolution418
437
Commentary
Biological Sciences
Evolution
Unchecked growth: Pushing the limits on RNA virus genome size in the absence of known proofreading
Glasner Dustin R. a
Daugherty Matthew D. mddaugherty@ucsd.edu
a 1 https://orcid.org/0000-0002-4879-9603

aDepartment of Molecular Biology, School of Biological Sciences, University of California, San Diego, La Jolla, CA 92093
1To whom correspondence may be addressed. Email: mddaugherty@ucsd.edu.
26 8 2024
3 9 2024
26 8 2024
121 36 e2414223121Copyright © 2024 the Author(s). Published by PNAS.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This open access article is distributed under Creative Commons Attribution-NonCommercial-NoDerivatives License 4.0 (CC BY-NC-ND).

HHS | NIH | National Institute of General Medical Sciences (NIGMS) 100000057 GM133633 Matthew D. Daugherty Burroughs Wellcome Fund (BWF) 100000861 PATH program Matthew D. Daugherty
==== Body
pmcRNA viruses have notoriously error-prone RNA-dependent RNA polymerases. As a result, substantial numbers of mutations are introduced to the virus during each round of genome replication, driving RNA viruses dangerously close to the so-called “error catastrophe” threshold in which the virus fails to produce sufficient infectious progeny (1–3). Mathematically speaking, smaller genomes will accumulate fewer total mutations and therefore are less likely to encounter an error catastrophe. Consistent with this premise, most nonsegmented RNA viruses have genomes that are ~10 kb or smaller (Fig. 1).

Fig. 1. Comparison of genome size and domain organization among nonsegmented RNA viruses. Most viruses, exemplified by dengue virus shown here, have a compact genome of ~10,000 bases (10 kb) or less. In contrast, many nidoviruses have large genomes ranging from 20 to 40 kb, including the ~30 kb SARS-CoV-2 genome shown here. Encoded within those nidovirus genomes is a 3′ to 5′ exonuclease (ExoN) domain (red) that is required for proofreading correction of errors made by the viral RNA-dependent RNA polymerase (RdRp, shown in blue). The ~40 kb genome of the newly discovered Maximus pesti-like virus is also shown, indicating positions of the RdRp and other domains with homology to other RNA viruses. Notably, there is no ExoN-like domain, but there are several domains (yellow) that are homologous to eukaryotic or bacterial proteins, which the authors suggest may represent novel mechanisms to facilitate replication and translation of such a large RNA viral genome.

However, for every trend in virology, there are exceptions. The best-known examples of large nonsegmented RNA virus genomes are found within the order Nidovirales. For example, coronaviruses, including SARS-CoV-2, have genomes that can be as large as 32 kb, while other nidoviruses have genomes >40 kb (4, 5). These viruses maintain genome sizes of >30 kb by encoding a polymerase error correcting mechanism that relies on a 3′ to 5′ exonuclease (ExoN) domain (6) (Fig. 1). This ExoN activity allows for proofreading during RNA genome synthesis, effectively lowering the rate of mutations for each round of genome replication and therefore avoiding error catastrophe (6–8). Other viruses, including arteriviruses within the Nidovirales, plant-infecting Closteroviridae, and recently described “large genome-flavi-like” (LGF) members of the Flaviviridae, have genomes that can span from ~15 to 27 kb but do not encode any known polymerase error correction mechanism (4, 9, 10). Thus, the size limit for a nonsegmented RNA virus genome that lacks an error correction mechanism appeared to be <30 kb, above which all known viruses were nidoviruses that contain an ExoN domain.

In violation of this previous genome size limit, Petrone et al. now describe the exciting discovery of a member of the Flaviviridae with a ~40 kb genome that contains no domains previously associated with polymerase error correcting (11) (Fig. 1). The authors discovered this viral genome while sampling the metagenomes of sea sponges and other invertebrates in Sydney Harbour, Sydney, Australia. The RNA viral genome, assembled with 139x sequence coverage, spans 39.8 kb and encodes one polyprotein of >12,500 amino acids. To determine the relationship between this virus and other RNA viruses, the authors used AlphaFold structural prediction and structure homology detection to identify domains within the viral polyprotein that are homologous to other viral proteins. Using this approach, they identified capsid and envelope proteins, a protease/helicase pair, and an RNA-dependent RNA-polymerase (RdRp), all of which are characteristic of the Flaviviridae family of viruses. Using the RdRp and protease/helicase domains, they demonstrate that this virus phylogenetically groups within the Flaviviridae family near the root of the clade that contains Pestiviruses and other LGF viruses, the largest of which was previously described to be 27 kb (10). Thus, this new virus, which they name “Maximus pesti-like virus,” represents the largest of all known viruses outside of the Nidovirales and falls within a basal portion of the Flaviviridae phylogeny, giving us a new view into the ancestral state of this important family of viruses.

Notably, using these sequence and structural homology analyses, the authors find no evidence for an ExoN-like proofreading domain. Indeed, much of the genome lacks obvious homology to domains found in any other RNA virus genomes (Fig. 1). However, the authors did find several domains with similarity to cellular (e.g., bacterial and eukaryotic) proteins, which they propose could encode activities unique to this virus to allow for the maintenance of its large RNA genome and polyprotein. For instance, one domain resembles cellular Tu elongation factors, which aid in accurate and rapid mRNA translation in eukaryotic and bacterial cells (12), and may therefore be beneficial for a virus with a >12,000 residue polyprotein. Even more intriguingly, they discovered a set of contiguous domains, which they dub a “nucleic acid metabolism cassette,” that they suggest might be involved in avoiding error catastrophe during genome replication. This cassette comprises adenylate kinase, nucleoside diphosphate kinase, phosphoprotein phosphatase, and photo-lyase-related domains that are homologous to proteins found in bacteria, eukaryotes, and in some cases, DNA viruses, but are found in no other RNA viruses. Due to the described roles of these domains in regulating levels of nucleoside phosphate pools, which are important for genome replication fidelity (13), as well as repairing UV-induced nucleotide damage (14), the co-option of these cellular domains by a virus provides an intriguing hypothesis for how a large RNA virus can avoid error catastrophe. Determining the function of these domains in viral replication will be an important next step in identifying whether they are required for genome replication fidelity. Regardless, the lack of any known viral error correcting proteins in this virus indicates that separate mechanisms must exist for RNA viruses to overcome the error catastrophe problem, whether through a higher-fidelity polymerase, a proofreading mechanism that is distinct from that found in nidoviruses, or some other mechanism that has yet to be described.

Petrone et al. now describe the exciting discovery of a member of the Flaviviridae with a ~40 kb genome that contains no domains previously associated with polymerase error correcting.

Beyond the implications of this newly discovered virus on novel mechanisms to circumvent error catastrophe, this study also sheds light on virus evolution more generally. These findings clearly indicate that >30 kb genomes have independently evolved in two separate nonsegmented RNA virus families, suggesting that other examples are likely to exist. Interestingly, while it has been proposed that nidoviruses acquired a proofreading mechanism prior to genome expansion (15), the phylogenetic placement of this novel virus has led the authors to suggest that members of the Flaviviridae may have evolved from large viruses such as the one they discovered toward the smaller genome viruses more commonly found in vertebrates such as dengue, Zika, and hepatitis C viruses. Finally, these data indicate that gene acquisition by RNA viruses from eukaryotes, and even from bacteria, may be more common than currently appreciated. While gene acquisition by viruses has been described, it has been overwhelmingly found in large (e.g., poxvirus and herpesvirus) or giant (e.g., pandoravirus and mimivirus) DNA virus genomes (16). The results here suggest that large RNA viruses may also employ gene acquisition to generate evolutionary novelty. Indeed, a recent paper described multiple independent acquisitions of an immune antagonizing enzymatic domain by distinct lineages of large genome nidoviruses (17). As the number of RNA viruses with large genomes increases, it is probable that we will also observe more instances of captured cellular genes in RNA viruses that aid in viral replication or immune evasion.

Discoveries such as the one described here highlight the importance of continued sampling and characterization of viruses from diverse hosts, especially nonvertebrate eukaryotic hosts. The so-called giant DNA viruses, which vastly exceed both the physical size as well as genome size of other known viruses, were first discovered in algae and amoeba and have up-ended many previously held assumptions about viral biology and evolution (18, 19). Likewise, individual viruses discovered in insects, flatworms, and now a sea sponge have each expanded our understanding of RNA virus genome size evolution (5, 10, 11, 15). As the pace of RNA virus discovery continues to grow (20–22), we are guaranteed to continue to need to revise our understanding of genome size constraints and the limits of virus error catastrophe.

The work was supported by grants from the NIH (R35GM133633) and Burroughs Wellcome Fund Investigators in the Pathogenesis of Infectious Disease program to M.D.D.

Author contributions

D.R.G. and M.D.D. wrote the paper.

Competing interests

The authors declare no competing interest.

See companion article, “A ~40kb flavi-like virus does not encode a known error-correcting mechanism,” 10.1073/pnas.2403805121.
==== Refs
1 S. Crotty, C. E. Cameron, R. Andino, RNA virus error catastrophe: Direct molecular test by using ribavirin. Proc. Natl. Acad. Sci. U.S.A. 98 , 6895–6900 (2001).11371613
2 J. W. Drake, J. J. Holland, Mutation rates among RNA viruses. Proc. Natl. Acad. Sci. U.S.A. 96 , 13910–13913 (1999).10570172
3 E. C. Holmes, Error thresholds and the constraints to RNA virus evolution. Trends Microbiol. 11 , 543–546 (2003).14659685
4 F. Ferron, B. Sama, E. Decroly, B. Canard, The enzymes for genome size increase and maintenance of large (+)RNA viruses. Trends Biochem. Sci. 46 , 866–877 (2021).34172362
5 A. Saberi, A. A. Gulyaeva, J. L. Brubacher, P. A. Newmark, A. E. Gorbalenya, A planarian nidovirus expands the limits of RNA genome size. PLoS Pathog. 14 , e1007314 (2018).30383829
6 E. Minskaia , Discovery of an RNA virus 3’->5’ exoribonuclease that is critically involved in coronavirus RNA synthesis. Proc. Natl. Acad. Sci. U.S.A. 103 , 5108–5113 (2006).16549795
7 L. D. Eckerle, X. Lu, S. M. Sperry, L. Choi, M. R. Denison, High fidelity of murine hepatitis virus replication is decreased in nsp14 exoribonuclease mutants. J. Virol. 81 , 12135–12144 (2007).17804504
8 E. C. Smith, H. Blanc, M. C. Surdel, M. Vignuzzi, M. R. Denison, Coronaviruses lacking exoribonuclease activity are susceptible to lethal mutagenesis: Evidence for proofreading and potential therapeutics. PLoS Pathog. 9 , e1003565 (2013).23966862
9 V. V. Dolja, J. F. Kreuze, J. P. Valkonen, Comparative and functional genomics of closteroviruses. Virus Res. 117 , 38–51 (2006).16529837
10 E. E. Matsumura, L. Nerva, J. C. Nigg, B. W. Falk, S. Nouri, Complete genome sequence of the largest known flavi-like virus, Diaphorina citri flavi-like virus, a novel virus of the Asian citrus psyllid, Diaphorina citri. Genome Announc. 4 , e00946-16 (2016).27609921
11 M. E. Petrone , A ~40kb flavi-like virus does not encode a known error-correcting mechanism. Proc. Natl. Acad. Sci. U.S.A. 121 , e2403805121 (2024).39018195
12 K. W. Ieong, U. Uzun, M. Selmer, M. Ehrenberg, Two proofreading steps amplify the accuracy of genetic code translation. Proc. Natl. Acad. Sci. U.S.A. 113 , 13744–13749 (2016).27837019
13 I. Kapoor, U. Varshney, Diverse roles of nucleoside diphosphate kinase in genome stability and growth fitness. Curr. Genet. 66 , 671–682 (2020).32249353
14 K. Brettel, M. Byrdin, Reaction mechanisms of DNA photolyase. Curr. Opin. Struct. Biol. 20 , 693–701 (2010).20705454
15 P. T. Nga , Discovery of the first insect nidovirus, a missing evolutionary link in the emergence of the largest RNA virus genomes. PLoS Pathog. 7 , e1002215 (2011).21931546
16 N. A. T. Irwin, A. A. Pittis, T. A. Richards, P. J. Keeling, Systematic evaluation of horizontal gene transfer between eukaryotes and viruses. Nat. Microbiol. 7 , 327–336 (2022).34972821
17 S. A. Goldstein, N. C. Elde, Recurrent viral capture of cellular phosphodiesterases that antagonize OAS-RNase L. Proc. Natl. Acad. Sci. U.S.A. 121 , e2312691121 (2024).38277437
18 H. A. M. Monttinen, C. Bicep, T. A. Williams, R. P. Hirt, The genomes of nucleocytoplasmic large DNA viruses: Viral evolution writ large. Microb. Genom. 7 , 000649 (2021).34542398
19 F. Schulz, C. Abergel, T. Woyke, Giant virus biology and diversity in the era of genome-resolved metagenomics. Nat. Rev. Microbiol. 20 , 721–736 (2022).35902763
20 M. Shi , Redefining the invertebrate RNA virosphere. Nature 540 , 539–543 (2016).27880757
21 Y. I. Wolf , Doubling of the known set of RNA viruses by metagenomic analysis of an aquatic virome. Nat. Microbiol. 5 , 1262–1270 (2020).32690954
22 A. A. Zayed , Cryptic and abundant marine viruses at the evolutionary origins of Earth’s RNA virome. Science 376 , 156–162 (2022).35389782
