
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

69896
10.1038/s41598-024-69896-1
Article
Unravelling aggregation propensity of rotavirus A VP6 expressed as E. coli inclusion bodies through in silico prediction
http://orcid.org/0000-0002-5972-0550
Kuri Pooja Rani
http://orcid.org/0000-0003-4098-043X
Goswami Pranab pgoswami@iitg.ac.in

https://ror.org/0022nd079 grid.417972.e 0000 0001 1887 8311 Department of Biosciences and Bioengineering, Indian Institute of Technology Guwahati, Guwahati, Assam 781039 India
13 9 2024
13 9 2024
2024
14 2146416 6 2024
9 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
The inner capsid protein of rotavirus, VP6, emerges as a promising candidate for next-generation vaccines against rotaviruses owing to its abundance in virion particles and high conservation. However, the formation of inclusion bodies during prokaryotic VP6 expression poses a significant hurdle to rotavirus research and applications. Here, we employed experimental and computational approaches to investigate inclusion body formation and aggregation-prone regions (APRs). Heterologous recombinant VP6 expression in Escherichia coli BL21(DE3) cells resulted in inclusion body formation, confirmed by transmission electron microscopy revealing amorphous aggregates. Thioflavin T assay demonstrated incubation temperature-dependent aggregation of VP6 inclusion bodies. Computational predictions of APRs in rotavirus A VP6 protein were performed using sequence-based tools (TANGO, AGGRESCAN, Zyggregator, Waltz, FoldAmyloid, ANuPP, Camsol intrinsic) and structure-based tools (SolubiS, CamSol structurally corrected, Aggrescan3D). A total of 24 consensus APRs were identified, with 21 of them being surface-exposed in VP6. All identified APRs display a predominance of hydrophobic amino acids, ranging from 33 to 100%. Computational identification of these APRs corroborates our experimental observation of VP6 inclusion body or aggregate formation. Characterization of VP6's aggregation propensity facilitates understanding of its behaviour during prokaryotic expression and opens avenues for protein engineering of soluble variants, advancing research on rotavirus VP6 in pathology, therapy, and diagnostics.

Keywords

Rotavirus
Protein aggregation
Aggregation prone regions
Inclusion bodies
Subject terms

Biotechnology
Expression systems
http://dx.doi.org/10.13039/501100001407 Department of Biotechnology, Ministry of Science and Technology, India BT/PR41449/NER/95/1687/2020 issue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Rotaviruses infect a wide range of vertebrates, including both birds and mammals. The serological classification of the Rotavirus genus encompasses ten distinct species (Rotavirus groups A–J) based on the antigenic specificity and genetic variability exhibited by the viral capsid protein VP61–3. Group A rotavirus (RVA) is the primary etiological agent responsible for acute gastroenteritis in humans, predominantly afflicting the under-five population, exerting a global impact4. RVA infections have accounted for a substantial annual mortality toll, surpassing half a million cases, preceding rotavirus vaccine implementation5,6.

Rotavirus VP6 has been acknowledged as a promising non-live next-generation vaccine candidate against rotavirus owing to its attributes of high abundance in the virion particle, high conservancy, and immunogenic nature conferring heterotypic protection7,8. Diverse vaccine modalities encoding rotavirus VP6 antigen, encompassing DNA vaccines 9–11, subunit vaccines incorporating recombinant VP6 protein12 and self-assembled structures have been documented to elicit immune responses or confer protection in animal models13–16. VP6 self-assembled structures, particularly VP6 nanotubes/Virus-Like Particles (VLPs), demonstrate superior immunogenicity than that of VP6 monomers or trimers, eliciting strong immune responses7. Additionally, they exhibit dual functionality as adjuvants and carriers for immunological applications17. Advancements in the production and purification of VP6 nanotubes/VLPs have enabled exploration of their multifaceted potential as immunogens, adjuvants, delivery vehicles, and nano-biomaterials18. Furthermore, the rotavirus VP6 antigen is utilized as a detection biomarker for the purpose of rotavirus diagnosis19,20.

VP6 is a structural protein, present in the middle of the triple layered capsids of the virion. Notably, rotavirus structural proteins govern host specificity, cellular entry, enzymatic machinery for viral transcript synthesis, and harbor immunological epitopes. Situated in the middle layer, 260 VP6 proteins engage in extensive interactions with both the outer layer (VP4 and VP7) and the inner layer (VP2)21–24. These interactions are essential for the assembly of the viral capsid and transcriptase activity. The integrity of these layers and the orchestrated interplay between their constituent proteins are fundamental for the overall rotavirus structure.

The VP6 protein, which constitutes 51% of the total rotavirus virion mass25 demonstrates a hierarchical arrangement when expressed in vitro in appropriate expression systems. The native VP6 protein exhibits self-assembly with structural polymorphism and forms trimers, hexagonal structures and further assembles into diverse architectures such as nanotubes, nanospheres (VLPs), or nanosheets25–29. Such polymorphic VP6 self-assembly is influenced by factors such as pH and ionic strength of divalent cations (Ca2+ and Mg2+)30,31.

The recombinant expression of VP6 has been of enormous importance to investigate rotavirus pathogenesis, development of vaccines, and diagnostic potentials. Several expression systems, including prokaryotic Escherichia coli13,32,33 and various eukaryotes such as fungi34, baculovirus-silkworm systems35, insect cell/baculovirus systems36,37, mammalian cells36, and plant cells38–40 have been explored for the production of VP6. However, the prokaryotic expression of heterologous proteins remains to be the most convenient and low cost method for laboratory as well as technological purposes41. Ironically, prokaryotic VP6 expression leads to the formation of inclusion bodies13,32,33, which represents a significant obstacle to VP6 protein research and applications. This necessitates an explicit structural understanding of aggregation propensity of rotavirus VP6 perhaps leading to inclusion body formation.

The establishment of protein's native conformation mostly involves intramolecular non-covalent interactions. However, the high translation rate in cells continuously generates a large number of unfolded polypeptides, exposing hydrophobic side chains and hydrogen bond donors/acceptors to the surrounding solvent. This enables both intramolecular and intermolecular interactions of protein segments, leading to the formation of either native proteins or aggregated proteins, respectively42. Consequently, protein aggregation, forming inclusion bodies, dominates in many cases of heterologous protein expression in prokaryotes, such as E. coli43. Because protein aggregation is associated with structural compromise of the otherwise native structure, the inclusion bodies are significantly rendered non-functional. Recovery of functional target protein necessitates unfolding and refolding procedures, which can be cumbersome and inefficient44. Protein aggregation is facilitated by specific sequences in the polypeptide chain that initiate the aggregation process, leading to the enrichment of intermolecular beta-sheet structures. While amyloid-like aggregates are characterized by cross-beta sheet structures, protein aggregates can encompass both amorphous and amyloid-like forms45,46. These sequential determinants of protein aggregation are called aggregation-prone regions (APRs) or aggregation hotspots46–49. The identification of APRs can aid in protein engineering through sequence modification of the target protein, achieved by mutating residues in aggregation hotspots. This approach aims to enhance soluble expression of the target protein by reducing its aggregation propensity and increasing its solubility44,50. Insights into the mechanisms of aggregation, derived from experimental and computational findings, may facilitate the development of reliable algorithms for prediction of APRs.

In this work, computational prediction of APRs in the VP6 protein of rotavirus A has been performed in response to the observation of VP6 inclusion body formation during heterologous recombinant expression in E. coli. The determination of the intrinsic aggregation propensity of VP6 can contribute to our understanding of its aggregation mechanisms and open avenues for future protein engineering efforts to generate more soluble forms during recombinant expression. This, in turn, can support fundamental and practical research concerning the pathology, therapy, or diagnostics of rotavirus VP6. A series of computational APR prediction tools were employed here to analyze both the sequence and structure of VP6. Various algorithms were selected to mitigate biases stemming from training sets, parameterization, and the unique characteristics inherent to each method. Utilizing this approach, diverse sequence-based as well as structure-based computational methods were applied. These computational techniques play a pivotal role in forecasting APRs within proteins, which is crucial for both protein characterization and biotechnological applications. Emphasis was placed on APRs identified by two or more prediction tools, underscoring their potential significance in elucidating protein behavior and devising strategies for protein solubilization and aggregation inhibition.

Results

Cell transformation and optimization of VP6 expression in E. coli BL21(DE3)

E. coli BL21(DE3) was successfully transformed with pET-28a(+)-RVAVP6 recombinant plasmid confirmed by single/double restriction digestion and PCR. VP6 was effectively overexpressed at various tested Isopropyl β-d-1-thiogalactopyranoside (IPTG) concentrations ranging from 0.0 mM to 1 mM, with comparable band intensities. Hence, an IPTG concentration of 0.4 mM was considered for the subsequent VP6 protein expression. However, despite successful overexpression, the VP6 protein consistently appeared in the pellet fraction without evident solubilization (Fig. S1).

To determine the optimum post-induction temperature for overexpression of VP6, the bacterial cultures were subjected to post-induction growth at temperatures of 18 °C, 25 °C, 30 °C and 37 °C. The overexpression was at a higher extent from 25 to 37 °C, almost equally. Henceforth, overexpression of VP6 was carried out at 25 °C (Fig. S2).

Post-induction incubation time was assessed to determine the optimal time for protein expression and PMSF, a serine protease inhibitor, was used to investigate possible protein degradation of overexpressed VP6 in the soluble fraction during protein extraction, in a time-dependent manner. The effect of post-induction incubation time on solubilizing the protein of interest was tested with and without PMSF treatment (0.2 mM). Protein extraction was carried out at 4 h, 6 h, 8 h, 10 h, and 12 h post-induction to assess solubilization efficiency. The overexpression of VP6 protein increased from 4 to 6 h post-induction, after which the amount of overexpressed protein remained stable until 12 h of incubation. Interestingly, the presence or absence of PMSF treatment did not affect VP6 overexpression. However, under all conditions tested, the protein consistently appeared in the insoluble fraction (Fig. S3).

To achieve protein solubilization, urea treatment was tested with concentrations ranging from 1 to 8 M in the lysis and resuspension buffers during protein extraction. Analysis of the urea-treated protein samples by SDS-PAGE revealed considerable partial solubilization of VP6 at urea concentrations of 6–8 M (Fig. S4).

VP6 purification and western blot

The urea-solubilized, overexpressed VP6 was subjected to Ni–NTA column chromatography and coupled with on-column renaturation. The purity of the protein sample eluate was analysed by SDS-PAGE in comparison to the crude lysate (Fig. 1a,b). The eluted sample containing purified VP6 was subjected to Western Blot, and a dark brown coloured band was observed at ~ 45 kDa, corresponding to the molecular weight of VP6 (Fig. 1c).Figure 1 Characterization of recombinant RVA VP6 protein. (a) SDS-PAGE analysis of crude cell lysates from induced E. coli BL21(DE3)/pET-28a(+)-RVAVP6. S represents the supernatant/soluble fraction while P represents pellet/insoluble fraction. (b) SDS-PAGE and corresponding (c) western blot analysis of purified RVA VP6 with VP6 specific monoclonal antibody. (d) FETEM images depict VP6 inclusion bodies at magnifications of 1 µm (I) and 100 nm (II), alongside purified VP6 at magnifications of 1 µm (III) and 100 nm (IV), showing amorphous aggregates.

Field emission transmission electron microscopy (FETEM) of VP6 inclusion bodies and purified VP6

FETEM analysis was conducted to assess the structural characteristics of purified VP6 samples (Fig. 1d). Intriguingly, the FETEM images revealed the presence of amorphous aggregates within the purified VP6 samples from prokaryotic expression system. These aggregates exhibited a lack of defined structure or crystallinity, appearing as irregularly shaped accumulations dispersed throughout the sample. The observation of amorphous aggregates suggests a propensity for VP6 to undergo non-specific interactions and self-association, which may have implications for its stability, functionality, and downstream applications. Such aggregation behaviour suggests potential challenges in the downstream processing and formulation of VP6-based products, necessitating further investigation into the factors influencing VP6 aggregation and strategies for mitigating this phenomenon.

ThT binding assay

The examination of VP6 inclusion bodies through ThT assay, at an excitation wavelength of 440 nm, unveiled a peak emission at approximately 487 nm, indicative of amyloidogenic aggregates. Intriguingly, fluorescence intensity analysis revealed a temperature-dependent modulation of amyloid character in VP6 inclusion body formation (Fig. 2a, b). The plot depicts a linear trend in ThT fluorescence intensity in VP6 inclusion bodies with increasing incubation temperature, suggesting a systematic relationship between incubation temperature and amyloid characteristic of VP6 inclusion bodies. This finding underscores the thermally sensitive nature of amyloid formation in VP6 inclusion bodies, aligning with the established phenomenon that higher temperatures enhance amyloid aggregation propensity51. This understanding can contribute to a more comprehensive characterization of the conformational dynamics and thermal stability of VP6 aggregates following expression at different temperatures.Figure 2 Aggregation profiling of VP6 inclusion bodies expressed at different temperatures using ThT binding assay. (a) Fluorescence Emission Spectra of ThT bound to VP6 inclusion bodies (excitation: 440 nm). (b) Temperature-dependent fluorescence response of ThT in VP6 inclusion bodies (excitation: 440 nm, emission: 487 nm).

Physicochemical characterization and structure prediction of rotavirus A VP6 protein

The retrieved VP6 sequence was submitted in FASTA format to ProtParam for physicochemical characterization. The VP6 protein is 397 amino acids in length with molecular weight of 44.873 kDa. and pI of 5.81. Instability index had been predicted to be 40.27, suggesting that the protein may be unstable. This indicates the protein's susceptibility to structural instability. This metric assesses the likelihood of the protein undergoing denaturation or loss of structure under physiological conditions. A higher instability index implies increased vulnerability to unfolding or aggregation, potentially affecting its functional integrity and experimental outcomes. Therefore, this value suggests that VP6 may exhibit characteristics predisposing it to instability, a factor important for understanding its behavior and applications in biological studies. The aliphatic index of a protein is defined as the relative volume occupied by aliphatic side chains (alanine, valine, isoleucine, and leucine). It may be regarded as a positive factor for the increase of thermostability of globular proteins. The aliphatic index of VP6 were found to be 88.66. The aliphatic index, indicative of the proportion of aliphatic side chains in the protein sequence, highlights its influence on stability, hydrophobicity, and thermal stability. With an aliphatic index of 88.66, VP6 exhibits a notable presence of hydrophobic residues, which can promote aggregation when exposed on the protein surface. Conversely, buried aliphatic residues contribute to stability but may hinder solubility if essential for interactions with aqueous environments. These characteristics underscore how aliphatic amino acids, crucial for stability and function, significantly impact the solubility profile of VP6. Engineering strategies often focus on modifying surface-exposed aliphatic residues to enhance solubility in aqueous solutions, balancing the need for stability with the challenges posed by hydrophobic interactions and potential aggregation pathways. GRAVY value of VP6 were predicted to be − 0.144. The GRAVY number of a protein is a measure of its hydrophobicity or hydrophilicity. A negative GRAVY value indicates that the protein is hydrophilic. The negative GRAVY index of VP6, indicating hydrophilicity, contrasts with the presence of numerous exposed hydrophobic patches identified in our study. This discrepancy is intriguing as it suggests a potential complexity in VP6's structural and functional roles. While hydrophobic patches typically correlate with aggregation propensity, the overall hydrophilic nature indicated by GRAVY may imply a specific mechanism of interaction or folding that counters the aggregative tendency of these patches. Understanding this dichotomy can be crucial for elucidating VP6's behaviour in both native and recombinant contexts, influencing strategies for protein engineering and therapeutic development. Its half-life is estimated to be 30 h (mammalian reticulocytes, in vitro), > 20 h (yeast, in vivo) and > 10 h (E. coli, in vivo).

The solubility of VP6 in the E. coli expression system was predicted by the Protein-Sol server. The population average for the experimental dataset (PopAvrSol) for E. coli expressed proteins is 0.45. Therefore, a scaled solubility value above 0.45 suggests higher predicted solubility than the average soluble E. coli protein in the dataset, whereas values below 0.45 indicate lower predicted solubility. Notably, the predicted solubility of VP6 in the E. coli expression system, assessed by the Protein-Sol server, was lower compared to other experimentally expressed proteins (Fig. S5).

While there is no available complete experimental structural data for rotavirus A VP6 sequence used in the current investigation, experimental structural data for bovine rotavirus A VP6 is available (PDB ID: 1QHD), which exhibit 96.98% amino acid sequence identity to the VP6 sequence used in the current study. The experimentally determined structure of bovine rotavirus A VP6 reveals the following secondary structure composition: alpha helix 27.7%, 3–10 helix 4.5%, pi helix 0.0%, extended strand 31.0%, isolated beta bridge 0.5%, turn 9.3%, bend 7.3%, and coil 19.6%. Given the high sequence identity, it is plausible that the human rotavirus A VP6 sequence exhibits a similar secondary structure composition.

The VP6 three-dimensional structure predicted by SWISS-MODEL underwent structural validation. Procheck analysis showed that 92.4% of residues were in the most favoured regions according to the Ramachandran Plot. ProSA indicated an overall Z-score of − 8.42, suggesting a structure within an acceptable range of native-like conformations. Protein structure analysis by Verify 3D demonstrated that 87.91% of residues achieved a 3D–1D score ≥ 0.1, exceeding the 80% threshold. Verify 3D scores assess the compatibility of a protein's 3D structure with its amino acid sequence, indicating higher scores for residues in environments typical of native proteins. Additionally, the ERRAT score was 96.063, surpassing the 95% threshold (Fig. 3). ERRAT evaluates overall structure quality by analyzing non-bonded interactions. The X-axis of the ERRAT graph represents the amino acid positions in the protein sequence. The Y-axis represents the error values, which are indicative of the quality of the non-bonded interactions for each segment of the protein. Lower error values suggest better quality and more reliable interactions. These structural validations collectively support the high quality and reliability of this predicted model of VP6. Further, to compare the quality of the SWISS-MODEL and AlphaFold2 predicted structures, we analysed their Ramachandran plots using the MolProbity server (Fig. S6A,B). Additionally, we performed pairwise structure alignment using the RCSB PDB alignment tool to assess the structural congruence with the reference structure (PDB ID: 1QHD) (Fig. S6C). The SWISS-MODEL predicted VP6 structure exhibited 97.2% (384/395) of all residues in favoured regions, 100.0% (395/395) in allowed regions, and no outliers. The AlphaFold2 predicted VP6 structure showed 97.5% (385/395) of all residues in favoured regions, 99.7% (394/395) in allowed regions, with one outlier (phi, psi): 130 Asp (55.0, 149.5). Pairwise structural alignment of models predicted by SWISS-MODEL and AlphaFold2 with the reference structure (PDB ID 1QHD) showed root mean square deviation (RMSD) values of 0.08 and 0.35, respectively, and a template modeling (TM)-score of 1 for both predicted structures. RMSD measures the alignment of backbone C-alpha atoms between superposed structures in Å, with lower values indicating better alignment. The TM-score assesses topological similarity between template and model structures, ranging from 0 to 1, where scores above 0.5 typically indicate a similar protein fold. These results indicate that both predicted models are comparable and of reliable quality. The SWISS-MODEL predicted VP6 structure was used for further analysis.Figure 3 Structure Validation of RVA VP6. (A) Predicted model of Human RVA VP6, (B) Ramachandran Plot. (C) ProSA with overall Z-score of − 8.42. (D) Verify 3D analysis with 87.91% of the residues averaging 3D–1D score ≥ 0.1. (E) ERRAT analysis with score 96.063.

Sequence-based and structure-based detection of APRs in VP6

TANGO utilizes a statistical algorithm that merges protein sequence information with pertinent physicochemical parameters. This combined methodology plays a pivotal role in discerning APRs within proteins. TANGO analysis identified five segments within VP6 as potential APRs, characterized by aggregation scores exceeding the predefined threshold (5%). These segments, as shown in Table 1 are indicative of their potential involvement in cross-β aggregation (Fig. 4a). Notably, the peptide VFTVASI exhibited the highest TANGO β-aggregation score, while the MIITM segment demonstrated the lowest score. To discern between buried and exposed APRs as predicted by TANGO, VP6 was analysed using the SolubiS server. The resulting stretch-plot, generated by SolubiS, depicted the aggregation propensity and local stability of APRs within VP6. Analysis revealed that none of the five APRs identified by TANGO exhibited a positive summed ΔG, indicating the absence of critical, surface-exposed APRs mediating native state aggregation. However, the segment spanning residues 245–249 (TTWFF), characterized by a least negative SolubiS score and therefore its low contribution to thermodynamic stability, was identified as a structural, buried APR implicated in non-native aggregation (Fig. 4b).Table 1 Summary of predicted APRs with amino acid residue positions and sequences in VP6 protein.

Tools	APR sequence	Tools	APR sequence	Tools	APR sequence	
APR region	APR region	APR region	
TANGO	FoldAmyloid	Camsol intrinsic	
3–7	VLYSL	1–6	MDVLYS	1–4	MDVL	
37–41	MIITM	35–40	NQMIIT	21	G	
217–225	VLTTATITL	55–65	PIRNWNFNFGL	25	S	
245–249	TTWFF	67–72	GTTLLN	33	Q	
385–391	VFTVASI	85–91	IDYFVDF	36–40	QMIIT	
AGGRESCAN	93–99	DNVCMDE	61–64	FNFG	
1–9	MDVLYSLSK	122–129	IKFKRINF	66–68	LGT	
32–39	QQFNQMII	137–141	ENWNL	86–90	DYFVD	
58–75	NWNFNFGLLGTTLLNLDA	178–183	GTMWLN	94	N	
86–94	DYFVDFVDN	212–219	VPLRRVLT	138	N	
119–125	LSAIKFK	232–236	FSFPR	159–164	PYSASF	
160–164	YSASF	246–255	TWFFNPVILR	178–182	GTMWL	
179–186	TMWLNAGS	261–265	VEFLL	196	S	
190–198	VAGFDYSCA	286–295	DTIRLSFQLM	198	A	
214–225	LRRVLTTATITL	305–309	AVLFP	211	I	
242–250	DGATTWFFN	322–326	LTLRI	220–224	TATIT	
257–270	NNVEVEFLLNGQII	366–370	NWTDL	235	P	
272–278	TYQARFG	383–390	QRVFTVAS	246–251	TWFFNP	
280–287	IVARNFDT	392–397	RSMLIK	262	E	
289–293	RLSFQ	AmylPred	266–267	NG	
303–308	AVAVLF	3–7	VLYSL	270–272	INT	
321–335	GLTLRIESAVCESVL	35–40	NQMIIT	278–280	GTI	
343–348	LANVTS	58–65	NWTFDFG	282	A	
350–355	RQEYAI	67–68	GT	288–293	IRLSFQ	
385–397	VFTVASIRSMLIK	85–93	IEYFIDFID	303–309	AVAVLFP	
Zyggregator	149–153	GFVFH	319–323	TVGLT	
11–15	LKDAR	178–185	GTMWLNAG	332–333	ES	
100–101	MV	188–189	IQ	344–346	ANV	
104–106	SQR	210–214	HIVQL	357–358	VG	
114–118	DSLRK	245–253	TTWFFNPII	385–389	VFTVA	
143–147	NRRQR	260–266	EVEFLLN	SolubiS	
170–171	QP	278–283	GTIVAR	245–249	TTWFF	
376–381	PSREDN	290–295	LSFQLM	CamSol structurally corrected	
18, 228, 230, 299	I, D, E, N	322–326	LTLRI	39, 163, 220, 235, 246, 248, 279	I, S, T, P, T, P, F	
Waltz	342–348	LLANVTA	Aggrescan3D	
1–7	MDVLYSL	367–373	WTDLITN	23–24	LY	
35–41	NQMIITM	383–394	QRVFTVASIRSM	55–56	PI	
57–67	RNWNFNFGLLG	ANuPP	64–65	GL	
73–92	LDANYVETARNTIDYFVDFV	18–43	IVEGTLYSNVSDLIQQFNQMIITMN	68–71	TTLL	
135–141	YIENWNL	84–97	TIEYFIDFIDNVC	121–122	GI	
207–212	QFEHIV	148–167	TGFVFHKPNIFPYSASFTL	159–164	PYSASF	
245–252	TTWFFNPV	173–215	HDNLMGTMWLNAGSEIQVAGFDYSCALNAPANIQQFEHIVQL	181–184	WLNA	
266–289	NGQIINTYQARFGTIVARNFDTIR	244–273	ATTWFFNPIILRPNNVEVEFLLNGQIINT	234–239	FPRVIT	
304–309	VAVLFP	301–339	TPAVNALFPQAQPFQHHATVGLTLRIESAVCESVLADA	252–253	VI	
370–374	LITNY	352–375	EYAIPVGPVFPPGMNWTDLITNYS	39, 117, 211, 226, 246, 248, 250, 278, 281	I, I, I, L, T, Y, N, G, I	
303–307	AVAAL	
356–358	PVG	

Figure 4 Prediction of APRs in VP6. (a) Schematic representation of APRs predicted by TANGO in VP6. TANGO aggregation propensity and SolubiS score were plotted against the protein sequence. (b) Stretch-plot showing APR aggregation propensity and thermodynamic stability in VP6. Positive summed ΔG indicates surface-exposed APRs, negative SolubiS score denotes structural, buried APRs.

AGGRESCAN predicts APRs by identifying primary sequences of query protein with protein fragments experimentally linked to disease-associated protein aggregation. According to the AGGRESCAN hot spot area data (Fig. 5a) prepared for the VP6, nineteen aggregation hotspots have been recognized in the VP6 protein (Table 1).Figure 5 Prediction of APRs in VP6. (a) AGGRESCAN normalized hot spot area plot. (b) Zyggregator scores (Ziagg). The brown line represents the aggregation tendency threshold. Residues above this threshold are predicted to contribute to fibrillar aggregate formation. (c) Amyloidogenic regions identified using WALTZ. (d) The amyloidogenic regions dentified using FoldAmyloid. (e) ANuPP showing predicted total aggregation scores of identified amyloidogenic regions. (f) Amyloidogenic regions identified by AmylPred.

Zyggregator considers physicochemical conditions like pH, temperature, ionic strength, and trifluoroethanol concentration on aggregation. It predicts local instabilities using the CamP program. Zyggregator also assesses “gatekeeper” residues flanking a sliding window for charged residues that could mitigate aggregation. Zyggregator identified 32 residues (Table 1) with the Ziagg score greater than the threshold value (Fig. 5b).

Analysis using Waltz algorithm revealed ten distinct regions as APRs within the VP6 sequence. These APRs, spanning residues as shown in Table 1, were identified as potential sites for amyloid formation (Fig. 5c). Prediction of APRs using FoldAmyloid revealed nineteen aggregation hotspots regions (Table 1) (Fig. 5d). ANuPP, utilizing atomic-level features derived from hexapeptides, predicted 7 APRs as shown in Table 1 (Fig. 5e). AmylPred server prediction relies on the concurrence of at least two out of five methods (refer “AmylPred” section). In VP6, Amylpred identified 17 consensus aggregation hits as shown in Table 1 (Fig. 5f).

The intrinsic solubility profile of VP6, assessed by CamSol, identified certain residues as poorly soluble (aggregation-prone) (Table 1) (Fig. 6). According to the Camsol structurally-corrected solubility profile, the residues I 39, S 163, T 220, P 235, T246, F 248, and T 279 have structurally corrected solubility scores (Fig. 6). These are the solvent-exposed poorly soluble residues of VP6. These residues represent potential positions for mutations aimed at enhancing solubility. The structurally-corrected solubility scores are derived from the intrinsic solubility scores, considering the proximity of the amino acids in the 3D structure and their solvent exposure. The Aggrescan3D (A3D) profile of VP6 provides a comprehensive assessment of residues implicated in aggregation propensity, utilizing spatial conformation and intermolecular interaction potentials within the protein structure. Notably, residues identified are shown in Table 1. Residues 39, 65, 70, 71, 160, 234, 237, 238, 248, 252, 253, and 357 demonstrate significant aggregation propensity as delineated by the A3D analysis (Fig. 7). A compiled information on all the predicted APRs have been provided in Table 1.Figure 6 Camsol intrinsic and Camsol structurally corrected solubility profiles predicted for VP6. The solvent-exposed poorly soluble amino acids in VP6 predicted by Camsol structurally corrected are shown with the black arrows and marked with the purple colour.

Figure 7 A3D structure-based analysis of VP6 aggregation propensity. Positive scores denote aggregation-prone residues, while negative scores indicate solubility-prone residues. The protein surface is color-coded based on the A3D score gradient: red for high aggregation propensity, white for negligible effect, and blue for high solubility regions.

A comprehensive investigation utilizing both sequence-based and structure-based approaches revealed a total of 24 consensus APRs within VP6 (Fig. 8a,b). These regions, defined as areas consistently predicted by two or more computational tools, offer insights into potential aggregation tendencies of the protein. Six APRs (1, 2, 3, 5, 7, 19, 21) were found to be part of helical structures, while five APRs (4, 6, 22, 23, 24) form a combination of random coil and helix. Additionally, five APRs (8, 9, 10, 12, 18) are located within beta sheet structures, and eight APRs (9, 11, 13, 14, 15, 16, 17, 20) are positioned in regions with a combination of beta sheet and random coil structures in the native VP6 3-D structure. Amino acid composition of APRs exhibited hydrophobic amino acid residues in the range of 33.33–100%, acidic amino acid residues from 0 to 30%, basic amino acid residues from 0 to 40% and neutral amino acid residues from 0 to 55.56% showing clear predominance of hydrophobic amino acid residues (Fig. 8c). The predominance of hydrophobic amino acid residues in the APRs suggests that these regions are likely to contribute to protein aggregation and reduced solubility. Hydrophobic interactions are known to drive protein aggregation, leading to the formation of insoluble aggregates such as inclusion bodies. Therefore, the high percentage of hydrophobic residues in the APRs indicates a propensity for reduced solubility of the protein. Notably, a substantial majority of these APRs, comprising 21 out of 24, were mapped to be surface-exposed in VP6 3-D structure (Fig. 8d).Figure 8 A comprehensive summary of all APRs identified by different tools in the RVA VP6 protein. (a) Summary of consensus APRs within RVA VP6, as predicted by various sequence-based and structure-based APR predictors. Each predictor's results are indicated by distinct colours, facilitating comparison. (b) Table displaying consensus APR sequences predicted by two or more tools and their positions in the RVA VP6 protein. Entries in white indicate totally surface-exposed APRs, those highlighted in green denote totally buried APRs, and light blue entries signify partially surface-exposed APRs. (c) Amino Acid Composition of APRs. (d) RVA VP6 structure is shown in both space-filling and cartoon models. Consensus APRs are highlighted in different colours, with their assigned APR number corresponding to their APR numbers from the table in (b).

Discussion

In this investigation, we evaluated the heterologous prokaryotic expression of codon-optimized rotavirus A VP6 in E. coli BL21(DE3) cells. Subsequent protein extraction and solubility analysis revealed an abundant presence of the recombinant protein within the insoluble fraction, contrasting with the preferred localization within the soluble fraction for efficient downstream processing, including protein purification, formulation, characterization, and functional studies. This phenomenon underscores a common challenge associated with high-level heterologous protein expression in bacterial systems, where the propensity for protein inclusion body formation presents a substantial impediment to protein research endeavours44. While protein recovery from inclusion bodies is indeed possible, the process entails labour-intensive and time-consuming procedures, resulting in inevitable protein loss at each stage. Additionally, the absence of a universal technique emphasizes the necessity for a trial-and-error approach to identify optimal strategies, thereby highlighting the intricate nature of protein handling in such instances. Therefore, to optimize soluble protein expression, a series of experimental conditions were systematically investigated for VP6 expression in the prokaryotic host. This included varying the concentration of the inducer (IPTG), post-induction temperature, post-induction incubation time with and without PMSF. These efforts were undertaken with the objective of enhancing protein solubility.

The inducer concentration profoundly influences the soluble protein fraction in E. coli expression systems. Higher concentrations typically yield elevated protein expression levels, potentially leading to a greater proportion of protein. Excessively high inducer concentrations can also lead to cellular stress or saturation of protein folding machinery, potentially causing misfolding, aggregation, and a decrease in solubility. Conversely, lower inducer concentrations may result in reduced protein expression but could enhance solubility by allowing for more controlled folding. However, excessively low inducer concentrations may not provide sufficient stimulation for protein expression, leading to inadequate levels of soluble protein. Thus, an optimal inducer concentration is crucial, striking a balance between robust protein expression and efficient folding to maximize soluble protein fraction while minimizing aggregation and cellular stress52. In the present study, the effect of varying IPTG concentrations on VP6 expression was examined, revealing localization in the insoluble fraction across all tested IPTG concentrations (Fig. S1).

Post-induction temperature in E. coli expression systems modulates protein solubility via its effects on protein folding kinetics, chaperone functionality, and potential protein denaturation. Optimal temperatures facilitate correct folding and enhanced solubility, while deviations from this range can induce protein misfolding or aggregation. Generally, protein solubility and activity are reported to be enhanced when expressed at lower temperatures in E. coli53. While VP6 expression level showed a slight increase with rising temperature (25 °C, 30 °C, and 37 °C) (Fig. S2).

Incubation times of 4, 6, 8, 10, and 12 h were assessed both with and without PMSF (0.2 mM) to investigate potential degradation during protein extraction of VP6 expressed in the soluble fraction. The results revealed that VP6 was not detectable in the soluble fraction under any of the assessed incubation times, regardless of the presence of PMSF (Fig. S3).

As a final approach, solubilization employing urea was undertaken, demonstrating efficacy solely at higher concentrations ranging from 6 to 8 M for the solubilization of VP6 (Fig. S4). However, given urea's protein-denaturing properties, a subsequent renaturation step was implemented utilizing on-column renaturation coupled with protein purification via Ni–NTA column chromatography. Subsequent visualization of the renatured protein via FETEM unveiled the formation of amorphous aggregates (Fig. 1). In contrast, the native VP6 protein demonstrates self-assembly characterized by structural polymorphism, resulting in the formation of trimers, hexagonal structures, and diverse architectures such as nanotubes, nanospheres, or nanosheets25–29.

Additionally, inclusion bodies generated at different temperatures underwent ThT binding assays. An augmentation in ThT fluorescence emission was observed corresponding to increasing incubation temperature, providing additional confirmation of aggregate formation within the inclusion bodies (Fig. 2). The analysis of ThT binding to VP6 inclusion bodies at varying incubation temperatures offers insights into the temperature sensitivity of protein aggregation processes associated with VP6 expression. An increase in ThT binding to VP6 inclusion bodies with incubation temperature suggests enhanced formation or stabilization of VP6 aggregates at higher incubation temperatures, indicating a temperature-dependent and potentially accelerated aggregation process. Conversely, a decrease in ThT binding with incubation temperature implies inhibition or destabilization of VP6 aggregates at lower temperatures. This suggests a less favorable or slower aggregation process, increasing the likelihood of proper folding of VP6 into its native structure. This finding is consistent with previous reports that have demonstrated how low expression temperatures promote the production of non-classical inclusion bodies, which exhibit significant biological activity54–56.

The observation of VP6 inclusion body formation during heterologous recombinant expression in E. coli, despite extensive optimization of experimental culture conditions, prompted us to investigate the potential root cause of aggregation: the APRs within the VP6 protein. Various computational methods have been developed to identify aggregation-prone regions in amyloidogenic proteins. While certain algorithms like TANGO, AGGRESCAN, AGGRESCAN3D, and CamSol primarily focus on predicting β-aggregation prone regions, others such as Waltz, Zyggregator, and Fold-amyloid target both β-aggregation and amyloid-fibril prone regions. In this study, we employed multiple tools utilizing diverse algorithms to analyse VP6, striving to alleviate biases due to training datasets, parameterization, and the distinct methodological traits inherent in each database. Finally, we shortlisted APRs that exhibited consensus identification by two or more prediction tools. The prediction tools utilized for this investigation have also been utilized for computational determination of APRs in human keratinocyte growth factor57, SARS‑CoV‑2 proteome58, tumour suppressor protein PTEN59, Naked Mole-Rat and Mouse proteome60 in previous reports.

The sequence-based APR prediction tools predicted multiple aggregation hotspot regions in VP6. A total of five, nineteen, ten, nineteen, seventeen, and seven aggregation hotspot sequences were identified by TANGO, AGGRESCAN, Waltz, FoldAmyloid, AmylPred, and ANuPP respectively. 32 residues were identified as aggregation prone residues by Zyggregator, while 30 regions involving residues as well as sequences were identified by Camsol intrinsic. The tools for predicting APRs based on structure restrict the quantity of anticipated APRs depending on the dynamically-exposed hydrophobic or aggregation-prone residues or regions. A majority of the total identified APRs by different tools exhibited overlaps with each other, and the ones in consensus with two or more tools were taken into consideration. In case of overlapping sequences, we considered the residues that were common within the overlapped region. This resulted in the consensus identification of total 24 APRs (Fig. 8). All the consensus epitopes were successfully mapped on the native VP6 three-dimensional structure to assess their surface accessibility. This revealed 21 APRs as solvent-exposed and 3 were found to be buried inside the protein. Furthermore, the consistent predominance of hydrophobic amino acids in all predicted APRs further elucidates their contribution to VP6 aggregation and, consequently, reduced solubility of the protein. Therefore, these 21 surface exposed APRs may promote misfolding or aggregation of VP6 during prokaryotic heterologous expression. Computational identification of these APRs corroborates our experimental observation of VP6 inclusion body or aggregate formation.

Conclusion

VP6 expression in E. coli BL21(DE3) led to its localizaiton in inclusion bodies despite exhaustive optimization of culture conditions. This indicates its inherent tendency for aggregation during prokaryotic expression. Computational investigation revealed 24 APRs, 21 of which are mapped as surface-exposed APRs in native VP6 structure, predominantly enriched in hydrophobic amino acid residues, corroborating our experimental observations. The identified regions can be considered as the candidate positions to create the VP6 mutant varieties with reduced aggregation propensity through site-directed mutagenesis in prokaryotic expression systems. It is noteworthy that the presence of exposed hydrophobic patches on VP6 is likely crucial for its intermolecular interactions that drive the formation of VP6 trimers and ultimately the assembly of the VP6 capsid structure. These patches may serve as critical interaction sites necessary for maintaining the structural integrity of the viral capsid. Concurrently, our identification of multiple APRs within VP6 provides a basis for potential mutagenesis strategies aimed at enhancing protein solubility, particularly for recombinant expression in E. coli. However, the decision to mutate these regions must carefully balance the goal of improving solubility with the imperative to preserve VP6's structural and immunological integrity within the viral capsid. This dual consideration is essential to avoid compromising VP6's functional roles in capsid assembly and its potential immunogenicity. The future scope of this research involves providing a detailed rationale for selecting mutation positions within APRs and experimentally validating these mutations. Specific mutation sites can be determined based on factors such as hydrophobicity, structural motifs, conservation analysis, surface accessibility relevant empirical evidence from prior studies. Subsequently, the study should progress to conduct mutagenesis, expression, and solubility analysis of VP6 mutants through recombinant expression in E. coli, aiming to ascertain the potential enhancement of VP6 solubility. Given the pivotal role of VP6 as a prospective rotavirus vaccine candidate owing to its high conservation and abundance in the rotavirus capsid structure, its successful solubilization promises to catalyse further advancements in pathological elucidation, therapeutic interventions, and diagnostic applications. This investigation represents the first comprehensive analysis of APRs within the rotavirus A VP6 capsid protein, marking a notable stride in elucidating its molecular complexities.

Materials and methods

Chemicals and reagents

The rotavirus A (RVA) VP6 gene sequence was codon optimized for bacterial expression system, synthesized, then cloned at the Xho I and BamH I restriction sites of the pET-28a(+) plasmid by Genscript®. Restriction enzymes were procured by New England Biolabs and Taq Red PCR mix was obtained from Sigma Aldrich respectively. The gene specific primers were synthesized from Bioserve (India). Nickel–Nitrilotriacetic acid (Ni–NTA) Hi-TRAP column was procured from GE healthcare. The mouse anti-RV inner capsid protein primary antibody and anti-mouse IgG specific-peroxidase secondary antibody was obtained from Novus Biologicals. The 3,3′-Diaminobenzidine was obtained from Amresco. The polyvinylidene fluoride (PVDF) membrane (Hybond P) was procured from Amershan. Thioflavin T was procured from Sigma Aldrich. All other chemicals used in the current work were of analytical grade. All solutions were prepared using nanopure water, MilliQ (18.2 MΩ cm), (Millipore Co., USA).

Cell transformation optimization of VP6 expression in E. coli BL21(DE3)

Competent E. coli BL21(DE3) cells were prepared by CaCl2 treatment. The pET-28a(+)-RVAVP6 recombinant plasmid was transformed into the prepared competent E. coli BL21(DE3) cells with a heat-shock treatment at 42 °C for 90 s. These transformed cells were grown in Luria Bertani (LB) broth (HiMedia, India) supplemented with kanamycin (Sigma) (50 μg/mL) on shaker-incubator at 37 °C (180 rpm) for 1 h followed by plating on Luria-Bertan agar plate supplemented with kanamycin (50 μg/mL), isopropyl beta-d-thiogalactoside (IPTG) (0.1 mM) and 5-bromo-4-chloro-3-indolyl-β-d-galactopyranoside (X-gal) (40 ug/mL) followed by incubation at 37 °C for about 16 h and observed for blue-white screening and validated with restriction digestion (NdeI and/or XhoI) and polymerase chain reaction (PCR) using gene specific primers.

The recombinant E. coli BL21(DE3) clones were subjected to optimization of RVAVP6 protein expression. Recombinant E. coli expression host, BL21(DE3), was inoculated in 5 mL (1:100) of LB medium supplemented 50 μg/mL of kanamycin and cultivated overnight at 37 °C in shaker-incubator. This preculture was used to inoculate (1:100) different sets of 10 mL LB medium, which were cultivated at 37 °C until reaching a 600 nm optical density (OD) of 0.6–0.8, followed by induction with IPTG (0.5 mM) and post induction temperature of 37 °C for 12 h. Cells were collected by centrifugation at 10,000g at 4 °C for 10 min, the culture medium was discarded and cells resuspended in lysis buffer containing 300 mM NaCl, 50 mM sodium phosphate buffer (SPB), pH 8. The homogenates were sonicated for 10 times for 60 s, with 30 s interval between each sonication, and then centrifuged at 12,000 rpm to separate supernatants and pellets. The crude supernatant and pellet fractions (pellets resuspended in resuspension buffer consisting of 300 mM NaCl, 50 mM SPB, pH 8) from bacterial protein extractions were confirmed by Sodium Dodecyl Sulphate–Polyacrylamide Gel Electrophoresis (SDS-PAGE) analysis.

To optimize the overexpression of the recombinant protein (VP6), several induction conditions—inducer concentration (0.0 mM to 1 M), post-induction temperature (18 °C, 25 °C, 30 °C and 37 °C), post-induction incubation time with and without PMSF treatment (0.2 mM, 0.4 mM, 0.6 mM and 0.8 mM) during protein extraction, and urea treatment were tested.

VP6 purification and western blot

The over-expressed recombinant VP6 protein localized as inclusion bodies in the pellet fraction or the insoluble fraction of the protein extract as revealed by SDS-PAGE analysis. Therefore, recombinant VP6 protein was purified by Ni–NTA affinity chromatography after solubilizing the inclusion bodies in a urea solubilization buffer (50 mM SPB, pH 8.0, 300 mM NaCl, 8 M Urea). Briefly, after equilibrating of the Ni–NTA column was carried out using lysis equilibration buffer (50 mM SPB, pH 8.0, 300 mM NaCl, 8 M Urea), followed by passage of an equal volume of lysis equilibration buffer with urea-treated cell lysate agarose column after equilibration. For on-column renaturation of VP6, the column was washed with five bed volumes of wash buffers with step-down urea concentrations (50 mM SPB, pH 8.0, 300 mM NaCl, 8–0 M Urea, 160 mM Imidazole). The recombinant VP6 protein was subsequently eluted with two bed volumes of elution buffer (50 mM SPB, pH 8.0, 300 mM NaCl, 500 mM Imidazole). The eluates collected at each step were analyzed by SDS-PAGE. The elution with purified recombinant VP6 was subjected to Western Blotting.

Briefly, recombinant VP6 protein was subjected to SDS-PAGE and protein bands were transferred over Polyvinylidene difluoride (PVDF) membrane (MDI, India) using constant voltage of 25 V, 300 mA, for 3–4 h at 4 °C. The unoccupied sites on PVDF membrane were blocked by 3% BSA at room temperature for 3 h. The protein bands were allowed to bind with mouse anti-RV inner capsid protein monoclonal antibody (NB110-37243) (1:1000 in 3% BSA) at 4 °C overnight, followed by incubation with anti-mouse IgG specific-peroxidase secondary antibody (1:5000 in 3% BSA) after washing with PBST (0.1% Tween-20 in PBS). The protein bands were detected using 3,3-Diaminobenzidine (DAB) (Sigma-Aldrich, USA).

Field emission transmission electron microscopy (FETEM) of VP6 inclusion bodies and purified VP6

Samples for FETEM analysis constituted the VP6 inclusion body and purified renatured VP6 protein. Each sample was diluted, and then 5 μL was placed on a carbon-coated copper grid, followed by staining with 2% uranyl acetate solution which was dried at room temperature using vacuum pump desiccator for 8 h. The grids with the sample were examined on JEOL JEM-2100F at an accelerating voltage of 100 kV.

Thioflavin T (ThT) binding assay

VP6 was expressed in BL21(DE3) at four different temperatures separately with optimized inducer concentration for 8 h. Cells were harvested by centrifugation at 10,000g for 10 min (4 °C). Cell pellets were resuspended in buffer 1 (50 mM SPB, pH 8.0, 300 mM NaCl) to form suspensions of equal OD600 nm. Inclusion bodies were purified using several sonication and washing steps. Purified VP6 inclusion bodies purified from equal quantity of cells were then subjected to ThT assay to study amyloid nature of VP6 inclusion bodies expressed at different temperatures. Briefly, inclusion bodies expressed at various temperatures were diluted in 50 mM SPB (pH 8.0) containing 300 mM NaCl to achieve uniform optical density suspensions (measured at 350 nm, OD350 nm = 0.3). Purified inclusion bodies were incubated with 50 mM ThT for 15 min. Fluorescence spectra were recorded on FluoroMax-4 (HORIBA Scientific). Samples were excited at 440 nm and emission spectra were recorded in the range of 450–650 nm with excitation and emission slit widths set to 5 nm. Data were smoothed using 5-point Savitzky-Golay smoothing. ThT without protein served as negative control. Difference spectra (after deducting spectra of negative control) were used for analysis. The final analysis was based on the average of three independent spectra.

Physicochemical characterization and structure prediction of rotavirus A VP6 protein

Human rotavirus A VP6 protein sequence (accession number YP_002302229.1, Fig. S7) in FASTA format was retrieved from the National Center for Biotechnology Information (NCBI) (https://www.ncbi.nlm.nih.gov/)61. The physicochemical characterization was performed using ProtParam Expasy tool (https://web.expasy.org/protparam/)62. The Protein-Sol webserver (https://protein-sol.manchester.ac.uk/) was used to predict protein solubility of VP6 in E. coli expression system63.

The three-dimensional structure of the chosen VP6 protein sequence has not been determined experimentally to date. However, crystal structure of VP6 isolated from Bovine rotavirus strain RF is available in RSCB PDB (1QHD)64. This structure exhibits 96.98% amino acid sequence identity, as confirmed by NCBI protein BLAST analysis and has therefore been used as a template for modeling the structure of VP6 under study using SWISS-MODEL65 (https://swissmodel.expasy.org). The generated three dimensional structure was subjected to structural validation using the tools Procheck67, ProSA68, Verify 3D69 and ERRAT70 and were visualized using the PyMOL Molecular Graphics System, Version 3.0 Schrödinger, LLC. Further, this structure was compared with the one generated using AlphaFold266 (https://cosmic-cryoem.org/tools/alphafold2/). To compare the quality of the predicted structures, we analyzed their Ramachandran plots using the MolProbity server (http://molprobity.biochem.duke.edu/). Additionally, we performed pairwise structure alignment using the RCSB PDB alignment tool (rcsb.org/alignment) to assess the structural congruence with the reference structure (PDB ID 1QHD).

Sequence-based detection of APRs in VP6

The amino acid sequence of VP6 was subjected to analysis using a range of sequence-based APR detection tools. These tools employ diverse prediction methodologies, which are described in the subsequent sections.

TANGO

TANGO employs a statistical mechanics-based algorithm to predict the nucleating regions for protein aggregation. Additionally, it analyzes the impact of mutations and environmental conditions on the propensity of these regions to undergo aggregation71. TANGO conducts computations to evaluate the occurrence of cross-beta aggregation in peptides and denatured proteins. It considers the relative propensities of different structural states, including b-turn, alpha-helix, and beta-sheet, which compete with each other71. To determine the propensity of each conformation, TANGO utilizes a partition function and predicts cross-beta aggregation. This prediction assumes that the core regions involved in the aggregation process are fully buried and fulfil their hydrogen-bond potential. The TANGO aggregation score considers environmental factors such as stability, pH, ionic strength, protein concentration, and the presence of denaturant trifluoroethanol (TFE). The default parameters of pH, temperature, and ionic strength were used for APR prediction of VP6. In terms of data interpretation, any residue exhibiting an aggregation score exceeding 5% over a span of 5–6 residues is considered a potential APR.

AGGRESCAN

AGGRESCAN functions as a computational framework aimed at the recognition of aggregation-nucleating segments, commonly referred to as "aggregation hot spots," in polypeptides. This platform relies on a scale of aggregation propensity, which is established based on empirically validated hotspots observed in a diverse assortment of proteins, encompassing both natively unfolded and pathogenic variants such as Ab42, synuclein, prion, and amylin. By harnessing this scale, AGGRESCAN enables the precise identification and characterization of specific regions within protein sequences that exhibit a heightened likelihood for aggregation. For APR utilizing AGGRESCAN, the sequence of VP6 was submitted to the AGGRESCAN database (http://bioinf.uab.es/aggrescan/)49. Various parameters were obtained, including the aggregation-propensity value (a3v) for each amino acid, the average of a3v within a sliding window (a4v), a graphical representation of the protein's aggregation profile, the area of the aggregation profile above the hotspot threshold (HST) in a specific hotspot, and a graphical depiction of the peak area. Additionally, putative aggregation hot spots were identified, defined as regions containing five or more residues with an a4v value surpassing the HST. The Hotspot Area (HSA) was calculated as the aggregation profile area divided by the number of residues in the entry sequence. Furthermore, the Normalized Hotspot Area (NHSA) per residue was determined by dividing the HSA by the number of residues in the entry sequence.

Zyggregator

Zyggregator is a method based on the protein's sequence that computes various characteristics, including the local stability of the individual units (ln Pi), the formation likelihood of β-rich oligomers (Zitox), and the propensity for forming fibrillar aggregates (Ziagg). These assessments aid in predicting the inherent amyloid aggregation tendency of the protein. The Ziagg score of 0 corresponds to an aggregation propensity equivalent to that of a random sequence at position i, while a score of 1 indicates a one standard deviation increase in the propensity for aggregation. The presence of a Ziagg score of 1 or higher indicates the involvement of aggregation-prone residues. For the analysis of VP6 using Zyggregator, the sequence of VP6 was submitted to the Zyggregator database, and the aforementioned parameters were computed. Subsequently, a graphical representation was created based on the Ziagg scores of the individual residues72.

Waltz

WALTZ is a computational algorithm designed to predict amyloids in proteins effectively. It demonstrates the ability to differentiate genuine amyloid structures from disordered amorphous aggregates. It is capable of discerning challenging amyloid aggregates, such as those found in prion diseases, which are typically difficult to predict73. WALTZ employs a comprehensive three-component scoring function. Firstly, it utilizes a position-specific scoring matrix (PSSM) focused on hexapeptides to identify regions prone to amyloid formation. The PSSM is derived from the AmylHex dataset, enriched with numerous experimentally determined amyloidogenic and non-amyloidogenic sequences. Secondly, WALTZ considers the physicochemical properties of amino acids, including hydrophobicity and β-structure-forming propensity. Lastly, it incorporates a position-specific pseudoenergy matrix, obtained from the Sup35 GNNQQNY peptide's crystal structure, which consist of cross-β spine of amyloid fibrils. This matrix, in combination with FoldX, estimates the relative energies involved in the process74,75. The training of the WALTZ algorithm involved an extensive dataset comprising experimentally validated amyloid-forming peptides. Approximately eighty percent of the dataset is validated using techniques such as circular dichroism, electron microscopy, infrared spectroscopy (FTIR), and Thioflavin-T binding assays75. The prediction of amyloid-forming peptides can be conducted using any of the three pre-existing thresholds available in the WALTZ database, namely “Best overall performance”, “High specificity”, and “High sensitivity”. Alternatively, customized threshold may also be set. To analyze VP6 using WALTZ, the amino acid sequence of VP6 in FASTA format was submitted to the online WALTZ web server (http://waltz.switchlab.org/) using “Best overall performance” as a threshold.

FoldAmyloid

The FoldAmyloid webserver predicts amyloidogenic regions in protein sequences based on two key characteristics: expected probability of hydrogen bond formation and expected packing density of residues. High probabilities of backbone-backbone hydrogen bond formation and dense packing are indicative of amyloidogenic potential. The FoldAmyloid server is accessible at http://antares.protres.ru/fold-amyloid/76. The primary protein sequence of VP6 was submitted to FoldAmyloid webserver with an averaging frame of 5 and threshold set to 21.4.

ANuPP

ANuPP is an ensemble classifier designed to discern potential APRs within peptides and proteins, employing atomic-level features derived from hexapeptides. This method demonstrated an accuracy rate of 83% in its predictions. The ANuPP server is accessible at https://web.iitm.ac.in/bioinfo2/ANuPP/77. RVAVP6 primary amino acid sequence was submitted in FASTA format with a default threshold of 0.52.

Camsol intrinsic

CamSol comprises two distinct algorithms, one based on sequence analysis and the other on protein structure analysis, to accurately determine protein solubility. The initial algorithm, referred to as CamSolintrinsic, focuses on predicting the inherent solubility of the unfolded state of the protein using the amino acid sequence of a protein. The intrinsic solubility analysis of VP6 protein was conducted by submitting its amino acid sequence to CamSolintrinsic web server (http://www-vendruscolo.ch.cam.ac.uk/camsolmethod.html)78.

Amylpred

Amylpred employs a consensus approach combining multiple strategies to predict amyloid fibril formation features associated with diseases like Alzheimer’s, Parkinson’s, type II diabetes, and prion disease. This tool also assists in understanding protein folding/misfolding properties and controlling aggregation/solubility in biopharmaceuticals. Amylpred integrates predictions from five methods: Average Packing Density, Possible Conformational Switches, Amyloidogenic Pattern, TANGO, and Hexapeptide Conformational Energy. Amylpred webserver is accessible at http://aias.biol.uoa.gr/AMYLPRED/. The primary amino acid sequence of the query protein served as the input79.

Structure based detection of APRs

Three tools were utilized for the structure-based prediction of APRs in VP6, as outlined in the following section.

SolubiS

SolubiS is a comprehensive method developed to identify APRs characterized by high aggregation propensity and low thermodynamic stability using the protein's three-dimensional structure. This approach integrates the TANGO algorithm for predicting β-aggregation-prone regions and utilizes the FoldX empirical force field to assess protein stability. SolubiS distinguishes between two types of APRs: structural APRs, which are protected from aggregation-triggering events by folding, and critical APRs, which can induce aggregation without major unfolding80. The three-dimensional structure of VP6 was submitted to SolubiS web-based tool (http://solubis.switchlab.org/) and stretch-plots representing APRs based on the TANGO summed score were generated.

Camsol structurally corrected

CamSol structurally corrected implements structural corrections on three-dimensional structure of protein to enhance intrinsic solubility and compute protein solubility based on amino acid proximity and solvent exposure within the structural context78. The structural analysis of the VP6 PDB file served as an input for CamSol. Further details regarding CamSol can be accessed at http://www-vendruscolo.ch.cam.ac.uk/camsolmethod.html.

Aggrescan3D (A3D)

AGGRESCAN3D is a method designed to predict APRs within natively folded proteins by utilizing structural analysis in both static and dynamic states. This approach focuses on identifying surface-exposed APRs based on amino acid positioning and structural characteristics. During the static mode calculation of AGGRESCAN3D (A3D), the input protein structure is first minimized using the FoldX force field. Following this, the intrinsic aggregation propensity score of each amino acid is adjusted based on its structural context, resulting in a structurally corrected aggregation score for each residue. The 3-D structure of VP6 in PDB format served as the input for AGGRESCAN3D server (http://biocomp.chem.uw.edu.pl/A3D/), using a radius of 10 Å. All the APRs were subjected to amino acid composition analysis using PEPTIDE 2.0 webserver (https://www.peptide2.com/N_peptide_hydrophobicity_hydrophilicity.php).

Supplementary Information

Supplementary Figures.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-69896-1.

Acknowledgements

We acknowledge the financial assistance (Grant No. BT/PR41449/NER/95/1687/2020) of DBT, India, and computational facility provided by Computer Centre, IIT Guwahati to carry out the work.

Author contributions

Conceptualization, P.G and P.R.K.; methodology, analysis, and visualization, P.R.K.; investigation and data curation, P.R.K.; writing—original draft preparation, and writing—review and editing, P.G. and P.R.K.; resources, supervision, project administration, and funding acquisition, P.G. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the DBT, India under Grant BT/PR41449/NER/95/1687/2020.

Data availability

All data generated or analyzed during this study are included in this published article [and its supplementary information files].

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Matthijnssens J Otto PH Ciarlet M Desselberger U Van Ranst M Johne R VP6-sequence-based cutoff values as a criterion for rotavirus species demarcation Arch. Virol. 2012 157 1177 1182 10.1007/s00705-012-1273-3 22430951
Matthijnssens, J. et al. VP6-sequence-based cutoff values as a criterion for rotavirus species demarcation. Arch. Virol. 157, 1177–1182 (2012).22430951 10.1007/s00705-012-1273-3
2. Dogs S Candidate new rotavirus species in sheltered dogs, Hungary Emerg. Infect. Dis. 2015 21 4 7
Dogs, S. et al. Candidate new rotavirus species in sheltered dogs, Hungary. Emerg. Infect. Dis. 21, 4–7 (2015).
3. Bányai K Candidate new rotavirus species in Schreiber’s bats, Serbia Infect. Genet. Evol. 2020 48 19 26 10.1016/j.meegid.2016.12.002
Bányai, K. et al. Candidate new rotavirus species in Schreiber’s bats, Serbia. Infect. Genet. Evol. 48, 19–26 (2020).10.1016/j.meegid.2016.12.002
4. Bishop RF Davidson GP Holmes IH Ruck BJ Virus particles in epithelial cells of duodenal mucosa from children with acute non-bacterial gastroenteritis Lancet 1973 302 1281 1283 10.1016/S0140-6736(73)92867-5
Bishop, R. F., Davidson, G. P., Holmes, I. H. & Ruck, B. J. Virus particles in epithelial cells of duodenal mucosa from children with acute non-bacterial gastroenteritis. Lancet 302, 1281–1283 (1973).10.1016/S0140-6736(73)92867-5
5. Parashar UD Gibson CJ Bresee JS Glass RI Rotavirus and severe childhood diarrhea Emerg. Infect. Dis. 2006 12 304 306 10.3201/eid1202.050006 16494759
Parashar, U. D., Gibson, C. J., Bresee, J. S. & Glass, R. I. Rotavirus and severe childhood diarrhea. Emerg. Infect. Dis. 12, 304–306 (2006).16494759 10.3201/eid1202.050006
6. Tate JE 2008 Estimate of worldwide rotavirus-associated mortality in children younger than 5 years before the introduction of universal rotavirus vaccination programmes: A systematic review and meta-analysis Lancet Infect. Dis. 2012 12 136 141 10.1016/S1473-3099(11)70253-5 22030330
Tate, J. E. et al. 2008 Estimate of worldwide rotavirus-associated mortality in children younger than 5 years before the introduction of universal rotavirus vaccination programmes: A systematic review and meta-analysis. Lancet Infect. Dis. 12, 136–141 (2012).22030330 10.1016/S1473-3099(11)70253-5
7. Tamminen K Lappalainen S Huhti L Vesikari T Blazevic V Trivalent combination vaccine induces broad heterologous immune responses to norovirus and rotavirus in mice PLoS ONE 2013 8 e70409 10.1371/journal.pone.0070409 23922988
Tamminen, K., Lappalainen, S., Huhti, L., Vesikari, T. & Blazevic, V. Trivalent combination vaccine induces broad heterologous immune responses to norovirus and rotavirus in mice. PLoS ONE 8, e70409 (2013).23922988 10.1371/journal.pone.0070409
8. Schwartz-cornil I Benureau Y Greenberg H Hendrickson BA Cohen J Heterologous protection induced by the inner capsid proteins of rotavirus requires transcytosis of mucosal immunoglobulins J. Virol. 2002 76 8110 8117 10.1128/JVI.76.16.8110-8117.2002 12134016
Schwartz-cornil, I., Benureau, Y., Greenberg, H., Hendrickson, B. A. & Cohen, J. Heterologous protection induced by the inner capsid proteins of rotavirus requires transcytosis of mucosal immunoglobulins. J. Virol. 76, 8110–8117 (2002).12134016 10.1128/JVI.76.16.8110-8117.2002
9. Afchangi A Jalilvand S Mohajel N Marashi SM Shoja Z Rotavirus VP6 as a potential vaccine candidate Rev. Med. Virol. 2019 29 2 e2027 10.1002/rmv.2027 30614135
Afchangi, A., Jalilvand, S., Mohajel, N., Marashi, S. M. & Shoja, Z. Rotavirus VP6 as a potential vaccine candidate. Rev. Med. Virol. 29(2), e2027 (2019).30614135 10.1002/rmv.2027
10. Chen SC Protective immunity induced by oral immunization with a rotavirus DNA vaccine encapsulated in microparticles J. Virol. 1998 72 5757 5761 10.1128/JVI.72.7.5757-5761.1998 9621034
Chen, S. C. et al. Protective immunity induced by oral immunization with a rotavirus DNA vaccine encapsulated in microparticles. J. Virol. 72, 5757–5761 (1998).9621034 10.1128/JVI.72.7.5757-5761.1998
11. Jalilvand S Mahdi S Shoja Z Rotavirus VP6 preparations as a non-replicating vaccine candidates Vaccine 2015 33 3281 3287 10.1016/j.vaccine.2015.05.026 26021725
Jalilvand, S., Mahdi, S. & Shoja, Z. Rotavirus VP6 preparations as a non-replicating vaccine candidates. Vaccine 33, 3281–3287 (2015).26021725 10.1016/j.vaccine.2015.05.026
12. Kuri PR Goswami P Current update on rotavirus in-silico multiepitope vaccine design ACS Omega 2023 8 190 207 10.1021/acsomega.2c07213 36643547
Kuri, P. R. & Goswami, P. Current update on rotavirus in-silico multiepitope vaccine design. ACS Omega 8, 190–207 (2023).36643547 10.1021/acsomega.2c07213
13. Afchangi A Immunization of mice by rotavirus NSP4-VP6 fusion protein elicited stronger responses compared to VP6 alone Viral Immunol. 2017 31 233 241 10.1089/vim.2017.0075 29185875
Afchangi, A. et al. Immunization of mice by rotavirus NSP4-VP6 fusion protein elicited stronger responses compared to VP6 alone. Viral Immunol. 31, 233–241 (2017).29185875 10.1089/vim.2017.0075
14. Feng H Oral administration of a seed-based bivalent rotavirus vaccine containing VP6 and NSP4 induces specific immune responses in mice Front. Plant Sci. 2017 8 910 10.3389/fpls.2017.00910 28620404
Feng, H. et al. Oral administration of a seed-based bivalent rotavirus vaccine containing VP6 and NSP4 induces specific immune responses in mice. Front. Plant Sci. 8, 910 (2017).28620404 10.3389/fpls.2017.00910
15. Lappalainen S Pastor AR Malm M López-Guerrero V Esquivel-Guadarrama F Palomares LA Vesikari T Blazevic V Protection against live rotavirus challenge in mice induced by parenteral and mucosal delivery of VP6 subunit rotavirus vaccine Arch. Virol. 2015 160 2075 2078 10.1007/s00705-015-2461-8 26016444
Lappalainen, S. et al. Protection against live rotavirus challenge in mice induced by parenteral and mucosal delivery of VP6 subunit rotavirus vaccine. Arch. Virol. 160, 2075–2078 (2015).26016444 10.1007/s00705-015-2461-8
16. Vega CG IgY antibodies protect against human rotavirus induced diarrhea in the neonatal gnotobiotic piglet disease model PLoS ONE 2012 7 e42788 10.1371/journal.pone.0042788 22880110
Vega, C. G. et al. IgY antibodies protect against human rotavirus induced diarrhea in the neonatal gnotobiotic piglet disease model. PLoS ONE 7, e42788 (2012).22880110 10.1371/journal.pone.0042788
17. Blazevic V Lappalainen S Nurminen K Huhti L Vesikari T Norovirus VLPs and rotavirus VP6 protein as combined vaccine for childhood gastroenteritis Vaccine 2011 29 8126 8133 10.1016/j.vaccine.2011.08.026 21854823
Blazevic, V., Lappalainen, S., Nurminen, K., Huhti, L. & Vesikari, T. Norovirus VLPs and rotavirus VP6 protein as combined vaccine for childhood gastroenteritis. Vaccine 29, 8126–8133 (2011).21854823 10.1016/j.vaccine.2011.08.026
18. Shoja Z Jalilvand S Latifi T Roohvand F Rotavirus VP6: Involvement in immunogenicity, adjuvant activity, and use as a vector for heterologous peptides, drug delivery, and production of nano-biomaterials Arch. Virol. 2022 167 1013 1023 10.1007/s00705-022-05407-9 35292854
Shoja, Z., Jalilvand, S., Latifi, T. & Roohvand, F. Rotavirus VP6: Involvement in immunogenicity, adjuvant activity, and use as a vector for heterologous peptides, drug delivery, and production of nano-biomaterials. Arch. Virol. 167, 1013–1023 (2022).35292854 10.1007/s00705-022-05407-9
19. Gautam R Detection of rotavirus antigen in stool specimens J. Clin. Virol. 2015 58 292 294 10.1016/j.jcv.2013.06.022
Gautam, R. et al. Detection of rotavirus antigen in stool specimens. J. Clin. Virol. 58, 292–294 (2015).10.1016/j.jcv.2013.06.022
20. Kaplon J Diagnostic accuracy of seven commercial assays for rapid detection of group a rotavirus antigens J. Clin. Microbiol. 2015 53 3670 3673 10.1128/JCM.01984-15 26378280
Kaplon, J. et al. Diagnostic accuracy of seven commercial assays for rapid detection of group a rotavirus antigens. J. Clin. Microbiol. 53, 3670–3673 (2015).26378280 10.1128/JCM.01984-15
21. Jayaram H Estes MK Prasad BVV Emerging themes in rotavirus cell entry, genome organization, transcription and replication Virus Res. 2004 101 67 81 10.1016/j.virusres.2003.12.007 15010218
Jayaram, H., Estes, M. K. & Prasad, B. V. V. Emerging themes in rotavirus cell entry, genome organization, transcription and replication. Virus Res. 101, 67–81 (2004).15010218 10.1016/j.virusres.2003.12.007
22. Crawford SUEE Trypsin cleavage stabilizes the rotavirus VP4 spike J. Virol. 2001 75 6052 6061 10.1128/JVI.75.13.6052-6061.2001 11390607
Crawford, S. U. E. E. et al. Trypsin cleavage stabilizes the rotavirus VP4 spike. J. Virol. 75, 6052–6061 (2001).11390607 10.1128/JVI.75.13.6052-6061.2001
23. Settembre EC Chen JZ Dormitzer PR Grigorieff N Harrison SC Atomic model of an infectious rotavirus particle EMBO J. 2010 30 408 416 10.1038/emboj.2010.322 21157433
Settembre, E. C., Chen, J. Z., Dormitzer, P. R., Grigorieff, N. & Harrison, S. C. Atomic model of an infectious rotavirus particle. EMBO J. 30, 408–416 (2010).21157433 10.1038/emboj.2010.322
24. Charpilienne A Lepault J Rey F Cohen J Identification of rotavirus VP6 residues located at the interface with VP2 that are essential for capsid assembly and transcriptase activity J. Virol. 2002 76 7822 7831 10.1128/JVI.76.15.7822-7831.2002 12097594
Charpilienne, A., Lepault, J., Rey, F. & Cohen, J. Identification of rotavirus VP6 residues located at the interface with VP2 that are essential for capsid assembly and transcriptase activity. J. Virol. 76, 7822–7831 (2002).12097594 10.1128/JVI.76.15.7822-7831.2002
25. Estes MK Rotaviruses 2007 Lippincott Williams & Wilkins
Estes, M. K. Rotaviruses (Lippincott Williams & Wilkins, 2007).
26. Estes MK Synthesis and immunogenicity of the rotavirus major capsid antigen using a baculovirus expression system J. Virol. 1987 61 1488 1494 10.1128/jvi.61.5.1488-1494.1987 3033276
Estes, M. K. et al. Synthesis and immunogenicity of the rotavirus major capsid antigen using a baculovirus expression system. J. Virol. 61, 1488–1494 (1987).3033276 10.1128/jvi.61.5.1488-1494.1987
27. Bugli F Caprettini V Cacaci M Martini C Paroni Sterbini F Torelli R Della Longa S Papi M Palmieri V Giardina B Posteraro B Sanguinetti M Arcovito A Synthesis and characterization of different immunogenic viral nanoconstructs from rotavirus VP6 inner capsid protein Int. J. Nanomed. 2014 9 2727 2739
Bugli, F. et al. Synthesis and characterization of different immunogenic viral nanoconstructs from rotavirus VP6 inner capsid protein. Int. J. Nanomed. 9, 2727–2739 (2014).
28. Choi AH Basu M Neal MMMC Clements JD Ward RL Antibody-independent protection against rotavirus infection of mice stimulated by intranasal immunization with chimeric VP4 or VP6 protein J. Virol. 1999 73 7574 7581 10.1128/JVI.73.9.7574-7581.1999 10438847
Choi, A. H., Basu, M., Neal, M. M. M. C., Clements, J. D. & Ward, R. L. Antibody-independent protection against rotavirus infection of mice stimulated by intranasal immunization with chimeric VP4 or VP6 protein. J. Virol. 73, 7574–7581 (1999).10438847 10.1128/JVI.73.9.7574-7581.1999
29. Kapikian AZ Hoshino Y Reoviridae 2001 Lippincott Williams & Wilkins
Kapikian, A. Z. & Hoshino, Y. Reoviridae (Lippincott Williams & Wilkins, 2001).
30. Lepault J Petitpas I Erk I Navaza J Bigot D Dona M Vachette P Cohen J Rey FA Structural polymorphism of the major capsid protein of rotavirus EMBO J. 2001 20 1498 1507 10.1093/emboj/20.7.1498 11285214
Lepault, J. et al. Structural polymorphism of the major capsid protein of rotavirus. EMBO J. 20, 1498–1507 (2001).11285214 10.1093/emboj/20.7.1498
31. Ready KFM Sabara M In vitro assembly of bovine rotavirus nucleocapsid protein Virology 1987 157 189 198 10.1016/0042-6822(87)90328-X 3029958
Ready, K. F. M. & Sabara, M. In vitro assembly of bovine rotavirus nucleocapsid protein. Virology 157, 189–198 (1987).3029958 10.1016/0042-6822(87)90328-X
32. Nanobiotechnol J A milk-based self-assemble rotavirus VP6—ferritin nanoparticle vaccine elicited protection against the viral infection J. Nanobiotechnol. 2019 10.1186/s12951-019-0446-6
Nanobiotechnol, J. et al. A milk-based self-assemble rotavirus VP6—ferritin nanoparticle vaccine elicited protection against the viral infection. J. Nanobiotechnol.10.1186/s12951-019-0446-6 (2019).10.1186/s12951-019-0446-6
33. Zhao Q Self-assembled virus-like particles from rotavirus structural protein VP6 for targeted drug delivery Bioconjug. Chem. 2011 22 346 352 10.1021/bc1002532 21338097
Zhao, Q. et al. Self-assembled virus-like particles from rotavirus structural protein VP6 for targeted drug delivery. Bioconjug. Chem. 22, 346–352 (2011).21338097 10.1021/bc1002532
34. Bredell H Smith JJ Prins WA Görgens JF van Zyl WH Expression of rotavirus VP6 protein: A comparison amongst Escherichia coli, Pichia pastoris and Hansenula polymorpha FEMS Yeast Res. 2016 16 foW001 10.1093/femsyr/fow001 26772798
Bredell, H., Smith, J. J., Prins, W. A., Görgens, J. F. & van Zyl, W. H. Expression of rotavirus VP6 protein: A comparison amongst Escherichia coli, Pichia pastoris and Hansenula polymorpha. FEMS Yeast Res. 16, foW001 (2016).26772798 10.1093/femsyr/fow001
35. Kato T Expression and purification of porcine rotavirus structural proteins in silkworm larvae as a vaccine candidate Mol. Biotechnol. 2023 65 401 409 10.1007/s12033-022-00548-3 35963985
Kato, T. et al. Expression and purification of porcine rotavirus structural proteins in silkworm larvae as a vaccine candidate. Mol. Biotechnol. 65, 401–409 (2023).35963985 10.1007/s12033-022-00548-3
36. da Silva Junior HC da Silva E Mouta Junior S de Mendonça MC de Souza Pereira MC da Rocha Nogueira A de Azevedo ML Leite JP de Moraes MT Comparison of two eukaryotic systems for the expression of VP6 protein of rotavirus species A: Transient gene expression in HEK293-T cells and insect cell-baculovirus system Biotechnol. Lett. 2012 34 1623 1627 10.1007/s10529-012-0946-z 22576283
da Silva Junior, H. C. et al. Comparison of two eukaryotic systems for the expression of VP6 protein of rotavirus species A: Transient gene expression in HEK293-T cells and insect cell-baculovirus system. Biotechnol. Lett. 34, 1623–1627 (2012).22576283 10.1007/s10529-012-0946-z
37. Lappalainen S Vesikari T Blazevic V Simple and efficient ultrafiltration method for purification of rotavirus VP6 oligomeric proteins Arch. Virol. 2016 161 3219 3223 10.1007/s00705-016-2991-8 27518400
Lappalainen, S., Vesikari, T. & Blazevic, V. Simple and efficient ultrafiltration method for purification of rotavirus VP6 oligomeric proteins. Arch. Virol. 161, 3219–3223 (2016).27518400 10.1007/s00705-016-2991-8
38. Zhou B Oral administration of plant-based rotavirus VP6 induces antigen-specific IgAs, IgGs and passive protection in mice Vaccine 2010 28 6021 6027 10.1016/j.vaccine.2010.06.094 20637305
Zhou, B. et al. Oral administration of plant-based rotavirus VP6 induces antigen-specific IgAs, IgGs and passive protection in mice. Vaccine 28, 6021–6027 (2010).20637305 10.1016/j.vaccine.2010.06.094
39. Chung IS Production of recombinant rotavirus VP6 from a suspension culture of transgenic tomato (Lycopersicon esculentum Mill.) cells Biotechnol. Lett. 2000 22 251 255 10.1023/A:1005626000329
Chung, I. S. et al. Production of recombinant rotavirus VP6 from a suspension culture of transgenic tomato (Lycopersicon esculentum Mill.) cells. Biotechnol. Lett. 22, 251–255 (2000).10.1023/A:1005626000329
40. O’Brien GJ Bryant CJ Voogd C Greenberg HB Gardner RC Bellamy A Rotavirus VP6 expressed by PVX vectors in nicotiana benthamiana coats PVX rods and also assembles into viruslike particles Virology 2000 270 444 453 10.1006/viro.2000.0314 10793003
O’Brien, G. J. et al. Rotavirus VP6 expressed by PVX vectors in nicotiana benthamiana coats PVX rods and also assembles into viruslike particles. Virology 270, 444–453 (2000).10793003 10.1006/viro.2000.0314
41. Yin J Li G Ren X Herrler G Select what you need: A comparative evaluation of the advantages and limitations of frequently used expression systems for foreign genes J. Biotechnol. 2007 127 335 347 10.1016/j.jbiotec.2006.07.012 16959350
Yin, J., Li, G., Ren, X. & Herrler, G. Select what you need: A comparative evaluation of the advantages and limitations of frequently used expression systems for foreign genes. J. Biotechnol. 127, 335–347 (2007).16959350 10.1016/j.jbiotec.2006.07.012
42. Jahn TRJ Radford SE Folding versus aggregation: Polypeptide conformations on competing pathways Arch. Biochem. Biophys. 2008 469 100 117 10.1016/j.abb.2007.05.015 17588526
Jahn, T. R. J. & Radford, S. E. Folding versus aggregation: Polypeptide conformations on competing pathways. Arch. Biochem. Biophys. 469, 100–117 (2008).17588526 10.1016/j.abb.2007.05.015
43. Jahn TR Radford SE The Yin and Yang of protein folding FEBS J. 2005 272 5962 5970 10.1111/j.1742-4658.2005.05021.x 16302961
Jahn, T. R. & Radford, S. E. The Yin and Yang of protein folding. FEBS J. 272, 5962–5970 (2005).16302961 10.1111/j.1742-4658.2005.05021.x
44. Ventura S Villaverde A Protein quality in bacterial inclusion bodies Trends Biotechnol. 2006 24 179 185 10.1016/j.tibtech.2006.02.007 16503059
Ventura, S. & Villaverde, A. Protein quality in bacterial inclusion bodies. Trends Biotechnol. 24, 179–185 (2006).16503059 10.1016/j.tibtech.2006.02.007
45. Willbold D Strodel B Schröder GF Hoyer W Heise H Amyloid-type protein aggregation and prion-like properties of amyloids Chem. Rev. 2021 121 8285 8307 10.1021/acs.chemrev.1c00196 34137605
Willbold, D., Strodel, B., Schröder, G. F., Hoyer, W. & Heise, H. Amyloid-type protein aggregation and prion-like properties of amyloids. Chem. Rev. 121, 8285–8307 (2021).34137605 10.1021/acs.chemrev.1c00196
46. Narayanan S Short amino acid stretches can mediate amyloid formation in globular proteins: The SRC homology 3 ( SH3) Case Proc. Natl. Acad. Sci. U.S.A. 2004 101 7258 7263 10.1073/pnas.0308249101 15123800
Narayanan, S. et al. Short amino acid stretches can mediate amyloid formation in globular proteins: The SRC homology 3 ( SH3) Case. Proc. Natl. Acad. Sci. U.S.A. 101, 7258–7263 (2004).15123800 10.1073/pnas.0308249101
47. Wang L Maji SK Sawaya MR Eisenberg D Riek R Bacterial inclusion bodies contain amyloid-like structure PLoS Biol. 2008 6 e195 10.1371/journal.pbio.0060195 18684013
Wang, L., Maji, S. K., Sawaya, M. R., Eisenberg, D. & Riek, R. Bacterial inclusion bodies contain amyloid-like structure. PLoS Biol. 6, e195 (2008).18684013 10.1371/journal.pbio.0060195
48. Rousseau F Schymkowitz J Serrano L Prediction of sequence-dependent and mutational effects on the aggregation of peptides and proteins Nat. Biotechnol. 2004 22 1302 1306 10.1038/nbt1012 15361882
Rousseau, F., Schymkowitz, J. & Serrano, L. Prediction of sequence-dependent and mutational effects on the aggregation of peptides and proteins. Nat. Biotechnol. 22, 1302–1306 (2004).15361882 10.1038/nbt1012
49. Conchillo-solé O AGGRESCAN: A server for the prediction and evaluation of “Hot spots” of aggregation in polypeptides BMC Bioinform. 2007 8 1 17 10.1186/1471-2105-8-65
Conchillo-solé, O. et al. AGGRESCAN: A server for the prediction and evaluation of “Hot spots” of aggregation in polypeptides. BMC Bioinform. 8, 1–17 (2007).10.1186/1471-2105-8-65
50. Ventura S Sequence determinants of protein aggregation: Tools to increase protein solubility Microb. Cell Fact. 2005 4 1 8 10.1186/1475-2859-4-11 15629064
Ventura, S. Sequence determinants of protein aggregation: Tools to increase protein solubility. Microb. Cell Fact. 4, 1–8 (2005).15629064 10.1186/1475-2859-4-11
51. Chiti F Dobson CM Protein misfolding, functional amyloid, and human disease Annu. Rev. Biochem. 2006 75 333 366 10.1146/annurev.biochem.75.101304.123901 16756495
Chiti, F. & Dobson, C. M. Protein misfolding, functional amyloid, and human disease. Annu. Rev. Biochem. 75, 333–366 (2006).16756495 10.1146/annurev.biochem.75.101304.123901
52. Sambrook JRD Molecular Cloning a Laboratory Manual 2001 Cold Spring Harbor Laboratory Press
Sambrook, J. R. D. Molecular Cloning a Laboratory Manual (Cold Spring Harbor Laboratory Press, 2001).
53. Sahdev S Khattar SK Saini KS Production of active eukaryotic proteins through bacterial expression systems: A review of the existing biotechnology strategies Mol. Cell. Biochem. 2008 307 249 264 10.1007/s11010-007-9603-6 17874175
Sahdev, S., Khattar, S. K. & Saini, K. S. Production of active eukaryotic proteins through bacterial expression systems: A review of the existing biotechnology strategies. Mol. Cell. Biochem. 307, 249–264 (2008).17874175 10.1007/s11010-007-9603-6
54. Ami D Natalello A Gatti-Lafranconi P Lotti M Doglia SM Kinetics of inclusion body formation studied in intact cells by FT-IR spectroscopy FEBS Lett. 2005 579 3433 3436 10.1016/j.febslet.2005.04.085 15949804
Ami, D., Natalello, A., Gatti-Lafranconi, P., Lotti, M. & Doglia, S. M. Kinetics of inclusion body formation studied in intact cells by FT-IR spectroscopy. FEBS Lett. 579, 3433–3436 (2005).15949804 10.1016/j.febslet.2005.04.085
55. Upadhyay AK Murmu A Singh A Panda AK Kinetics of inclusion body formation and its correlation with the characteristics of protein aggregates in Escherichia coli PLoS ONE 2012 7 e33951 10.1371/journal.pone.0033951 22479486
Upadhyay, A. K., Murmu, A., Singh, A. & Panda, A. K. Kinetics of inclusion body formation and its correlation with the characteristics of protein aggregates in Escherichia coli. PLoS ONE 7, e33951 (2012).22479486 10.1371/journal.pone.0033951
56. Peternel Š Grdadolnik J Gaberc-Porekar V Komel R Engineering inclusion bodies for non denaturing extraction of functional proteins Microb. Cell Fact. 2008 7 1 9 10.1186/1475-2859-7-34 18211716
Peternel, Š, Grdadolnik, J., Gaberc-Porekar, V. & Komel, R. Engineering inclusion bodies for non denaturing extraction of functional proteins. Microb. Cell Fact. 7, 1–9 (2008).18211716 10.1186/1475-2859-7-34
57. Shahbazi Dastjerdeh M Shokrgozar MA Rahimi H Golkar M Potential aggregation hot spots in recombinant human keratinocyte growth factor: A computational study J. Biomol. Struct. Dyn. 2022 40 8169 8184 10.1080/07391102.2021.1908912 33843469
Shahbazi Dastjerdeh, M., Shokrgozar, M. A., Rahimi, H. & Golkar, M. Potential aggregation hot spots in recombinant human keratinocyte growth factor: A computational study. J. Biomol. Struct. Dyn. 40, 8169–8184 (2022).33843469 10.1080/07391102.2021.1908912
58. Gour S Yadav JK Aggregation hot spots in the SARS-CoV-2 proteome may constitute potential therapeutic targets for the suppression of the viral replication and multiplication J. Proteins Proteom. 2021 12 1 13 10.1007/s42485-021-00057-y 33613009
Gour, S. & Yadav, J. K. Aggregation hot spots in the SARS-CoV-2 proteome may constitute potential therapeutic targets for the suppression of the viral replication and multiplication. J. Proteins Proteom. 12, 1–13 (2021).33613009 10.1007/s42485-021-00057-y
59. Palumbo E Zhao B Xue B Uversky VN Davé V Analyzing aggregation propensities of clinically relevant PTEN mutants: A new culprit in pathogenesis of cancer and other PTENopathies J. Biomol. Struct. Dyn. 2020 38 2253 2266 10.1080/07391102.2019.1630005 31232187
Palumbo, E., Zhao, B., Xue, B., Uversky, V. N. & Davé, V. Analyzing aggregation propensities of clinically relevant PTEN mutants: A new culprit in pathogenesis of cancer and other PTENopathies. J. Biomol. Struct. Dyn. 38, 2253–2266 (2020).31232187 10.1080/07391102.2019.1630005
60. Besse S Poujol R Hussin JG Comparative study of protein aggregation propensity and mutation tolerance between naked mole-rat and mouse Genome Biol. Evol. 2022 14 evac057 10.1093/gbe/evac057 35482036
Besse, S., Poujol, R. & Hussin, J. G. Comparative study of protein aggregation propensity and mutation tolerance between naked mole-rat and mouse. Genome Biol. Evol. 14, evac057 (2022).35482036 10.1093/gbe/evac057
61. Sayers EW Database resources of the national center for biotechnology information Nucleic Acids Res. 2022 50 D20 D26 10.1093/nar/gkab1112 34850941
Sayers, E. W. et al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 50, D20–D26 (2022).34850941 10.1093/nar/gkab1112
62. Gasteiger E Protein Identification and Analysis Tools on the Expasy Server. The Proteomics Protocols Handbook 2005 Humana Press
Gasteiger, E. et al. Protein Identification and Analysis Tools on the Expasy Server. The Proteomics Protocols Handbook (Humana Press, 2005). 10.1385/1592598900.
63. Hebditch M Carballo-Amador MA Charonis S Curtis R Warwicker J Protein-Sol: A web tool for predicting protein solubility from sequence Bioinformatics 2017 33 3098 3100 10.1093/bioinformatics/btx345 28575391
Hebditch, M., Carballo-Amador, M. A., Charonis, S., Curtis, R. & Warwicker, J. Protein-Sol: A web tool for predicting protein solubility from sequence. Bioinformatics 33, 3098–3100 (2017).28575391 10.1093/bioinformatics/btx345
64. Berman HM The protein data bank Nucleic Acids Res. 2000 28 235 242 10.1093/nar/28.1.235 10592235
Berman, H. M. et al. The protein data bank. Nucleic Acids Res. 28, 235–242 (2000).10592235 10.1093/nar/28.1.235
65. Waterhouse A SWISS-MODEL: Homology modelling of protein structures and complexes Nucleic Acids Res. 2018 46 296 303 10.1093/nar/gky427
Waterhouse, A. et al. SWISS-MODEL: Homology modelling of protein structures and complexes. Nucleic Acids Res. 46, 296–303 (2018).10.1093/nar/gky427
66. Jumper J Highly accurate protein structure prediction with AlphaFold Nature 2021 596 583 589 10.1038/s41586-021-03819-2 34265844
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).34265844 10.1038/s41586-021-03819-2
67. Laskowski RA MacArthur MW Moss DS Thornton JM PROCHECK: A program to check the stereochemical quality of protein structures J. Appl. Crystallogr. 1993 26 283 291 10.1107/S0021889892009944
Laskowski, R. A., MacArthur, M. W., Moss, D. S. & Thornton, J. M. PROCHECK: A program to check the stereochemical quality of protein structures. J. Appl. Crystallogr. 26, 283–291 (1993).10.1107/S0021889892009944
68. Wiederstein M Sippl MJ ProSA-web: Interactive web service for the recognition of errors in three-dimensional structures of proteins Nucleic Acids Res. 2007 35 407 410 10.1093/nar/gkm290
Wiederstein, M. & Sippl, M. J. ProSA-web: Interactive web service for the recognition of errors in three-dimensional structures of proteins. Nucleic Acids Res. 35, 407–410 (2007).10.1093/nar/gkm290
69. Eisenberg D Lüthy R Bowie JU VERIFY3D: Assessment of protein models with three-dimensional profiles Methods Enzymol. 1997 277 396 404 10.1016/S0076-6879(97)77022-8 9379925
Eisenberg, D., Lüthy, R. & Bowie, J. U. VERIFY3D: Assessment of protein models with three-dimensional profiles. Methods Enzymol. 277, 396–404 (1997).9379925 10.1016/S0076-6879(97)77022-8
70. Colovos C Yeates T Verification of protein structures: Patterns of nonbonded atomic interactions Protein Sci. 1993 2 1511 1519 10.1002/pro.5560020916 8401235
Colovos, C. & Yeates, T. Verification of protein structures: Patterns of nonbonded atomic interactions. Protein Sci. 2, 1511–1519 (1993).8401235 10.1002/pro.5560020916
71. Buck PM Voynov V Caravella JA Computational methods to predict therapeutic protein aggregation Therapeutic Proteins 2012 Springer 425 451
Buck, P. M. et al. Computational methods to predict therapeutic protein aggregation. In Therapeutic Proteins (eds Voynov, V. & Caravella, J. A.) 425–451 (Springer, 2012). 10.1007/978-1-61779-921-1.
72. Tartaglia GG Vendruscolo M The Zyggregator method for predicting protein aggregation propensities Chem. Soc. Rev. 2008 37 1395 1401 10.1039/b706784b 18568165
Tartaglia, G. G. & Vendruscolo, M. The Zyggregator method for predicting protein aggregation propensities. Chem. Soc. Rev. 37, 1395–1401 (2008).18568165 10.1039/b706784b
73. Oliveberg M Waltz, an exciting new move in amyloid prediction aMAZe-ing tools for mosaic analysis in zebrafish Nat. Methods 2010 7 187 188 10.1038/nmeth0310-187 20195250
Oliveberg, M. Waltz, an exciting new move in amyloid prediction aMAZe-ing tools for mosaic analysis in zebrafish. Nat. Methods 7, 187–188 (2010).20195250 10.1038/nmeth0310-187
74. Ahmed AB Kajava AV Breaking the amyloidogenicity code: Methods to predict amyloids from amino acid sequence FEBS Lett. 2013 587 1089 1095 10.1016/j.febslet.2012.12.006 23262221
Ahmed, A. B. & Kajava, A. V. Breaking the amyloidogenicity code: Methods to predict amyloids from amino acid sequence. FEBS Lett. 587, 1089–1095 (2013).23262221 10.1016/j.febslet.2012.12.006
75. Meric G Robinson AS Roberts CJ Driving forces for nonnative protein aggregation and approaches to predict aggregation-prone regions Annu. Rev. Chem. Biomol. Eng. 2017 8 139 159 10.1146/annurev-chembioeng-060816-101404 28592179
Meric, G., Robinson, A. S. & Roberts, C. J. Driving forces for nonnative protein aggregation and approaches to predict aggregation-prone regions. Annu. Rev. Chem. Biomol. Eng. 8, 139–159 (2017).28592179 10.1146/annurev-chembioeng-060816-101404
76. Garbuzynskiy SO Lobanov MY Galzitskaya OV FoldAmyloid: A method of prediction of amyloidogenic regions from protein sequence Bioinformatics 2009 26 326 332 10.1093/bioinformatics/btp691 20019059
Garbuzynskiy, S. O., Lobanov, M. Y. & Galzitskaya, O. V. FoldAmyloid: A method of prediction of amyloidogenic regions from protein sequence. Bioinformatics 26, 326–332 (2009).20019059 10.1093/bioinformatics/btp691
77. Prabakaran R Rawat P Kumar S Michael Gromiha M ANuPP: A versatile tool to predict aggregation nucleating regions in peptides and proteins J. Mol. Biol. 2021 433 166707 10.1016/j.jmb.2020.11.006 33972019
Prabakaran, R., Rawat, P., Kumar, S. & Michael Gromiha, M. ANuPP: A versatile tool to predict aggregation nucleating regions in peptides and proteins. J. Mol. Biol. 433, 166707 (2021).33972019 10.1016/j.jmb.2020.11.006
78. Sormanni P Aprile FA Vendruscolo M The CamSol method of rational design of protein mutants with enhanced solubility J. Mol. Biol. 2015 427 478 490 10.1016/j.jmb.2014.09.026 25451785
Sormanni, P., Aprile, F. A. & Vendruscolo, M. The CamSol method of rational design of protein mutants with enhanced solubility. J. Mol. Biol. 427, 478–490 (2015).25451785 10.1016/j.jmb.2014.09.026
79. Tsolis AC Papandreou NC Iconomidou VA Hamodrakas SJ A consensus method for the prediction of ‘aggregation-prone’ peptides in globular proteins PLoS One 2013 8 1 6 10.1371/journal.pone.0054175
Tsolis, A. C., Papandreou, N. C., Iconomidou, V. A. & Hamodrakas, S. J. A consensus method for the prediction of ‘aggregation-prone’ peptides in globular proteins. PLoS One 8, 1–6 (2013).10.1371/journal.pone.0054175
80. Van Durme J Solubis: A webserver to reduce protein aggregation through mutation Protein Eng. Des. Sel. 2016 29 285 289 10.1093/protein/gzw019 27284085
Van Durme, J. et al. Solubis: A webserver to reduce protein aggregation through mutation. Protein Eng. Des. Sel. 29, 285–289 (2016).27284085 10.1093/protein/gzw019
