
==== Front
Sci Rep
Sci Rep
Scientific Reports
2045-2322
Nature Publishing Group UK London

70096
10.1038/s41598-024-70096-0
Article
Improving rigor and reproducibility in western blot experiments with the blotRig analysis
Omondi Cleopa 1
Chou Austin 1
Fond Kenneth A. 1
Morioka Kazuhito 1
Joseph Nadine R. 1
Sacramento Jeffrey A. 1
Iorio Emma 1
Torres-Espin Abel 134
Radabaugh Hannah L. 1
Davis Jacob A. 1
Gumbel Jason H. 1
Huie J. Russell Russell.huie@ucsf.edu

12
Ferguson Adam R. adam.ferguson@ucsf.edu
Russell.huie@ucsf.edu

12
1 grid.266102.1 0000 0001 2297 6811 Weill Institute for Neurosciences, University of California, San Francisco, CA USA
2 https://ror.org/049peqw80 grid.410372.3 0000 0004 0419 2775 San Francisco Veterans Affairs Medical Center, San Francisco, CA USA
3 https://ror.org/01aff2v68 grid.46078.3d 0000 0000 8644 1405 School of Public Health Sciences, Faculty of Health Sciences, University of Waterloo, Waterloo, ON Canada
4 https://ror.org/0160cpw27 grid.17089.37 Department of Physical Therapy, Faculty of Rehabilitation Medicine, University of Alberta, Edmonton, AB Canada
17 9 2024
17 9 2024
2024
14 2164419 12 2023
13 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by/4.0/ Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Western blot is a popular biomolecular analysis method for measuring the relative quantities of independent proteins in complex biological samples. However, variability in quantitative western blot data analysis poses a challenge in designing reproducible experiments. The lack of rigorous quantitative approaches in current western blot statistical methodology may result in irreproducible inferences. Here we describe best practices for the design and analysis of western blot experiments, with examples and demonstrations of how different analytical approaches can lead to widely varying outcomes. To facilitate best practices, we have developed the blotRig tool for designing and analyzing western blot experiments to improve their rigor and reproducibility. The blotRig application includes functions for counterbalancing experimental design by lane position, batch management across gels, and analytics with covariates and random effects.

Keywords

Western blot
Analytical chemistry
Antibodies
Biostatistics
Computational biology
Computational chemistry
Subject terms

Biochemistry
Biological techniques
Computational biology and bioinformatics
Neuroscience
National Institutes of Health/National Institute of Neurological Disorders and Stroke grantR01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 R01NS088475 Omondi Cleopa Chou Austin Fond Kenneth A. Morioka Kazuhito Joseph Nadine R. Sacramento Jeffrey A. Iorio Emma Torres-Espin Abel Radabaugh Hannah L. Davis Jacob A. Gumbel Jason H. Huie J. Russell Ferguson Adam R. NIH NINDSR01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 R01NS122888 Omondi Cleopa Chou Austin Fond Kenneth A. Morioka Kazuhito Joseph Nadine R. Sacramento Jeffrey A. Iorio Emma Torres-Espin Abel Radabaugh Hannah L. Davis Jacob A. Gumbel Jason H. Huie J. Russell Ferguson Adam R. US Veterans Affairs (VA)I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 I01RX002245 Omondi Cleopa Chou Austin Fond Kenneth A. Morioka Kazuhito Joseph Nadine R. Sacramento Jeffrey A. Iorio Emma Torres-Espin Abel Radabaugh Hannah L. Davis Jacob A. Gumbel Jason H. Huie J. Russell Ferguson Adam R. US Veterans Affairs (VAI50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 I50BX005878 Omondi Cleopa Chou Austin Fond Kenneth A. Morioka Kazuhito Joseph Nadine R. Sacramento Jeffrey A. Iorio Emma Torres-Espin Abel Radabaugh Hannah L. Davis Jacob A. Gumbel Jason H. Huie J. Russell Ferguson Adam R. Wings for Life Foundation, Craig H. Neilsen Foundationissue-copyright-statement© Springer Nature Limited 2024
==== Body
pmcIntroduction

Proteomic technologies such as protein measurement with folin phenol reagent were first introduced by Lowry et al. in 19511. The resulting qualitative data are typically confirmed by a second, independent method such as western blot (WB)2,3. The WB method, first described by Towbin et al.4 and Burnette5 in 1979 and 1981, respectively, uses specific antibody-antigen interactions to confirm the protein present in the sample mixture. Quantitative WB (qWB assay) is a technique to measure protein concentrations in biological samples with four main steps: (1) protein separation by size, (2) protein transfer to a solid support, (3) marking a target protein using proper primary and secondary antibodies for visualization, and (4) semi-quantitative analysis6. Importantly, qWB data is considered semi-quantitative because methods to control for experimental variability ultimately yield relative comparisons of protein levels rather than absolute protein concentrations2,3, 7, 8. Similarly, western blotting applying ECL (enhanced chemiluminescence) is considered a semi-quantitative method because it lacks cumulative luminescence linearity and offers limited quantitative reproducibility9. However, the emergence of highly sensitive fluorescent labeling techniques, which exhibit a wider quantifiable linear range, greater sensitivity, and improved stability when compared to the conventional ECL detection method, now permits the legitimate characterization of protein expression as linearly quantitative10. Current methodologies do not sufficiently account for diverse sources of variability, producing highly variable results between different laboratories and even within the same lab11–13. Indeed, qWB data exhibits more variability compared to other experimental techniques such as enzyme linked immunosorbent assay (ELISA)14. For example, results have shown that qWB can produce significant variability in detecting host cell proteins and lead to researchers missing or overestimating true biological effects15. This in turn results in publication of irreproducible qWB interpretations, which leads to loss of its credibility13. In the serious cases, qWB results may even provide clinical misdiagnosis16 that could impact on a larger public health concern due to the prevalence of WB in biomedical research, such as diagnosis of SARS-CoV2 infection17.

The process of recognizing and accounting for variability in WB analyses will ultimately improve reproducibility between experiments. A growing body of studies has shown that this requires a fundamental shift in the experimental methodology across data acquisition, analysis, and interpretation to achieve precise and accurate results2,3,11–13.

Here we highlight experimental design practices that enable a statistics-driven approach to improve the reproducibility of qWBs. Specifically, we discuss major sources of variability in qWB including the non-linearity in antibody signal2,3; imbalanced experimental design13; lack of standardization in the treatment of technical replicates3,18; and variability between protein loading, lanes, and blots2,7,19. To address these issues, we provide new comprehensive suggestions for quantitative evaluation of protein expression by combining linear range characterization for antibodies, appropriate counterbalancing during gel loading, running technical replicates across multiple gels, and by taking careful consideration of the analysis method. By applying these experimental practices, we can then account for more sources of variability by running analysis of covariance (ANCOVA) or generalized linear mixed models (LMM). Such approaches have been shown to successfully improve reproducibility compared to other methods13.

Good options for qWB protein bands analysis using free, downloadable tools are available for researchers. Amongst others, LI-COR Image Studio Lite can be used to measure the intensity of protein bands in western blots and calculate their relative abundance. Likewise, ThermoFisher ImageQuant Lite offers features such as the ability to perform background subtraction and normalization. However, to date, no specific tools are freely available to provide a map to counterbalance samples, which overcome imperfect uniform protein electrophoresis/transfer and perform statistical analysis. Here, we present blotRig, a tool for researchers with functionalities to counterbalance samples and perform statistical analysis.

To help improve WB rigor we developed the blotRig protocol and application harnessing a database of 6000 + western blots from N = 281 subjects (rats and mice) collected by multiple UCSF labs on core equipment. To demonstrate blotRig best practices in a real-world experiment, we carried out prospective multiplexed WB analysis of protein lysate from lumbar cord in rodent models of spinal cord injury (SCI) (N = 29 rats) in 2 groups (experimental group & control group). In order to show that these experimental suggestions could improve qWB reproducibility, we compared different statistical approaches to handling loading controls and technical replicates. Specifically, we applied two strategies to integrate loading controls: (i) normalizing the target protein levels by dividing by the loading control or (ii) treating the loading control as a covariate in a LMM. Additionally, we analyzed technical replicates in four ways: (1) assume each sample was only run once without replication, (2) treat each technical replicate as an independent sample, (3) use the mean of the three technical replicate values, and 4) treat the replicate as a random effect in a LMM. Altogether, we found that the statistical power of the experiment was significantly increased when we used loading control as a covariate with technical replicates as a random effect during analysis. In addition, the effect size was increased, and the p-value of our analysis decreased when using this LMM, suggesting the potential for greater sensitivity in our WB experiment when using this approach20. Through rigorous experimental design and statistical analysis we show that we can account for greater variability in the data and more clearly identify underlying biological effects.

Materials and methods

Animals

All experiments protocol were approved by the University Laboratory Animal Care Committee at University of California, San Francisco (UCSF, CA, USA) and followed the animal guidelines of the National Institutes of Health Guide for the Care and Use of Laboratory animals (National Research Council (US) Committee for the Update of the Guide for the Care and Use of Laboratory Animals, 2011). We followed The ARRIVE guidelines (Animal Research: Reporting In Vivo Experiments) to describe our in vivo experiments.

Male Simonsen Long Evans rats (188–385 g; Gilroy (Santa Clara, CA, USA), (N = 29) aged 3 weeks were housed under standard conditions with a 12-h light–dark cycle (6:30 am to 6:30 pm) and were given food and water ad libitum. The animals were housed mostly in pairs in 30 × 30 × 19-cm isolator cages with solid floors covered with a 3 cm layer of wood chip bedding. The experimenters were blind to the identity of treatments and experimental conditions, and all experiments were designed to minimize suffering and limit the number of animals required.

Anesthesia and surgery

We performed non-survival spinal cord injury and spared nerve injury surgeries on animals. Specifically, 3 week old female rats were anesthetized with continuous inhalation of isoflurane (1–5% mg/kg) while on oxygen (0.6–1 mg/kg) in accordance with the IACUC surgical and anesthesia guidelines. Preoperative 0.5% lidocaine local infiltration was applied once at surgical site, avoiding injection into muscle. Fur over the T7–T9 thoracic level was shaved. The dorsal skin was aseptically prepared with surgical iodine or chlorhexidine and 70% ethanol. A small longitudinal incision was made along the spine through the skin, fascia, and muscle to expose the T7-T9 vertebrae. Animals undergoing sham procedure did not undergo laminectomy and immediately proceeded to wound closure. Overlying muscle and subcutaneous tissue was sutured closed using an absorbable suture in a layered fashion. External skin was reinforced using monofilament suture or tissue glue as needed. Animals were euthanized after 30 min to extract spinal cord tissue through fluid expulsion.

Experimental methodology

In accordance with established quality standards for preclinical neurological research21, experimenters were kept blind to experimental group conditions throughout the entire study. Western blot loading order was determined a priori by a third-party coder, who ensured that a representative sample from each condition was included on each gel in a randomized block design. The number of subjects per condition was kept consistent across groups for each experiment to ensure that proper counterbalancing could be achieved across independent western runs. All representative western images presented in the figures represent lanes from the same gel. Sometimes, the analytical comparisons of interest were not available on adjacent lanes even though they come from the same gel because of our randomized counterbalancing procedure.

Western blot

The example western blot data used in this paper are taken from a model of spared nerve injury in animals with spinal cord injury. The nerve injury model used is based on models from pain literature22, where two of the three branches of the sciatic nerve are transected, sparing the sural nerve (SNI)23. Two surgeons perform the procedure simultaneously, with injuries occurring 5 min apart. The spinal cord of animals was obtained based on fluid expulsion model24 and a 1 cm section of the lumbar region was excised at the lumbar enlargement section. The tissue was then preserved in a -80 degree freezer until it was needed for an experiment, at which point it was thawed and used to run a Western blot. We conducted a Western blot analysis on 29 samples from animals using standard biochemical methods. We measured the protein levels of the AMPA receptor subunit GluA2 and used beta-actin as a loading control. The data from these experiments was then aggregated and used for statistical analysis.

Protein assay

We assayed sample protein concentration using a bicinchoninic acid (BCA assay (Pierce) for reliable quantification of total protein using a plate reader (Tecan; GeNios) with triplicate samples (technical replicates) detected against a Bradford Assay (BSA) standard curve. Technical replicates are multiple measurements that are performed under the same conditions in order to quantify and correct for technical variability and improve the accuracy and precision of the results (48). We ran the same WB loading scheme three times (technical replicates of the entire gel) and measured the protein levels of AMPA receptors.

Polyacrylamide gel electrophoresis and multiplexed near-infrared immunoblotting

The approach involved performing serial 1:2 dilutions with cold Laemmli sample buffer in room temperature; 15 μg of total protein per sample was loaded into separate lanes on a precast 10–20% electrophoresis gel (Tris–HCl polyacrylamide, BioRad) to establish linear range (Fig. 1). The blotRig software helps counterbalance sample positions across the gel by treatment condition. (Fig. 2). A kaleidoscope ladder was loaded on the first lane of each gel to confirm molecular weight (Fig. 2). The gel was electrophoresed for 30 min at 200 V in SDS buffer (25 mm Tris, 192 mm glycine, 0.1% SDS, pH 8.3; BioRad). Protein was transferred to a nitrocellulose membrane in cold transfer buffer (25 mm Tris, 192 mm glycine, 20% ethanol, pH 8.3). Membrane transfer was confirmed using Ponceau S stain (67) followed by a quick rinse and blocking in Odyssey blocking buffer (Li-Cor) containing Tween-20.Figure 1 Determining linear range of antibodies to optimize parametric analysis of Western blot data. When small or large protein concentrations are loaded, there is often a possibility that their representation on western blot band density may become non-linear. If there is a disconnect between the observed and expected protein concentrations, results may be inaccurate. Thus determining the linear range wherein, a one-unit increase in protein is reflected in a linear increase in band density for each western blot antibody is a crucial initial step to ensure confidence in reproducibility of the linear models commonly applied to western blot data analysis.

Figure 2 Counterbalancing to reduce bias. (A) Experimental design. A simple hypothetical experimental design for illustrating counterbalancing. Two experimental groups (Wild Type vs Transgenic), with two treatments (Drug vs Vehicle) analyzed within each individual. This 2 (Experimental Condition) by 2 (Tissue Area) design yields four groups. (B) Counter-balanced Gel Loading. The goal of appropriate counterbalancing is to optimize the sequence in which samples are loaded such that groups are represented equally across the gel. Those with red X have with the experimental groups and treatment condition grouped in the same area of the gel, and thus variability across the gel may be conflated with group differences. In contrast, those with the green check are organized so that experimental condition and treatment condition are better placed to reduce the possibility of any single group being over-represented in a particular area of the gel.

The membrane was blocked for 1 h in Odyssey Blocking Buffer (Li-Cor) containing 0.1% Tween-20, followed by an overnight incubation in primary antibody solution at 4 °C. Membrane incubation was done in a primary antibody solution containing Odyssey blocking buffer, Tween-20, appropriate primary antibody receptor targeting1:2000 mouse PSD-95 (cat # MA1-046,Thermofisher), 1:200 rabbit GluA1 (cat # AB1504, Millipore), 1:200 rabbit GluA2 (cat # AB1766, Millipore), 1:200 rabbit pS831(cat # 04–823, Millipore), 1:200 p880 (cat#07–294, Millipore) or 1:1,500 mouse actin loading control (cat # 612,857, BD Transduction)]. Following incubation, the membrane was washed 4 × 5 min with Tris-buffered saline containing 0.1% Tween 20 (TTBS) and incubated in fluorescent-labeled secondary antibody (1:30 K LiCor IRdye appropriate goat anti-rabbit in Odyssey blocking buffer plus 0.2% Tween 20) for 1 h in the dark. Subsequent to 4 × 5 min washes in TTBS, followed by a 5 min wash in TBS.

Membrane incubation was used to detect the presence of a specific protein or antigen on a membrane. In this case, the membrane was incubated with a fluorescently labeled secondary antibody solution that was specifically tuned to the emission spectra of the laser lines used by the Li-Cor Odyssey quantitative near-infrared molecular imaging system instrument. This allows for specific detection of the protein of interest on the membrane. The sample is then imaged using an infrared imaging system that is optimized for detecting the specific wavelengths of light emitted by the fluorescent label. Additional rounds of incubation and imaging are performed to detect additional proteins using the multiplexing functionality of the Li-Cor instrument, with each round adding new bands at different molecular weight ranges. This allows for the detection of multiple proteins in the same sample, maximizing the proteomic detection.

Quantitative near-IR densitometric analysis

Using techniques optimized in the our lab25,26, we established near-infrared labeling and detection techniques (Odyssey Infrared Imaging System, Li-Cor) to quantify linear intensity detection of fluorescently labeled protein bands. The biochemistry is performed in a blinded, counterbalanced fashion, and three independent replications of the assay are run on different days27. Fluorescent Western blotting utilizes fluorescent-labeled secondary antibodies to detect the target protein, which allows for more sensitive and specific detection compared to chemiluminescence11,28,29. Additionally, fluorescence imaging allows multiple detection of a target protein and internal loading control in the same blot, which enables more accurate correction of sample-to-sample and lane-to-lane variation11,30,31. This provides a more accurate and reliable quantification of the target protein, making it a popular choice for quantitative analysis of WB data.

Blinding

It is good practice for the pipetting experimenter to remain blind to experimental conditions during gel loading, transfer, and densitometric quantification. We achieved this using de-identified tube codes and a priori gel loading sequences that were developed by an outside experimenter using the method implemented in the blotRig software.

Statistical analyses

Statistical analyses were performed using the R statistical software. Our WB data was analyzed using parametric statistics. The WB was run using three independent replications and covariance corrected by beta-actin loading control, with replication statistically controlled as a random factor. Significance was assessed at p < 0.0525,26,32,33,34. We report estimated statistical power and standardized regression coefficient effect sizes in the results section.

All ANOVAs were run using the stats R package; standardized effect size was calculated using the parameters R package35. Linear mixed models were run using the lme4 R package. Observed power was calculated by Monte Carlo simulation (1000x) run on the fitted model (either ANOVA or LMM) using the simR package36. For the development of the blotRig interface, the R packages used included: shiny, tidyverse, DT, shinythemes, shinyjs, and sortable)37–42. You can access the blotRig analysis software, which includes code for inputting experimental parameters for all Western blot analysis, through the following link: https://atpspin.shinyapps.io/BlotRig/.

Results

Designing reproducible western blot experiments

Determining linear range for each primary antibody

Most WB analyses assume semi-quantitatively that the relationship between qWB assay optical density data (i.e. western band signal) and protein abundance is linear2,3,11,18. Accordingly, most qWB analyses use statistical tests (t-test; ANOVA) that assume a linear effect. However, recent studies have shown that the relationship can potentially be highly non-linear19 As Fig. 1 illustrates, the WB band signal can become non-linearly correlated with protein concentrations at low and high values. This may result in inaccurate quantification of relative target protein amount in the experiment and violates the assumptions for linear model which can lead to false inferences. To address the assumption of linearity, it is important to first determine the optimal linear range for each protein of interest so that one can be confident that a unit change in band density reflects a linear change in protein concentration. This enables an experimenter to accurately quantify the protein of interest and apply linear statistical methods appropriately for hypothesis testing.

Counterbalancing during experimental design

Counterbalancing is the practice of having each experimental condition represented on each gel and evenly distributing them to prevent overrepresentation of the same experimental groups in consecutive lanes. For example, imagine an experimental design in which we are studying two experimental groups (wild type and transgenic animals) and are also looking at two treatment conditions (Drug and Vehicle). The best way to determine the effects and interactions between our experimental and treatment groups would be to create a balanced factorial design. A factorial design is one in which all combinations of levels across factors are represented. For the current example, a balanced factorial design would produce four groups, covering each possible combination (Drug-treated Wild Type, Vehicle-treated Wild Type, Drug-Treated Transgenic and Vehicle-treated Transgenic) (Fig. 2A). During WB gel loading, experimenters often distribute their samples unevenly such that certain experimental conditions may be missing on some gels or samples from the same experimental condition are loaded adjacently on a gel. This is problematic because we know that polyacrylamide gel electrophoresis (PAGE) gels are not perfectly uniform, reflecting a source of technical variability43; in the worst case, if we have only loaded a single experimental group on a gel and found a significant effect of the group, we cannot conclude if the effect is due to the experimental condition or a technical problem of the gel. At minimum, experimenters should ensure that every group in a factorial design is represented on each gel to avoid confounding technical gel effects with experimental differences. If the number of combinations is too large to represent on a single gel because of the number of factors or the number of levels of the factors, then a smaller "fractional factorial" design will provide maximal counterbalancing to ensure unbiased estimates of all factor effects and the most important interactions.

In addition, experimenters can further counter technical variability by arranging experimental groups on each gel to ensure adequately counterbalanced design assuming the uniformed protein concentration and fluid volume of all samples. This importantly addresses the variability due to physical effects within an individual gel. In our example, this means alternating the tissue areas and experimental conditions as much as possible to minimize similar samples from being loaded next to one another (Fig. 2B). By spreading the possibility of technical variability across all samples by counterbalancing across and within gels, we can mitigate potential technical effects that can bias our results. Proper counterbalancing also enables us to implement more rigorous statistical analysis to account for and remove more technical variability25,26,32,33. Overall, this will help to ensure that experimenters can find the same result in the future and improve reproducibility.

Technical replication

Technical replicates are used to measure the precision of an assay or method by repeating the measurement of the same sample multiple times. The results of these replicates can then be used to calculate the variability and error of the assay or method13. This is important to establish the reliability and accuracy of the results. Most experimenters acknowledge the importance of running technical replicates to avoid false positives and negatives due to technical error13. Even beyond extreme results, technical replicates can account for the differences in gel makeup, human variability in gel loading, and potential procedural discrepancies. In fact, most studies run at least duplicates; however, the experimental implementation of replicates (e.g., running replicates on the same gel or separate gels) as well as the statistical analysis of replicates (e.g., dropping “odd-man-out” or taking the mean or standard deviation) can differ greatly44,45. This experimental variability ultimately impedes our ability to meaningfully compare results. For experimenters to establish accuracy and advance reproducibility in WB experiments, it is important to implement standardized and rigorous protocols to handle technical replicates11,13. In doing so, we can further reduce the technical variability with statistical methods during analysis.

As underscored previously, we recommend that technical replicates are counterbalanced on separate gels to mitigate any possible gel effect. Additionally, by running triplicates, we can treat replicates as a random effect in a LMM during statistical analysis. Importantly, triplicates provide more values to measure the distribution of technical variance to ensure the robustness of the LMM than only running duplicates. This approach isolates and removes technical variance from biological variation which ultimately improves our sensitivity for true experimental effects46.

In the following demonstration of statistical methods, we replicated all WB analyses in triplicate with a randomized counterbalanced design. We then explore how the way in which technical replicates and loading controls are incorporated into analysis can have a significant impact on both the sensitivity of our results and the interpretation of the findings. An example mockup of a dataset illustrating the various ways in which western blot data are typically prepared for analysis can be found in Fig. 3.Figure 3 Western Blot Gel and Replication Strategies. (A) Illustration of Western Blot Gel. This depiction of a typical multiplexed western blot gel highlights the antibody-labeled target protein bands of interest (green/yellow) and housekeeping protein loading control that is always run and quantified in the same sample and lane as the target of interest. Total protein stain (fluorescent ponceau stain) is shown in red can can be used as an alternative loading control. Specific, quantification is typically executed on a single antibody-labeled channel for the target protein and housekeeping protein loading control (gray scale image). (B) Balanced Factorial Technical Replicate Strategy. Here we show the western blot data for the first 3 subjects from an example dataset. In a balanced factorial design, an equal number of samples from all possible experimental groups are represented on each gel. This table shows the subject number, the technical replicate, experimental group, and the band quantifications for both the target protein and the loading control. A ratio of target protein and loading control is also calculated. (C) Other Common Technical Replicate Strategies. In this example table are two of the other ways western blot data are typically formatted. Some experimenters choose to not include technical replicates, with only one sample from each subject quantified. In another replication strategy, technical replicates are averaged. Averaging may bias or skew the data. We recommend running technical replicates on separate gels or batches, and using gel/batch as a random factor when analyzing western blot data.

Statistical methodology to improve western blot analysis

Loading control as a covariate

Most qWB assay studies use loading controls (either a housekeeping protein or total protein within lane) to ensure that there are no biases in total protein loaded in a particular lane2,11,27. The most common way that loading controls are used to account for variability between lanes is by normalizing the target protein expression values by dividing it by the loading control values (Fig. 3), resulting in a ratio between target protein to loading control2,47,48. However, ratios may violate assumptions of common statistical test used to analyze qWB (e.g., t-test, ANOVA, etc.)49 This ultimately hinders the ability to statistically account for the variance in qWB outcomes and have a reliable estimate of the statistics. An alternative approach to improve the parametric properties would be to include loading control values as a covariate—a variable that is not our experimental factors but that may affect the outcome of interest and presents a source of variance that we may account for50. For instance, we know the amount of protein loaded is a source of variability in WB quantification, so we can use the loading control as a covariate to adjust for that variance. In doing so, we extend the method of ANOVA into that of ANCOVA51. This approach accounts for the technical variability present between lanes while meeting the necessary assumptions for parametric statistics which helps curb bias and averts false discoveries.

Replication and subject as a random effect

Most WB studies use ANOVA, a test that allows comparison of the means of three or more independent samples, for quantitative analysis of WB data49. One of the assumptions in ANOVA is the independence of observations49. This is problematic because we often collect multiple observations from the same analytical unit, for example different tissue samples from a single subject, or technical replicates. As a result, those observations don’t qualify as independent and should be analyzed using models controlling for variability within units of observations (e.g., the animal) to mitigate inferential errors (false positives and negatives)52 caused by what is known as pseudoreplication. This arises when the quantity of measured values or data points surpasses the number of actual replicates, and the statistical analysis treats all data points as independent, resulting in their full contribution to the final result53.

In addition, when conducting experiments, it is important to consider the randomness of the conditions being observed. Treating both subjects and conditions as fixed effects can lead to inaccurate p-values. Instead, subjects/ animals should be treated as random effects and the conditions should be considered as a sample from a larger population54. This is especially important when collecting data from different replicates or gels, as the separate technical replicate runs should be considered as random.

In Fig. 4 we use a simple experimental design comparing the difference in a target protein between two experimental groups to demonstrate four of the most common ways researchers tend to analyze western blot data: (1) running each sample once without replication, (2) treating each technical replicate as an independent sample, (3) taking the mean of technical replicate values, and (4) treating subject and replication as a random effect (Fig. 4). We then tested how effect size, power, and p value are affected by each of these strategies to get a sense of how much these estimates vary between analyses. For each of these strategies, we also tested the difference between using the ratio of target protein to loading controls versus using loading control as a statistical covariate. For further exploration of the way these data are prepared and analyzed, see the data workup in Supplementary Figs. 1 and 2.Figure 4 Effect of different replication and loading control strategies on statistical outcomes. Eight possible strategies are shown, representing the most common ways in which replication and loading controls are treated in a typical Western blot analysis. Four replication strategies: either no replication at all, 3 technical replicate gels treated as independent, mean of three replicates, or replicate treated as a random effect in a linear mixed model. These are crossed with two loading control strategies: either target protein is divided by loading control, or loading control is treated as a covariate in a linear mixed model. (A) Effect Size: Standardized effect size coefficient is generally improved when loading control is treated as a covariate, compared to using a ratio of the target protein and loading control values. (B) Power: By treating each replication as independent the statistical power is increased (due to the inaccurate assumption that technical replicates are not related, thus artificially tripling the n). Conversely, including the variability inherent in technical replicates as a part of the statistical model, we work to identify and account for a major source of variability, thus improving power in a more appropriate way. (C) P value: As expected, when each replication is inaccurately treated as independent the p value is low (due to artificially inflated n). We found that using the mean of replications and loading controls as covariates also resulted in a p value below 0.05. The smallest p value was found when including replication as a random factor. Across each of these statistical measures, only when replication is included as a random factor and loading control as a covariate do we see a strong effect size, high power, and low p value.

In the first scenario, we imagined that no technical replication was run at all (by using only the first replication). With this strategy, we found that standardized effect size is weak, power is low, and the p value was high (Fig. 4). Second, we demonstrate how analytical output would be different if we did run three technical replicates, but treated each as independent. As discussed above, this strategy does not take into account the fact that each sample is being run three times, and consequently the overall n of your experiment is artificially tripled! As one might expect, observed power is quite high, and our p value is low (< 0.05). Power is increased by an increase in sample size, so it is not surprising that the power is much higher if we erroneously report that we have a 3X larger sample size (i.e., pseudoreplication)53. In this case, the observed power is inflated and an artifact of inappropriate statistics, and the probability of a false positive is considerably increased with respect to the expected 5%.

So, what would be a more appropriate way to handle technical replicates? One method that researchers often use is to take the mean of their technical replicates. This does ensure that we are not artificially inflating our sample size, which is certainly an improvement over the previous strategy. With this strategy, we do find that our p value is less than 0.05 (when loading control is treated as a covariate). But we also see that our power is still low. We have effectively taken our replicates into account by collapsing across them within each sample, but this can be dangerous. If there is wide variation across replicates of a particular sample, then taking the mean of three replicates could produce an inaccurate estimate of the ‘true’ sample value. Ideally, we want to find a solution where instead of collapsing this variation, we add it to our statistical model so that we can better understand what amount of variation is randomly coming from within technical replicates, and in turn what amount of variation is actually due to potential differences in our experimental groups.

To achieve this, we need to model both the fixed effect of all groups in a full factorial design, and the random effect of replication across western blot gels. When we use both fixed and random effects, this is referred to as a linear mixed model (LMM). When using this strategy, we find that our effect size remains strong, and our p value is low. But importantly, we now have strong observed power (Fig. 4). This suggests that we can achieve greater sensitivity in our WB experiment when using this approach. Specifically, if we implement careful counterbalancing while designing our experiments, then we can use the variability between gels to our advantage during analysis using linear mixed effects model55.

LMM is recommended because it takes into account both the multiple observations within a single subject/animal in a given condition and differences across subjects observed in multiple conditions. This reduces chances of inaccurate p-values and improves reliability56. Further, treating both subjects and replication as random effects generalizes the results to the population of subjects and also to the population of conditions57.

Real world application of blotRig software for western blot experimental design, technical replication, and statistical analysis

We have designed a user interface that is designed to facilitate appropriate counterbalancing and technical replication for western blot experimental design. The ‘blotRig’ application is run through RStudio, and can be found here: https://atpspin.shinyapps.io/BlotRig/ Upon starting the blotRig application, the user is prompted to upload a comma separated values (CSV) spreadsheet. This spreadsheet should include separate columns for subject ID and experimental group. The user is then prompted to enter the total number of lanes that are available on their particular western blot gel apparatus. The blotRig software will first run a quality check to confirm that each subject ID (unique sample or subject) is only found in one experimental group. If duplicates are found, a warning will be shown that specifies which subjects are repeated across groups. If no errors are found, a centered gel map will be generated that illustrates the western blot gel lanes into which each subject should be loaded (Fig. 5A). The decision for each lane loading is based on two main principles outlined above: (1) each western blot gel should hold a representative sample of each experimental group (2) samples from the same experimental group are not loaded in adjacent lanes whenever possible. This ensures that proper counterbalancing is achieved so that we can limit the chances that the inherent variability within and across western blot gels is confounded with the experimental groups that we are interested in experimentally testing.Figure 5 Example of the blotRig Gel Creator interface. (A) Illustration of the blotRig interface. User has entered their sample IDs, experimental groups, and the number of lanes per western blot gel. (B) The blotRig system then creates a counterbalanced gel map that ensures each gel contains a representative from each experimental group. This illustration shows the exact lane for each gel in which each sample should be run.

Once the gel map has been generated, the user can then select to export this gel map to a CSV spreadsheet. This sheet is designed to clearly show which gel each sample is on, which lane on each gel a sample is found, what experimental group each sample belongs to, and importantly, a repetition of each of these values for three technical replicates (Fig. 5B). User will also see columns for Target Protein and Loading Control. These are the cells where the user can then input their densitometry values upon completing their western blot runs. Once this spreadsheet is filled out, it is then ready to go for blotRig analysis.

To analyze western blot data, users can upload the completed template that was exported in the blotRig experimental design phase or their own CSV file under the ‘Analysis’ tab (Fig. 6). The blotRig software will first ask the user to identify which columns from the spreadsheet represent Subject/SampleID, Experimental Group, Protein Target, Loading Control, and Replication. The blotRig software will again run a quality check to confirm that there are no subject/sample IDs that are duplicated across experimental groups. If no errors are found, the data will then be ready to analyze. The blotRig analysis will then be run, using the principles discussed above. Specifically, a linear mixed-model runs using the lmer R package, with Experimental Group as a fixed effect, Loading Control as a covariate, and Replication (nested within Subject/Sample ID) as a random factor. Analytical output is then displayed, giving a variety of statistical results from the linear mixed model output table, including fixed and random effects and associated p values (Fig. 6). A bar graph of group means and 95% confidence interval error bars will also be generated, along with a summary of the group means, standard error of the mean, and upper/lower 95% confidence intervals. These outputs can be directly reported in the results sections of papers, improve the statistical rigor of published WB reports. In addition, since the entire pipeline is opensource, the blotRig code itself can be reported to support transparency and reproducibility.Figure 6 Workflow for running statistical analysis of replicate western blot data using blotRig. First, fill out spreadsheet with subject ID, experimental group assignment, number of technical replication, the densitometry values for your target proteins and loading controls. After saving this spreadsheet as a.csv file, the file can be uploaded to blotRig. Tell blotRig the exact names of each of your variables, then click ‘Run Analysis’. This will produce a statistical output using linear mixed model testing for group differences using loading control as a covariate and replication as a random effect. Bar graph with error bars and summary statistics can then be exported.

Discussion

Although the western blot technique has proven to be a workhorse for biological research, the need to enhance its reproducibility is critical13,19,27. Current qWB assay methods are still lacking for reproducibly identifying true biological effects13. We provide a systematic approach to generate quantitative data from western blot experiments that incorporates key technical and statistical recommendations which minimize sources of error and variability throughout the western blot process. First, our study shows that experimenters can improve the reproducibility of western blots by applying the experimental recommendations of determining the linear range for each primary antibody, counterbalancing during experimental design, and running technical triplicates13,27. Furthermore, these experimental implementations allow for application of the statistical recommendations of incorporating loading controls as covariates and analyzing gel and subject as random effects58,59. Altogether, these enable more rigorous statistical analysis that accounts for more technical variability which can improve the effect size, observed power, and p-value of our experiments and ultimately better identify true biological effects.

Biomedical research has continued to rely on p-values for determining and reporting differences between experimental groups, despite calls to retire the p-value60. Power (sensitivity) calculations have also become increasingly common. In brief, p-values and the related alpha value are associated with Type I error rate—the probability of rejecting the null hypothesis (i.e., claiming there is an effect) when there is no true effect61. On the other hand, power effectively measures the probability of rejecting the null hypothesis (i.e. stating there is not effect) when there is indeed a true underlying effect—a concept that is closely related to reducing the Type II error rate59,62. Critically, empirical evidence estimates that the median statistical power of studies in neuroscience is between ∼8% and ∼31%, yet best practices suggest that an experimenter should aim to achieve a power of 80% with an alpha of 0.0520. By being underpowered, experiments are at higher likelihood of producing a false inference. If an underpowered experiment is seeking to reproduce a previous observation, the resulting false negative may throw into question the original findings and directly exacerbate the reproducibility crisis59. Even more alarmingly, a low power also increases the likelihood that a statistically significant result is actually a false positive due to small sample size problems61. In our analyses, we show that our technical and statistical recommendations lower the p-value (indicating that the observed relationship between variables is less likely to be due to chance) as well as observed power of our experiments. This translates into the ability to better avoid false negatives when there is a true effect as well as reduce the likelihood of false positives when there is not a true experimental effect, both of which will ultimately improve the reproducibility of qWB assay experiments.

Another useful component of statistical analyses that is not as commonly reported but is critically related to p-value and power is effect size. Effect size is a statistical measure that describes the magnitude of the difference between two groups in an experiment63. It is used to quantify the strength of the relationship between the variables being studied63. The estimated effect size is important because it answers the most frequent question that researchers ask: how big is the difference between experimental groups, or how strong is the relationship or association?63. The combination of standardized effect size, p-value and power reflect crucial experimental results that can be broadly understood and compared with findings from other studies62, thus improving comparability of qWB experiments49,63. In particular, studies with large effect sizes have more power: we are more likely to detect a true positive experimental effect and avoid the false negative if the underlying difference between experimental groups is large46. In some cases, the calculated effect size is greatly influenced by how sources of variance are handled during analysis13. Our results demonstrate that by reducing the residual variance (by modeling the random effect of replication) the estimated effect size of our experiment increases. This could mean that the magnitude of the difference between the groups in our experiment is larger than it was originally thought to be. This could be due to a variety of factors such as improving the experimental design, sample size, or the measurement of the variables13. Likewise, conducting a power analysis is an essential step in experimental design that should be done before collecting data to ensure that the study is adequately powered to detect an effect of a certain size64.

Increasingly, power analysis is becoming a requirement for publications and grant proposals65. This is because a study with low statistical power is more likely to produce false negative results, which means that the study may fail to detect a real effect that actually exists. This can lead to the rejection of true hypotheses, wasted resources, and potentially harmful conclusions. In brief, given an experimental effect size and variance, we can calculate the sample size needed to achieve an alpha of 0.05 and power of 0.8; an increased sample size reduces the standard error of mean (SEM), which is the measured spread of sample means and consequently increases the power of the experiment66. We have demonstrated that our experimental and statistical recommendations lead to a lower p value (Fig. 3C) and effect size (Fig. 3B) without changing the sample size. This may be of greatest interest to researchers: more rigorous analytics ultimately improves experimental sensitivity without relying solely on increasing the sample size.

Reducing the sample size of an experiment can be beneficial for several reasons, one of which is cost-effectiveness. A smaller sample size can lead to a reduction in the number of animals or other resources that are needed for the study, which can result in lower costs. Additionally, it can also save time and reduce the duration of the experiment, as fewer subjects need to be recruited, and the data collection process can be completed more quickly. However, it is important to note that reducing the sample size can also lead to decreased statistical power. As a result, reducing sample size too much can increase the risk of a type II error, failing to detect significance when there is a true effect62.Therefore, it is important to consider the trade-off between sample size and power when designing an experiment, and to use statistical techniques like power analysis to ensure that the sample size is sufficient to detect an effect of a certain size. Moreover, when using animals in research, it's always important to consider the ethical aspect and the 3Rs principles of reduction, refinement, and replacement55.

Despite our best efforts in creating a balanced, full factorial experimental design, there will always be random variation in biological experiments. Fixed effects such as experimental group differences are expected to be generalizable if the experiment is replicated. Random effects (such as gel variation) on the other hand are unpredictable across experiments. Western blot analyses are particularly susceptible to this random gel variation, as different values may be observed for technical replicates run on different gels. By using a linear mixed model paired with rigorous full factorial design, we can ensure that we account for as much of that random variation as possible. When we acknowledge, identify, and model random effects we enhance the possibility of discovering our fixed effect of experimental treatment, if one exists.

The linear mixed model framework discussed above assumes that our western blot outcome measures are on a linear scale. As described above, parametric work to identify the linear range of a protein of interest is critical for ensuring that the results of a LMM (or ANOVA and t-test) are accurate and interpretable. While we recommend using loading control (or total protein control) as a covariate in a linear mixed model, many bench researchers may prefer to use the within-lane loading control (or total protein) to normalize target protein values. It is important to consider that in doing so, one creates a ratio value that is multiplicative instead of linear. This property has the side effect of artificially distorting the variance. To account for this non-linearity, we recommend that one uses semi-parametric mixed models such as generalized estimating equations with a gamma distribution link function that appropriately represents ratio data.

There has been recent recognition that an appropriate study design can be achieved by balancing sample size (n), effect size, and power31. The experimental and statistical approach presented in this study provide insight into how more rigorous planning for western blot experimental design and corresponding statistical analysis without depending on p-values only can acquire precise data resulting in true biological effects. Using blotRig as a standardized, integrated western blot methodology, quantitative western blot may become highly reproducible, reliable, and a less controversial protein measurement technique18,28,67.

Study reporting

This study is reported in accordance with ARRIVE guidelines.

Supporting information

This article contains supporting information. You can access the blotRig analysis software, which includes code for inputting experimental parameters for all Western blot analysis, through the following link: https://atpspin.shinyapps.io/BlotRig/

Supplementary Information

Supplementary Figure 1.

Supplementary Figure 2.

Abbreviations

AAALAC American association for accreditation of laboratory animal care

ARRIVE Animal research reporting of in vivo experiments

AVMA American veterinary medical association

IACUC Institutional animal care and use committee

qWB Quantitative western blot

ELISA Enzyme linked immunosorbent assay

SARS-CoV2 Severe acute respiratory syndrome coronavirus 2

ANCOVA Analysis of covariance

ANOVA Analysis of variance

SCI Spinal cord injury

SNI Spared nerve injury

AMPA α-Amino-3-hydroxy-5-methyl-4-isoxazoleproprionic acid

GluA1 Glutamate receptor 1

GluA2 Glutamate receptor 2

LMM Linear mixed models

TTBS Tris-buffered saline containing 0.1% Tween 20

PAGE Polyacrylamide gel electrophoresis

SEM Standard error of mean

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-024-70096-0.

Acknowledgements

The authors would like to thank Alexys Maliga Davis for data librarian services.

Author contributions

C.O: Writing-original draft preparation, Investigation, Validation, Data Curation, Visualization, Formal Analysis, Writing-Review & Editing; A. C: Formal analysis, Writing-Review & Editing; K. A. F: Software, Writing-Review & Editing; K. M: Methodology, Writing-Review & Editing; N. R. J: Writing—Review & Editing; J. A. S: Investigation, Project Administration, Writing—Review & Editing; E. I: Resources, Writing-Review & Editing; A.T.E: Software, Writing—Review & Editing; H. L. R: Software, Writing—Review & Editing; J. A. D: Investigation, Writing—Review & Editing; J. H. G: Investigation, Writing – Review & Editing; J. R. H: Conceptualization, Methodology, Validation, Formal Analysis, Investigation, Data Curation, Writing- Review & Editing, Visualization, Supervision; A. R. F:Conceptualization, Methodology, Validation, Formal Analysis, Resources, Investigation, Data Curation, Writing-Review & Editing, Visualization, Supervision, Project Administration, Funding Acquisition.

Funding

This work was supported by a National Institutes of Health/National Institute of Neurological Disorders and Stroke grant (R01NS088475) to A. R. F. NIH NINDS: R01NS122888, UH3NS106899, U24NS122732, US Veterans Affairs (VA): I01RX002245, I01RX002787, I50BX005878, Wings for Life Foundation, Craig H. Neilsen Foundation. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Correspondence and requests for materials should be addressed to A.R.F.

Data availability

The datasets and computer code generated or used in this study are accessible in a public, open-access repository at 10.34945/F51C7B and https://github.com/ucsf-ferguson-lab/blotRig/ respectively.

Competing interests

The authors declare no competing interests.

Publisher's note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
==== Refs
References

1. Lowry O Rosebrough N Farr AL Randall R Protein measurement with the Folin phenol reagent J. Biol. Chem. 1951 193 265 275 14907713
Lowry, O., Rosebrough, N., Farr, A. L. & Randall, R. Protein measurement with the Folin phenol reagent. J. Biol. Chem. 193, 265–275. 10.1016/S0021-9258(19)52451-6 (1951).14907713
2. Aldridge GM Podrebarac DM Greenough WT Weiler IJ The use of total protein stains as loading controls: An alternative to high-abundance single protein controls in semi-quantitative immunoblotting J. Neurosci. Methods 2008 172 250 254 18571732
Aldridge, G. M., Podrebarac, D. M., Greenough, W. T. & Weiler, I. J. The use of total protein stains as loading controls: An alternative to high-abundance single protein controls in semi-quantitative immunoblotting. J. Neurosci. Methods 172, 250–254. 10.1016/j.jneumeth.2008.05.00 (2008).18571732
3. McDonough AA Veiras LC Minas JN Ralph DL Considerations when quantitating protein abundance by immunoblot Am. J. Physiol. Cell Physiol. 2015 308 C426 433 25540176
McDonough, A. A., Veiras, L. C., Minas, J. N. & Ralph, D. L. Considerations when quantitating protein abundance by immunoblot. Am. J. Physiol. Cell Physiol. 308, C426-433. 10.1152/ajpcell.00400.2014 (2015).25540176
4. Towbin H Staehelin T Gordon J Electrophoretic transfer of proteins from polyacrylamide gels to nitrocellulose sheets: Procedure and some applications PNAS 1979 76 4350 4354 388439
Towbin, H., Staehelin, T. & Gordon, J. Electrophoretic transfer of proteins from polyacrylamide gels to nitrocellulose sheets: Procedure and some applications. PNAS 76, 4350–4354. 10.1073/pnas.76.9.4350 (1979).388439
5. Burnette WN “Western blotting”: Electrophoretic transfer of proteins from sodium dodecyl sulfate-polyacrylamide gels to unmodified nitrocellulose and radiographic detection with antibody and radioiodinated protein A Anal. Biochem. 1981 112 195 203 6266278
Burnette, W. N. “Western blotting”: Electrophoretic transfer of proteins from sodium dodecyl sulfate-polyacrylamide gels to unmodified nitrocellulose and radiographic detection with antibody and radioiodinated protein A. Anal. Biochem. 112, 195–203. 10.1016/0003-2697(81)90281-5 (1981).6266278
6. Mahmood T Yang P-C Western blot: Technique, theory, and trouble shooting N. Am. J. Med. Sci. 2012 4 429 434 23050259
Mahmood, T. & Yang, P.-C. Western blot: Technique, theory, and trouble shooting. N. Am. J. Med. Sci. 4, 429–434. 10.4103/1947-2714.100998 (2012).23050259
7. Alegria-Schaffer A Lodge A Vattem K Performing and optimizing Western blots with an emphasis on chemiluminescent detection Methods Enzymol. 2009 463 573 599 19892193
Alegria-Schaffer, A., Lodge, A. & Vattem, K. Performing and optimizing Western blots with an emphasis on chemiluminescent detection. Methods Enzymol. 463, 573–599. 10.1016/S0076-6879(09)63033-0 (2009).19892193
8. Khoury MK Parker I Aswad DW Acquisition of chemiluminescent signals from immunoblots with a digital SLR camera Anal. Biochem. 2010 397 129 131 19788886
Khoury, M. K., Parker, I. & Aswad, D. W. Acquisition of chemiluminescent signals from immunoblots with a digital SLR camera. Anal. Biochem. 397, 129–131. 10.1016/j.ab.2009.09.041 (2010).19788886
9. Zellner M Babeluk R Diestinger M Pirchegger P Skeledzic S Oehler R Fluorescence-based western blotting for quantitation of protein biomarkers in clinical samples Electrophoresis 2008 29 3621 3627 18803224
Zellner, M. et al. Fluorescence-based western blotting for quantitation of protein biomarkers in clinical samples. Electrophoresis 29, 3621–3627. 10.1002/elps.200700935 (2008).18803224
10. Gingrich JC Davis DR Nguyen Q Multiplex detection and quantitation of proteins on western blots using fluorescent probes Biotechniques 2000 29 636 642 10997278
Gingrich, J. C., Davis, D. R. & Nguyen, Q. Multiplex detection and quantitation of proteins on western blots using fluorescent probes. Biotechniques 29, 636–642. 10.2144/00293pf02 (2000).10997278
11. Janes KA An analysis of critical factors for quantitative immunoblotting Sci. Signal 2015 8 rs2 25852189
Janes, K. A. An analysis of critical factors for quantitative immunoblotting. Sci. Signal 8, rs2. 10.1126/scisignal.2005966 (2015).25852189
12. Mollica JP Oakhill JS Lamb GD Murphy RM Are genuine changes in protein expression being overlooked? Reassessing western blotting Anal. Biochem. 2009 386 270 275 19161968
Mollica, J. P., Oakhill, J. S., Lamb, G. D. & Murphy, R. M. Are genuine changes in protein expression being overlooked? Reassessing western blotting. Anal. Biochem. 386, 270–275. 10.1016/j.ab.2008.12.029 (2009).19161968
13. Pillai-Kastoori L Schutz-Geschwender AR Harford JA A systematic approach to quantitative western blot analysis Anal. Biochem. 2020 593 113608 32007473
Pillai-Kastoori, L., Schutz-Geschwender, A. R. & Harford, J. A. A systematic approach to quantitative western blot analysis. Anal. Biochem. 593, 113608. 10.1016/j.ab.2020.113608 (2020).32007473
14. Aydin S A short history, principles, and types of ELISA, and our laboratory experience with peptide/protein analyses using ELISA Peptides 2015 72 4 15 25908411
Aydin, S. A short history, principles, and types of ELISA, and our laboratory experience with peptide/protein analyses using ELISA. Peptides 72, 4–15. 10.1016/j.peptides.2015.04.012 (2015).25908411
15. Seisenberger C Graf T Haindl M Wegele H Wiedmann M Wohlrab S Questioning coverage values determined by 2D western blots: A critical study on the characterization of anti-HCP ELISA reagents Biotechnol. Bioeng. 2021 118 1116 1126 33241851
Seisenberger, C. et al. Questioning coverage values determined by 2D western blots: A critical study on the characterization of anti-HCP ELISA reagents. Biotechnol. Bioeng. 118, 1116–1126. 10.1002/bit.27635 (2021).33241851
16. Edwards VM Mosley JW Reproducibility in quality control of protein (western) immunoblot assay for antibodies to human immunodeficiency virus Am. J. Clin. Pathol. 1989 91 75 78 2910017
Edwards, V. M. & Mosley, J. W. Reproducibility in quality control of protein (western) immunoblot assay for antibodies to human immunodeficiency virus. Am. J. Clin. Pathol. 91, 75–78. 10.1093/ajcp/91.1.75 (1989).2910017
17. Matschke J Lütgehetmann M Hagel C Sperhake JP Schröder AS Edler C Mushumba H Fitzek A Allweiss L Dandri M Dottermusch M Heinemann A Pfefferle S Schwabenland M Sumner Magruder D Bonn S Prinz M Gerloff C Püschel K Krasemann S Aepfelbacher M Glatzel M Neuropathology of patients with COVID-19 in Germany: A post-mortem case series Lancet Neurol. 2020 19 919 929 33031735
Matschke, J. et al. Neuropathology of patients with COVID-19 in Germany: A post-mortem case series. Lancet Neurol. 19, 919–929. 10.1016/S1474-4422(20)30308-2 (2020).33031735
18. Murphy RM Lamb GD Important considerations for protein analyses using antibody based techniques: Down-sizing western blotting up-sizes outcomes J. Physiol. 2013 591 5823 5831 24127618
Murphy, R. M. & Lamb, G. D. Important considerations for protein analyses using antibody based techniques: Down-sizing western blotting up-sizes outcomes. J. Physiol. 591, 5823–5831. 10.1113/jphysiol.2013.263251 (2013).24127618
19. Butler TAJ Paul JW Chan E-C Smith R Tolosa JM Misleading westerns: Common quantification mistakes in western blot densitometry and proposed corrective measures Biomed. Res. Int. 2019 2019 5214821 30800670
Butler, T. A. J., Paul, J. W., Chan, E.-C., Smith, R. & Tolosa, J. M. Misleading westerns: Common quantification mistakes in western blot densitometry and proposed corrective measures. Biomed. Res. Int. 2019, 5214821. 10.1155/2019/5214821 (2019).30800670
20. Button KS Ioannidis JPA Mokrysz C Nosek BA Flint J Robinson ESJ Munafò MR Power failure: Why small sample size undermines the reliability of neuroscience Nat. Rev. Neurosci. 2013 14 365 376 23571845
Button, K. S. et al. Power failure: Why small sample size undermines the reliability of neuroscience. Nat. Rev. Neurosci. 14, 365–376. 10.1038/nrn3475 (2013).23571845
21. Landis SC Amara SG Asadullah K Austin CP Blumenstein R Bradley EW Crystal RG Darnell RB Ferrante RJ Fillit H Finkelstein R Fisher M Gendelman HE Golub RM Goudreau JL Gross RA Gubitz AK Hesterlee SE Howells DW Huguenard J Kelner K Koroshetz W Krainc D Lazic SE Levine MS Macleod MR McCall JM Moxley RT Narasimhan K Noble LJ Perrin S Porter JD Steward O Unger E Utz U Silberberg SD A call for transparent reporting to optimize the predictive value of preclinical research Nature 2012 490 187 191 23060188
Landis, S. C. et al. A call for transparent reporting to optimize the predictive value of preclinical research. Nature 490, 187–191. 10.1038/nature11556 (2012).23060188
22. Shields SD Eckert WA Basbaum AI Spared nerve injury model of neuropathic pain in the mouse: A behavioral and anatomic analysis J. Pain 2003 4 465 470 14622667
Shields, S. D., Eckert, W. A. & Basbaum, A. I. Spared nerve injury model of neuropathic pain in the mouse: A behavioral and anatomic analysis. J. Pain 4, 465–470. 10.1067/s1526-5900(03)00781-8 (2003).14622667
23. Decosterd I Woolf C Spared nerve injury: An animal model of persistent peripheral neuropathic pain Pain 2000 87 149 158 10924808
Decosterd, I. & Woolf, C. Spared nerve injury: An animal model of persistent peripheral neuropathic pain. Pain 87, 149–158. 10.1016/S0304-3959(00)00276-1 (2000).10924808
24. Richner M Jager SB Siupka P Vaegter CB Hydraulic extrusion of the spinal cord and isolation of dorsal root ganglia in rodents J. Vis. Exp. 2017 28190031
Richner, M., Jager, S. B., Siupka, P. & Vaegter, C. B. Hydraulic extrusion of the spinal cord and isolation of dorsal root ganglia in rodents. J. Vis. Exp.10.3791/55226 (2017).28190031
25. Ferguson AR Christensen RN Gensel JC Miller BA Sun F Beattie EC Bresnahan JC Beattie MS Cell death after spinal cord injury is exacerbated by rapid TNFα-induced trafficking of GluR2-lacking AMPARS to the plasma membrane J Neurosci 2008 28 11391 11400 18971481
Ferguson, A. R. et al. Cell death after spinal cord injury is exacerbated by rapid TNFα-induced trafficking of GluR2-lacking AMPARS to the plasma membrane. J Neurosci 28, 11391–11400. 10.1523/JNEUROSCI.3708-08.2008 (2008).18971481
26. Ferguson AR Huie JR Crown ED Grau JW Central nociceptive sensitization vs. spinal cord training: Opposing forms of plasticity that dictate function after complete spinal cord injury Front. Physiol. 2012 3 1 22275902
Ferguson, A. R., Huie, J. R., Crown, E. D. & Grau, J. W. Central nociceptive sensitization vs. spinal cord training: Opposing forms of plasticity that dictate function after complete spinal cord injury. Front. Physiol. 3, 1. 10.3389/fphys.2012.00396 (2012).22275902
27. Taylor SC Berkelman T Yadav G Hammond M A defined methodology for reliable quantification of western blot data Mol. Biotechnol. 2013 55 217 226 23709336
Taylor, S. C., Berkelman, T., Yadav, G. & Hammond, M. A defined methodology for reliable quantification of western blot data. Mol. Biotechnol. 55, 217–226. 10.1007/s12033-013-9672-6 (2013).23709336
28. Bakkenist CJ Czambel RK Hershberger PA Tawbi H Beumer JH Schmitz JC A quasi-quantitative dual multiplexed immunoblot method to simultaneously analyze ATM and H2AX phosphorylation in human peripheral blood mononuclear cells Oncoscience 2015 2 542 554 26097887
Bakkenist, C. J. et al. A quasi-quantitative dual multiplexed immunoblot method to simultaneously analyze ATM and H2AX phosphorylation in human peripheral blood mononuclear cells. Oncoscience 2, 542–554. 10.18632/oncoscience.162 (2015).26097887
29. Wang YV Wade M Wong E Li Y-C Rodewald LW Wahl GM Quantitative analyses reveal the importance of regulated Hdmx degradation for p53 activation Proc. Natl. Acad. Sci. USA 2007 104 12365 12370 17640893
Wang, Y. V. et al. Quantitative analyses reveal the importance of regulated Hdmx degradation for p53 activation. Proc. Natl. Acad. Sci. USA 104, 12365–12370. 10.1073/pnas.0701497104 (2007).17640893
30. Bass J Wilkinson D Rankin D Phillips B Szewczyk N Smith K Atherton P An overview of technical considerations for western blotting applications to physiological research Scand. J. Med. Sci. Sports 2017 27 4 25 27263489
Bass, J. et al. An overview of technical considerations for western blotting applications to physiological research. Scand. J. Med. Sci. Sports 27, 4–25. 10.1111/sms.12702 (2017).27263489
31. Lazzeroni LC Ray A The cost of large numbers of hypothesis tests on power, effect size and sample size Mol. Psychiatry 2012 17 108 114 21060308
Lazzeroni, L. C. & Ray, A. The cost of large numbers of hypothesis tests on power, effect size and sample size. Mol. Psychiatry 17, 108–114. 10.1038/mp.2010.117 (2012).21060308
32. Huie JR Stuck ED Lee KH Irvine K-A Beattie MS Bresnahan JC Grau JW Ferguson AR AMPA receptor phosphorylation and synaptic colocalization on motor neurons drive maladaptive plasticity below complete spinal cord injury eNeuro 2015 26668821
Huie, J. R. et al. AMPA receptor phosphorylation and synaptic colocalization on motor neurons drive maladaptive plasticity below complete spinal cord injury. eNeuro10.1523/ENEURO.0091-15.2015 (2015).26668821
33. Stück ED Christensen RN Huie JR Tovar CA Miller BA Nout YS Bresnahan JC Beattie MS Ferguson AR Tumor necrosis factor alpha mediates GABAA receptor trafficking to the plasma membrane of spinal cord neurons in vivo Neural Plast 2012 22530155
Stück, E. D. et al. Tumor necrosis factor alpha mediates GABAA receptor trafficking to the plasma membrane of spinal cord neurons in vivo. Neural Plast10.1155/2012/261345 (2012).22530155
34. Krzywinski M Altman N Points of significance: Power and sample size Nat. Method. 2013 10 1139 1140
Krzywinski, M. & Altman, N. Points of significance: Power and sample size. Nat. Method. 10, 1139–1140. 10.1038/nmeth.2738 (2013).
35. R Core Team R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/ (2021).
36. Green, P. & MacLeod C. J. “simr: An R package for power analysis of generalised linear mixed models by simulation.” Meth. Ecol. Evolut. 7(4), 493–498. 10.1111/2041-210X.12504, https://CRAN.R-project.org/package=simr (2016).
37. Attali, D. shinyjs: Easily Improve the User Experience of Your Shiny Apps in Seconds. R package version 2.1.0, https://deanattali.com/shinyjs/ (2022).
38. Chang, W. et al. shiny: Web Application Framework for R. R package version 1.9.1.9000, https://github.com/rstudio/shiny, https://shiny.posit.co/ (2024).
39. Chang, W. shinythemes: Themes for Shiny. R package version 1.2.0, https://github.com/rstudio/shinythemes (2024).
40. de Vries, A., Schloerke, B., Russell, K. sortable: Drag-and-Drop in ‘shiny’ Apps with ‘SortableJS’. R package version 0.5.0, https://github.com/rstudio/sortable (2024).
41. Wickham H Averick M Bryan J Chang W McGowan L François R Grolemund G Hayes A Henry L Hester J Kuhn M Pedersen T Miller E Bache S Müller K Ooms J Robinson D Seidel D Spinu V Takahashi K Vaughan D Wilke C Woo K Yutani H Welcome to the tidyverse JOSS 2019 4 43 1686
Wickham, H. et al. Welcome to the tidyverse. JOSS 4(43), 1686. 10.21105/joss.01686 (2019).
42. Xie, Y., Cheng, J., Tan, X. DT: A Wrapper of the JavaScript Library ‘DataTables’. R package version 0.33.1, dt. https://github.com/rstudio/ (2024).
43. Krzywinski M Altman N Points of significance: Analysis of variance and blocking Nat Methods 2014 11 699 700 25110779
Krzywinski, M. & Altman, N. Points of significance: Analysis of variance and blocking. Nat Methods 11, 699–700. 10.1038/nmeth.3005 (2014).25110779
44. Heidebrecht F Heidebrecht A Schulz I Behrens S-E Bader A Improved semiquantitative western blot technique with increased quantification range J. Immunol. Methods 2009 345 40 48 19351538
Heidebrecht, F., Heidebrecht, A., Schulz, I., Behrens, S.-E. & Bader, A. Improved semiquantitative western blot technique with increased quantification range. J. Immunol. Methods 345, 40–48. 10.1016/j.jim.2009.03.018 (2009).19351538
45. Huang Y-T van der Hoorn D Ledahawsky LM Motyl AAL Jordan CY Gillingwater TH Groen EJN Robust comparison of protein levels across tissues and throughout development using standardized quantitative western blotting J. Vis. Exp. 2019 31904744
Huang, Y.-T. et al. Robust comparison of protein levels across tissues and throughout development using standardized quantitative western blotting. J. Vis. Exp.10.3791/59438 (2019).31904744
46. Krzywinski M Altman N Points of view: Designing comparative experiments Nat. Methods 2014 11 597 598 25019145
Krzywinski, M. & Altman, N. Points of view: Designing comparative experiments. Nat. Methods 11, 597–598. 10.1038/nmeth.2974 (2014).25019145
47. Thacker JS Yeung DH Staines WR Mielke JG Total protein or high-abundance protein: Which offers the best loading control for western blotting? Anal. Biochem. 2016 496 76 78 26706797
Thacker, J. S., Yeung, D. H., Staines, W. R. & Mielke, J. G. Total protein or high-abundance protein: Which offers the best loading control for western blotting?. Anal. Biochem. 496, 76–78. 10.1016/j.ab.2015.11.022 (2016).26706797
48. Zeng L Guo J Xu H-B Huang R Shao W Yang L Wang M Chen J Xie P Direct blue 71 staining as a destaining-free alternative loading control method for western blotting Electrophoresis 2013 34 2234 2239 23712695
Zeng, L. et al. Direct blue 71 staining as a destaining-free alternative loading control method for western blotting. Electrophoresis 34, 2234–2239. 10.1002/elps.201300140 (2013).23712695
49. Jaeger TF Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models J. Mem. Lang. 2008 59 434 446 19884961
Jaeger, T. F. Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. J. Mem. Lang. 59, 434–446. 10.1016/j.jml.2007.11.007 (2008).19884961
50. Mefford J Witte JS The covariate’s dilemma PLoS Genet. 2012 8 e1003096 23162385
Mefford, J. & Witte, J. S. The covariate’s dilemma. PLoS Genet. 8, e1003096. 10.1371/journal.pgen.1003096 (2012).23162385
51. Schneider BA Avivi-Reich M Mozuraitis M A cautionary note on the use of the analysis of covariance (ANCOVA) in classification designs with and without within-subject factors Front. Psychol. 2015 6 474 25954230
Schneider, B. A., Avivi-Reich, M. & Mozuraitis, M. A cautionary note on the use of the analysis of covariance (ANCOVA) in classification designs with and without within-subject factors. Front. Psychol. 6, 474. 10.3389/fpsyg.2015.00474 (2015).25954230
52. Nieuwenhuis S Forstmann BU Wagenmakers E-J Erroneous analyses of interactions in neuroscience: A problem of significance Nat. Neurosci. 2011 14 1105 1107 21878926
Nieuwenhuis, S., Forstmann, B. U. & Wagenmakers, E.-J. Erroneous analyses of interactions in neuroscience: A problem of significance. Nat. Neurosci. 14, 1105–1107. 10.1038/nn.2886 (2011).21878926
53. Freeberg TM Lucas JR Pseudoreplication is (still) a problem J. Com. Psychol. 2009 123 450 451
Freeberg, T. M. & Lucas, J. R. Pseudoreplication is (still) a problem. J. Com. Psychol. 123, 450–451. 10.1037/a0017031 (2009).
54. Judd CM Westfall J Kenny DA Treating stimuli as a random factor in social psychology: A new and comprehensive solution to a pervasive but largely ignored problem J. Pers. Soc. Psychol. 2012 103 54 69 22612667
Judd, C. M., Westfall, J. & Kenny, D. A. Treating stimuli as a random factor in social psychology: A new and comprehensive solution to a pervasive but largely ignored problem. J. Pers. Soc. Psychol. 103, 54–69. 10.1037/a0028347 (2012).22612667
55. Lee OE Braun TM Permutation tests for random effects in linear mixed models Biometrics 2012 68 486 493 21950470
Lee, O. E. & Braun, T. M. Permutation tests for random effects in linear mixed models. Biometrics 68, 486–493. 10.1111/j.1541-0420.2011.01675.x (2012).21950470
56. Baayen RH Davidson DJ Bates DM Mixed-effects modeling with crossed random effects for subjects and items J. Mem. Lang 2008 59 390 412
Baayen, R. H., Davidson, D. J. & Bates, D. M. Mixed-effects modeling with crossed random effects for subjects and items. J. Mem. Lang 59, 390–412. 10.1016/j.jml.2007.12.005 (2008).
57. Barr DJ Levy R Scheepers C Tily HJ Random effects structure for confirmatory hypothesis testing: Keep it maximal J. Mem. Lang. 2013 24403724
Barr, D. J., Levy, R., Scheepers, C. & Tily, H. J. Random effects structure for confirmatory hypothesis testing: Keep it maximal. J. Mem. Lang.10.1016/j.jml.2012.11.001 (2013).24403724
58. Blainey P Krzywinski M Altman N Points of significance: Replication Nat. Methods 2014 11 879 880 25317452
Blainey, P., Krzywinski, M. & Altman, N. Points of significance: Replication. Nat. Methods 11, 879–880. 10.1038/nmeth.3091 (2014).25317452
59. Drubin DG Great science inspires us to tackle the issue of data reproducibility Mol. Biol. Cell 2015 26 3679 3680 26515968
Drubin, D. G. Great science inspires us to tackle the issue of data reproducibility. Mol. Biol. Cell 26, 3679–3680. 10.1091/mbc.E15-09-0643 (2015).26515968
60. Amrhein V Greenland S McShane B Scientists rise up against statistical significance Nature 2019 567 305 307 30894741
Amrhein, V., Greenland, S. & McShane, B. Scientists rise up against statistical significance. Nature 567, 305–307. 10.1038/d41586-019-00857-9 (2019).30894741
61. Cohen J The earth is round (p <.05) Am. Psychol. 1994 49 997 1003
Cohen, J. The earth is round (p <.05). Am. Psychol. 49, 997–1003. 10.1037/0003-066X.49.12.997 (1994).
62. Ioannidis JPA Tarone R McLaughlin JK The false-positive to false-negative ratio in epidemiologic studies Epidemiology 2011 22 450 456 21490505
Ioannidis, J. P. A., Tarone, R. & McLaughlin, J. K. The false-positive to false-negative ratio in epidemiologic studies. Epidemiology 22, 450–456. 10.1097/EDE.0b013e31821b506e (2011).21490505
63. Sullivan GM Feinn R Using effect size-or why the P value Is not enough J. Grad. Med. Educ. 2012 4 279 282 23997866
Sullivan, G. M. & Feinn, R. Using effect size-or why the P value Is not enough. J. Grad. Med. Educ. 4, 279–282. 10.4300/JGME-D-12-00156.1 (2012).23997866
64. Brysbaert M Stevens M Power analysis and effect size in mixed effects models: A tutorial J. Cogn. 2018 1 9 31517183
Brysbaert, M. & Stevens, M. Power analysis and effect size in mixed effects models: A tutorial. J. Cogn. 1, 9. 10.5334/joc.10 (2018).31517183
65. Kline, R. B. Beyond significance testing: Reforming data analysis methods in behavioral research. Am. Psychol. Associat. 10.1037/10693-000 (2024).
66. Rosner, Bernard (Bernard A.). Fundamentals of biostatistics. (Boston, Brooks/Cole, Cengage Learning, 2011).
67. Bromage E Carpenter L Kaattari S Patterson M Quantification of coral heat shock proteins from individual coral polyps Mar. Ecol. Progress Ser. 2009 376 123 132
Bromage, E., Carpenter, L., Kaattari, S. & Patterson, M. Quantification of coral heat shock proteins from individual coral polyps. Mar. Ecol. Progress Ser. 376, 123–132 (2009).
