
==== Front
Proc Natl Acad Sci U S A
Proc Natl Acad Sci U S A
PNAS
Proceedings of the National Academy of Sciences of the United States of America
0027-8424
1091-6490
National Academy of Sciences

39236231
202404035
10.1073/pnas.2404035121
persPerspectivesoc-scienceSocial Sciences432
447
Perspective
Social Sciences
Social Sciences
Toward a more credible assessment of the credibility of science by many-analyst studies
Auspurg Katrin Katrin.auspurg@lmu.de
a 1 https://orcid.org/0000-0003-4504-0391

Brüderl Josef a https://orcid.org/0000-0001-8636-9922

aDepartment of Sociology, Ludwig-Maximilians-Universität (LMU) Munich, Munich 80801, Germany
1To whom correspondence may be addressed. Email: Katrin.auspurg@lmu.de.
Edited by Douglas Massey, Princeton University, Princeton, NJ; received May 13, 2024; accepted July 5, 2024

5 9 2024
17 9 2024
5 9 2024
121 38 e2404035121Copyright © 2024 the Author(s). Published by PNAS.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This open access article is distributed under Creative Commons Attribution-NonCommercial-NoDerivatives License 4.0 (CC BY-NC-ND).

We discuss a relatively new meta-scientific research design: many-analyst studies that attempt to assess the replicability and credibility of research based on large-scale observational data. In these studies, a large number of analysts try to answer the same research question using the same data. The key idea is the greater the variation in results, the greater the uncertainty in answering the research question and, accordingly, the lower the credibility of any individual research finding. Compared to individual replications, the large crowd of analysts allows for a more systematic investigation of uncertainty and its sources. However, many-analyst studies are also resource-intensive, and there are some doubts about their potential to provide credible assessments. We identify three issues that any many-analyst study must address: 1) identifying the source of variation in the results; 2) providing an incentive structure similar to that of standard research; and 3) conducting a proper meta-analysis of the results. We argue that some recent many-analyst studies have failed to address these issues satisfactorily and have therefore provided an overly pessimistic assessment of the credibility of science. We also provide some concrete guidance on how future many-analyst studies could provide a more constructive assessment.

meta-science
many-analyst projects
replicability
credibility of science
meta-analysis
German Research Foundation AU 394/5-1 Katrin Auspurg
==== Body
pmcCentral to the nature of science is the ability of the scientific community to scrutinize scientific claims (1). The better scientific findings are repeated through independent replications, the higher their certainty and credibility. Replications thus help to confirm existing research, but also to test its robustness to small changes in design, such as modifications in the statistical model used to analyze data (2–4).

In this Perspective, we discuss an increasingly used replication design: many-analyst studies. In these studies, several teams of analysts examine the same research question using the same data. Beginning with the seminal project by Silberzahn and Uhlmann published in 2015 (5) (hereafter SU), many-analyst studies have become a popular design for evaluating research with large-scale observational data in the social sciences, but also in other fields such as the neuro and life sciences (see, for example, refs. 6–10). In the SU study, 29 teams analyzed the same sports data to answer the question of whether soccer players with darker skin tones were more likely to receive a red card than soccer players with lighter skin tones. Another prominent example is the study by Breznau, Rinke, and Wuttke et al. (11) (hereafter BRW), in which 73 teams analyzed the same survey data to test the hypothesis that immigration reduces public support for social policies. Both studies found wide variation in analysts’ results, with some estimates being statistically significant, others not, or with different signs. The studies were able to explain only a very small fraction of the variation in results, leaving one to speculate about the source of the large amount of uncertainty. The conclusions of these studies were therefore alarming: No single study result should be taken seriously; there appears to be a large “hidden universe” of uncertainty. These studies are now a major source of evidence that (social) science research has a credibility problem, cited in hundreds of articles and discussed in several commentaries on the need to improve science. Other many-analyst studies have also found substantial cross-analyst variation in results (for some reviews: refs. 7 and 12).

However, many-analyst studies themselves have proved controversial, with some seeing them as an essential tool for detecting the potential fragility of findings (13, 14) and others seeing them as exaggerated in their overall conclusions (15–17). It is also an open question whether the method is worth the effort, given that studies typically involve the research work of dozens or even hundreds of researchers, which is then lost to other (original) research. [For example, a recent many-analyst study estimated the total time required at 27 full-time equivalent person-years (9).]

To date, systematic evaluations and practical guidelines are lacking (for an exception on few methodological aspects: ref. 12). First, the lack of standardization makes it difficult to synthesize the results of different studies. Second, variations in results can occur for a variety of reasons that do not necessarily indicate that something is wrong (1, 8). For example, differences in estimates may simply indicate substantively meaningful variation that is due to different target populations or estimands that are studied. Constructive many-analyst studies allow these sources of variation to be separated from more problematic sources such as model misspecification or uncertainties in data weighting and preparation. Third, due to the lack of adherence to meta-analytic standards, some of the diagnosed uncertainty is likely not caused by analysts conducting standard research, but by Principal Investigators (PIs) introducing artificial uncertainty when pooling and meta-analyzing analysts’ results. By overlooking these issues, many of the existing many-analyst studies may have overestimated the amount of uncertainty, leading to an overly pessimistic picture of the credibility of science.

We aim to advance the use of many-analyst studies by discussing how they can best unfold their potential. We begin with a brief discussion of why many-analyst studies are in general a promising research design. We then discuss and provide guidance on three issues that any many-analyst study must address. Our intended audience is researchers planning to conduct a many-analyst study, as well as scholars interested in interpreting the results.

1. Why Many-Analyst Studies?

Standard research articles express uncertainty in estimates by SEs or CIs. These show the uncertainty caused by sampling: When observing only a random subsample of the population of interest, there is inevitably some random deviation from the true population parameters. However, especially when analyzing large-scale observational data, there are many other sources of uncertainty. For example, one must decide how to operationalize treatment and outcome variables, which control variables to use, whether to specify linear or nonlinear relationships between variables, and make many other decisions that can introduce variation in results. This has been termed a “garden of forking paths” (18), “modeling uncertainty” (19, 20), or “excess heterogeneity” (i.e., heterogeneity that exceeds sampling error, 21, 22).

Meta-research attempts to quantify this uncertainty. To identify variation beyond the sampling error, one must not resample data, but rather reanalyze existing data with different approaches to data preparation or analysis (which by design keeps the sampling error constant; see ref. 23 for why this replication design is especially useful with observational data). Nevertheless, there are different designs for quantifying the amount of modeling uncertainty. First, there is the classical approach of performing a reanalysis (also called “robustness reproduction”) of an original study (24). However, similar to original research, a single reanalysis may suffer from “cherry-picked” results. This is because results that contradict existing research are more likely to be published than those that confirm it (25). Researchers involved in a reanalysis may therefore also show only few extreme findings that result from fishing in the garden of forking paths (for evidence: ref. 26). This incentive structure can lead to a literature that alternates between extreme research claims and refutations (25, 27, 28). It is then difficult for readers to judge which results are more credible (the original result or the reanalysis).

Second, it has recently been suggested that “multiverse” analyses (also called “specification curves,” “multimodel” or “multistrategy” analyses, 20, 29–31) can be used to quantify the amount of modeling uncertainty. The idea here is to show not just one main estimate, but rather the distribution of estimates that result from running all defensible data processing and analysis pipelines (i.e., the whole garden of “forking paths”). But similar incentive problems apply here as with a single reanalysis. Multiverse analysts may choose questionable paths (e.g., include collider variables that should be avoided, but which produce extreme results; 32).

Third, many-analyst studies have recently been introduced. By having analysts use the same data, one can abstract from sampling error and identify uncertainty beyond that source. Unlike a single reanalysis or a multiverse study, the definition of “defensible” research decisions is no longer in the hands of one or a few authors. Instead, it is the “wisdom of the crowd” that defines the appropriate data preparation and identification strategies. In addition, the problematic incentive structure affecting a single reanalysis could be overcome by guaranteeing analysts coauthorship independent of the delivery of a “significant” or “remarkable” finding (9, 11). Thus, if done properly, many-analyst studies hold the promise of better understanding modeling uncertainty inherent in standard scientific work with observational data (1).

2. Three Issues That Any Many-Analyst Study Must Address

As a measure of the amount of modeling uncertainty in standard research, many-analyst studies use variation in the results produced by the many analysts. In doing so, any many-analyst study must address three issues:1) Meta-estimand. There are multiple sources of variation in research results. Not all variation indicates problematic “uncertainty.” Therefore, any many-analyst study should clearly define what uncertainty it wants to estimate (its meta-estimand). Second, to be productive, it should try to separate different sources of variation.

2) External validity. The many-analyst design creates an artificial research environment. To be externally valid, the incentive structure should be similar to standard research.

3) Internal validity. For correctly identifying the amount and sources of variation in results, many-analyst studies must use appropriate meta-analytic tools.

2.1. Many-Analyst Studies Should Define Uncertainty and Separate Sources of Variation.

2.1.1. “Uncertainty” vs. other causes of variation.

Existing many-analyst studies tend to interpret any variation in findings as problematic uncertainty hampering progress in science. According to theoretical meta-science work (1), this interpretation needs to be qualified depending on the source of the variation. We propose to conceptualize these sources as shown in Table 1.

Table 1. Sources of variation in research findings when exploring a broad research question

Problematic “uncertainty” (impeding progress of science)

	a) Model uncertainty due to visible decisions in

i. Data processing (coding of variables, weighting, imputing, etc.)

ii. Modeling assumptions (different identification assumptions, different estimation models)

b) Coding errors and other typically “hidden” sources of variation (modeling decisions not reported by researchers)

	
Helpful variation (advancing theories)

	c) Substantively meaningful variation across

i. Different target populations (variation over time or space)

ii. Different estimands (different definitions of outcome variables, treatment variables, conceptualizations of causal effects and counterfactuals: e.g., total vs. direct or indirect effect)

	

According to recent discussions (1), only a) and b) can be considered “problematic” because they make scientific claims more uncertain and thus less credible. a) reflects a lack of certainty about the correct approach to identifying an estimand of interest (e.g., it is often not clear which sample weight would allow obtaining the best unbiased population parameter). b) introduces random or systematic variation due to sloppiness or lack of transparency (errors, failures to report consequential research decisions). Both sources of uncertainty can be considered problematic: They prevent clear answers to research questions and thus slow down or even prevent scientific progress. b) is additionally problematic because this source of variance is not directly visible to readers of a study. One might call variation caused by b) “hidden uncertainty.” Especially this source of variation impacts certainty and trust in science.

In contrast, the variation in results due to c) is productive for scientific progress because it provides information about effect heterogeneity and scope conditions for our theories. This variation is not due to misleading research processes but is variation that exists in the real world. Science is the more credible, the better one is able to assess and predict this variation across different populations or estimands. Therefore, the goal of good scientific practice is not to remove this variation but to correctly identify, theoretically conceptualize, and confirm this variation in reanalyses.

Most many-analyst studies have not clearly defined their meta-estimand. Instead, they have implicitly equated all variation with uncertainty. An example is the BRW study (11). In this study, analysts used different subsamples of the data to focus on different target populations (e.g., countries). However, it is already known that the impact of immigration on support for social policies depends on many factors, such as the size of the immigrant population living in a region and the nature and size of the welfare state (see, e.g., refs. 33 and 34). In addition, the BRW study pooled various estimands. All analysts were asked to produce estimates of two treatments (immigrant stock and flow) on the support of six different outcomes (six different social policies; resulting in a total of 2 × 6 = 12 different estimands). However, according to the state of research in this area, sudden surges in immigration (i.e., flows) create a greater sense of instability and competition among natives than stable proportions of immigrants (stocks; see, e.g., ref. 35). Natives are then likely to favor an increase in social policies that protect them from competition with immigrants (e.g., unemployment benefits restricted to the labor-force), while at the same time favoring a decrease in social policies that would also benefit new immigrants (e.g., providing a job for anyone who wants one). Differences in empirical estimates for these various estimands are substantial and not due to flawed research designs. The variation induced by these 12 different estimands is therefore not “problematic” uncertainty; rather, it informs (and in this case confirms) theories.

Another illustration of why pooling different estimands is problematic is the SU study. In the case of very general research questions (such as the effects of skin color on the likelihood of receiving a red card in soccer), researchers may choose to identify different causal effects, each of which has its merit, but each of which implies different estimands and corresponding identification strategies (reasonable counterfactuals; 15, 36). For example, one might be interested in (i) identifying the descriptive difference in the probability of receiving a red card by skin color, or one might be interested in (ii) identifying the causal direct effect of skin color (“racial bias”). For a valid identification, such a direct effect requires the removal of all mediating pathways, which might also explain a higher likelihood of receiving a red card but not be considered an instance of discrimination (such as player position). In this example, one would expect the estimate of the direct effect to be smaller than the estimate of the descriptive difference (15). This is not uncertainty, but it is theoretically productive variation.

2.1.2. How to separate sources of variation.

To become more productive, many analyst-studies should do a better job of separating the sources of variation. We see two approaches to achieving this: i) by designing the study accordingly or ii) by using a meta-regression.

Ad (i): Design-based approach

A design-based approach to separating sources of variation was proposed by Huntington-Klein (8). First, he used the data preparation of one analyst team to produce “estimation data” on which to run the analyses of all the other teams; and second, he ran the data analysis used by one team on all the different estimation data produced by the other teams. This allowed him to disentangle the variation caused by the two analytical steps (data preparation, including operationalization of variables and possible coding errors, which is typically not visible to readers; vs. specification of regression models, which is typically reported to readers of an article). However, as the author himself admits, this approach requires choosing a “reference” team for both steps, which is always somewhat arbitrary. This limitation would also apply to other design-based approaches in which some sources of variation are held constant ex ante by explicitly restricting the analysts’ degrees of freedom (e.g., by requesting all analysts to use some preprepared variables instead of having analysts create their own; for an example, see ref. 37). Another limitation of these designs is that they allow only a few broad categories of sources of uncertainty/variation to be contrasted (here: data preparation vs. data analysis), which may not yet be indicative of possible interventions. However, especially for studies with too few participants/estimates to allow stable meta-regressions focusing on many different sources of variation, a design-based approach may be helpful.

Ad (ii): Meta-regression

A meta-regression approach (MRA) allows sources of variation to be teased out without limiting the analysts’ degrees of freedom. The analysts’ results are regressed on all known sources of variation as explanatory factors in a linear regression. This allows for easy interpretation: The effects represent estimates of the absolute deviation of the effect sizes from the mean estimate, and the R² value can be interpreted as the proportion of variance explained. R²-increments provide information about the relative importance of the different sources of variation. Unexplained variance indicates “hidden uncertainty” due to unknown sources of variation.

It is important to note that all meta-regressions are a type of second-stage regression (22, 38): The observations that form the outcome variable are themselves estimates (usually regression coefficients or some derivates). The consequence is that the observations typically vary more in their precision than would be the case in a standard first-stage regression, where observations are a random draw from a common population with unique variance (such as survey respondents from a general population). The precision may differ because the estimates may be based on different subsamples or different model specifications that suffer from more or less confounding bias or multicollinearity. The consequence is that the homoscedasticity assumption of a standard regression (assuming that the observations have a constant variance) does not hold, and therefore, a standard regression will not provide “best” (i.e., most precise) linear unbiased estimators (BLUE).

To overcome this problem, it is common practice to use some variant of a weighted least squares (WLS) regression (39). Typically, some function of the precision of the estimates is used as the weight (22, 40, 41). This addresses the problem of heteroscedasticity. At the same time, such weighting by precision prevents that a few noisy (low-precision) outliers dominate the results (which is likely to happen in meta-analyses). One also does not need to define “outliers” by some arbitrary thresholds. Finally, MRA allows one to estimate the extent to which the variation in the outcome is greater than would be expected from sampling error alone (e.g., using the I² value in random effects meta-regressions; see ref. 22 for tools to achieve this with WLS-MRA). Note that if analysts are allowed to use subsamples of data, some of the variation in results may still be sampling error, see ref. 17.

To date, only a few many-analyst studies have used a WLS-MRA. We recommend a multiple WLS-MRA with the square of the inverse of the SE (1/SE2) as weights (22, 38, 42). This approach covers all the sources of variation we distinguish and is underpinned by extensive methodological research (21). One should include all measured sources of variation (43). Some many-analyst studies have stuck to bivariate analyses (44) or have limited the number of predictors to a small number for fear of “overfitting” (11). However, when the goal is to explain variance, overfitting is not an issue.

On the contrary, by missing predictors, one misclassifies substantively meaningful variation as uncertainty and/or overestimates the hidden portion of uncertainty. Indeed, the underuse of multiple, weighted MRAs may have led to massive overestimates of hidden uncertainty. For example, in the BRW study, the PIs were able to explain only 5% of the variance in the analysts’ estimates by visible researcher decisions. However, in a reanalysis of this study, we found that when we switched to a WLS-MRA and included more predictors, the amount of explained variance increased substantially to 64% (see ref. 45).

Box 1. In summary.

Many-analyst studies should be clear about their meta-estimand and identification strategy. What kind of variation or uncertainty does the study want to identify? By which methods? We recommend that studies be designed to separate the sources of variation shown in Table 1. This can be achieved by coding all relevant analyst decisions and including them as predictors in a WLS-MRA. Unexplained variance then points to hidden sources of variance (but also to possible misspecification of the MRA). Careful interpretation is required.

2.2. Many-Analyst Studies Should Mimic a Standard Research Process.

Many-analyst studies that seek to assess uncertainty in standard research must replicate that research in at least all elements related to uncertainty. Any systematic deviation from i) the composition of researchers, including their incentives to conduct research, ii) the research process, or iii) peer quality control could add hidden variation and compromise the transferability of results to standard research (i.e., compromise the external validity of many-analyst studies).

Ad (i): Composition and incentives of researchers

We recommend that many-analyst studies opt for samples and incentives that are in line with standard research. Some studies used very inexperienced researchers, such as undergraduate students or scholars who were not trained in the methods used (see supplemental information to refs. 6 and 7). Participation was incentivized by emphasizing the importance of replication research and/or by financial rewards, and often also by guaranteeing coauthorship unconditionally on the results provided. The latter was justified by freeing researchers from the “perverse” incentives (e.g., p-hacking) prevalent in standard research (11). However, by guaranteeing coauthorship, many-analyst studies may have artificially introduced another source of uncertainty: if coauthorship is guaranteed, analysts may have adopted a kind of “satisficing” rather than “optimizing” strategy (46): they may have tried to achieve an acceptable estimate with little effort. This may have led to exceptionally sloppy research.

In addition, analysts may have used some “freaky” methods to get some extreme results. Particularly alarming results that diagnose a very low health of science may help to attract attention and boost citation scores for many-analyst studies, thus helping analysts to gain more impact and reputation as coauthors (for some evidence that replicators tend to produce alarming [outlier] results: 26). In fact, the “justifications for the methods chosen” collected in some studies show that some analysts deliberately choose exotic models; simply because it allows them to try out something new, to position themselves in a “unique way,” or to add some variation through a method that would otherwise be missing (see, for example, supplemental material to ref. 44). Combined with lax quality controls (nobody is rejected), such designs obviously overestimate problematic uncertainty.

Ad (ii): Research process

PIs of several many-analyst projects argued that one should use a very generic research question (e.g., whether immigration reduces support for social policies, whether dark skin color increases the risk of receiving a red card in soccer) but without saying anything about the exact estimand to be tested or the target population to be studied. A more precise specification would make the studies more realistic. In standard studies, the questions are not so broad because the authors are typically trying to test only one or a few specific hypotheses derived from a particular theory. Focusing on different estimands or target populations is typically done in different studies, but not in one original study. [Or researchers may at least not pool the estimates, but deliberately keep them separate, as was the case, for example, in the study of Brady and Finnigan (47) to be reanalyzed in the BRW study]. Artificial uncertainty is at least added when many-analyst studies pool estimates that standard researchers would not consider interchangeable (15, 16, the meta-analytic literature warns against this with the metaphor of pooling “apples and oranges,” see ref. 48).

Analyses may also have suffered from time constraints due to strict deadlines in the study design: In some studies, analysts indicated that their results may not be reliable because they would have performed different analyses had time permitted (see, e.g., analysts’ “method justification” reported in supplemental information to ref. 44). Another common departure from standard research (that we also recommend be avoided) is that analysts are typically asked to provide results with only brief notes on, e.g., the statistical model used, but no detailed elaboration on their target estimand and identification strategy (see, e.g., refs. 7 and 11; for an exception: ref. 9). As a result, analysts essentially just provide an empirical number without explaining what question that number is trying to answer. This is presumably done to make it easier to recruit analysts by reducing the time required. However, it also increases the risk of sloppiness and misspecification errors.

Ad (iii): Quality control by peer review

The standard solution for avoiding sloppy and biased research is peer review. Peer review helps filter out poor quality research and can also motivate proper research at the first stage, when authors know that they must convince reviewers that their work is credible. So far, however, most many-analyst studies have used only internal review processes (analysts reviewing each other) without any real consequences. We know of only one many-analyst study that used a standard review with external reviewers. This study found a substantial reduction in uncertainty once the low-quality results identified by the reviewers were removed (9).

Box 2. In summary.

Our recommendation is to work with researchers who are experienced in the area under study and to be precise about the definition of the research question (estimand). Otherwise, it is very important to control this source of uncertainty in the analyses (see Section 2.1). One should ask teams to provide some justification for the design chosen, as is standard in journal articles. Implementing peer review is probably the only way to get a realistic picture of uncertainty in standard research. You may not have the resources to meet all these standards. Deviations from standard research should then at least be justified and transparently discussed as a possible limitation.

2.3. Many-Analyst Studies Should Do a Proper Meta-Analysis.

After the many analysts have provided their results (estimates), the PIs must analyze them. What is often overlooked is that this is a meta-analysis: A statistical analysis of research results. To obtain efficient and unbiased meta-estimates, it is very important to adhere to at least two standards that are well established in the meta-analysis literature. [A third one (efficient meta-regression) has already been discussed in Section 2.1.]

2.3.1. Do not get lost in translation: Effects must be on the same metric before they are pooled.

It is well established in the meta-analysis literature that effect sizes must be measured on the same scale (i.e., using the same effect size metric) before they can be pooled (see, e.g., refs. 22, 40, and 41). If analysts are allowed to use different metrics, they must be brought to a common metric post hoc (by rescaling). Different metrics can result from different units in which treatment and/or outcome variables were measured (e.g., immigration levels measured in units of one per hundred or one per thousand), but also from the use of different statistical models. For example, logistic regression results may be reported as odds ratios or logit coefficients; group differences may be quantified as risk ratios, odds ratios, correlation coefficients, or average marginal effects (AMEs). Converting different effects into a common effect size metric is complex, and in some cases impossible, because no formulas exist. In this case, one should not pool them.

Errors in this rescaling process are consequential as they always introduce artificial variance. Effects vary more than they would if the rescaling transformations were done correctly. The result is (i) a downward bias in the effects of all predictors in a meta-regression because the random noise caused by inconsistent measures biases all association coefficients toward zero; see, e.g., ref. 39 for this “attenuation bias,” also called “regression dilution.” The consequence is (ii) an overestimation of uncertainty. This means that one is per se less successful in identifying and decomposing sources of uncertainty. Accordingly, there appears to be more unexplained, hidden uncertainty.

Unsuccessful rescaling operations may indeed have led to erroneous conclusions in many-analyst studies. For example, in the BRW study, several missing or arbitrary rescaling operations added massive variance to the estimates (e.g., effects were erroneously multiplied by up to a factor of 10, with no clear justification in the main text as to why these transformations made sense; see ref. 45). As a consequence, the most extreme effect sizes were caused by the PI’s incorrect rescaling operations. This created problematic “hidden” variation that was then falsely attributed to the research practices of the analysts. At the same time, the PIs of this study overlooked some necessary transformations (for details: ref. 45), which introduced additional, unexplained noise into the data.

To avoid such errors, all rescaling attempts should be carefully checked for possible errors, e.g., by organizing an internal peer review process (49). If it is questionable whether one can get them to the same effect size metric, it is better not to pool the estimates (50). Detailed guidance on how to perform which transformations can be found in standard textbooks on meta-analysis (e.g., ref. 51).

2.3.2. Do not be fooled by a few outliers.

After the estimates have been pooled, the meta-analysis begins by reporting descriptive statistics, such as the range of effect sizes or their variance. Often, the proportions of effects in different categories are also reported (e.g., positive vs. negative, nonsignificant vs. statistically significant estimates). It should be noted that simply counting estimates in different categories can make marginal discrepancies look like meaningful ones (e.g., two null effects with effect sizes of −0.001 vs. +0.001; P-values of 0.049 vs. 0.051). In standard research, reported point estimates or exact P-values (or CIs) can make these marginal differences visible, whereas meta-analyses that report only categories hide them. For metric differences, one should opt for easy-to-interpret effect sizes and discuss them focusing not only on statistical but also substantive significance (17, 22, 52). Especially when true effects are very close to zero, small absolute differences can lead to large relative differences (e.g., two null effects of 0.001 vs. 0.0001; the first effect would be 10 times larger). This must be kept in mind when using variance to quantify uncertainty: Variance is inherently a relative measure since it measures deviations relative to the mean.

When interpreting the range or variance, care must also be taken to ensure that these measures are not driven by a few outliers. Even if all of the effects fall within a narrow range (replicate well), one strong outlier could give the false impression that there is a lot of variation (estimates span a wide range). More guidelines on how to assess the variation in results by means of descriptive statistics and visual representations can be found in ref. 24.

We illustrate the importance of these issues again with a recent many-analyst study. In Fig. 1, the Left panel shows the 1,252 AMEs provided by the analysts in the BRW study, along with 95% CIs. The AMEs report the effect of a one percent increase in immigration on the probability [0;1] that respondents support social policies such as the government providing jobs for the unemployed or a decent standard of living. [Unfortunately, not all AMEs are measured on this probability scale. About 40% are measured on other scales (ordinal or continuous), violating the same-metric standard (45)].

Fig. 1. Distribution of the AMEs in the BRW study. Left panel: With 95% CIs. Right panel: As funnel plot (scattergram of the AMEs and their precision). Left panel: N = 1,252 AMEs (red) and their 95% CIs (blue) as they entered the BRW study (one AME dropped because it had a missing CI). The CIs are trimmed to the interval [−1.5, +1.5] to better show the variability of the point estimates. Right panel: Scattergram (funnel plot) of the AMEs and their precision (1/SE). N=1,126, because the 10% most precise estimates (precision > >894) have been dropped to better show the association with precision. The blue lines denote the borders of the range of the estimates reported in the original study (47) that the teams replicated.

It can be seen that most of the results are consistent and very close to zero, confirming the main conclusions of the original study to be replicated (47; which concluded that immigration has only a very weak effect, if any, on support for social policies). There are also some outliers on both tails that, however, are not plausible at all. For example, the most extreme effect is an AME of 1.3, meaning that the probability of support would increase by +1.3, which is mathematically not defined. These outliers were used by BRW as indication for high uncertainty [“teams results varied greatly;” c.f. BRW (11), p. 1]. What BRW did not discuss, however, was that most results were in a narrow range. They also did not show the CIs of the estimates. As can be seen in Fig. 1 (Left), the outliers had very large CIs (i.e., low measurement precision). One should, however, avoid letting a few low-quality outliers distort the overall picture (48).

To better see whether only a few noisy outliers are destroying an otherwise consistent picture of research findings, we recommend plotting analysts’ estimates in a “funnel plot” (see the Right panel of Fig. 1). This descriptive graph is standard in meta-analyses in order to detect possible bias from low-quality estimates. The effect sizes (here: AMEs) are plotted on the x-axis, while their precision (measured by the inverse of the SE) is plotted on the y-axis. The highly precise estimates are at the top of the graph, and the imprecise estimates are at the bottom. For the BRW study, it can clearly be seen that the vast majority of the findings (82%) fall within a very narrow range close to zero [−0.02; +0.03]. This is the range reported in the original study (47). Only a minority of the results fall outside this range, and almost all of them are associated with very low measurement precision (large SE).

By not performing such diagnostics for outliers and influential data points, BRW overlooked that their main conclusion of a large “hidden universe of uncertainty” was driven by only a few very imprecisely estimated outliers.

Box 3. In summary.

Before pooling estimates, one has to define a consistent effect metric to be used for the analyses. This can be done before the teams do their analyses. Otherwise, the PIs will need to bring effects to a common metric. Meta-analysis methods are also required to adequately describe and analyze the data. This includes using meaningful metrics to measure and interpret variation in estimates, and accounting for their varying quality (precision), e.g., by plotting CIs. Be careful not to allow a small number of noisy outliers to distort the result of otherwise consistent (i.e., certain) findings.

3. Concluding Remarks

By not adhering to the three methods standards discussed in this Perspective (defining a meta-estimand and separating uncertainty from other sources of variation; replicating the standard research process; and using meta-analytic tools to analyze the data), existing many-analyst studies have likely provided a too pessimistic picture of the health of science. Overestimation of uncertainty in scientific results is likely to threaten public confidence in science. Other scientists, the public, or politicians may think we are just “poking around in the fog” and thus may be reluctant to invest in further research or evidence-based interventions. For these and other reasons, a credible assessment of credibility is important (1). We hope that our guidelines will help to make future many-analyst studies a more reliable and constructive tool for advancing knowledge about uncertainty.

However, even with better adherence to meta-analytic standards, many-analyst studies will be challenging undertakings: As explained in this Perspective, much uncertainty can be artificially introduced by the meta-analytic step itself. We therefore also recommend i) to double-check all data preparations and analyses and ii) to use transparent reporting standards. One should give readers and reviewers the best possible insight into all methods used, e.g., by reporting formulas used for transforming effect sizes, informing on outliers that were dropped or “capped,” and any other “manipulations” of the data. Existing many-analyst studies were often surprisingly opaque in their methods, either not reporting key methodological decisions or relegating them to supplements or thousands of lines of statistical syntax code.

Finally, to make many-analyst studies even more credible, they should share all data and syntax code used by the teams and PIs and also provide raw data on analysts’ estimates before PIs applied rescaling transformations. That way, many-analyst studies can provide an even more credible assessment of the credibility of standard research.

Author contributions

K.A. and J.B. performed formal analysis, methodology, and visualization; K.A. conceptualized the manuscript; and K.A. and J.B. wrote the paper.

Competing interests

The authors declare no competing interest.

Data, Materials, and Software Availability

Previously published data were used for this work (https://github.com/nbreznau/CRI) (53). Stata code for reproducing Figure 1 can be found on our OSF project (54).

This article is a PNAS Direct Submission.
==== Refs
1 National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (The National Academies Press, Washington, DC, 2019), 10.17226/25303.
2 J. Freese, D. Peterson, Replication in social science. Annu. Rev. Sociol. 43 , 147–165 (2017).
3 M. Duvendack, R. Palmer-Jones, W. R. Reed, What is meant by “replication” and why does it encounter resistance in economics? Am. Econ. Rev. 107 , 46–51 (2017).
4 G. Christensen, J. Freese, E. Miguel, Transparent and Reproducible Social Science Research. How to Do Open Science (University of California Press, ed. 1, 2019), 10.2307/j.ctvpb3xkg.
5 R. Silberzahn, E. L. Uhlmann, Crowdsourced research: Many hands make tight work. Nature 526 , 189–191 (2015).26450041
6 E. Gould , Same data, different analysts: Variation in effect sizes due to analytical decisions in ecology and evolutionary biology (OSF, 14 May, 2023). https://osf.io/us3jh/.
7 M. Schweinsberg , Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis. Organ. Behav. Hum. Decis. Process. 165 , 228–249 (2021).
8 N. Huntington-Klein , The influence of hidden researcher decisions in applied microeconomics. Econ. Inq. 59 , 944–960 (2021).
9 A. J. Menkveld , Non-standard errors. J. Finance 79 , 2339–2390 (2024).
10 R. Botvinik-Nezer , Variability in the analysis of a single neuroimaging dataset by many teams. Nature 582 , 84–88 (2020).32483374
11 N. Breznau , Observing many researchers using the same data and hypothesis reveals a hidden universe of uncertainty. Proc. Natl. Acad. Sci. U.S.A. 119 , e2203150119 (2022).36306328
12 B. Aczel , Consensus-based guidance for conducting and reporting multi-analyst studies. Elife 10 , e72185 (2021).34751133
13 C. F. Camerer, The apparent prevalence of outcome variation from hidden “dark methods” is a challenge for social science. Proc. Natl. Acad. Sci. U.S.A. 119 , e2216020119 (2022).36538485
14 E. J. Wagenmakers, A. Sarafoglou, B. Aczel, One statistical analysis must not rule them all. Nature 605 , 423–425 (2022).35581494
15 K. Auspurg, J. Brüderl, Has the credibility of the social sciences been credibly destroyed? Reanalyzing the “many analysts, one data set” project. Socius 7 , 23780231211024421 (2021).
16 P. Engzell, A universe of uncertainty hiding in plain sight. Proc. Natl. Acad. Sci. U.S.A. 120 , e2218530120 (2023).36595682
17 M. B. Mathur, C. Covington, T. J. VanderWeele, Variation across analysts in statistical significance, yet consistently small effect sizes. Proc. Natl. Acad. Sci. U.S.A. 120 , e2218957120 (2023).36623183
18 A. Gelman, E. Loken, The statistical crisis in science. Data-dependent analysis—A ‘garden of forking paths’—Explains why many statistically significant comparisons don’t hold up. Am. Sci. 102 , 460–466 (2014).
19 C. Young, Model uncertainty in sociological research: An application to religion and economic growth. Am. Sociol. Rev. 74 , 380–397 (2009).
20 C. Young, K. Holsteen, Model uncertainty and robustness: A computational framework for multimodel analysis. Sociol. Methods Res. 46 , 3–40 (2017).
21 T. D. Stanley, C. Doucouliagos, Practical significance, meta-analysis and the credibility of economics. IZA DP No. 12458 [Preprint]. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3427595 (Accessed 5 May 2024).
22 T. Stanley, H. Doucouliagos, Meta-Regression Analysis in Economics and Business (Routledge Advances in Research Methods, Routledge Taylor and Francis Group, London and New York, 2012), vol. 5 , pp. 1–190.
23 K. Auspurg, J. Brüderl, “How to increase reproducibility and credibility of sociological research” in Handbook of Sociological Science: Contributions to Rigorous Sociology, K. Gërxhani, N. de Graaf, W. Raub, Eds. (Edward Elgar Publishing, 2022), chap. 26, pp. 512–527, 10.4337/9781789909432.
24 A. Dreber Almenberg, M. Johannesson, A framework for evaluating reproducibility and replicability in economics. Econ. Inq., 1–19 (2024).
25 N. Malhotra, “Threats to the scientific credibility of experiments. Publication bias and p-hacking” in Advances in Experimental Political Science, J. Druckman, D. P. Green, Eds. (Cambridge University Press, Cambridge, 2021), chap. 19, pp. 354–368.
26 C. J. Bryan, D. S. Yeager, J. M. O’Brien, Replicator degrees of freedom allow publication of misleading failures to replicate. Proc. Natl. Acad. Sci. U.S.A. 116 , 25535–25545 (2019).31767750
27 J. Ankel-Peters, N. Fiala, F. Neubauer, Is economics self-correcting? Replications in the American economic review. Econ. Inq., 1–23 (2024).
28 J. P. A. Ioannidis, T. A. Trikalinos, Early extreme contradictory estimates may appear in published research: The Proteus phenomenon in molecular genetics research and randomized trials. J. Clin. Epidemiol. 58 , 543–549 (2005).15878467
29 S. Steegen, F. Tuerlinckx, A. Gelman, W. Vanpaemel, Increasing transparency through a multiverse analysis. Perspect. Psychol. Sci. 11 , 702–712 (2016).27694465
30 U. Simonsohn, J. P. Simmons, L. D. Nelson, Specification curve analysis. Nat. Hum. Behav. 4 , 1208–1214 (2020).32719546
31 L. T. Droy, J. Goodwin, H. O’Connor, Methodological uncertainty and multi-strategy analysis: Case study of the long-term effects of government sponsored youth training on occupational mobility. Bull. Sociol. Methodol. 147–148 , 200–230 (2020).
32 M. Del Giudice, S. W. Gangestad, A traveler’s guide to the multiverse: Promises, pitfalls, and a framework for the evaluation of analytic decisions. Adv. Methods Pract. Psychol. Sci. 4 , 2515245920954925 (2021).
33 A. Alesina, E. Murard, H. Rapoport, Immigration and preferences for redistribution in Europe. J. Econ. Geogr. 21 , 925–954 (2021).
34 S. L. Schneider, Anti-immigrant attitudes in Europe: Outgroup size and perceived ethnic threat. Eur. Sociol. Rev. 24 , 53–67 (2007).
35 D. J. Hopkins, Politicized places: Explaining where and when immigrants provoke local opposition. Am. Polit. Sci. Rev. 104 , 40–60 (2010).
36 I. Lundberg, R. Johnson, B. M. Stewart, What is your estimand? Defining the target quantity connects statistical evidence to theory. Am. Sociol. Rev. 86 , 532–565 (2021).
37 N. Huntington-Klein, Many-economists project (2022). https://nickch-k.github.io/ManyEconomists/. Accessed 5 May 2024.
38 T. D. Stanley, S. B. Jarrell, Meta-regression analysis: A quantitative method of literature surveys. J. Econ. Surv. 3 , 161–170 (1989).
39 M. Wooldridge, Introductory Econometrics: A Modern Approach (South-Western Cengage Learning, Mason, OH, ed. 5, 2012).
40 M. Borenstein, L. V. Hedges, J. P. T. Higgins, H. R. Rothstein, Introduction to Meta-Analysis (Wiley, 2009).
41 J. Gurevitch, J. Koricheva, S. Nakagawa, G. Stewart, Meta-analysis and the science of research synthesis. Nature 555 , 175–182 (2018).29517004
42 T. D. Stanley , Meta-analysis of economic research reporting guidelines. J. Econ. Surv. 27 , 390–394 (2013).
43 B. B. McShane, J. L. Tackett, U. Böckenholt, A. Gelman, Large-scale replication projects in contemporary psychological research. Am. Stat. 73 , 99–105 (2019).
44 R. Silberzahn , Many analysts, one data set: Making transparent how variations in analytic choices affect results. Adv. Methods Pract. Psychol. Sci. 1 , 337–356 (2018).
45 K. Auspurg, J. Brüderl, Claims about “a hidden universe of uncertainty” in social research are exaggerated: How many-analyst studies risk overestimating uncertainty. MetaArXiv [Preprint] (2024). 10.31222/osf.io/uc84k (Accessed 31 July 2024).
46 H. A. Simon, A behavioral model of rational choice. Q. J. Econ. 69 , 99–118 (1955).
47 D. Brady, R. Finnigan, Does immigration undermine public support for social policy? Am. Sociol. Rev. 79 , 17–42 (2014).
48 D. Sharpe, Of apples and oranges, file drawers and garbage: Why validity issues in meta-analysis will not go away. Clin. Psychol. Rev. 17 , 881–901 (1997).9439872
49 P. C. Gøtzsche, A. Hróbjartsson, K. Maric, B. Tendal, Data extraction errors in meta-analyses that use standardized mean differences. JAMA 298 , 430–437 (2007).17652297
50 M. A. L. M. van Assen, A. H. Stoevenbelt, R. C. M. van Aert, The end justifies all means: Questionable conversion of different effect sizes to a common effect size measure. Relig. Brain Behav., 13 , 345–347 (2022), 10.1080/2153599X.2022.2070249.
51 M. Borenstein, L. V. Hedges, J. P. T. Higgins, H. Rothstein, Introduction to Meta-Analysis (Wiley, Somerset, ed. 2, 2011).
52 F. Bernardi, L. Chakhaia, L. Leopold, ‘Sing me a song with social significance’: The (mis)use of statistical significance testing in European sociological research. Eur. Sociol. Rev. 33 , 1–15 (2016).
53 N. Breznau, Data from “The Hidden Universe of Data-Analysis”. https://github.com/nbreznau/CRI/blob/master/data/cri.csv. Accessed 30 November 2022.
54 K. Auspurg, J. Brüderl, Replication Files for “Is Social Research Really Not Better Than Alchemy? How Many-Analysts Studies Produce ‘a Hidden Universe of Uncertainty’ by Not Following Meta-Analytical Standards”. BRW Analyses – Reproduction by AB.do. https://osf.io/dtv2p/files/osfstorage/647cf174bf3d0f09d9d872be. Deposited 4 June 2023.
