
==== Front
PLoS Genet
PLoS Genet
plos
PLOS Genetics
1553-7390
1553-7404
Public Library of Science San Francisco, CA USA

39241053
PGENETICS-D-24-00217
10.1371/journal.pgen.1011391
Research Article
Research and Analysis Methods
Research Assessment
Research Errors
Biology and Life Sciences
Genetics
Single Nucleotide Polymorphisms
Research and Analysis Methods
Simulation and Modeling
Research and Analysis Methods
Mathematical and Statistical Techniques
Statistical Methods
Instrumental Variable Analysis
Physical Sciences
Mathematics
Statistics
Statistical Methods
Instrumental Variable Analysis
Biology and Life Sciences
Genetics
Medicine and Health Sciences
Nephrology
Renal Diseases
Chronic Kidney Disease
Biology and Life Sciences
Anatomy
Renal System
Kidneys
Medicine and Health Sciences
Anatomy
Renal System
Kidneys
Biology and Life Sciences
Genetics
Heredity
MR-SPLIT: A novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies
Robust one-sample MR analysis with sample splitting
Shi Ruxin Data curation Formal analysis Investigation Methodology Software Validation Visualization Writing – original draft Writing – review & editing 1
https://orcid.org/0000-0003-2452-6691
Wang Ling Formal analysis Resources 2
Burgess Stephen Formal analysis Investigation 3
https://orcid.org/0000-0001-8099-1753
Cui Yuehua Conceptualization Investigation Methodology Project administration Resources Supervision Writing – original draft Writing – review & editing 1 *
1 Department of Statistics and Probability, Michigan State University, East Lansing, Michigan, United States of America
2 Department of Medicine, Michigan State University, East Lansing, Michigan, United States of America
3 Biostatistics Unit, University of Cambridge, Cambridge, United Kingdom
Morrison Jean Editor
University of Michigan School of Public Health, UNITED STATES OF AMERICA
The authors have declared that no competing interests exist.

* E-mail: cuiy@msu.edu
9 2024
6 9 2024
20 9 e101139123 2 2024
9 8 2024
© 2024 Shi et al
2024
Shi et al
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Mendelian Randomization (MR) is a widely embraced approach to assess causality in epidemiological studies. Two-stage least squares (2SLS) method is a predominant technique in MR analysis. However, it can lead to biased estimates when instrumental variables (IVs) are weak. Moreover, the issue of the winner’s curse could emerge when utilizing the same dataset for both IV selection and causal effect estimation, leading to biased estimates of causal effects and high false positives. Focusing on one-sample MR analysis, this paper introduces a novel method termed Mendelian Randomization with adaptive Sample-sPLitting with cross-fitting InstrumenTs (MR-SPLIT), designed to address bias issues due to IV selection and weak IVs, under the 2SLS IV regression framework. We show that the MR-SPLIT estimator is more efficient than its counterpart cross-fitting MR (CFMR) estimator. Additionally, we introduce a multiple sample-splitting technique to enhance the robustness of the method. We conduct extensive simulation studies to compare the performance of our method with its counterparts. The results underscored its superiority in bias reduction, effective type I error control, and increased power. We further demonstrate its utility through the application of a real-world dataset. Our study underscores the importance of addressing bias issues due to IV selection and weak IVs in one-sample MR analyses and provides a robust solution to the challenge.

Author summary

Mendelian randomization (MR) is a method used in genetic epidemiology to determine whether a specific exposure has a causal effect on a health outcome. Ensuring the accuracy of this method is crucial for its reliability and for making informed decisions that can enhance public health and medical practices. Typically, researchers employ the two-stage least squares (2SLS) method which involves selecting a set of valid instrumental variables (IVs) to estimate and infer the causal effect. However, 2SLS can produce biased results when the effects of the IVs are weak, known as weak instrument bias. Additionally, the “winner’s curse” problem may occur when using the same dataset for both IV selection and causal effect estimation, introducing additional bias. Here we introduce a novel approach called MR-SPLIT, which addresses these two bias issues by randomly splitting the data into two parts: one for IV selection and the other for IV construction and causal effect estimation. Through effective integration, this strategy enhances the power, reduces bias, and provides more precise estimates. Our approach is validated through extensive simulation studies, and its effectiveness is demonstrated by an application to a real-world dataset.

http://dx.doi.org/10.13039/100000968 American Heart Association 24TPA1288424 https://orcid.org/0000-0001-8099-1753
Cui Yuehua http://dx.doi.org/10.13039/100000968 American Heart Association 24TPA1288424 https://orcid.org/0000-0003-2452-6691
Wang Ling This work was supported by a grant from the American Heart Association (24TPA1288424 to LW and YC). The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. PLOS Publication Stagevor-update-to-uncorrected-proof
Publication Update2024-09-18
Data AvailabilityThe Chronic Renal Insufficiency Cohort (CRIC) Study was conducted by the CRIC Study Investigators and supported by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). The data from the CRIC Study reported here were supplied by the CRIC investigators and were downloaded from dbGaP with accession number phs000524.v1.p1. This manuscript was not prepared in collaboration with Investigators of the CRIC Study and does not necessarily reflect the opinions or views of the CRIC Study, or the NIDDK. R codes for MR-SPLIT can be freely accessed at: https://github.com/RuxinShi/MRSPLIT.
Data Availability

The Chronic Renal Insufficiency Cohort (CRIC) Study was conducted by the CRIC Study Investigators and supported by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). The data from the CRIC Study reported here were supplied by the CRIC investigators and were downloaded from dbGaP with accession number phs000524.v1.p1. This manuscript was not prepared in collaboration with Investigators of the CRIC Study and does not necessarily reflect the opinions or views of the CRIC Study, or the NIDDK. R codes for MR-SPLIT can be freely accessed at: https://github.com/RuxinShi/MRSPLIT.
==== Body
pmcIntroduction

Mendelian Randomization (MR) utilizes genetic variants as instruments to detect causal effects between an exposure and an outcome variable [1–4], and has emerged as a pivotal method in the realm of causal inference. In observational studies, establishing a causal relationship between two variables exhibiting an observed association can be challenging, because of unknown confounding factors that may influence both variables simultaneously. However, MR harnesses Single Nucleotide Polymorphisms (SNPs) as instrumental variables (IVs) to infer causative associations between exposures and outcomes, circumventing the confounding biases inherent in conventional observational studies. Given that SNPs, as genetic variants, are randomly segregated during meiosis, they offer a unique opportunity for robust causal inference.

This innovative utilization of SNPs as IVs rests on a triad of fundamental assumptions, which are indispensable for ensuring the validity of MR results [4, 5]: 1) Relevance assumption; 2) Independence assumption; and 3) Exclusion-restriction assumption. Using SNPs as IVs can also avoid reverse causation because genetic variants are randomly assigned at conception and remain constant throughout an individual’s life. However, violations [6] of the triadic assumptions underpinning MR analysis can produce biased and unreliable estimates. Specifically, when the relevance assumption is contravened, we encounter what is termed a “weak instrument” [7, 8] phenomenon. Conversely, breaches of the latter two assumptions often manifest as “pleiotropy” [9–11], where the instrumental variable exerts effects on the outcome via pathways other than the exposure of interest. Throughout this work, the primary focus is to address the challenge of IV selection and weak instrument bias in one-sample MR analyses.

In MR analysis, two main frameworks are commonly used: two-sample MR analysis with GWAS summary statistics and one-sample MR analysis with individual-level data. While two-sample MR analysis has gained popularity due to easy access to public datasets, it comes with a couple of limitations. Firstly, it relies on marginal estimates of SNP statistics, which can be biased when not accounting for linkage disequilibrium (LD) properly. Secondly, it lacks the flexibility to model other causal mechanisms, such as nonlinear causal effects. As a result, there continues to be a significant interest in the advancement of statistical methods for one-sample MR analysis. The most popular method used in one-sample MR analysis is the two-stage least squares (2SLS) approach [12], which is relatively straightforward to implement and can yield consistent estimates of causal effects. However, the 2SLS estimate can be biased in the presence of weak instruments [8]. The bias is in the direction of the confounded association and can cause inflated false positive rates, particularly when more than one IV is included in the analysis [13]. To date, weak instrument bias still remains one of the significant concerns in one-sample MR analysis [13].

A potential solution to mitigate the impact of weak IVs is to opt for a two-sample MR analysis. While this approach might mitigate some biases, it does not eliminate them entirely. Specifically, bias due to weak instruments in two-sample MR tends to be directed towards the null [14]. Limited information maximum likelihood (LIML) method [15, 16] was introduced as an alternative to 2SLS when dealing with weak instruments. Burgess et al. [17] showed that LIML could provide a less biased estimate compared to 2SLS in the presence of weak instruments, but at the expense of incurring larger variance. Nevertheless, LIML is still subject to weak instrument problems and its finite sample performance can be poor. Angrist et al. [18] proposed two jackknife instrumental variables estimators (JIVE) as alternatives to 2SLS and LIML to reduce the bias with many weak instruments. However, Sören and Matz [19] showed that neither LIML nor the JIVE estimators perform uniformly better than the 2SLS does in terms of root mean square error.

In one-sample MR analysis, when the same dataset is used for both IV selection and causal effect estimation, the “winner’s curse” or IV selection bias emerges as another notable concern in addition to the weak IV bias issue [13, 20]. This bias could lead to biased causal effect estimates and hence inflate false positive rates under the 2SLS IV regression framework. This is evident in Section 2 of the S1 Text, where it is shown that using the same data (the whole sample) for both IV selection and causal effect estimation, both LIML and 2SLS methods amplify bias compared to using half data for selection and the other half for causal effect estimation. Thus, it is critical to address the IV selection bias issue in one-sample MR analysis. Typically, in one-sample MR analysis, IVs are typically chosen based on a p-value threshold. However, the usage of a p-value threshold criterion in the selection of IVs is somewhat arbitrary and lacks robust justification. The 2SLS approach relies on the fitted values from the first stage for estimating causal effects in the second stage, highlighting the critical role of prediction accuracy and thus questioning the robustness of models that depend solely on p-value thresholds for validation. Given the typically vast dimensionality of SNP data, the use of penalized shrinkage methods can effectively mitigate the winner’s curse effect in one-sample MR analysis. This strategy prioritizes prediction accuracy and hence provides a potentially more dependable and robust framework for causal inference.

Denault et al. [21] introduced a method called ‘Cross-Fitting for Mendelian Randomization’ (CFMR) to handle the weak instrument issue in one-sample MR analysis, which consolidates information from multiple IVs into a single IV, termed the Cross-Fitted Instrument (CFI). CFMR randomly splits a sample into K subgroups {I1, ⋯, IK} and define the complement of the partition Ik as Ikc={1,⋯,N∉Ik}. In each subset {Ikc}, it first selects γk independent variants {Z1,k,⋯,Zγk,k}, and then defines a CFI of the exposure X on Ik, which is the prediction of X on Ik trained using data with indexes in Ikc.

This predicted value can be viewed as a polygenic risk score in risk prediction analysis. Then, the new CFI is used as the IV to fit the 2SLS model for further causal inference. CFMR consolidates all the IVs into one single IV (CFI), thus it produces less biased results. However, CFMR does not completely solve the selection bias issue. Taking K = 10 as an example, CFMR employs 9 folds of data for selecting IVs and applies the estimated effects to construct the composite IV in a separate fold of data. By iterating this process 10 times, the composite IVs across any pair of folds are constructed with 80% of data in common. Thus, the composite IVs are not constructed using completely independent data. This could lead to a new manifestation of the winner’s curse problem. Furthermore, by relying on one CFI as the only IV to represent the collective information of all IVs, there could be potential information loss which further leads to variance inflation and consequently reduced power (as shown in our theorem and simulation studies).

In general, the selection of IVs involves a bias and variance trade-off when estimating the causal effect. Using more IVs tends to introduce a larger bias but smaller variance, whereas employing too few IVs results in a smaller bias but larger variance. Pierce et al. [22] did intensive simulations to evaluate the power and IV strength requirements for MR analyses based on 2SLS. They employed four strategies to combine information across IVs and evaluated the consequences of these strategies on power and overall IV strength, as measured by the first-stage F statistic in 2SLS. The results suggest that categorizing IVs into major and weak ones and then consolidating the weak ones into a single IV based on the knowledge of the genetic architecture underlying the exposure, can mitigate the issue of weak IVs. However, the study identifies a gap in current methodologies: it does not provide a clear approach for differentiating between major and weak IVs, nor does it offer a strategy for combining weak IVs in the context of one-sample MR analysis. This highlights an area for further research and methodology development in the field.

In this work, we propose an adaptive Sample-sPLitting method with cross-fitting InstrumenTs (MR-SPLIT) to address the bias issue of IV selection and weak instruments. This approach can effectively reduce the number of weak IVs without the loss of much information, thereby enhancing the performance of causal inference in MR studies by improving the power of causal inference. Our method has two advantages over the existing ones: 1) It adaptively selects major and weak IVs, subsequently creating a composite IV from the weaker ones. We theoretically proved that the variance of the MR-SPLIT estimate is smaller than that of the CFMR estimate under the condition of one sample split. Simulation results also show that MR-SPLIT can always achieve higher power and lower RMSE than CFMR; and 2) A multi-sample splitting strategy is further employed to enhance the robustness of estimation and testing. Extensive simulation studies were conducted to assess the performance of our method in comparison to its counterparts, including 2SLS, LIML, and CFMR. Our method offers an efficient and powerful solution for one-sample MR analysis by addressing two primary sources of bias: IV selection bias and the bias associated with weak instruments.

Statistical method

Assume the following structural equation model, yi=xiβ+εyixi=Gi·α+εxi

where xi is the exposure, and yi denotes the outcome of the ith individual. Gi ⋅ is a p-dim vector of SNP IVs, where Gi·={Gi1,Gi1,…,Gip}∈Rp. The error term is denoted by εi = (εxi, εyi) ∼ N(0, σ2R) where R12(= ρ) is the correlation due to confounding. β is the interested causal effect which needs to be estimated. Suppose we have N independent individuals, and denote Y=(y1,…,yN)′∈RN×1, X=(x1,…,xN)′∈RN×1, G={G1,…,Gp}∈RN×p, where the jth IV denoted as Gj = (G1j, …, GNj)′, j = 1, …, p, then we have Y=Xβ+εyX=Gα+εx (1)

MR-SPLIT: Mendelian Randomization with adaptive Sample-sPLitting with cross-fitting InstrumenTs

Cross-fitting instruments with sample split

Given the observed data {X, Y, G}, we first need to select a valid IV subset from the existing SNP pool, where the number of SNPs can be much larger than the sample size. To reduce potential biases and enhance the accuracy of estimates in MR analysis, one can use one sample for the selection of appropriate IVs and a separate, independent sample for the 2SLS estimation. By doing so, over-fitting and biases stemming from sample-specific peculiarities, such as the double dipping issue, can be minimized, leading to more robust and credible causal effect estimates. When only one sample is available, one simple idea is to randomly split the data into two equal subsets {I1, I2}, each containing roughly N/2 samples. Then, one can use one subset (say I1) to select the IVs and use the other (say I2=I1c) to get the estimates of β.

For the IV selection, if no prior information about specific SNPs is available, researchers usually regress the exposure variable on each SNP, and then select those SNPs that yield marginal p-values smaller than a preset threshold (e.g., 5 × 10−8) followed by LD pruning or LD clumping. However, such a threshold is quite ad hoc and sometimes can be too stringent, prompting the need for relaxation, as advocated in some studies [23]. Such strictness can lead to the exclusion of valid IVs and the loss of valuable information. Conversely, if the threshold is too lenient, it may result in the selection of an excessive number of SNP IVs, potentially introducing challenges associated with weak IVs [17]. We suggest using some high-dimensional screening methods such as sure independence screening (SIS) [24] to first reduce the SNP dimension from ultra-high to high dimension. Methods like SIS have the sure screening property in which they ensure that, as the sample size increases, the probability of including all relevant variables becomes close to one. After this step, shrinkage methods such as LASSO or adaptive LASSO [25, 26] can be employed to select and estimate non-zero SNP effects. Other penalized methods with different penalty functions such as MCP or SCAD can also be applied.

After the SNP selection, directly employing these IVs in 2SLS might lead to the issue of weak instruments, potentially resulting in biased estimate. To mitigate this, we group the IVs into two groups, major IVs and weak IVs, based on their association strength with the exposure. Conventionally, the validity of IVs is assessed using F statistics. A common benchmark used in econometrics and statistical literature suggests that an F-statistic exceeding 10 is indicative of strong instruments, particularly when assessing the strength of a collective set of IVs [27, 28]. However, the determination of the weakness of an individual IV lacks a widely recognized standard. In this analysis, we employed partial F-statistics with different thresholds as criteria for selecting major IVs. Generally, the partial F statistic is defined as: F=(RSSr-RSSf)/p(RSSf)/(N-k-1) (3)

where RSSr and RSSf are the residual sums of squares for the reduced and full model, respectively; N is the total number of observations; k and p are the numbers of variables in the reduced and full model, respectively. This statistic measures how much the addition of p variables improves the model, compared to the increase in complexity these variables bring. It is a good statistic for calculating the strength of each IV and is consistent with the commonly used F-statistic for evaluating IV strength. In our model, p = 1 because we calculate the partial F statistics for each IV. The thresholds were set at partial F-statistics greater than 10, 30, and 50. We conducted a simulation study to compare these three statistics for the purpose of identifying weak IVs, as detailed in section Evaluation of Major IV identification. Based on the simulation results, it is recommended to use a threshold of partial F > 30 to define major IVs.

Consolidating weak IVs to form a composite IV

Following the separation of major and weak IVs, we propose consolidating the weak ones into a composite IV. Then, the major IVs and the composite IV are included in the 2SLS model to infer the causal effect (see Fig 1 for the flowchart of MR-SPLIT). By only consolidating the weak IVs into a single instrument, we can substantially reduce the number of IVs in the model while retaining most of the information they carry.

10.1371/journal.pgen.1011391.g001 Fig 1 The flow chart of MR-SPLIT with one random split.

The original data is randomly split into two parts indexed by I1 and I2. We use data in I1 to select major and weak IVs, then form the composite IV for the weak ones in I2 with the weight ω being calculated based on Eq (3), then get the fitted X^2 in I2. Similarly, we use data in I2 to select major and weak IVs, then use data in I1 to form the composite IV and get the fitted values X^1. Next, we combine X^1 and X^2 to get X^=(X^1T,X^2T)T and fit the 2nd stage regression model Y∼X^ to get the causal estimate (β^) and its testing p-value.

Denote the selected index of weak IVs as Sk,W, k = 1, 2, where |Sk,W| = p2 represents the selected numbers of weak IV using data in Ik. Taking sample I1 as an example, let the estimated effects for the weak IVs on the exposure be denoted as α^1,W={α^1,j;j∈S1,W} for data in I1. Here, the subscript 1 indicates that this parameter is estimated from subsample I1, and the subscript W signifies that it corresponds to the direct effect of weak IVs on X from Eq (1). The new composite IV constructed in sample I2, G^2,W, is then defined as G^2,W=∑j∈S1,WωjG2,j (2)

where ωj=sign(α^1,j)|α^1,j|∑j∈S1,W|α^1,j| (3)

In other word, we use weak IVs selected from sample I1 to construct the new composite IV in sample I2. Then, we can use the major IVs and the new composite IV, i.e., {G2,M,G^2,W}, to get the cross-fitted exposure in I2. Here, the subscript M represents that G2,M is identified as the major IV, while the subscript W signifies that G^2,W is estimated from the weak IV. The subscript 2 indicates that these values are obtained from subsample I2.

To clarify logic, for each k = 1, 2 we use the subset Ik to identify the SNP IVs and obtain the estimated effects α^ for the selected IVs. Then, we categorize them into two groups, major IVs and weak IVs, using the partial F-statistic criterion defined earlier. We then combine the weak IVs in Ikc using the estimated weights from Ik. This approach enables us to avoid overfitting by selecting IVs and estimating the causal effect using different samples.

Estimating the causal effect

Once we get the IVs {GM,G^W} in each Ik, k = 1, 2, we can then perform the first stage of the 2SLS regression on these IVs to get the cross-fitted exposures X^k which are then aggregated, i.e., X^=(X^1X^2)∈RN×1.

The causal effect can be estimated by regressing Y on X^ using the whole sample, which is given by β^=(X^′X^)-1X^′Y (4)

Cross-fitting allows for the utilization of the entire dataset in estimating causal effects, thereby circumventing the winner’s curse problem that arises when the same data is employed for both IV selection and causal effect estimation.

Remark 1: Both MR-SPLIT and CFMR implement a cross-fitting idea with sample splitting, but the analysis is fundamentally different. MR-SPLIT combines the cross-fitted exposures for further causal inference, while CFMR combines cross-fitting instruments for further causal inference. CFMR first calculates the composite IV in the kth split sample (denoted as G˜k,nk×1), where nk denotes the sample size in the kth split sample, then combines these composite IVs to form the final composite IV (denoted as G˜N×1 by stacking all G˜k), and finally uses the full data (X,G˜,Y) to perform 2SLS analysis for causal inference. Each composite IV G˜k,nk×1 can be regarded as a polygenic risk score based on the kth split sample. As the dimensions of major and weak IVs identified from different sample splits are different, CFMR is infeasible to separate the two components and incorporate them in downstream causal inference.

Remark 2: CFMR uses data in Ikc to select IVs, then calculates the composite IV based on data in Ik. Typically, a 10-fold split suffices. On the other hand, MR-SPLIT benefits from a 2-fold sample splitting. It uses data in I1 to select and separate major and weak IVs, then forms the cross-fitting composite IV in I2. After that, it calculates the cross-fitted exposures based on the major IV(s) and the cross-fitting composite IV for further causal inference. More data for IV selection leads to less data to fit the cross-fitted exposure and vice versa. To balance the two components, a 2-fold sample split is recommended.

Remark 3: By combining the cross-fitted exposures, Theorem 1 shows that MR-SPLIT produces estimates with a variance no larger than that of CFMR if both approaches implement a 2-fold sample split. The proof is given in S1 Text. This finding extends to scenarios involving a k(>2)-fold sample split for CFMR. Although providing theoretical proof for this result poses a challenge, we have demonstrated its validity through simulations.

Theorem 1 Let β^CFMR and β^MR-SPLIT be the 2SLS estimates obtained respectively by the CFMR and MR-SPLIT method with a 2-fold random split. Then, β^MR-SPLIT is more efficient than β^CFMR in the sense that var(β^MR-SPLIT)≤var(β^CFMR)

The proof of Theorem 1 is given in S1 Text.

Multiple sample splitting to improve robustness

Given the inherent uncertainty in single-sample splitting, particularly in cases of limited sample size, we propose a multiple-splitting strategy to improve the robustness of the approach. We randomly split data (into two halves) L times. For each random split, the same estimation and testing procedure as described before are conducted. Let pvall denote the p-value at the lth random split. There are different ways to aggregate these L p-values. One approach involves employing the aggregation method for p-values proposed by Wasserman and Roeder [29, 30]. However, this method has proven to be overly conservative in our simulations. Another simple way is to use the Cauchy combination rule for correlated p-values [31], which is similar to the minimum p-value method but does not require an intensive resampling procedure to assess the null distribution of the minimum p-value. Given its computational efficiency, we adopt the Cauchy combination rule to aggregate p-values obtained from multiple sample splitting. Following [31], the test statistics is defined as: Tcauchy=∑l=1Lωltan((0.5-pvall)π)

where the weights ωl are non-negative and ∑l=1Lωl=1. If no further information is available, the weight ωl can be simply chosen as 1/L. The p-value of Tcauchy can be simply approximated by p-value=12-arctan(Tcauchy)/π (5)

In essence, augmenting the number of sample splits improves result robustness. However, this enhancement comes with the trade-off of requiring increased computational resources. To provide general guidance on the number of splitting times, we conducted a simulation study (see section Simulation study). The results suggest that conducting multiple splits about 50–60 times is sufficient to achieve a robust outcome in terms of controlling type I errors and maintaining stable statistical power. In the case of a large sample size and strong SNP heritability, the splitting time can be dramatically reduced (see the simulation results).

Algorithm

The detailed algorithm of the MR-SPLIT is given below:

For each l = 1, ⋯, L random split, repeat the following steps:

Split the sample into two equal subsets {I1, I2}, i.e., {1, ⋯, N} = I1 ∪ I2 with I1 ∩ I2 = ∅ and |I1| = [N/2] and |I2| = N − [N/2], and denote the complementary sets as {I1c,I2c} accordingly.

For each k = 1, 2, we use Ikc to select IVs and get the estimated effect size for each IV. Then, categorize the selected IVs into two distinct groups, major IV(s) and weak IVs, based on the partial F > 30 criterion.

In each subset Ik, combine the weak IVs using the effect size estimated from Ikc following Eq (2). Then regress the exposure variable X on the new IVs (major IV(s) + composite IV) to get the fitted value X^k.

Denote X^=(X^1X^2), and do the second stage regression of Y on X^ to get the causal effect estimate β^l and the p-value pvall.

Calculate the Cauchy combination statistics Tcauchy=1L∑l=1Ltan((0.5-pvall)π), and the aggregated p-value as pval=12-arctan(Tcauchy)/π. The final aggregated causal effect estimate can be calculated as β^=1L∑l=1Lβ^l.

Results

Simulation study

We conducted simulations to assess the performance of our method and provided guidance on the identification of major IVs and selecting an efficient number of sample splitting. Subsequently, we compared the proposed MR-SPLIT with the existing approaches, including 2SLS, LIML, and CFMR, across various settings.

Evaluation of Major IV identification

We applied 3 criteria, F > 10, F > 30, and F > 50, to distinguish the major and weak IVs under various settings. We randomly generated 300 independent SNPs each with MAF = 0.3, and assumed only 5 SNP had effects on the exposure. The effects of these SNPs were set to be β = (0.4, 0.4, 0.1, 0.05, 0.05)ω0, where ω0 was chosen to ensure that these SNPs account for h2 = {0.15, 0.30, 0.50} of the variation in exposure (h2 can be interpreted as the exposure heritability). The error term was assumed to follow the standard normal distribution with mean 0 and variance 1. The rest 295 SNPs were assumed to be noises with no effect on the exposure (i.e., β = 0). Then, we followed model (1) to simulate the exposure. In this setting, the initial two SNPs may be regarded as the major IVs, whereas the remaining three are categorized as weaker ones. However, this differentiation can also be contingent on the signal-to-noise ratio, meaning that the first two SNPs may not be deemed as the major ones when h2 is low, say h2 = 0.15. And when the IVs are strong enough, say h2 = 0.5, the three weaker IVs may be regarded as strong IVs. After applying SIS screening and LASSO estimation on these 300 SNPs, we then used these three criteria to distinguish the major and weak IVs. Part of the results can be seen in Table 1. As we mentioned before, there were 295 noise SNPs in total. It is possible that some of these noise SNPs may be incorrectly identified as major IVs. We also summarized these results in the last column of Table 1. More detailed information can be found in Table A and Fig A in S1 Text. Our analysis indicates that employing a partial F > 10 threshold to define major IVs is excessively lenient, leading to misidentifying noises as major IVs, particularly in scenarios with small sample sizes (e.g., N = 500). Conversely, a threshold of F > 50 proves overly stringent, failing to recognize SNP 1 and 2 as major IVs in conditions characterized by low sample sizes and heritability. A threshold of F > 30 emerges as a balanced criterion for defining major IVs, effectively mitigating the aforementioned issues. Thus, we propose to use a partial F > 30 threshold in the selection of major IVs.

10.1371/journal.pgen.1011391.t001 Table 1 Mean numbers of being identified as major IV using different criteria in 1,000 simulations.

h 2	N	Criteria	SNP1	SNP2	SNP3	SNP4	SNP5	Noises (×295)*	
0.15	500	F>10	0.55	0.5	0.15	0	0.1	1.25	
F>30	0.05	0.05	0	0	0	0	
F>50	0	0	0	0	0	0	
1000	F>10	0.95	0.95	0.35	0.1	0	0.65	
F>30	0.35	0.35	0	0	0	0	
F>50	0.15	0	0	0	0	0	
2000	F>10	1	1	0.75	0.35	0.55	0.6	
F>30	1	0.95	0.1	0	0	0	
F>50	0.75	0.75	0	0	0	0	
0.3	500	F>10	0.95	0.8	0.25	0.25	0.1	1.15	
F>30	0.5	0.55	0	0	0	0	
F>50	0.25	0.25	0	0	0	0	
1000	F>10	1	1	0.75	0.45	0.3	0.65	
F>30	1	1	0.2	0	0	0	
F>50	0.8	0.7	0.05	0	0	0	
2000	F>10	1	1	1	0.75	0.9	0.6	
F>30	1	1	0.6	0.15	0.1	0	
F>50	1	1	0.1	0	0.05	0	
0.5	500	F>10	1	1	0.8	0.35	0.4	1.8	
F>30	1	1	0.35	0	0.05	0	
F>50	0.9	0.9	0	0	0	0	
1000	F>10	1	1	1	1	0.95	0.6	
F>30	1	1	0.8	0.5	0.15	0	
F>50	1	1	0.4	0.05	0	0	
2000	F>10	1	1	1	1	1	0.4	
F>30	1	1	1	0.9	0.9	0	
F>50	1	1	1	0.4	0.3	0	
*The total number of noise SNPs incorrectly identified as major IVs out of the 295 noise SNPs.

Comparison with 2SLS and LIML

We compared the proposed MR-SPLIT with the widely-used 2SLS approach and the LIML method which is particularly designed to address the weak instruments bias issue. We simulated 300 SNPs independently and randomly selected 5 SNPs as the IVs to generate the exposure variable X. We set h2 = {0.15, 0.3, 0.5} which respectively represent weak, moderate, and strong overall effect, and ρ = (0.1, 0.2) where ρ = cor(εxi, εyi) controls the unknown confounding effect.

We set the sample size (N) to 1000. To ensure a fair comparison with 2SLS and LIML, we only split the sample once (i.e., no multiple splitting). We then used one subset for selecting the IVs and incorporated the other subset with the selected IVs for estimation. Both 2SLS and LIML followed the same process but did not differentiate between major and weak IVs for further causal inference. To check the impact of selection bias for 2SLS and LIML, we also did the analysis using the whole dataset for both IV selection and causal effect estimation. The simulation settings are the same as what we previously described. The only difference is that we do not split the sample and use the whole sample to do the IV selection and estimation. Results for this analysis were given in S1 Text. The respective boxplots, illustrating the distribution of estimations across 1000 simulation iterations, are provided in Figs B, C and D in S1 Text. It is evident from the results that using the entire sample for both IV selection and effect estimation results in estimates with smaller variance but larger bias, leading to a significantly higher type I error rate. In the following, we only show the results based on sample splitting.

Table 2 presents a comparative analysis of the estimation accuracy among MR-SPLIT, LIML, and 2SLS. It shows that MR-SPLIT can provide estimates with a significantly small bias. In contrast, the estimates from 2SLS exhibit large bias, especially under weak IV and substantial confounding effects (e.g., ρ = 0.2). In some of the cases, LIML gives a smaller bias than MR-SPLIT does, but it has consistently larger variance than MR-SPLIT, leading to a conservative coverage probability (CP) compared to MR-SPLIT. The variance of 2SLS is uniformly smaller than the other two methods. However, given its large bias, it has the most poor coverage probability among the three methods. On the other hand, MR-SPLIT shows consistently good coverage probabilities under different scenarios, showcasing its robust performance under different conditions.

10.1371/journal.pgen.1011391.t002 Table 2 Simulation comparison between M* (MR-SPLIT), LIML and 2SLS.

h 2	ρ	β	Bias(|β-β^|×100)	Est. SE	CP*	
M*	LIML	2SLS	M*	LIML	2SLS	M*	LIML	2SLS	
0.15	0.1	-0.08	0.18	0.17	4.86	0.1252	0.1776	0.0788	0.955	0.827	0.895	
0.08	0.82	0.23	5.05	0.1253	0.1884	0.0795	0.957	0.832	0.886	
0.2	-0.08	0.17	1.19	9.45	0.1230	0.1737	0.0782	0.949	0.843	0.740	
0.08	0.31	1.01	9.44	0.1320	0.1770	0.0801	0.952	0.825	0.732	
0.30	0.1	-0.08	0.02	0.11	2.4	0.0520	0.0844	0.0605	0.958	0.895	0.929	
0.08	0.13	0.59	2.94	0.0515	0.0831	0.0600	0.959	0.898	0.919	
0.2	-0.08	0.43	0.67	4.77	0.0501	0.0840	0.0612	0.946	0.904	0.837	
0.08	0.47	0.31	5.03	0.0524	0.0840	0.0621	0.947	0.905	0.845	
0.50	0.1	-0.08	0.32	0.08	1.12	0.0329	0.0482	0.0430	0.938	0.943	0.945	
0.08	0.11	0.22	0.82	0.0328	0.0474	0.0424	0.948	0.931	0.944	
0.2	-0.08	0.14	0.02	2.08	0.0335	0.0513	0.0457	0.942	0.909	0.902	
0.08	0.2	0.03	2.07	0.0318	0.0469	0.0423	0.954	0.938	0.924	
CP*=coverage probability

Fig 2 shows the results of the type I error of the three methods. We can observe that MR-SPLIT can effectively control Type I errors, even in the presence of strong unknown confounding. As depicted in Fig 2, both LIML and 2SLS methods exhibit much poorer performance than MR-SPLIT. Notably, 2SLS suffers from poor type I error control when the confounding effect is strong (i.e., ρ = 0.2), leading to inflated error rates. LIML has high false positive rates when the SNP effects are weak (i.e., weak instruments with low h2), especially under ρ = 0.2. As the SNP effects increase, its performance improves; however, it can only effectively control type I errors when the instrumental variables are strong, as demonstrated in the scenario with h2 = 0.5. Conversely, MR-SPLIT consistently demonstrates robust type I error control under all conditions, even under h2 = 0.15 and ρ = 0.2, where 2SLS and LIML exhibit their poorest performance. The inflated type I error rates lead to inflated statistical power for 2SLS and LIML. Consequently, comparing power between MR-SPLIT and these two methods may not be a fair comparison; thus, we did not show the detailed power comparison here. Nevertheless, in the scenario where h2 = 0.5 and ρ = 0.1, MR-SPLIT still attains the highest power, reaching 0.683 compared to 0.545 for 2SLS and 0.419 for LIML.

10.1371/journal.pgen.1011391.g002 Fig 2 Type I error comparison between MR-SPLIT, 2SLS and LIML.

The horizontal dashed line denotes the 0.05 level.

Comparison with CFMR

We compared our method with CFMR under different simulation scenarios. To ensure a fair comparison with CFMR, we applied 10-fold CFMR as recommended in the CFMR work, and 2-fold MR-SPLIT with 50 random sample splits. We applied the same procedure for selecting IVs. While CFMR combined all the selected IVs into a single composite one, our method differentiated between major and weak IVs using the partial F > 30 criterion and only weak IVs were combined into a composite one. We also followed the simulation settings described in the CFMR work to ensure a fair comparison. We generated a set of 300 SNPs, and the minor allele frequency is fixed as 0.3 for all the SNPs. We randomly chose 5 SNP IVs to generate the exposure variable with the model X=∑j=15Gjαj+εx, and the outcome with the model Y = Xβ + εy, where (εxεy)∼N(0,5(10.160.161))

We set two scenarios to comprehensively compare MR-SPLIT and CFMR:

Scenario I: The effect sizes of the 5 SNPs are different, i.e., α = (0.4, 0.4, 0.1, 0.05, 0.05). Potentially, SNPs with the effect of 0.4 can be regarded as major IVs and the rest can be considered as weak ones. This also depends on the SNP heritability level h2.

Scenario II: The effect sizes of the 5 SNPs are the same, i.e., α = (0.2, 0.2, 0.2, 0.2, 0.2). In this case, differentiating between major and weak IVs can be challenging, presenting a less favorable condition for our method.

In each scenario, we compared the two methods in different aspects by changing the sample size (N = {1000, 3000, 5000}), variation in the exposure explained by the SNP IVs (h2 = {0.15, 0.2, 0.3}) and the exposure’s effect size for β.

Fig 3 depicts the type I error control of the two methods in scenario I and scenario II. In general, the control of type I error for the two methods is highly comparable across different settings characterized by distinct sample sizes and SNP heritability levels. Though the type I error is a little inflated for MR-SPLIT under a small sample size (N = 1000), particularly in scenario II, it controls the type I error well as the sample size increases.

10.1371/journal.pgen.1011391.g003 Fig 3 Comparison of type I error between MR-SPLIT and CFMR in Scenario I (top) and II (bottom).

The horizontal dashed line denotes the 0.05 level.

Fig 4 shows the results of the power for the two methods in these two scenarios. Regardless of the settings, MR-SPLIT consistently exhibits higher power than CFMR. This discrepancy becomes especially noticeable when the IVs are relatively weak (i.e., h2 = 0.15).

10.1371/journal.pgen.1011391.g004 Fig 4 Power comparison between MR-SPLIT and CFMR in Scenario I (top) and II (bottom).

In S1 Text, we also presented the estimation performance of both methods when β = 0 and 0.08 in Figs E-J. The results reveal minimal difference in the causal effect estimation between the two methods. However, a noticeable distinction is the smaller standard error observed in MR-SPLIT across nearly all the scenarios, resulting in a smaller Root Mean Square Error (RMSE) (see Fig K in S1 Text) and higher statistical power when compared to CFMR. This aligns well with the theoretical finding in Theorem 1, though the result is proved under a 2-fold sample split. The findings further underscore the advantages of MR-SPLIT.

In summary, MR-SPLIT consistently demonstrates robust type I error control when compared to 2SLS and LIML across various simulation settings. In comparison to CFMR, MR-SPLIT exhibits superior performance, yielding smaller RMSE and higher statistical power. Even under a less favorable condition for MR-SPLIT, the type I error can be controlled when the sample size is reasonably large. The simulation results further corroborate our theoretical finding, consistently showing that MR-SPLIT results in smaller standard errors for causal effect estimation compared to CFMR, which leads to higher statistical power when testing for the causal effect.

Evaluation of Multiple data splitting to improve robustness

Intuitively, more data splitting should yield more robust results, which, however, would entail higher computational resource usage. We implemented our methods under different splitting times, different sample size N and different h2 values, to check if we can find an efficient number of splitting. In our simulations, the true causal effect of the exposure on the outcome is set to equal 0.2 (β = 0.2), and the sample size ranges from 500 to 2000 (N = 500, 1000, 2000). We did simulations in Scenario I as described in section Comparison with CFMR. Fig 5 demonstrates how the type I error fluctuates with an increasing number of splits. To obtain a smoother estimate of the type I error rate, we repeated the simulation 5,000 times Under a small sample size, the type I error rates get stable as the number of sample splits increases. Though the type I error increases as the sample split times increase under small sample sizes, this increase is considered acceptable, particularly in light of the associated boost in power (see Fig 5), which is especially pertinent for smaller sample sizes.

10.1371/journal.pgen.1011391.g005 Fig 5 Type I error under different sample sizes: N = 500(left), 1000(middle), 2000(right), and under different h2: 0.15 (top) and 0.2 (bottom).

Fig 6 shows the empirical power under different sample sizes and h2. The type I error and power results when h2 = 0.3 can be found in Figs L and M in S1 Text. When the sample size is small (N = 500) and the IVs are relatively weak (h2 = 0.15), the power gets stabilized after 50 splits. As the sample size increases, there is a decrease in the need for the number of sample splits to maintain stable power. This indicates that in practical data analysis, it is possible to estimate the exposure heritability based on the selected SNP IVs, and thereafter determine the appropriate number of sample splits. In any case, opting for 50 sample splits represents a highly conservative option.

10.1371/journal.pgen.1011391.g006 Fig 6 Empirical power under different sample sizes: N = 500(left), 1000(middle), 2000(right), and under different h2: 0.15 (top) and 0.2 (bottom).

Case study

We demonstrated the effectiveness of our method by applying it to the Chronic Renal Insufficiency Cohort (CRIC) dataset, to understand the progression of chronic kidney disease (CKD).

CKD is evaluated utilizing two straightforward tests: a blood test known as the estimated glomerular filtration rate (eGFR) and a urine test, the urine albumin-creatinine ratio (uACR). Both eGFR and uACR measure kidney function, with low eGFR and high uACR values indicating impaired kidney function. In this application, we are interested in evaluating the causal relationship between CKD and apparent Treatment-Resistant Hypertension (aTRH). aTRH is a condition where a patient’s blood pressure remains above target levels despite using three different classes of antihypertensive drugs at optimal doses, typically including a diuretic. The definition of aTRH also extends to cases where four or more medications are required to effectively control blood pressure [32]. In a recent two-sample MR analysis using summary statistics, Yu et al. [33] identified the causal effect of higher kidney function (measured by eGFR estimated from serum creatinine) on lower systolic blood pressure. To date, the causal relationship between CKD and aTRH and the causal link between them remains to be established [34–37]. To this end, we utilized eGFR and uACR as the exposure variable and aTRH as the outcome, applying MR-SPLIT for our analysis. For comparative purposes, we also employed CFMR and 2SLS on the same dataset. Given that the outcome variable (aTRH) is binary (0/1) in nature, the LIML method is not suitable in this analysis.

Genetic data processing

The original data have 3,541 samples containing 970,342 SNPs. Our initial step involved removing SNPs with missing rate larger than 10%, resulting in 886,384 SNPs. After excluding SNPs with minor allele frequency (MAF) lower than 0.05, 762,664 SNPs were left. The next phase entailed the elimination of SNPs with p-values less than 1e-5 in the Hardy-Weinberg equilibrium test, which narrowed our SNP count down to 693,848. To ensure the robustness of our genetic instruments, we then implemented LD pruning. SNPs were filtered out in close LD by considering pairs of SNPs within a window of 100 kb. If a pair of SNPs has an LD measure (r2) exceeding 0.64, one SNP from the pair is removed. After completing all these steps, we were left with 467,597 SNPs.

Causal effect of eGFR on aTRH

In the initial dataset, eGFR values were obtained on multiple occasions. For consistency and relevance, we selected the eGFR measurements corresponding to visit number 3, which also represents the baseline assessment. Following the exclusion of samples with missing values for either eGFR or aTRH, and then combined with the SNP data, our analysis proceeded with a total of N = 1, 353 samples. A simple logistic regression shows there is a strong association between aTRH and eGFR (p < 2 × 10−16). We would like to evaluate if this association is causal. Fig S20 shows the boxplots of eGFR in aTRH positive and negative groups.

Next, we proceeded with the MR-SPLIT and used SIS for preliminary screening, reducing the number of SNPs from ultra-high to high. To optimize computational efficiency in the analysis, we first conducted univariate regression of each SNP against the exposure before applying sample split, using the whole data set. A total of 4,580 SNPs (p < 0.01) remained for further analysis. The removed SNPs would most likely be screened out by the SIS procedure in subsequent steps even after the sample split if not discarded at this stage. For each of the 50 sample splits, we used the ‘screening’ function from the R package ‘screening’ with the SIS option. The number of SNP variables retained post-screening adhered to the default setting, which is half the size of the sample. In this real data analysis, instead of applying the LASSO algorithm to select and estimate SNP effects, we employed a high-dimensional inference procedure, specifically a LASSO-projection method which provides debiased coefficient estimates and hence a valid p-value for each coefficient. This is done by using the ‘lasso.proj’ function in the R ‘hdi’ package [30]. As the regular LASSO estimates are biased, this approach can give debiased estimates and further provide p-values for testing each coefficient. To compare the performance of the LASSO-projection with the regular LASSO, We conducted a simulation (detailed in Section 6 in S1 Text). The results show that the LASSO-projection method slightly outperforms LASSO, exhibiting higher power and better control of the type I error rate and smaller RMSE. After getting the p-values for each SNP, we retained those with a p-value less than or equal to 0.05. This resulted in an average of 98 IVs out of 50 sample splits. We used the partial F > 30 as the criterion to declare major IVs. And the weak IVs were then combined into a composite IV. Finally, we used both the composite IV and the major IV(s) to obtain the causal effect estimate and the p-value.

Fig 7 shows the p-value distribution and the causal effect estimates out of 50 sample splits. The majority of p-values obtained from MR-SPLIT are below 0.05, and the majority of the estimated causal effects β^ is centered around -0.0343 (indicated by the black dashed line). In these 50 sample splits, there was an average of 98.06 IVs incorporated into the model for the causal effect estimate and the majority were classified as weak IVs. Among these, an average of 0.54 IVs were identified as major IVs each time. After aggregating all the results using Cauchy’s combination rule, our method provided an estimate of β^=-0.0343 (OR = 0.9663), with an aggregated p-value of 5.96 #x00D7; 10−5. We also tried lowering the partial F threshold to 20, which yielded slightly more major IVs than the F > 30 threshold (see Fig V in S1 Text). Among the 50 sample splits, an average of 4.26 IVs were identified as major IVs each time. The results show that the p-value for MR-SPLIT improved slightly (from 5.9 × 10−5 to 2.9 × 10−6), but the estimates remained nearly the same (β^=-0.0342).

10.1371/journal.pgen.1011391.g007 Fig 7 Histogram of p-values and causal effect estimates from 50 sample splits when eGFR is treated as the exposure.

We also applied the CFMR method with a 10-fold split. The CFMR method yielded an average estimate of β^=-0.0378 (OR = 0.9629), with a p-value of <1 × 10−5. For reference, simply conducting the 2SLS method yields an estimate of β^=-0.0407 (OR = 0.9601), with a p-value of <1 × 10−7. The three methods established a consistent causal relationship between eGFR and aTRH.

Causal effect of uACR on aTRH

Following a similar procedure, we excluded samples with incomplete data for either uACR or aTRH. After merging the remaining data with the SNP data, the dataset was reduced to 1,324 samples. The distribution of uACR is very skewed (to the right) (See Fig W in S1 Text). We opted to do a logarithmic transformation of uACR, denoted as log(uACR). A simple logistic regression shows there is a strong association between log(uACR) and aTRH (p = 4.6 × 10−12). Similar procedures as described before were followed for further analysis.

Fig 8 shows the p-value distribution as well as the causal effect estimate out of 50 sample splits with MR-SPLIT. In these 50 sample splits, there were on average 74.3 SNPs selected as IVs with the majority as weak ones for the causal effect estimate. Among them, an average 0.22 IVs were identified as major IV each time. After aggregating all the results, the final causal estimate was β^=0.1675 (OR = 1.186), with a p-value of 1.9 × 10−3. For comparison, CFMR provided an estimate of β^=0.1584 (OR = 1.1716), with a p-value of 5.2 × 10−5. The two methods yielded statistically significant results and presented comparable estimates. While applying 2SLS on the same dataset, we also observed significant results (p-value = 1.3 × 10−5), albeit with a different causal effect estimate of β^=0.0363.

10.1371/journal.pgen.1011391.g008 Fig 8 Histogram of p-values and causal effect estimates from 50 sample splits when log(uACR) is treated as the exposure.

Integrating the results from the two analyses that utilized eGFR and uACR separately as exposures, we infer that there exists a causal relationship between CKD function and aTRH. Specifically, a lower eGFR and a higher uACR tend to contribute to an increased risk of aTRH. However, we recognize the limited sample size of this study, which necessitates cautious interpretation of the causal relationship identified. To assess the possibility of a reverse causal effect, we require a method capable of accommodating a binary exposure variable, such as aTRH in this context. This will be explored in our future studies.

Discussion

MR analysis has been an instrumental means in epidemiology studies, enabling the assessment and revelation of causal connections between exposures or interventions and particular outcomes, leveraging genetic variants as IVs to mitigate confounding factors. In this study, we introduced an innovative adaptive sample splitting method known as MR-SPLIT, designed to address the issue of IV selection bias and weak instruments in the context of one-sample MR analysis using individual-level data. By a random sample split, we use half sample to select IVs and another independent half to estimate the causal effect, hence avoiding the winner’s curse problem by using the same data for IV selection and causal effect estimation. Additionally, we presented a multi-sample splitting strategy to further enhance the robustness of causal estimation and testing. Our approach involves the adaptive identification of major and weak IVs and further aggregate weak IVs to form a composite IV. The final set of IVs comprises the major IV(s) and the composite IV. Such a strategy, as shown in the theoretical evaluation and simulation results, yields a more efficient causal estimate than CFMR, thereby enhancing testing power. In addition, MR-SPLIT shows consistently superior performance in terms of coverage probability. Therefore, MR-SPLIT offers significant improvements over existing methods by effectively handling weak instruments in one-sample MR analysis and providing robust results with enhanced statistical power.

In comparison to the traditional 2SLS and LIML methods, MR-SPLIT yields less biased results and effectively controls type I error, under different simulation settings. Compared to the CFMR approach, which is designed to tackle weak IV issues, our approach provides estimates with smaller variance and higher statistical power. In the application to the CRIC dataset, both MR-SPLIT and CFMR produce highly comparable results. We established the causal impact of kidney function, as assessed by eGFR and uACR, on aTRH. It is worth noting that both CFMR and MR-SPLIT not only address the issue of weak instrument bias (i.e. finite-sample bias from IV analysis with a given set of IVs), they also solve the problem of “winner’s curse” (i.e. bias due to variant selection in the same dataset as the analysis is performed, in particular under a high-dimensional scenario). The two sets of biases are related but are conceptually distinct. By employing sample splitting strategies, both methods tackle the two bias issues and offer a solution to one-sample MR analysis. On the other hand, as shown in our theoretical evaluation as well as the intensive simulation studies, MR-SPLIT demonstrates superior performance compared to CFMR. Within the proposed sample splitting strategy, additional tasks such as nonlinear causal estimation can also be executed using one-sample individual-level data.

In the process of selecting IVs, CFMR recommends employing predictive methodologies, such as LASSO regression, for their efficacy in enhancing prediction accuracy through variance minimization. However, this approach often introduces bias in effect estimates, as it may incorporate SNPs without significant association with the exposure—potentially compromising the relevance assumption for IVs. On the other hand, 2SLS analysis prioritizes the use of predicted exposure values in its secondary causal inference phase, underlining the importance of prediction accuracy for causal estimation. Recent advancements in the realm of high-dimensional statistical inference offer a promising solution by enabling the evaluation of estimation uncertainty for LASSO-derived estimates [30]. This is achieved through a de-biasing step that facilitates the calculation of p-values, thereby presenting an innovative approach for SNP IV selection within the context of high-dimensional SNP-exposure regressions. This technique allows for the derivation of p-values for individual SNPs, enabling the validation of IV suitability through a p-value based method. Unlike traditional practices that determine p-values by fitting each SNP individually in marginal regressions, this approach fits all SNPs (after the SIS step) in a multiple regression model. This yields partial SNP effect estimates and hence partial p-values, offering a nuanced perspective compared to conventional methods. By adopting a p-value threshold criterion (e.g., p < 0.05), the selected SNPs meet the relevance assumption, providing a more robust framework for IV selection. In our analysis, we observed that the LASSO variable selection technique typically identifies a greater number of IVs compared to the debiased LASSO method. If computational resources are not a limiting factor, we recommend the implementation of the debiased LASSO approach in practical applications.

While MR-SPLIT offers notable advantages, there is still considerable potential for further enhancement and refinement. In this work, we applied the partial F statistics for identifying major IVs, which does not rule out the application of other measures such as those studied in James and Motohiro [27]. Any statistical measure capable of ranking the effect sizes of the selected IVs could be considered for enhancing the robustness and effectiveness of our approach. It is essential to devise robust methods for discerning between major and weak IVs. This represents a promising direction for future research. It is worth mentioning that we do not specify the ratio of major IVs to weak IVs; their quantities are entirely determined by the data itself, that is, based on the strength calculated from the IVs. On the other hand, as revealed by the simulation studies, the declaration of major IVs may vary under different F thresholds and under different sample sizes and SNP heritability levels. In real applications, the F > 30 threshold can be relaxed under a small sample size and low heritability level. The genomewide SNP heritability can be estimated with software such as GCTA [38].

In addition, we employed a straightforward weighted combination approach to aggregate the information from all weak IVs into a single composite IV. Other advanced machine learning techniques could also be borrowed by minimizing information loss which could potentially yield improved results.

An additional constraint of our approach is that we did not take the pleiotropic effects into consideration. Addressing pleiotropy can be a complex endeavor, but there are several test statistics available to identify its presence [39, 40]. Integrating these statistics into our method presents a challenge, as it requires the repeated selection of new IVs at each sample split and the detection of pleiotropic effects associated with the chosen IVs. We plan to tackle this issue in our future investigation. Studies also show that incorporating the invalid IVs with uncorrelated and correlated horizontal pleiotropic effects can potentially increase power and decrease bias [41–43]. We will investigate this in our future work.

The concept of sample splitting and cross-fitting instruments introduced in this study has potential applications beyond the scope of traditional one-sample MR analyses using individual-level data. For example, this framework can be adapted for use in multiple exposure MR analyses, where it would involve adapting the existing approach to handle multiple sets of selected IVs simultaneously. For another example, the proposed framework enables the investigation of potential non-linear causal relationships through a control function approach while effectively addressing the two bias issues previously mentioned. Accomplishing this task is not feasible with summary statistics, highlighting the framework’s capability to provide more nuanced insights into causal mechanisms that cannot be captured by summary-level data. In essence, the expansion of our methodology to encompass various types of MR analyses could facilitate innovative research into causal relationships, opening new avenues for investigation.

Supporting information

S1 Text Supplementary Materials for “MR-SPLIT: A novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies”.

(PDF)

The Chronic Renal Insufficiency Cohort (CRIC) Study was conducted by the CRIC Study Investigators and supported by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). This manuscript was not prepared in collaboration with Investigators of the CRIC Study and does not necessarily reflect the opinions or views of the CRIC Study, or the NIDDK.

10.1371/journal.pgen.1011391.r001
Decision Letter 0
Zhu Xiaofeng Section Editor
Morrison Jean Guest Editor
© 2024 Zhu, Morrison
2024
Zhu, Morrison
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version0
4 Apr 2024

Dear Dr Cui,

Thank you very much for submitting your Research Article entitled 'MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies' to PLOS Genetics.

The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important topic but identified some concerns that we ask you address in a revised manuscript.

We therefore ask you to modify the manuscript according to the review recommendations. Your revisions should address the specific points made by each reviewer.

In addition we ask that you:

1) Provide a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript.

2) Upload a Striking Image with a corresponding caption to accompany your manuscript if one is available (either a new image or an existing one from within your manuscript). If this image is judged to be suitable, it may be featured on our website. Images should ideally be high resolution, eye-catching, single panel square images. For examples, please browse our archive. If your image is from someone other than yourself, please ensure that the artist has read and agreed to the terms and conditions of the Creative Commons Attribution License. Note: we cannot publish copyrighted images.

We hope to receive your revised manuscript within the next 30 days. If you anticipate any delay in its return, we would ask you to let us know the expected resubmission date by email to plosgenetics@plos.org.

If present, accompanying reviewer attachments should be included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist.

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org.

Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process.

To resubmit, you will need to go to the link below and 'Revise Submission' in the 'Submissions Needing Revision' folder.

Please let us know if you have any questions while making these revisions.

Yours sincerely,

Jean Morrison

Guest Editor

PLOS Genetics

Xiaofeng Zhu

Section Editor

PLOS Genetics

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: The manuscript presented, MR-SPLIT, a novel method to address selection and weak instrument bias in one-sample MR analysis. The authors first provided the theoretical justification in terms of the efficiency for the causal effect estimate, and then conducted various simulations, parallelized with real data applications, to illustrate the advantage of MR-SPILIT over other competing methods. Overall, the manuscript is of great interest and well-written, I just have the following comments for further improvement.

1: Given that including more correlated IVs, rather than only independent IVs through LD clumping or pruning, can increase the power of MR analysis. I wonder whether MR-SPLIT can be developed for correlated IVs. I suggested the authors to discuss this.

2. It seems that there is no gold standard to determine either the major or the weak IVs. For fixed number of IVs, dose the performance of MR-SPLIT rely on the ratio of the number of major IVs to the number of weak IVs.

3: In the real data analysis, it seems that the power advantage of MR-SPLIT over CFMR is not that obvious, and CFMR can even produce much smaller p value than MR-SPLIT. I suggested the authors to investigate the possible reasons. For example, are there any IV associated with the outcome? If so, what happened if we remove them from the analysis. Indeed, this can partly alleviate the issue of pleiotropy.

4. In the real data analysis, what is the criteria to determine the major IVs or the weak IVs? Again, what happened if we changed the ratio of the number of major IVs to the number of weak IVs.

5. There are some typos in the manuscript, For example, in line 4 of the 4.3 section, p=4.6*10(-12). I suggested the authors to double check the whole manuscript again.

Reviewer #2: This manuscript proposed a novel method named as Mendelian Randomization with adaptive Sample-sPLitting with cross-fitting InstrumenTs (MR-SPLIT) to address IV selection issue and weak instrument bias under the 2SLS IV regression framework.

Overall, the paper is well-written and the methods are nice outlined both in the main text and in supplementary materials. Below are my comments and questions:

Introduction:

P2. The author explained the assumptions of IVs in the MR analysis. The violation of the relevance assumption leads to a 'weak instrument' issue. Please describe why the challenge of IV selection should also be considered when this problem was first mentioned.

P5 and P7. The author only introduced methods that handled the weak instrument issue. Are there any methods that consider the IV selection issue? Please add some citations.

P8. Please summarize how this method addressed the bias of IV selection.

Methods:

In section 2.1.1 P3. Please describe the definition of partial F-statistics and explain why to use partial F-statistics as criteria.

In section 2.1.2. Please explain how to use I_1 and I_2 to select major and weak IVs, respectively. Could you please explain whether the weak IVs used to construct composite IV are the same as the IVs obtained in sample I_1?

In Remark 2. The author described that MR-SPLIT benefits from a 2-fold sample splitting. It selects and separates major and weak IVs in I_1, then forms the cross-fitting composite IV in I_2. Please explain why 2-fold sample splitting can deal with the weak instrument bias and IV selection issue.

In section 2.1.5 algorithm. How to define the value of L random split in the simulations studies and real data analysis?

Simulation Study:

In section 3.1. The author can determine what criteria to distinguish the major and weak IVs by conducting simulations. Please describe the value of the threshold used in the real data analysis and how to determine this threshold.

In sections 3.1 and 3.2. Please add more details to explain how to randomly generate 300 SNPs. Do you generate the simulated datasets following the model (1) in section 2? Are these SNPs independent with each other? Do you consider the LD matrix when generating the SNPs?

P2. Please add more details on conducting simulations to check the impact of selection bias, including specific methods for sample division.

Case Study:

P2. Please explain the reason for using a LASSO-projection method. It would be helpful to provide some numerical results to show that this one is better than LASSO algorithm to select and estimate SNP effects.

Please describe the value of the partial F statistic threshold and the reason.

Could you provide any references that confirm the causal relationship between eGFR and aTRH? If available, please include the citations.

In real data analysis, the author only compares the MR-SPLIT with CFMR. Please add more results to compare with LIML and 2SLS.

Reviewer #3: The author introduces MR-Split, an interesting extension of the work of Denault and colleagues on CFMR. The limitations of the CFMR are well presented, and the approach proposed by the author is sound and seems to address correctly the limitations of CFMR. They even provide theoretical arguments for MR split superiority compared to CFMR.

I am positive about accepting this paper, but I have a minor revision.

Comments:

1) The paper needs some polishing/sharpening. For example, at the start of section two, the authors say, “assume the following 2SLS” model. 2SLS is an estimation method for this kind of model but not a model by itself. This model is an SEM. Another example in the supplementary material on page 14 is that the authors wrote, “To prove S14, we need to show ???” This is obviously a latex mistake regarding the references. It would be good if this kind of typos/impreciseness were corrected for the next version.

2) Regarding Figure 5, I appreciate that the authors point out the lack of control of type one error when increasing the number of splits in a small sample size. I suspect that some simple bootstrapping method could be used to correct this bias. If this is not possible, I still agree that the power gain is worth a slight increase in type I error. Furthermore, it would be good to have a smoother estimate of the type I error. Could the authors perform more simulations(given that this is for a small sample, I suspect that should be doable.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: None

Reviewer #2: Yes

Reviewer #3: Yes

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

Reviewer #3: Yes: William R.P. Denault

10.1371/journal.pgen.1011391.r002
Author response to Decision Letter 0
Submission Version1
24 Apr 2024

Attachment Submitted filename: Response to the comments.pdf

10.1371/journal.pgen.1011391.r003
Decision Letter 1
Zhu Xiaofeng Section Editor
Morrison Jean Guest Editor
© 2024 Zhu, Morrison
2024
Zhu, Morrison
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version1
18 Jul 2024

Dear Dr Cui,

Thank you very much for submitting your Research Article entitled 'MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies' to PLOS Genetics.

The manuscript was fully evaluated at the editorial level and by independent peer reviewers. The reviewers appreciated the attention to an important topic but identified some concerns that we ask you address in a revised manuscript.

We therefore ask you to modify the manuscript according to the review recommendations. Your revisions should address the specific points made by each reviewer.

In addition we ask that you:

1) Provide a detailed list of your responses to the review comments and a description of the changes you have made in the manuscript.

2) Upload a Striking Image with a corresponding caption to accompany your manuscript if one is available (either a new image or an existing one from within your manuscript). If this image is judged to be suitable, it may be featured on our website. Images should ideally be high resolution, eye-catching, single panel square images. For examples, please browse our archive. If your image is from someone other than yourself, please ensure that the artist has read and agreed to the terms and conditions of the Creative Commons Attribution License. Note: we cannot publish copyrighted images.

We hope to receive your revised manuscript within the next 30 days. If you anticipate any delay in its return, we would ask you to let us know the expected resubmission date by email to plosgenetics@plos.org.

If present, accompanying reviewer attachments should be included with this email; please notify the journal office if any appear to be missing. They will also be available for download from the link below. You can use this link to log into the system when you are ready to submit a revised version, having first consulted our Submission Checklist.

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email us at figures@plos.org.

Please be aware that our data availability policy requires that all numerical data underlying graphs or summary statistics are included with the submission, and you will need to provide this upon resubmission if not already present. In addition, we do not permit the inclusion of phrases such as "data not shown" or "unpublished results" in manuscripts. All points should be backed up by data provided with the submission.

To enhance the reproducibility of your results, we recommend that you deposit your laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols

Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

PLOS has incorporated Similarity Check, powered by iThenticate, into its journal-wide submission system in order to screen submitted content for originality before publication. Each PLOS journal undertakes screening on a proportion of submitted articles. You will be contacted if needed following the screening process.

To resubmit, log into your Editorial Manager account and select the option 'Revise Submission' in the 'Submissions Needing Revision' folder.

Please let us know if you have any questions while making these revisions.

Yours sincerely,

Jean Morrison

Guest Editor

PLOS Genetics

Xiaofeng Zhu

Section Editor

PLOS Genetics

Thank you for submitting this revision. Both reviewers had positive comments. Reviewer 2 has requested some clarification of notation and methods. If these minor changes are addressed, we should be able to move forward with accepting your manuscript.

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: The authors have addressed all my comments.

Reviewer #2: Review of “MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies” by Ruxin Shi, Ling Wang, Stephen Burgess and Yuehua Cui

Summary:

This manuscript proposed a novel method named as Mendelian Randomization with adaptive Sample-sPLitting with cross-fitting InstrumenTs (MR-SPLIT) to address IV selection issue and weak instrument bias under the 2SLS IV regression framework.

Overall, the paper is well revised, both in the main text and in supplementary materials based on the major comments, but I have the following minor suggestions:

Methods

1. Please include brief descriptions in Figure 1 to outline the workflow of MR-SPLIT and refine this flow chart. Specifically, please adjust the size of the rectangles (the third row).

2. In Sections 2.1.2 and 2.1.3, please describe the notations α ^ and {G_M,G ^_W} clearly.

3. In Section 2.1.3 Remark 1, please explain the notation n_k in G ~_(k,n_k×1).

Simulations

1. In Section 3.1.1, please add more details and explain the processes for generating independent SNPs, the exposure, and the outcome clearly (i.e., the value of β and error terms).

2. Please add the explanations of the last column (Noises) in the caption of Table 1 and Section 3.1.1.

3. In Section 3.1.2, the first line has an extra right parentheses on page 11.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: None

Reviewer #2: None

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

Reviewer #2: No

10.1371/journal.pgen.1011391.r004
Author response to Decision Letter 1
Submission Version2
29 Jul 2024

Attachment Submitted filename: Response to the comments_R2.pdf

10.1371/journal.pgen.1011391.r005
Decision Letter 2
Zhu Xiaofeng Section Editor
Morrison Jean Guest Editor
© 2024 Zhu, Morrison
2024
Zhu, Morrison
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Submission Version2
9 Aug 2024

Dear Dr Cui,

We are pleased to inform you that your manuscript entitled "MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies" has been editorially accepted for publication in PLOS Genetics. Congratulations!

Before your submission can be formally accepted and sent to production you will need to complete our formatting changes, which you will receive in a follow up email. Please be aware that it may take several days for you to receive this email; during this time no action is required by you. Please note: the accept date on your published article will reflect the date of this provisional acceptance, but your manuscript will not be scheduled for publication until the required changes have been made.

Once your paper is formally accepted, an uncorrected proof of your manuscript will be published online ahead of the final version, unless you’ve already opted out via the online submission form. If, for any reason, you do not want an earlier version of your manuscript published online or are unsure if you have already indicated as such, please let the journal staff know immediately at plosgenetics@plos.org.

In the meantime, please log into Editorial Manager at https://www.editorialmanager.com/pgenetics/, click the "Update My Information" link at the top of the page, and update your user information to ensure an efficient production and billing process. Note that PLOS requires an ORCID iD for all corresponding authors. Therefore, please ensure that you have an ORCID iD and that it is validated in Editorial Manager. To do this, go to ‘Update my Information’ (in the upper left-hand corner of the main menu), and click on the Fetch/Validate link next to the ORCID field.  This will take you to the ORCID site and allow you to create a new iD or authenticate a pre-existing iD in Editorial Manager.

If you have a press-related query, or would like to know about making your underlying data available (as you will be aware, this is required for publication), please see the end of this email. If your institution or institutions have a press office, please notify them about your upcoming article at this point, to enable them to help maximise its impact. Inform journal staff as soon as possible if you are preparing a press release for your article and need a publication date.

Thank you again for supporting open-access publishing; we are looking forward to publishing your work in PLOS Genetics!

Yours sincerely,

Jean Morrison

Guest Editor

PLOS Genetics

Xiaofeng Zhu

Section Editor

PLOS Genetics

www.plosgenetics.org

Twitter: @PLOSGenetics

----------------------------------------------------

Comments from the reviewers (if applicable):

Reviewer's Responses to Questions

Comments to the Authors:

Please note here if the review is uploaded as an attachment.

Reviewer #1: I have no further comments, the manuscript are acceptable.

**********

Have all data underlying the figures and results presented in the manuscript been provided?

Large-scale datasets should be made available via a public repository as described in the PLOS Genetics data availability policy, and numerical data that underlies graphs or summary statistics should be provided in spreadsheet form as supporting information.

Reviewer #1: None

**********

PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.

Reviewer #1: No

----------------------------------------------------

Data Deposition

If you have submitted a Research Article or Front Matter that has associated data that are not suitable for deposition in a subject-specific public repository (such as GenBank or ArrayExpress), one way to make that data available is to deposit it in the Dryad Digital Repository. As you may recall, we ask all authors to agree to make data available; this is one way to achieve that. A full list of recommended repositories can be found on our website.

The following link will take you to the Dryad record for your article, so you won't have to re‐enter its bibliographic information, and can upload your files directly: 

http://datadryad.org/submit?journalID=pgenetics&manu=PGENETICS-D-24-00217R2

More information about depositing data in Dryad is available at http://www.datadryad.org/depositing. If you experience any difficulties in submitting your data, please contact help@datadryad.org for support.

Additionally, please be aware that our data availability policy requires that all numerical data underlying display items are included with the submission, and you will need to provide this before we can formally accept your manuscript, if not already present.

----------------------------------------------------

Press Queries

If you or your institution will be preparing press materials for this manuscript, or if you need to know your paper's publication date for media purposes, please inform the journal staff as soon as possible so that your submission can be scheduled accordingly. Your manuscript will remain under a strict press embargo until the publication date and time. This means an early version of your manuscript will not be published ahead of your final version. PLOS Genetics may also choose to issue a press release for your article. If there's anything the journal should know or you'd like more information, please get in touch via plosgenetics@plos.org.

10.1371/journal.pgen.1011391.r006
Acceptance letter
Zhu Xiaofeng Section Editor
Morrison Jean Guest Editor
© 2024 Zhu, Morrison
2024
Zhu, Morrison
https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
29 Aug 2024

PGENETICS-D-24-00217R2

MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies

Dear Dr Cui,

We are pleased to inform you that your manuscript entitled "MR-SPLIT: a novel method to address selection and weak instrument bias in one-sample Mendelian randomization studies" has been formally accepted for publication in PLOS Genetics! Your manuscript is now with our production department and you will be notified of the publication date in due course.

The corresponding author will soon be receiving a typeset proof for review, to ensure errors have not been introduced during production. Please review the PDF proof of your manuscript carefully, as this is the last chance to correct any errors. Please note that major changes, or those which affect the scientific understanding of the work, will likely cause delays to the publication date of your manuscript.

Soon after your final files are uploaded, unless you have opted out or your manuscript is a front-matter piece, the early version of your manuscript will be published online. The date of the early version will be your article's publication date. The final article will be published to the same URL, and all versions of the paper will be accessible to readers.

Thank you again for supporting PLOS Genetics and open-access publishing. We are looking forward to publishing your work!

With kind regards,

Anita Estes

PLOS Genetics

On behalf of:

The PLOS Genetics Team

Carlyle House, Carlyle Road, Cambridge CB4 3DN | United Kingdom

plosgenetics@plos.org | +44 (0) 1223-442823

plosgenetics.org | Twitter: @PLOSGenetics
==== Refs
References

1 Davey Smith G , Hemani G . Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Human Molecular Genetics. 2014;23 (R1 ):R89–R98. doi: 10.1093/hmg/ddu328 25064373
2 Lawlor DA , Harbord RM , Sterne JA , Timpson N , Davey Smith G . Mendelian randomization: using genes as instruments for making causal inferences in epidemiology. Statistics in Medicine. 2008;27 (8 ):1133–1163. doi: 10.1002/sim.3034 17886233
3 Greenland S . An introduction to instrumental variables for epidemiologists. International Journal of Epidemiology. 2000;29 (4 ):722–729. doi: 10.1093/ije/29.4.722 10922351
4 Davey Smith G , Ebrahim S . ‘Mendelian randomization’: can genetic epidemiology contribute to understanding environmental determinants of disease?*. International Journal of Epidemiology. 2003;32 (1 ):1–22. doi: 10.1093/ije/dym289 12689998
5 Burgess S , Small DS , Thompson SG . A review of instrumental variable estimators for Mendelian randomization. Statistical Methods in Medical Research. 2017;26 (5 ):2333–2355. doi: 10.1177/0962280215597579 26282889
6 Glymour MM , Tchetgen Tchetgen EJ , Robins JM . Credible Mendelian randomization studies: approaches for evaluating the instrumental variable assumptions. American Journal of Epidemiology. 2012;175 (4 ):332–339. doi: 10.1093/aje/kwr323 22247045
7 Staiger D , Stock JH . Instrumental Variables Regression with Weak Instruments. National Bureau of Economic Research; 1994. 151 .
8 Bound J , Jaeger DA , Baker RM . Problems with Instrumental Variables Estimation When the Correlation Between the Instruments and the Endogeneous Explanatory Variable is Weak. Journal of the American Statistical Association. 1995;90 (430 ):443–450. doi: 10.2307/2291055
9 Bowden J , Davey Smith G , Burgess S . Mendelian randomization with invalid instruments: effect estimation and bias detection through Egger regression. International Journal of Epidemiology. 2015;44 (2 ):512–525. doi: 10.1093/ije/dyv080 26050253
10 Kolesár M , Chetty R , Friedman J , Glaeser E , Imbens GW . Identification and inference with many invalid instruments. Journal of Business & Economic Statistics. 2015;33 (4 ):474–484. doi: 10.1080/07350015.2014.978175
11 Hemani G , Bowden J , Davey Smith G . Evaluating the potential role of pleiotropy in Mendelian randomization studies. Human Molecular Genetics. 2018;27 (R2 ):R195–R208. doi: 10.1093/hmg/ddy163 29771313
12 Angrist JD , Krueger AB . Does compulsory school attendance affect schooling and earnings? The Quarterly Journal of Economics. 1991;106 (4 ):979–1014. doi: 10.2307/2937954
13 Burgess S , Smith GD , Davies NM , Dudbridge F , Gill D , Glymour MM , et al . Guidelines for performing Mendelian randomization investigations: update for summer 2023. Wellcome Open Research. 2019;4 . doi: 10.12688/wellcomeopenres.15555.3 32760811
14 Angrist JD , Krueger AB . Split-Sample Instrumental Variables Estimates of the Return to Schooling. Journal of Business & Economic Statistics. 1995;13 (2 ):225–235. doi: 10.1080/07350015.1995.10524597
15 Anderson TW , Rubin H . Estimation of the Parameters of a Single Equation in a Complete System of Stochastic Equations. The Annals of Mathematical Statistics. 1949;20 (1 ):46–63. doi: 10.1214/aoms/1177730090
16 Anderson TW . Origins of the limited information maximum likelihood and two-stage least squares estimators. Journal of Econometrics. 2005;127 (1 ):1–16. doi: 10.1016/S0304-4076(97)00076-6
17 Burgess S , Thompson SG , Collaboration CCG . Avoiding bias from weak instruments in Mendelian randomization studies. International Journal of Epidemiology. 2011;40 (3 ):755–764. doi: 10.1093/ije/dyr036 21414999
18 Angrist JD , Imbens GW , Krueger AB . Jackknife instrumental variables estimation. Journal of Applied Econometrics. 1999;14 (1 ):57–67. doi: 10.1002/(SICI)1099-1255(199901/02)14:1<57::AID-JAE501>3.0.CO;2-G
19 Blomquist S , Dahlberg M . Small sample properties of LIML and jackknife IV estimators: Experiments with weak instruments. Journal of Applied Econometrics. 1999;14 (1 ):69–88. doi: 10.1002/(SICI)1099-1255(199901/02)14:1<69::AID-JAE521>3.0.CO;2-7
20 Jiang T , Gill D , Butterworth AS , Burgess S . An empirical investigation into the impact of winner’s curse on estimates from Mendelian randomization. International Journal of Epidemiology. 2023;52 (4 ):1209–1219. doi: 10.1093/ije/dyac233 36573802
21 Denault WRP , Bohlin J , Page CM , Burgess S , Jugessur A . Cross-fitted instrument: A blueprint for one-sample Mendelian randomization. PLOS Computational Biology. 2022;18 (8 ):1–21. doi: 10.1371/journal.pcbi.1010268 36037248
22 Pierce BL , Ahsan H , VanderWeele TJ . Power and instrument strength requirements for Mendelian randomization studies using multiple genetic variants. International Journal of Epidemiology. 2010;40 (3 ):740–752. doi: 10.1093/ije/dyq151 20813862
23 Panagiotou OA , Ioannidis JPA , for the Genome-Wide Significance Project. What should the genome-wide significance threshold be? Empirical replication of borderline genetic associations. International Journal of Epidemiology. 2011;41 (1 ):273–286. doi: 10.1093/ije/dyr178 22253303
24 Fan J , Lv J . Sure independence screening for ultrahigh dimensional feature space. Journal of the Royal Statistical Society: Series B (Statistical Methodology). 2008;70 (5 ):849–911. doi: 10.1111/j.1467-9868.2008.00674.x
25 Tibshirani R . Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society: Series B (Methodological). 1996;58 (1 ):267–288. doi: 10.1111/j.2517-6161.1996.tb02080.x
26 Zou H . The adaptive lasso and its oracle properties. Journal of the American statistical association. 2006;101 (476 ):1418–1429. doi: 10.1198/016214506000000735
27 Stock JH, Yogo M. Testing for weak instruments in linear IV regression; 2002.
28 Shea J . Instrument relevance in multivariate linear models: A simple measure. Review of Economics and Statistics. 1997;79 (2 ):348–352. doi: 10.1162/rest.1997.79.2.348
29 Wasserman L , Roeder K . High dimensional variable selection. Annals of statistics. 2009;37 (5A ):2178. doi: 10.1214/08-aos646 19784398
30 Dezeure R , Bühlmann P , Meier L , Meinshausen N . High-dimensional inference: confidence intervals, p-values and R-software hdi. Statistical science. 2015; p. 533–558.
31 Liu Y , Xie J . Cauchy Combination Test: A Powerful Test With Analytic p-Value Calculation Under Arbitrary Dependency Structures. Journal of the American Statistical Association. 2019;115 :1–29.34012183
32 Judd E , Calhoun D . Apparent and true resistant hypertension: definition, prevalence and outcomes. Journal of human hypertension. 2014;28 (8 ):463–468. doi: 10.1038/jhh.2013.140 24430707
33 Yu Z , Coresh J , Qi G , Grams M , Boerwinkle E , Snieder H , et al . A bidirectional Mendelian randomization study supports causal effects of kidney function on blood pressure. Kidney international. 2020;98 (3 ):708–716. doi: 10.1016/j.kint.2020.04.044 32454124
34 Chen J , Bundy JD , Hamm LL , Hsu Cy , Lash J , Miller ER III , et al . Inflammation and apparent treatment-resistant hypertension in patients with chronic kidney disease: the results from the CRIC study. Hypertension. 2019;73 (4 ):785–793. doi: 10.1161/HYPERTENSIONAHA.118.12358 30776971
35 Thomas G , Xie D , Chen HY , Anderson AH , Appel LJ , Bodana S , et al . Prevalence and prognostic significance of apparent treatment resistant hypertension in chronic kidney disease: report from the chronic renal insufficiency cohort study. Hypertension. 2016;67 (2 ):387–396. doi: 10.1161/HYPERTENSIONAHA.115.06487 26711738
36 Kaboré J , Metzger M , Helmer C , Berr C , Tzourio C , Drueke TB , et al . Hypertension control, apparent treatment resistance, and outcomes in the elderly population with chronic kidney disease. Kidney international reports. 2017;2 (2 ):180–191. doi: 10.1016/j.ekir.2016.10.006 29142956
37 Kabore J , Metzger M , Helmer C , Berr C , Tzourio C , Massy ZA , et al . Kidney function decline and apparent treatment-resistant hypertension in the elderly. PLoS One. 2016;11 (1 ):e0146056. doi: 10.1371/journal.pone.0146056 26807712
38 Yang J , Lee SH , Goddard ME , Visscher PM . GCTA: a tool for genome-wide complex trait analysis. The American Journal of Human Genetics. 2011;88 (1 ):76–82. doi: 10.1016/j.ajhg.2010.11.011 21167468
39 Greco M FD , Minelli C , Sheehan NA , Thompson JR . Detecting pleiotropy in Mendelian randomisation studies with summary data and a continuous outcome. Statistics in Medicine. 2015;34 (21 ):2926–2940. doi: 10.1002/sim.6522 25950993
40 Sargan JD . The estimation of economic relationships using instrumental variables. Econometrica: Journal of the Econometric Society. 1958; p. 393–415. doi: 10.2307/1907619
41 Yuan Z , Liu L , Guo P , Yan R , Xue F , Zhou X . Likelihood-based Mendelian randomization analysis with automated instrument selection and horizontal pleiotropic modeling. Science Advances. 2022;8 (9 ):eabl5744. doi: 10.1126/sciadv.abl5744 35235357
42 Qi G , Chatterjee N . Mendelian randomization analysis using mixture models for robust and efficient estimation of causal effects. Nature communications. 2019;10 (1 ):1941. doi: 10.1038/s41467-019-09432-2 31028273
43 Burgess S , Foley CN , Allara E , Staley JR , Howson JM . A robust and efficient method for Mendelian randomization with hundreds of genetic variants. Nature communications. 2020;11 (1 ):376. doi: 10.1038/s41467-019-14156-4 31953392
