
==== Front
Crit Care Explor
Crit Care Explor
CC9
Critical Care Explorations
2639-8028
Lippincott Williams & Wilkins Hagerstown, MD

39302988
CCE-D-24-00097
00001
10.1097/CCE.0000000000001152
3
Methodology
Statistical Power and Performance of Strategies to Analyze Composites of Survival and Duration of Ventilation in Clinical Trials
Chen Ziming MSc 1
Harhay Michael O. PhD mharhay@pennmedicine.upenn.edu
2
Fan Eddy MD, PhD eddy.fan@uhn.ca
3
Granholm Anders MD anders.granholm@regionh.dk
4
McAuley Daniel F. MD d.f.mcauley@qub.ac.uk
56
Urner Martin MD martin.urner@uhn.ca
78
Yarnell Christopher J. MD 3910
Goligher Ewan C. MD, PhD ewan.goligher@uhn.ca
281112
https://orcid.org/0000-0002-7263-4251
Heath Anna PhD 11314
1 Child Health Evaluative Sciences, Peter Gilgan Centre for Research and Learning, The Hospital for Sick Children, Toronto, ON, Canada.
2 Department of Biostatistics, Epidemiology and Informatics Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA.
3 Department of Medicine, Division of Respirology, University Health Network, Toronto, ON, Canada.
4 Department of Intensive Care, Copenhagen University Hospital–Rigshospitalet, Copenhagen, Denmark.
5 School of Medicine, Dentistry and Biomedical Sciences, Wellcome-Wolfson Institute for Experimental Medicine, Queen’s University Belfast, Belfast, United Kingdom.
6 Regional Intensive Care Unit, Royal Victoria Hospital, Belfast, United Kingdom.
7 Department of Anesthesiology and Pain Medicine, University of Toronto, Toronto, ON, Canada.
8 Interdepartmental Division of Critical Care Medicine, University of Toronto, Toronto, ON, Canada.
9 Department of Critical Care Medicine, Scarborough Health Network, Toronto, ON, Canada.
10 Institute of Health Policy, Management, and Evaluation, University of Toronto, Toronto, ON, Canada.
11 Department of Physiology, University of Toronto, Toronto, ON, Canada.
12 Toronto General Hospital Research Institute, Toronto, ON, Canada.
13 Division of Biostatistics, Dalla Lana School of Public Health, University of Toronto, Toronto, ON, Canada.
14 Department of Statistical Science, University College London, London, United Kingdom.
For information regarding this article, E-mail: anna.heath@sickkids.ca
20 9 2024
10 2024
6 10 e1152Copyright © 2024 The Authors. Published by Wolters Kluwer Health, Inc. on behalf of the Society of Critical Care Medicine.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution-Non Commercial-No Derivatives License 4.0 (CCBY-NC-ND), where it is permissible to download and share the work provided it is properly cited. The work cannot be changed in any way or used commercially without permission from the journal.

BACKGROUND:

Patients with acute hypoxemic respiratory failure are at high risk of death and prolonged time on the ventilator. Interventions often aim to reduce both mortality and time on the ventilator. Many methods have been proposed for analyzing these endpoints as a single composite outcome (days alive and free of ventilation), but it is unclear which analytical method provides the best performance. Thus, we aimed to determine the analysis method with the highest statistical power for use in clinical trials.

METHODS:

Using statistical simulation, we compared multiple methods for analyzing days alive and free of ventilation: the t, Wilcoxon rank-sum, and Kryger Jensen and Lange tests, as well as the proportional odds, hurdle-Poisson, and competing risk models. We compared 14 scenarios relating to: 1) varying baseline distributions of mortality and duration of ventilation, which were based on data from a registry of patients with acute hypoxemic respiratory failure and 2) the varying effects of treatment on mortality and duration of ventilation.

RESULTS AND CONCLUSIONS:

All methods have good control of type 1 error rates (i.e., avoid false positive findings). When data are simulated using a proportional odds model, the t test and ordinal models have the highest relative power (92% and 90%, respectively), followed by competing risk models. When the data are simulated using survival models, the competing risk models have the highest power (100% and 92%), followed by the t test and a ten-category ordinal model. All models struggled to detect the effect of the intervention when the treatment only affected one of mortality and duration of ventilation. Overall, the best performing analytical strategy depends on the respective effects of treatment on survival and duration of ventilation and the underlying distribution of the outcomes. The evaluated models each provide a different interpretation for the treatment effect, which must be considered alongside the statistical power when selecting analysis models.

clinical trial outcomes
critical care
hospital-free days
statistical models
statistical power
ventilator-free days
Patient-Centered Outcomes Research Institute 10.13039/100006093 ME-2020C1-19220 Michael O HarhayNational Heart, Lung, and Blood Institute 10.13039/100000050 R00-HL141678 Michael O HarhayNational Heart, Lung, and Blood Institute 10.13039/100000050 R01-HL168202 Michael O HarhayInterdepartmental Division of Critical Care MedicineScholarship Martin UrnerCanada Research Chairs 10.13039/501100001804 Statistical Trial Design Anna HeathNatural Sciences and Engineering Research Council of Canada 10.13039/501100000038 RGPIN-2021-03366 Anna HeathNational Sanitarium Association 10.13039/100013329 Early Career Investigator Award Ewan GoligherOPEN-ACCESSTRUE
SDCT
==== Body
pmcSelecting the appropriate primary outcome for clinical trials investigating interventions for acute hypoxemic respiratory failure (AHRF) is challenging (1, 2). As mortality rates are high (3), interventions often aim to reduce mortality, but the sample sizes required to detect small but clinically relevant differences in mortality can be infeasible (4, 5). Additionally, reducing the time spent on the ventilator or in hospital is important for patients (1, 2) and may improve long-term recovery (6, 7). Thus, investigators are often interested in effects on both mortality and length of time-on-ventilator or in hospital (1, 2, 8). These outcomes can also be adjusted, for example, to account for days alive and out of hospital, in less acute settings (9).

One common strategy for combining these endpoints is to compute ventilator-free days (VFDs) (10–14). While the exact definition of VFDs varies among studies (1), VFDs for survivors are generally calculated as the difference between the total follow-up time (e.g., 28 d) and the number of days spent on the ventilator. Patients remaining on the ventilator at the end of follow-up are given a VFD outcome of 0, and deaths are assigned the worst possible outcome, either 0 or –1, depending on the definition (1, 13, 15).

VFDs are sometimes criticized as difficult to interpret (16), and there is no universally recognized method for their analysis. One major challenge is that an intervention reducing mortality might result in longer durations of mechanical ventilation, as patients survive longer on ventilators instead of dying. This makes it difficult to determine the efficacy of interventions on mortality and ventilation time when combined as a composite endpoint. Previous studies have evaluated different methods for analyzing effects on mortality and duration of ventilation (9) including standard statistical tests (1, 17), hurdle models (15), proportional odds models (18, 19), and bespoke statistical tests (20). Mortality and duration of ventilation can also be analyzed using competing risk models and joint models (17).

Investigators must consider various criteria when choosing the optimal analysis method for clinical trials (21), with high statistical power to detect treatment effects being crucial. Hence, we conducted a simulation study to inform the selection of the most powerful strategy for designing trials assessing treatment effects on mortality and ventilation duration.

METHODS

This statistical simulation study was undertaken using the Aims, Data-Generating Mechanisms, Estimands, Methods, and Performance Measures framework (22).

Aims

This simulation aimed to compare methods for analyzing outcomes combining data on mortality and ventilation duration to detect differences between two treatment groups in a randomized clinical trial.

Data-Generating Mechanisms

We used two methods to generate data from a simulated AHRF trial (23) randomizing 1500 patients (750 per arm) in a 1:1 ratio to evaluate the effect of an intervention on mortality and duration of ventilation in comparison to standard care (control). We assumed that all individuals were enrolled while receiving ventilatory support and that we were interested in the final liberation from the ventilator, that is, if patients were reintubated, the days they were not receiving ventilatory support were discarded. The first data generation method generated data using an observed distribution of mortality and duration of mechanical ventilation from a registry-derived dataset of patients admitted to nine ICUs from across the Greater Toronto Area (n = 326). These data provided the proportion of patients who were liberated on each day between days 1 and 28 after intubation and the proportion of patients who died by day 28. The distribution of VFDs in the treated population in the simulated trial was generated by applying a proportional odds model to this control group data (24).

Different data generation approaches influence simulation study outcomes (25). Thus, we also used Weibull distributions to simulate patient times of death and ventilator liberation, matching the proportions of deaths and ventilator-free to the registry data. It is common to use a longer follow-up time for mortality (26), set at 90 days, with patients assumed to survive if their simulated death time exceeded this period. We also generated the data such that individuals who were expected to spend longer on the ventilator were also more likely to die (Supplementary Material, http://links.lww.com/CCX/B398). Treatment effects on time-to-death and time-to-ventilator liberation were simulated by adjusting proportions of patients who died by day 90 or remaining on the ventilator at day 28.

For both data-generating processes, we prespecified seven different scenarios for how the treatment could impact the two outcomes:

1) Null: The treatment has no effect on either mortality or duration of ventilation.

2) Overall benefit: The treatment reduces the number of deaths and decreases time-on-ventilator for survivors.

3) Overall harm: The treatment increases the number of deaths and prolongs the time-on-ventilator.

4) Mortality benefit with increased duration of ventilation: The treatment reduces the number of deaths but prolongs the time-on-ventilator.

5) Duration of ventilation benefit with increased mortality: The treatment increases the number of deaths but reduces the number of days on a ventilator.

6) Mortality benefit alone: The treatment reduces the number of deaths but has no effect on the duration of ventilation.

7) Duration of ventilation benefit alone: The treatment decreases the duration of ventilation but does not impact the number of deaths.

For the first data generation approach, the proportional odds ratios used to specify benefit and harm were 0.75 and 1.25, respectively, to represent clinically meaningful differences in outcomes. For the second data generation approach, the treatment effect for mortality was generated by adjusting the 90-day survival to 0.32 for benefit and 0.42 for harm from 0.37. Similarly, the treatment effect for liberation adjusted the proportion of individuals liberated from the ventilator at day 28 to 0.93 for benefit and 0.86 for harm from 0.89. When the treatment effect differed across the two outcomes, these values were used to represent benefit and harm but only for the outcome affected. Table 1 displays the assumed proportion of deaths and duration of ventilation for each of the scenarios and two data-generating approaches.

TABLE 1. Simulated Scenarios of Treatment Effect on Mortality and Duration of Ventilation Analyzed in This Study

Scenarios	Data Generation Approach 1	Data Generation Approach 2	
Proportion of Patients Who Die	Duration of Ventilation for Survivors, Mean (sd)	Proportion of Patients Who Die	Duration of Ventilation for Survivors, Mean (sd)	
Null	0.37	10 (8.6)	0.37	10 (8.0)	
Overall benefit	0.30	9 (8.3)	0.32	9 (7.3)	
Overall harm	0.42	11 (8.8)	0.42	11 (8.3)	
Mortality benefit with extended duration of ventilation	0.30	11 (8.8)	0.32	11 (8.5)	
Duration of ventilation benefit with increased mortality	0.42	9 (8.3)	0.42	9 (7.2)	
Mortality benefit alone	0.30	10 (8.6)	0.32	11 (8.1)	
Duration of ventilation benefit alone	0.37	9 (8.3)	0.37	9 (7.3)	
For data generation approach 1, benefit is generated as a proportional odds ratio of 0.75 and harm is generated as a proportional odds ratio of 1.25. For data generation approach 2, benefit and harm for mortality were generated as a 5% absolute risk difference in 90-day survival and benefit and harm for duration of ventilation were generated as a 4% and 3% absolute risk difference, respectively, in 28-day ventilation.

Target of Analysis: Estimands

The target of this analysis is the probability that a positive (beneficial) treatment effect is declared. For the null scenario, the probability of a positive treatment effect is the type 1 error. For the methods using Bayesian analyses, we used the simulations from the null scenario to determine the threshold(s) used to declare a positive treatment effect to ensure that the type 1 error rate was 5%.

Analysis Methods

Endpoints combining mortality and duration of ventilation have been treated as continuous, count, ordinal, or time-to-event outcomes, with selection of the analytic method matching the outcome type. Note that each proposed analysis method provides an alternative estimand, which must be considered when selecting an appropriate analysis method (27).

Ventilator-Free Days As a Continuous/Count Outcome

VFDs can be defined as the number of days alive and free of mechanical ventilation to day 28 among survivors (ranging from 0 to 28), with nonsurvivors assigned the worst possible outcome of 0 (18). When considered as a continuous or count outcome, VFDs can be analyzed with the following analysis methods:

A one-tailed two-sample t test, which rejects the null hypothesis of no “increase” in VFDs with treatment if the p value is less than 0.05. This method aligns with a hypothesis of no difference between the mean VFDs but risks bias, and lack of flexibility, especially in our case where the data are not approximately normally distributed.

A Wilcoxon rank-sum test, which rejects the null hypothesis of no “difference” in VFDs between treatment and control if the p value is less than 0.05. In contrast to the t test, the Wilcoxon rank-sum cannot differentiate treatment harm from treatment benefit but does not require approximate normality.

A semi-parametric Kryger Jensen and Lange test (KJL) that separates the effect of treatment on death and number of days free from a ventilator but computes a single p value for hypothesis of no treatment effect (28). We reject the null hypothesis of no difference if the p value is less than 0.05. This method offers robustness against non-normality and unequal variances in small sample sizes but is more complex to implement and interpret and currently lacks recognition.

A logit-Poisson hurdle model that analyses mortality and the number of days free from a ventilator (as a count outcome) separately (15) but provides two treatment effects; for the effect of treatment on mortality and the mean number of days survivors spend off the ventilator. We were interested in whether treatment affects morality or the duration of time-on-ventilator, which required a novel method to declare a “positive” effect of treatment. We specified that the treatment was beneficial using a Bayesian hurdle model if it demonstrated a positive effect for either outcome and did not demonstrate a negative effect for both outcomes. Thresholds to declare positive and negative effects were chosen to control type 1 error (for positive effects) and to prevent harm (for negative effects). Specifically, a positive treatment effect was declared if: 1) the posterior probability of a reduction in mortality or a reduction in time-on-ventilator was greater than 99.8% and 2) the posterior probability of an increase in mortality and an increase in time-on-ventilator were both less than 20%. This method handles zero-inflation and overdispersion well but risks overfitting and is complex to implement compared with the other tests.

Ventilator-Free Days As an Ordinal Outcome

VFDs can be analyzed as ordinal endpoints, where deaths are usually assigned a value of –11 to distinguish them from patients remaining on the ventilator at day 28. Proportional odds models are commonly used for ordinal outcomes (24), measuring treatment effects as odds ratios of transitioning between outcome categories. We fitted ordinal models with three and ten categories to assess whether the number of categories affects performance. In the three-category outcome, patients who died were assigned –1, those requiring over a week for recovery were 0, and those ventilated for less than a week were 1. For the ten-category outcome, patients who died were assigned –1, patients who remained on the ventilator at day 28 were 0 and VFD values were recoded as in Table S1 (http://links.lww.com/CCX/B398) based on quantiles of the distribution, ensuring balanced category distribution. Due to computational constraints, models with more than ten categories were not considered in this simulation (29).

These ordinal models were fit using a Bayesian framework, with positive treatment effects declared when the posterior probability of an odds ratio greater than 1 exceeded 95.5% for the three-category outcome for both data-generating processes, and 96.5% and 98% for the ten-category outcome in data-generating processes one and two, respectively, to maintain a type 1 error rate below 5%.

Time-to-Liberation and Time-to-Death

Finally, mortality and ventilation duration can be jointly analyzed via “competing risk” analysis, fitting separate survival models for death and ventilator liberation (30). This method provides separate treatment effects for both outcomes, moving away from the composite VFDs measure2. We applied two Bayesian Cox proportional hazards models to compute these treatment effects as hazard ratios. Treatment efficacy was determined if the posterior probability of increasing time-to-death or decreasing time-on-ventilator exceeded 96% and the posterior probability of decreasing time-to-death or increasing time-on-ventilator was below 20% to control type 1 error below 5%.

In many settings, researchers may be interested in reducing overall mortality, rather than reducing time-to-death. Thus, we also considered an alternative to this competing risk analysis that fits a logistic model for mortality and a Bayesian Cox proportional hazards model for time-to-liberation, similar to the KJL and hurdle model approaches. This models the effect of treatment on the “probability” of dying within the study period, rather than the time-to-death. The treatment was then deemed effective if the posterior probability of reducing the probability of death or the time spent on the ventilator exceeded 97.5% and the posterior probability of increasing the mortality and increasing the time-on-ventilator were both below 20%. Again, these thresholds were chosen to maintain type 1 error below 5%.

A challenge for these competing risk methods is that the first data-generating approach only generated whether a patient has died (as a –1) and not the “time” they died. Thus, a traditional competing risk analysis could not be performed, as the time of death must be incorporated into the analysis. We considered two methods to adjust the second analysis method if the time of death was not observed in the available data. First, we only evaluated the time-to-liberation for survivors. Second, we generated pseudo-death times from a uniform distribution between 0 and 90 for patients who died, which is equivalent to assuming that we know the patient died but not when. For these two alternative methods, we adjusted the efficacy threshold from 97.5% to 96% and 96.5%, respectively, to maintain a 5% type 1 error.

Performance Measures

The probability that each method declares a positive treatment effect was estimated by the proportion of simulated trials that declare a positive treatment effect. We used 2000 simulations for each scenario (22).

RESULTS

Data Generation Approach 1 (Ordinal Model Applied to Registry Data)

Figure 1; and Table S2 (http://links.lww.com/CCX/B398) summarized the probability of detecting a positive treatment effect under each analytic approach for the first data generation process. We observed a slightly reduced type 1 error for the hurdle-Poisson model under these data generating process. When the treatment improves both mortality and time-to-liberation, the t test and the ordinal models provided the highest power (92% and 90%, respectively) with the two competing risk models also providing high power at 89% and 88%, respectively. The hurdle-Poisson model has the lowest power at 66%, under this treatment effect. When the treatment worsens both outcomes, the Wilcoxon rank-sum test was likely to conclude statistical significance (67% probability) and the three-category ordinal outcome incorrectly declared a positive treatment effect in 9% of the simulations.

Figure 1. Probability of each analysis method declaring that the intervention results in a positive effect of treatment, which is defined differently for each analysis method, for each of the seven treatment effect scenarios using the first data generating process based on simulating from a multinomial distribution and proportional odds model. Note that for the Wilcoxon rank-sum test, we plot the probability of declaring statistical significance, which is not equivalent to a positive treatment effect. KJL = Kryger Jensen and Lange test.

All analysis methods were more likely to detect differences when mortality is improved vs. when duration of ventilation was shortened when treatment effects are opposite for the two outcomes, especially for the competing risk model with uniform censoring and the KJL test. When the treatment improved mortality, most methods had a relatively high probability to detect this, except for the hurdle-Poisson model. The hurdle-Poisson model was most able to detect true differences in the time-on-ventilator, with a power of 42%, followed by the KJL test.

Data Generation Approach 2 (Weibull Distributions)

Figure 2 and Table S3 (http://links.lww.com/CCX/B398) summarized the probabilities of obtaining a positive treatment effect using each analytic approach for the second data-generating process. When the treatment benefited both outcomes, the competing risk model exhibited the highest power (100%), followed by the binary competing risk model for death (92%), t test (90%), and the ten-category ordinal outcome (88%). Apart from the Wilcoxon rank-sum test, which indicates there is a difference between the two interventions, none of the methods declared a positive treatment effect when the treatment was harmful for both outcomes.

Figure 2. Probability of each analysis method declaring that the intervention results in a positive effect of treatment, which is defined differently for each analysis method, for each of the seven treatment effect scenarios using the second data generating process based on simulating from two correlated Weibull distributions. Note that for the Wilcoxon rank-sum test, we plot the probability of declaring statistical significance, which is not equivalent to a positive treatment effect. KJL = Kryger Jensen and Lange test.

Most methods struggled to detect a positive treatment effect when the treatment had opposite impacts on mortality and time-to-liberation. The KLJ test was mostly likely to declare benefit, declaring a positive effect in approximately 30% and 38% of simulations for mortality and time-to-liberation improvements, respectively.

Competing risk models were effective in detecting positive effects when the treatment was null on one outcome but improved the other, especially when reducing time-on-ventilator. Ordinal models were sensitive to detecting mortality benefits, while hurdle models were sensitive to identifying differences in ventilation duration. The t test and Wilcoxon rank-sum tests generally had lower detection rates in these scenarios.

Across both data generation approaches, the ten-category ordinal outcome is preferred over the three-category outcome, resulting in higher power when there was a true benefit and typically lower errors when the treatment was harmful. Overall, across the two data generation approaches, the t test performs adequately. As expected, the proportional odds model and survival models perform better under their respective data generation frameworks but also perform adequately when their assumptions are not met.

DISCUSSION

We evaluated methods to analyze the composite VFDs endpoint that combines information on mortality and duration of ventilation. Across two data generation approaches, t tests, ordinal models, and competing risk models had the highest statistical power. Crucially, the power of the proportional odds model was well maintained when the data did not respect the proportional odds assumption. The competing risk analyses, which decompose the VFDs endpoint, had the highest power when information on the time-to-death for patients who die was available, suggesting these data should be collected. All methods performed poorly when the treatment only improves one outcome. Thus, VFDs should only be used if we believe that the treatment would improve both outcomes. These findings emphasize that the choice of primary endpoint in clinical trials in AHRF should be guided by the investigators’ prior beliefs about the likely effects of the intervention on outcome.

Another key consideration for selecting an analysis method is the estimand of interest (27), which precisely defines the treatment effect. For many, the treatment effect estimated for a t test, which values death and lengthy stays on the ventilator the same, is not clinically appropriate (1). Although we have included survival models to explore the analyses methods without hard coding deaths to mitigate this issue. For the proportional odds model, the treatment effect measures the odds of moving from a lower category to a higher category across the entire ordinal scale, assuming that death is the worst possible outcome. Finally, competing risk analyses separate the effect of treatment into an effect on mortality and on the length of stay on the ventilator. Our two proposed methods differ in their definition of the treatment effect on mortality, either an effect on the time-to-death or on the chance of death, meaning that the most appropriate definition should be selected. Researchers should also consider whether covariate adjustment is required, which can increase power (31) but can also change the estimand and the analysis method, for example, the t test must be replaced by a regression model to undertake a covariate adjusted analysis.

An important consideration when interpreting the results of this study is the potential disparity between the results derived from prespecified primary analyses and those from pre-planned sensitivity analyses, which explore alternative statistical methods. Prespecified analyses, while planned a priori and thus protected from various biases, might not always use the most powerful or appropriate statistical methods if new evidence or understanding emerges after the trial’s initiation. However, they usually provide the primary evidence for the effectiveness of an intervention in clinical trials. This article aimed to support researchers to select the most appropriate analysis method to include in a prespecified analysis plan. In contrast, post hoc analyses could optimize the analyses based on the observed distribution in the collected data. However, this risks introducing bias and overfitting. Post hoc analysis can be viewed as supplementary, offering insights that can refine understanding and guide future research but not as definitive evidence on their own. Therefore, when choosing a method to analyze VFDs at the design stage, it is important to consider the nature and the objectives of the analyses.

A key limitation of this study is that we used simulated data to explore these methods. While simulated data allowed us to calculate type 1 error and power, it means that the data may not completely represent realistic data. We aimed to mitigate this concern by generating data using two approaches and mimicking outcomes from a ventilated critical care population. Separately, we could not analyze the full 30-level ordinal outcome due to the computational constraints of running 2000 simulations (29). Finally, we did not consider methods that drew trial conclusions from one of the outcomes alone, which has been suggested to overcome the drawbacks with VFDs (30). This is because our focus was the primary analysis of a clinical trial; however, these separate analyses should always be presented as secondary analyses.

Future work should extend to alternative data-generating approaches to ensure the conclusions remain robust. For example, a multistate transition model can be considered when the data generating mechanism is inherently longitudinal and a method for specifying when the novel intervention is preferred is clearly defined (32). Research should also improve the understanding and communication of these analytic approaches to improve their use. A key challenge is specifying minimally important clinical differences, which are required to guide clinical trial design. Finally, a more complete exploration of the underlying assumptions and the estimands from each of these proposed models would further support the selection of the optimal analysis method for VFDs outcomes in critical care trials.

ACKNOWLEDGMENTS

We acknowledge Elizabeth Colantuoni for her insightful comments on a previous version of this article and Yanara Marks for her support in article preparation.

Supplementary Material

1 It is also possible, but less common, to assign individuals who die a value of –1 and treat the outcome as continuous, analyzing it with either the t test or the Wilcoxon rank-sum test. We ran this scenario as a sensitivity analysis and saw minimal differences to the presented results.

2 An extension of the competing risk models is a multistate model that can provide additional treatment effects, if they are not of interest and requires detailed longitudinal data.

Drs. Goligher and Heath contributed equally.

Drs. Goligher and Heath conceived the research idea. Ms. Chen, Dr. Goligher, and Dr. Heath drafted the research plan. Ms. Chen performed the analyses and prepared the tables. Ms. Chen, Dr. Goligher, and Dr. Heath drafted the article. Drs. Fan, Harhay, Granholm, Urner, and Yarnell contributed to the research plan and reviewed the article. All authors approved the final version for publication.

Dr. Harhay is funded by the Patient-Centered Outcomes Research Institute (Award ME-2020C1-19220) and the U.S. National Institutes of Health/National Heart, Lung, and Blood Institute (grant numbers R00-HL141678 and R01-HL168202). Dr. Urner is supported by a Scholarship of the Interdepartmental Division of Critical Care Medicine, University of Toronto. Dr. Heath is funded by a Canada Research Chair in Statistical Trial Design; Natural Sciences and Engineering Research Council of Canada (award No. RGPIN-2021 03366). Dr. Goligher is funded by an Early Career Investigator Award from the National Sanitarium Association. The remaining authors have disclosed that they do not have any potential conflicts of interest.

All statements in this report, including its findings and conclusions, are solely those of the authors and do not necessarily represent the views of the National Institutes of Health or Patient-Centered Outcomes Research Institute or its Board of Governors or Methodology Committee.

All code to undertake these simulations is available on GitHub in repository impactlab23/analyzing ventilator-free days.

Supplemental digital content is available for this article. Direct URL citations appear in the printed text and are provided in the HTML and PDF versions of this article on the journal’s website (http://journals.lww.com/ccejournal).
==== Refs
REFERENCES

1. Yehya N Harhay MO Curley MAQ : Reappraisal of ventilator-free days in critical care research. Am J Respir Crit Care Med 2019; 200 :828–836 31034248
2. Auriemma CL Taylor SP Harhay MO : Hospital-free days: A pragmatic and patient-centered outcome for trials among critically and seriously ill patients. Am J Respir Crit Care Med 2021; 204 :902–909 34319848
3. Unal AU Kostek O Takir M : Prognosis of patients in a medical intensive care unit. North Clin Istanb 2015; 2 :189–195 28058366
4. Harhay MO Wagner J Ratcliffe SJ : Outcomes and statistical power in adult critical care randomized trials. Am J Respir Crit Care Med 2014; 189 :1469–1478 24786714
5. Schoenfeld DA Bernard GR ; ARDS Network: Statistical evaluation of ventilator-free days as an efficacy measure in clinical trials of treatments for acute respiratory distress syndrome. Crit Care Med 2002; 30 :1772–1777 12163791
6. Herridge MS Chu LM Matte A ; RECOVER Program Investigators (Phase 1: towards RECOVER): The RECOVER program: Disability risk groups and 1-year outcome after 7 or more days of mechanical ventilation. Am J Respir Crit Care Med 2016; 194 :831–844 26974173
7. Combes A Costa MA Trouillet JL : Morbidity, mortality, and quality-of-life outcomes of patients requiring ≥14 days of mechanical ventilation. Crit Care Med 2003; 31 :1373–1381 12771605
8. Haviland K Tan KS Schwenk N : Outcomes after long-term mechanical ventilation of cancer patients. BMC Palliat Care 2020; 19 :42 32228554
9. Granholm A Kaas-Hansen BS Lange T : Use of days alive without life support and similar count outcomes in randomised clinical trials—an overview and comparison of methodological choices and analysis methods. BMC Med Res Methodol 2023; 23 :139 37316785
10. Matthay MA Brower RG Carson S ; National Heart, Lung, and Blood Institute Acute Respiratory Distress Syndrome (ARDS) Clinical Trials Network: Randomized, placebo-controlled clinical trial of an aerosolized β2-agonist for treatment of acute lung injury. Am J Respir Crit Care Med 2011; 184 :561–568 21562125
11. Rice TW Wheeler AP Thompson BT ; National Heart, Lung, and Blood Institute Acute Respiratory Distress Syndrome (ARDS) Clinical Trials Network: Initial trophic vs full enteral feeding in patients with acute lung injury: The EDEN randomized trial. JAMA 2012; 307 :795–803 22307571
12. Rice TW Wheeler AP Thompson BT ; NIH NHLBI Acute Respiratory Distress Syndrome Network of Investigators: Enteral omega-3 fatty acid, γ-linolenic acid, and antioxidant supplementation in acute lung injury. JAMA 2011; 306 :1574–1581 21976613
13. Angus DC Berry S Lewis RJ : The REMAP-CAP (randomized embedded multifactorial adaptive platform for community-acquired pneumonia) study rationale and design. Ann Am Thorac Soc 2020; 17 :879–891 32267771
14. Harhay MO Ratcliffe SJ Small DS : Measuring and analyzing length of stay in critical care trials. Med Care 2019; 57 :e53–e59 30664613
15. Verghis RM McDowell C Blackwood B : Re-analysis of ventilator-free days (VFD) in acute respiratory distress syndrome (ARDS) studies. Trials 2023; 24 :183 36907882
16. Varadhan R Weiss CO Segal JB : Evaluating health outcomes in the presence of competing risks: A review of statistical methods and clinical applications. Med Care 2010; 48 (6 Suppl ):S96–S105 20473207
17. Yehya N : Ventilator-free days should be presented as subdistribution hazard ratios using competing risk regression. In: C49. Critical Care: Every Breath You Take-Acute Respiratory Failure and Mechanical Ventilation. San Diego, American Thoracic Society, 2018, p A7768
18. Lawler PR Goligher EC Berger JS ; ATTACC Investigators; ACTIV-4a Investigators; REMAP-CAP Investigators: Therapeutic anticoagulation with heparin in noncritically ill patients with Covid-19. N Engl J Med 2021; 385 :790–802 34351721
19. Goligher EC Bradbury CA McVerry BJ ; REMAP-CAP Investigators; ACTIV-4a Investigators; ATTACC Investigators: Therapeutic anticoagulation with heparin in critically ill patients with Covid-19. N Engl J Med 2021; 385 :777–789 34351722
20. Andersen-Ranberg NC Poulsen LM Perner A ; AID-ICU Trial Group: Haloperidol for the treatment of delirium in ICU patients. N Engl J Med 2022; 387 :2425–2435 36286254
21. Popchak AJ Lynch AD Irrgang JJ : Framework for selecting clinical outcomes for clinical trials. In: Basic Methods Handbook for Clinical Orthopaedic Research. Musahl V Karlsson J Hirschmann MT Ayeni OR. Marx RG Koh JL Nakamura N (Eds). Berlin, Heidelberg, Springer Berlin Heidelberg, 2019, pp 133–142
22. Morris TP White IR Crowther MJ : Using simulation studies to evaluate statistical methods. Stat Med 2019; 38 :2074–2102 30652356
23. Saha R Assouline B Mason G : The impact of sample size misestimations on the interpretation of ARDS trials. Chest 2022; 162 :1048–1062 35643115
24. Peterson B Harrell FE : Partial proportional odds models for ordinal response variables. Appl Stat 1990; 39 :205
25. Pawel S Kook L Reeve K : Pitfalls and potentials in simulation studies: Questionable research practices in comparative simulation studies allow for spurious claims of superiority of any method. Biom J 2023; 66 :e2200091 36890629
26. Florescu S Stanciu D Zaharia M : Effect of antiplatelet therapy on survival and organ support–free days in critically ill patients with COVID-19. JAMA 2022; 327 :1247 35315874
27. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH): Addendum on Estimands and Sensitivity Analysis in Clinical Trials to the Guideline on Statistical Principles for Clinical Trials E9(R1). 2019. Available at: https://database.ich.org/sites/default/files/E9-R1_Step4_Guideline_2019_1203.pdf. Accessed January 16, 2022
28. Jensen AK Lange T : A novel high-power test for continuous outcomes truncated by death. arXiv:1910.12267v1
29. Rue H Riebler A Sørbye SH : Bayesian computing with INLA: A review. Annu Rev Stat Appl 2017; 4 :395–421
30. Bodet-Contentin L Frasca D Tavernier E : Ventilator-free day outcomes can be misleading. Crit Care Med 2018; 46 :425–429 29227369
31. Lee PH : Covariate adjustments in randomized controlled trials increased study power and reduced biasedness of effect size estimation. J Clin Epidemiol 2016; 76 :137–146 26921693
32. Cheung LC Albert PS Das S : Multistate models for the natural history of cancer progression. Br J Cancer 2022; 127 :1279–1288 35821296
