
==== Front
9111633
32751
J Stroke Cerebrovasc Dis
J Stroke Cerebrovasc Dis
Journal of stroke and cerebrovascular diseases : the official journal of National Stroke Association
1052-3057
1532-8511

38703875
10.1016/j.jstrokecerebrovasdis.2024.107750
nihpa2016990
Article
Multicenter comparison using two AI stroke CT perfusion software packages for determining thrombectomy eligibility
Alwood Benjamin T. MD ab*
Meyer Dawn M. FNP-C, PhD b
Ionita Chip PhD c
Snyder Kenneth V. MD, PhD c
Santos Roberta MD a
Perrotta Lindsey DNP a
Crooks Ryan MD a
Van Orden Kimberlee MD b
Torres Dolores MD b
Poynor Briana MD b
Pham Nhan BS d
Kelly Sophie d
Meyer Brett C. MD b
Bolar Divya S. MD, PhD de
a Department of Vascular Neurology, University of Florida, Jacksonville, FL, United States
b University of California San Diego Stroke Center, University of California San Diego, San Diego, CA, United States
c Department of Biomedical Engineering and Neurosurgery, University at Buffalo, Buffalo NY, United States
d Department of Radiology, University of California San Diego, San Diego, CA, United States
e Center for Functional MRI, University of California San Diego, San Diego, CA, United States
* Corresponding author. benjamin.alwood@jax.ufl.edu (B.T. Alwood).
28 8 2024
7 2024
02 5 2024
01 9 2024
33 7 107750107750
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).
Background:

Stroke AI platforms assess infarcted core and potentially salvageable tissue (penumbra) to identify patients suitable for mechanical thrombectomy. Few studies have compared outputs of these platforms, and none have been multicenter or considered NIHSS or scanner/protocol differences. Our objective was to compare volume estimates and thrombectomy eligibility from two widely used CT perfusion (CTP) packages, Viz.ai and RAPID.AI, in a large multicenter cohort.

Methods:

We analyzed CTP data of acute stroke patients with large vessel occlusion (LVO) from four institutions. Core and penumbra volumes were estimated by each software and DEFUSE-3 thrombectomy eligibility assessed. Results between software packages were compared and categorized by NIHSS score, scanner manufacturer/model, and institution.

Results:

Primary analysis of 362 cases found statistically significant differences in both software’s volume estimations, with subgroup analysis showing these differences were driven by results from a single scanner model, the Canon Aquilion One. Viz.ai provided larger estimates with mean differences of 8cc and 18cc for core and penumbra, respectively (p<0.001). NIHSS subgroup analysis also showed systematically larger Viz.ai volumes (p<0.001). Despite volume differences, a significant difference in thrombectomy eligibility was not found. Additional subgroup analysis showed significant differences in penumbra volume for the Phillips Ingenuity scanner, and thrombectomy eligibility for the Canon Aquilion One scanner at one center (7 % increased eligibility with Viz.ai, p=0.03).

Conclusions:

Despite systematic differences in core and penumbra volume estimates between Viz.ai and RAPID. AI, DEFUSE-3 eligibility was not statistically different in primary or NIHSS subgroup analysis. A DEFUSE-3 eligibility difference, however, was seen on one scanner at one institution, suggesting scanner model and local CTP protocols can influence performance and cause discrepancies in thrombectomy eligibility. We thus recommend centers discuss optimal scanning protocols with software vendors and scanner manufacturers to maximize CTP accuracy.

CT Perfusion
Ischemic stroke
Acute stroke
Interventional neuroradiology
DEFUSE-3, Mechanical thrombectomy
==== Body
pmcIntroduction

Since the publication of the extended window thrombectomy treatment trials DAWN and DEFUSE-3, CT perfusion (CTP) imaging has become a standardized, widely accepted method for determining eligibility for endovascular thrombectomy for large vessel occlusion (LVO) in the anterior circulation1,2. The DAWN trial found benefit of thrombectomy up to 24 hours from last known well using perfusion imaging to estimate core volume for the purposes of determining eligibility. DEFUSE-3 showed a similar benefit up to 16 hours from last known well using perfusion imaging to estimate both core and penumbra volumes for the purposes of determining eligibility. While MRI is considered the gold standard for evaluating ischemia, CTP is more widely available, significantly faster, and more practical. This is critical when shorter times from symptom onset to revascularization directly correlate with better functional outcomes3.

The 2019 update to American Heart Association/American Stroke Association stroke guidelines recommend that CTP should be used in the extended treatment window (6–24 hours after symptom onset or last known well) for evaluation of thrombectomy eligibility (Class 1 Level A)4. DEFUSE-3 requires an objective measurement of CBF < 30 % (reflecting infarcted core tissue) and time to maximum of the residue function (Tmax) of > 6s (reflecting penumbra tissue) to determine thrombectomy eligibility. CTP data is post-processed using sophisticated algorithms to estimate these volumes. The DEFUSE-3 criteria require a patient to have < 70cc of infarcted core, ≥ 15cc penumbra, and a mismatch ratio of penumbra to core of ≥ 1.8 to be considered eligible for thrombectomy1.

Volume estimates of core and penumbra are obtained by analyzing CTP imaging data using CTP post-processing software. Two industry-leading CTP processing programs widely utilized clinically are RAPID. AI (IschemaView, Inc., Menlo Park, CA) and Viz.ai (Viz.ai, Inc., San Francisco, CA). RAPID.AI was utilized in the landmark extended window trials1,2. Viz.ai was the first to develop a fully-functioned mobile application for image viewing, LVO detection, and efficient multidisciplinary communication. Both have since developed similar features and mobile applications, and are effectively used in clinical practice for stroke management. Differences in the proprietary algorithms between the software, however, can result in divergent volume estimates when processing identical CTP data, which could potentially affect DEFUSE-3 thrombectomy eligibility and alter clinical management.

CTP involves continuously imaging the brain to capture movement of an iodinated contrast bolus from the arterial circulation, through the brain parenchyma, and into the venous circulation5. RAPID.AI and Viz. ai analyze bolus delivery dynamics to estimate CBV, CBF, MTT, and Tmax. Perfusion results can be affected by a variety of factors, such as cardiac output, arrhythmia, previous stroke, patient motion, and scanner protocol6. Scanner protocol parameters can vary substantially between scanner and institution; these include image acquisition time, voltage, radiation dose, contrast bolus size, number of detectors in array, and image slice thickness7. These variables can affect image quality, for example, by changing spatial resolution, temporal resolution, and signal-to-noise ratio (SNR), and thereby alter CTP performance.

Recently, there has been a call for increased transparency and post-market monitoring of clinical artificial intelligence (AI) models8. Only a few retrospective studies have compared outputs between RAPID.AI and other CTP processing software packages9,10,11,12,13. Some studies have shown similar core and penumbra volumes between packages, while others have reported significant differences in these parameters, with mixed results on whether these differences affect thrombectomy eligibility9,10,11,12. To our knowledge, only one study has compared DEFUSE-3 thrombectomy eligibility specifically between Viz.ai to RAPID.AI and showed a 10.6 % difference. While an important finding, this was based on a small sample size of only 46 patients from a single institution, and information about scanner model/manufacturer was not provided13. A recent study of 242 patients showed substantial agreement for core and excellent agreement for penumbra between Viz. ai and RAPID.AI, though did not compare DEFUSE-3 eligibility14.

Moreoever, no studies have evaluated whether the discrepancies mentioned above could be affected by NIHSS score or variability in scanner model and manufacturer. Patients with lower NIHSS scores typically have smaller core and penumbra volumes, which means that minor discrepancies in volume estimates from RAPID.AI and Viz.ai will disproportionately affect DEFUSE-3 criteria for these patients, relative to those with higher scores. Consequently, this could lead to more frequent eligibility differences between software packages for patients with milder strokes. Similarly, variations in scanner model and/or manufacturer can lead to differential volume estimates and thrombectomy eligibility due to each software’s unique algorithms and AI training sets. Given the potential for these factors to impact software performance, a detailed subgroup analysis would be valuable to the stroke community.

Our study utilizes CTP data from acute stroke patients from four comprehensive academic stroke centers that use both Viz.ai and RAPID. AI to estimate tissue volumes for the purposes of determining thrombectomy eligibility. We specifically compared perfusion volume estimates by applying each software to identical CTP data sets across the four centers and assessing for differences in core and penumbra volume estimates and DEFUSE-3 thrombectomy eligibility. Secondary objectives were to assess whether NIHSS score (i.e., greater or less than 10), scanner manufacturer, scanner model, or institution could affect RAPID. AI and Viz.ai agreement. According to our knowledge and literature review, this study is the largest of its type with 362 cases.

Methods

We reviewed CTP data from patients who were scanned during acute stroke evaluation at the four participating centers. IRB approval/exemption was obtained locally at all centers. The first author had full access to all data in the study and takes responsibility for its integrity and the data analysis.

RAPID.AI and Viz.ai were used to analyze CTP data and report tissue volumes with CBF < 30 % and Tmax > 6s, corresponding to core and penumbra, respectively. Mismatch ratio is also reported and reflects the ratio of penumbra to core. Volume estimates and mismatch were used to determine and report thrombectomy eligibility based on DEFUSE-3 criteria. We opted to use DEFUSE-3 criteria instead of DAWN, as these criteria were the basis for thrombectomy decisions at all participating institutions. Arterial input function (AIF) and venous outflow function (VOF) input parameters used to estimate CBF and Tmax were selected automatically by the respective processing packages. NIHSS score was collected from each patient.

Screening

Our goal was to include patients who had undergone CTP imaging and were positive for LVO. At Center A, RAPID.AI was used from 2018–2021 to process CTP images on patients with CTA-diagnosed LVO in the 6–24 hour window. The original imaging data was retrospectively reprocessed with Viz.ai for the purposes of this study.

The remaining three centers utilized Viz.ai and RAPID.AI simultaneously to process CTP volumes. Center B routinely scans all stroke codes with CTP; data from Jan 2021 to July 2022 were screened. Center C used CTP in acute stroke patients with NIHSS ≥ 4 or for CTA-confirmed LVO; data from Dec 2022 to March 2023 were screened. Center D used CTP in acute stroke patients with NIHSS ≥5 or for CTA-confirmed LVO; data from Jan 2022 to March 2022 were screened.

As noted in Fig. 1 (screening methodology), 767 patients underwent CTP at the four stroke centers. The LVO screening process was dependent on the local CTP protocol and database documentation. Center A utilized CTP only for patients with confirmed LVO. Centers B and C did not document LVO presence in their local stroke databases; therefore, scans were screened for LVO by using Viz.ai’s automatic detection algorithm, which has been demonstrated to have high specificity for detecting LVO15. Cases with NIHSS ≥ 10 but negative for LVO by Viz.ai underwent manual review for LVO as a secondary verification, since prior studies have shown that NIHSS ≥ 10 is highly sensitive and specific for LVO16. Center D documented confirmed LVO presence in their local stroke database. This screening resulted in 362 acute stroke patients with confirmed LVO.

Scanner hardware and protocols

Studies were performed on ten scanners from four different manufacturers: four Canon Aquilion ONE, three GE Discovery, one GE Revolution, one Phillips Ingenuity, and one Siemens Somatom Definition Flash. Each scanner utilizes a slightly different scanning technique to optimize performance and radiation dose, resulting in slightly different image quality and temporal resolution.

Canon Aquilion ONE has 320 detector rows and uses a Dynamic Volume technique with five phases. This technique utilizes the increased detector rows to expand brain coverage without losing temporal resolution or the need to move the table. Parameters include: 15 passes, 1.95s per pass, 24 cm Display field-of-view (DFOV), 300mA, contrast dose 40cc at 5cc/s, 80 kV, and 0.5 mm slice thickness. For one of the scanners, data was reconstructed into 5 mm slices for Viz.ai and 10mm slices for RAPID.AI processing (based on software vendor recommendations). The other three used 5 mm thick reconstructions for both software packages. Total acquisition time was 49.75s.

GE Discovery has 64 detector rows and uses the shuttle technique. This technique moves the table between two cranial levels to increase brain coverage without requiring more radiation, though at the expense of lowering temporal resolution. Parameters include: 17 passes, 2.6 seconds per pass, 25 cm DFOV, 150mA, contrast dose 40cc at 5cc/s, 80 kV, 0.625 mm nominal slice thickness reconstructed into 5 mm slices, 45s scan duration.

GE Revolution has 256 detector rows and utilizes the shuttle technique. Parameters include: 21 passes at 2.2s, DFOV 25cm, 350mA, contrast dose 40cc at 5cc/sec, 80kv, slice thickness 5 mm (not resampled for post processing). Scan duration is 57.8 seconds.

Siemens Somatom Definition Flash is a dual source CT scanner and utilizes a dual source scan technique that accelerates the acquisition of scans using an adaptive 4D spiral approach that includes table movement. This technique utilizes two 128 detector rows which increases brain coverage with less radiation and without losing temporal resolution. Parameters include: 36 passes at 1.5s, DFOV 20cm, 200mA, contrast dose 40cc at 5cc/sec, 80kv, slice thickness 1.2 mm which were resampled to 5mm for Viz.ai and 8mm for RAPID.AI. Scan duration is 67.9 seconds.

Philips Ingenuity has 128 detector rows and uses the jog technique. This technique increases the covered scan range by moving the table back and forth, similar to the shuttle technique. Parameters include: 15 passes, 3.8 s per pass, 16 cm DFOV, 400 mA, contrast dose 40cc at 5cc/sec, 80 kV, slice thickness 1.0 mm and reconstructed 5mm slice thickness, 57s scan duration.

Statistical Analysis

Viz.ai and RAPID.AI data were compared with SPPS 28.0.1.1 (IBM, Chicago, IL) using pairwise non-parametric statistical approaches. A Wilcoxon signed rank test was performed on paired continuous data of core volume, penumbra volume, and mismatch ratio. A McNemar test was performed on paired nominal data including infinite mismatch, DEFUSE-3 thrombectomy eligibility, and the parameters used to determine eligibility (i.e., core ≤ 70 cc, penumbra ≥ 15 cc, and mismatch ≥ 1.8). Linear regression was used to assess correlation between DEFUSE-3 eligibility from both packages. Probability density estimates of the Viz.ai and RAPID.AI core and penumbra data distributions were calculated in MATLAB (Mathworks, Natick, MA) and assessed for similarity by calculating the Jensen-Shannon (JS) divergence (with a JS =0 reflecting perfect similarity and 1 reflecting maximum dissimilarity). The same statistical approaches were used for subgroup analysis, which stratified data by NIHSS score (greater or less than 10), scanner model/manufacturer, and institution.

Results

Primary analysis

In total, 362 CTP cases were included in the primary analysis (Table 1). Viz.ai core volumes were significantly higher than RAPID.AI, with average core volumes of 25.9 cc and 18.2 cc, respectively (Z =9.47, p < 0.001). The Viz.ai penumbra volumes were also significantly higher than RAPID.AI with an average penumbra volume of 102.4 cc and 84.6 cc, respectively (Z = 6.3, P < 0.001). These differences were primarily driven by Canon Aquilion One scanners at two separate institutions, which together comprised the majority of the data (see subgroup analyses, Table 3). Fig. 2a provides violin plots to visualize distributions of core and penumbra volumes, noting similar morphology of the estimated densities for both Viz.ai and RAPID.AI. Jensen-Shannon divergence was calculated as less than 0.02, suggesting a high degree of similarity. No significant difference was found in DEFUSE-3 thrombectomy eligibility (p = 0.68) determined by RAPID.AI or Viz.ai, though a significant difference was found in two of the three metrics used to determine DEFUSE-3 eligibility: core < 70cc and penumbra ≥ 15cc, though not for mismatch > 1.8. We also found a statistically significant difference in the percentage of cases with an infinite mismatch (p < 0.001), as well as in the percentage of cases with large core (p < 0.001).

Fig. 3 displays correlation analyses and Bland-Altman plots for core and penumbra volumes estimated by each software. A strong linear relationship (R2 = 0.83) was observed for core volumes and a moderate linear relationship for penumbra volumes (R2 = 0.56). The Bland-Altman plot reveals that the 95 % limits of agreement between Viz.ai and RAPID.AI spans a 75.3 cc range for core volume and 330.1 cc range for penumbra volume.

NIHSS subgroup analysis

Table 2 stratifies cases with NIHSS < 10 and NIHSS ≥ 10. In both groups, Viz.ai had significantly larger core and penumbra volumes with less infinite mismatch (all p < 0.001). The NIHSS < 10 group, however, demonstrated significant differences in both penumbra ≥ 15 and mismatch ≥ 1.8 (two DEFUSE-3 criteria), which were not seen with the NIHSS ≥ 10 group. Despite this, neither group demonstrated significant differences in thrombectomy eligibility (similar to the primary analysis).

Fig. 2b & c provide violin plots to visualize core and penumbra distributions from Viz.ai and RAPID.AI, again noting similar morphology of the estimated densities. Mean, median, interquartile range, and density envelope are all shifted to slightly higher values for NIHSS ≥ 10 data relative to NIHSS < 10 data. Jensen-Shannon divergence was calculated as less than 0.025 when comparing densities, also suggesting a high degree of similarity between the estimated densities.

Institution and scanner model subgroup analysis

Table 3 shows results separated by institution and secondarily by scanner model/manufacturer. Data from Canon Aquilion One models at Centers A and D demonstrated significant differences in both core and penumbra estimates, with Viz.ai demonstrating systematically larger volumes, though without a statistical difference in thrombectomy eligibility. Data from Center B, acquired on Phillips Ingenuity scanners, demonstrated a significant difference in penumbra estimates, with RAPID.AI estimating a larger penumbra, though this also did not result alter thrombectomy eligibility.

A significant difference in DEFUSE-3 eligibility was only seen with data from Center D, acquired on Canon Aquilion One scanners (53 % versus 46 % of patients eligible for Viz.ai and RAPID.AI, respectively, p = 0.03). No other significant differences in thrombectomy eligibility were seen in this subgroup analysis.

Discussion

Primary Analysis

Our analysis of all 362 patients showed statistically significant differences in CT perfusion volumes estimated by Viz.ai and RAPID.AI, with both core and penumbra systematically larger in Viz.ai. However, it is important to note that these differences were driven by data from a single scanner model, the Canon Aquilion One, acquired at two separate institutions, which accounted for the majority of the dataset. In theory, such differences can have significant implications for DEFUSE-3 thrombectomy eligibility by potentially altering both volume and mismatch values to such an extent that the thresholds for eligibility are crossed. For example, a larger core volume could cause the core > 70 cc threshold to be exceeded, whereas a smaller penumbra could result in the penumbra ≥ 15 cc threshold not being met. Changes in core and penumbra will also alter mismatch ratio, potentially altering eligibility based on the 1.8 mismatch threshold. Despite these possibilities, no significant difference was found in DEFUSE-3 thrombectomy eligibility.

Bland Altman plots quantitatively assess the two methods for agreement in the absence of a gold standard17. Agreement in core volume between software packages was stronger than penumbra, as seen in Fig. 3, with the limits of agreement for core and penumbra volume spanning a 75.3 cc and 330.1 cc range, respectively. Given these large ranges relative to the threshold values, a significant change in DEFUSE-3 eligibility may be expected. On the contrary, it was found that that regardless of the volume changes, mismatch was not significantly changed and the number of cases crossing either volume threshold was small (as seen in Table 1).

Violin plots allow visualization of the core and penumbra distribution and estimated probability density envelopes, which qualitatively appear similar for Viz.ai and RAPID.AI. This similarity is quantitatively affirmed by JS divergence scores lower than 0.03 (on a scale from 0 to 1, with 0 reflecting identical distributions). These plots do highlight the slight systematic bias between Viz.ai and RAPID.AI volume estimates, evidenced by the slightly higher median, mean, and interquartile values for Viz.ai. Similar density morphology, however, suggests comparable performance apart from this systematic bias.

Subgroup analysis

Subgroup analysis based on NIHSS score showed the same pattern of core and penumbra volume discrepancy for both NIHSS < 10 and NIHSS ≥ 10 groups as seen in the primary analysis, again without change in thrombectomy eligibility. The core and penumbra volume distributions for both software packages are shifted to higher values for the NIHSS ≥ 10 group compared to NIHSS < 10 group, which is unsurprising since a larger volume of tissue is expected to be affected in patients with higher stroke scales.

Given the lower volumes in the NIHSS < 10 group, we expected that differences between Viz.ai and RAPID.AI volume estimates would have more impact on DEFUSE-3 criteria (specifically, penumbra ≥ 15 ml and mismatch ≥ 1.8) for this group than for those with higher NIHSS scores. This expectation was based on the disproportionate effects of volume discrepancies when computing DEFUSE-3 criteria, when volumes are overall smaller. Indeed, our results in Table 2 confirm this hypothesis. Despite these observations, however, no differences in overall thrombectomy eligibility were found in either group.

Findings revealed by institution and scanner manufacturer/model stratification were heterogeneous. Systematically larger volumes were seen for Viz.ai relative to RAPID.AI in Canon subgroups, though again without a significant difference in thrombectomy eligibility. Given the disproportionate contribution of Canon data (246 cases out of 362), the volume discrepancies between Viz.ai and RAPID.AI seen in primary analysis were almost certainly driven by the Canon data. While data acquired on Center B’s Phillips scanner did show a significant difference in penumbra between Viz.ai and RAPID.AI, the discrepancy was in the opposite direction, i.e. RAPID.AI had the larger penumbra volume (with mean difference of 17 cc).

Only the data from Center D’s Canon scanners showed significant difference in DEFUSE-3 eligibility (53 % by Viz.ai, 46 % by RAPID.AI, p <0.03), a meaningful result particularly given that over half the cases are from this center (204/362).

Differences in the volume and eligibility parameters between Viz.ai and RAPID may be attributed to a variety of factors. For example, each center utilized different CTP protocols, including variable slice thickness, radiation dose, scan duration, CTP indication, and CT perfusion technique (jog, shuttle, dynamic, etc), amongst other factors. Center A’s Canon Aquilion One scanner sent 5mm cuts to Viz.ai and 10mm slice thickness to RAPID.AI (based on recommendations from the software vendors), whereas Center D’s Canon Aquilion One scanners (with a similar scanning protocol) sent 5mm cuts to both software platforms. Alterations in slice thickness will change image characteristics such as smoothness and SNR, both of which can affect volume estimation. For example, use of thinner slices will result in higher spatial resolution but lower SNR, potentially causing underestimation of core/penumbra volumes.

Since each software is trained using different datasets, a bias could theoretically develop from training datasets that lack scanner and/or protocol diversity. This could explain, for example, why RAPID.AI has larger volumes on Philips and Viz.ai has larger volumes on Canon.

Both Viz.ai and RAPID.AI utilize fundamentally different algorithms for processing CTP data, including for AIF/VOF selection and motion correction, though little information is provided to the end user given the proprietary nature. Differences in AIF/VOF selection will assuredly alter volume estimates, independent of CTP algorithm. A variety of approaches for estimating perfusion parameters from CTP data exist (for example, convolutional, deconvolutional, and machine learning) and will also likely lead to slightly different volume estimates18. As mentioned, choice of slice thickness directly affects the input data and each software may be optimized for a specific slice thickness, for example, to balance spatial resolution and SNR.

A practical interpretation of differences noted above is that each software has been designed and optimized for a set of specific scanning protocols, which may slightly differ depending on scanner model and CTP indication. Deviation from recommended parameter sets could result in less accurate volume estimates. Centers that utilize CTP should therefore routinely discuss optimal scan protocol parameters with both their CTP software vendor and scanner manufacturer to maximize accuracy of core and penumbra volume estimates.

Limitations

While our study encompasses a significant number of patients, we must acknowledge that our primary analysis is influenced by a large number of cases from Canon Aquilion One scanners, accounting for 246 out of 362 cases – 204 from Center D and 42 from Center A. Notably, Center D was the only one among all four centers to show significant DEFUSE-3 eligibility differences between Viz.ai and RAPID.AI. Interestingly, despite Center A observing similar core and penumbra volume differences as Center D, it showed identical DEFUSE-3 eligibility (64 % for both software packages). This suggests that the DEFUSE-3 eligibility variations observed at Center D may stem from its larger sample size, greater detection power, and other factors (such as different screening methods between Centers A and D). Conversely, the capacity to detect such differences at the other centers may have been limited by smaller sample sizes and lower detection power. A larger, more extensive study that includes more subjects per scanner model could potentially uncover differences not apparent in this initial analysis.

We did not manually choose AIF or VOF for individual cases and allowed them to be automatically chosen by Viz.ai and RAPID.AI, since this would be the approach used in actual practice. Since both software platforms use different methodology for choosing AIF and VOF, this likely contributes to the observed volume discrepancies.

While we focused on identifying volume differences between Viz.ai and RAPID.AI, we did not have access to criterion-standard reference values that would have allowed us to determine which platform provides more accurate estimates. Incorporating such reference values, for example via concurrently performed MRI, would be valuable in future studies to gain a more comprehensive understanding of the accuracy and reliability of both platforms.

It is important to note that DEFUSE-3 eligibility is not the only factor that determines whether a patient is taken for thrombectomy, and by extension, discrepancies in DEFUSE-3 eligibility do not necessarily imply missed thrombectomy opportunities. Ultimately, decision to treat with thrombectomy is multifactorial, individualized, and interdisciplinary, based on additional variables outside of the DEFUSE-3 volume criteria such as baseline functional status, goals of care, and other risk versus benefit discussions. While we focused on DEFUSE-3 criteria (since these were used for thrombectomy decisions at participating centers), we acknowledge that our findings may have been different if DAWN criteria were used instead for eligibility determination.

Our results highlight the inherent challenges and uncertainties of using CTP for determining thrombectomy eligibility. At Center D, using RAPID.AI over Viz.AI would have changed the thrombectomy strategy for 14 patients – a substantial impact. Furthermore, recent studies indicate that DAWN and DEFUSE-3 eligibility criteria may be overly restrictive, particularly concerning maximum core infarction size19,20,21. Comparable outcomes of patients selected using ASPECTS scores from non-contrast CT20,22,23 and/or collateral status from CTA21 suggest that alternative methods to CTP may be equally effective, if not preferable, in some scenarios. While RAPID.AI and Viz.ai are two of the most widely used packages, emerging software such as syngo.via, IntelliSpace Portal, and Olea are also gaining popularity and will likely introduce further variability. Although CTP remains the guideline-recommended standard4, mounting evidence challenges this idea and may prompt reevaluation of its role in determining thrombectomy eligibility in the future23.

Conclusion

This multicenter, retrospective study demonstrated significant discrepancy between core and penumbra volumes estimated by Viz.ai and RAPID.AI, which is driven by a single scanner model (Canon) and persists after stratifying for NIHSS. Despite these differences in volume estimates, no significant difference was seen for thrombectomy eligibility, which is reassuring. Subgroup analysis after stratifying data by institution and scanner model, however, revealed a statistically significant difference in thrombectomy eligibility for scans from a specific institution and scanner model. Our study demonstrates that performance of CTP software can be affected by parameters such as scanner model/manufacturer and local CTP protocol. These results support a recommendation that centers discuss optimal CTP parameter settings with both stroke AI software vendors and scanner manufacturers to adjust parameters for maximum accuracy.

Sources of Funding

BA is funded in part through NIH U24 NS107225-05.

Abbreviations:

AI Artificial Intelligence

AIF Arterial Input Function

CBF Cerebral Blood Flow

CBV Cerebral Blood Volume

CTP Computed Tomography Perfusion

DFOV Display Field of View

IRB Institutional review board

LVO Large vessel occlusion

MTT Mean Transit Time

NIHSS National Institutes of Health Stroke Scale

SNR Signal-to-noise Ratio

Tmax Time to maximum of the Tissue Residue Function

VOF Venous Outflow Function

Fig. 1. Study population flow chart and screening methodology.

Four different centers were included in the study. Each center’s indication for CTP differed, requiring a center-specific database query to identify patients with LVO. At centers that did not document LVO presence (Center B and C), patients were included if LVO detected on Viz.ai or by manual review in patients with NIHSS ≥ 10.

Fig. 2. Violin plots of core and penumbra for all 362 patients (a), NIHSS < 10 patients (b), and NIHSS ≥ 10 patients (c). White dot, horizontal lines, and vertical blue bar represent median, mean, and interquartile range, respectively. Colored envelopes reflect the kernel density estimate of the data distribution. Data for Viz.ai and RAPID.AI had similar distributions for all three categories, as evidenced by similar kernel density morphology and low Jensen-Shannon divergence scores.

Fig. 3. Correlation (left column) and Bland-Altman (right column) analysis of core and penumbra volumes between RAPID.AI and Viz.ai. Strong agreement for core volume (A) and moderate agreement for penumbra (B) is seen. Dotted lines in Bland-Altman plots indicate the mean difference of the Viz.AI-RAPID.AI volumes, with solid lines estimating the agreement interval in which 95 % of the differences between methods reside.

Table 1: Volume and thrombectomy eligibility results for all patients.

Group	Data	Viz.ai	RAPID.AI	z	P value	
	
	Core (cc)	25.9	18.2	9.5	<0.001**	
	Penumbra (cc)	102.4	84.6	6.3	<0.001**	
	Mismatch Ratio	9.5	5.7	0.4	0.71	
All patients	Infinite mismatch	26%	46%		<0.001**	
n=362	DEFUSE3 eligible	55%	54%		0.68	
	Core <70cc	88%	93%		<0.001**	
	Penumbra ≥15cc	67%	62%		<0.01**	
	Mismatch ≥1.8	73%	77%		0.13	
	Large core >70cc	11%	7%		<.001**	

Table 2: Volume and thrombectomy eligibility results stratified by NIHSS.

Group	Data	Viz.ai	RAPID.AI	z	P value	Avg NIHSS	
	
	Core (cc)	16.2	8.5	6.5	<0.001**		
NIHSS<10	Penumbra (cc)	61.9	49.4	4.6	<0.001**		
n=179	Mismatch Ratio	9.6	3.2	2.4	0.18		
Canon:139	Infinite mismatch	34%	60%		<0.001**		
GE:27	DEFUSE3 eligible	50%	45%		0.11	3.4	
Philips:8	Core <70cc	93%	97%		0.02**		
Siemens:5	Penumbra ≥15cc	57%	47%		0.01**		
	Mismatch ≥1.8	68%	75%		0.04**		
	Large core >70cc	7%	3%		0.02**		
	
	Core (cc)	37.3	29.2	7.0	<0.001**		
NIHSS≥10	Penumbra (cc)	149.0	123.4	4.9	<0.001**		
n=171	Mismatch Ratio	11.2	7.3	1.5	0.13		
Canon:108	Infinite mismatch	19%	34%		<0.001**	18.9	
GE:41	DEFUSE3 eligible	61%	64%		0.26		
Philips:19	Core <70cc	82%	88%		0.01**		
Siemens:3	Penumbra ≥15cc	80%	78%		0.61		
	Mismatch ≥1.8	81%	80%		0.81		
	Large core >70cc	16%	12%		0.04**		

Table 3: Volume and thrombectomy eligibility results stratified by institution and scanner model.

Center	Scanner	Data	Viz.ai	RAPID.AI	z	P value	
	
	Canon	Core (cc)	32.5	25.9	3.5	<0.001**	
	Aquilion	Penumbra (cc)	161.2	124.4	2.8	0.005**	
	One	Mismatch Ratio	17.1	10.8	1.0	0.31	
	n=42	Infinite mismatch	26%	40%		0.11	
A		DEFUSE3 eligible	64%	64%		1	
		
		Core (cc)	14.1	12.8	0.4	0.67	
	GE	Penumbra (cc)	102.0	98.2	0.7	0.47	
	Discovery	Mismatch Ratio	10.9	6.1	0.6	0.55	
	n=60	Infinite mismatch	27%	42%		0.02**	
		DEFUSE3 eligible	63%	73%		0.07	
	
		Core (cc)	23.2	19.2	1.8	0.08	
	Philips	Penumbra (cc)	60.7	78	2.6	0.009**	
B	Ingenuity	Mismatch Ratio	11.3	3.3	1.9	0.89	
	n=38	Infinite mismatch	13%	24%		0.29	
		DEFUSE3 eligible	37%	45%		0.38	
	
		Core (cc)	47.7	42.9	1.3	0.19	
	GE	Penumbra (cc)	121.1	130.2	0.8	0.41	
	Revolution	Mismatch Ratio	3.1	5.2	1.4	0.16	
	n=10	Infinite mismatch	10%	10%		1	
C		DEFUSE3 eligible	60%	70%		1	
		
	Siemens	Core (cc)	46.0	41.6	0.7	0.5	
	Somatom	Penumbra (cc)	91.4	88.6	0.1	0.89	
	Flash	Mismatch Ratio	3.4	3.8	0.3	0.75	
	n=8	Infinite mismatch	13%	25%		1	
		DEFUSE3 eligible	63%	63%		1	
	
	Canon	Core (cc)	26.7	16.0	9.2	<0.001**	
	Aquilion	Penumbra (cc)	97.7	71.2	8.1	<0.001**	
D	One	Mismatch Ratio	7.8	5.1	0.8	0.48	
	n=204	Infinite mismatch	29%	55%		<0.001**	
		DEFUSE3 eligible	53%	46%		0.03**	

Declaration of competing interest

BA has nothing to disclose is funded in part through NIH U24 NS107225–05. DMM is on Speaker’s Bureau for Chiesi and AstraZeneca. CI, KVS, RS, LP, RC, KVO, DT, BP, BH, SK, have nothing to disclose. BCM is an advisor to Sevarro, Inc. DSB has industry-funded research grants from GE Healthcare and Siemens Medical Solutions USA.

CRediT authorship contribution statement

Benjamin T. Alwood: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. Dawn M. Meyer: Conceptualization, Data curation, Formal analysis, Methodology, Resources, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. Chip Ionita: Data curation, Writing – review & editing. Kenneth V. Snyder: Data curation, Supervision, Writing – review & editing. Roberta Santos: Data curation, Writing – review & editing. Lindsey Perrotta: Data curation, Writing – review & editing. Ryan Crooks: Data curation, Supervision, Writing – review & editing. Kimberlee Van Orden: Data curation, Writing – review & editing. Dolores Torres: Data curation, Writing – review & editing. Briana Poynor: Data curation, Writing – review & editing. Nhan Pham: Data curation, Validation, Writing – review & editing. Sophie Kelly: Data curation, Software, Validation, Writing – review & editing. Brett C. Meyer: Conceptualization, Data curation, Formal analysis, Methodology, Resources, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. Divya S. Bolar: Conceptualization, Formal analysis, Investigation, Methodology, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing, Data curation, Resources, Software.
==== Refs
References

1. Albers GW , Lansberg MG , Kemp S , Tsai JP . A multicenter randomized controlled trial of endovascular therapy following imaging evaluation for ischemic stroke (DEFUSE-3). Int J Stroke: Off J In. Stroke Soc. 2017;12 (8 ):896–905.
2. Nogueira RG , Jadhav AP , Haussen DC . Thrombectomy 6 to 24 hours after stroke with a mismatch between deficit and infarct. N Engl J Med. 2018;378 (1 ):11–21.29129157
3. Saver JL , Goyal M , Lugt A . Time to treatment with endovascular thrombectomy and outcomes from ischemic stroke: a meta-analysis. JAMA. 2016;(316 ):1279–1289.27673305
4. Powers WJ , Rabinstein AA , Ackerson T . Guidelines for the early management of patients with acute ischemic stroke: 2019 update to the 2018 guidelines for the early management of acute ischemic stroke: A guideline for healthcare professionals from the american heart association/american stroke association. Stroke; A J Cerebral Circulation. 2019;50 (12 ):e344–e418.
5. Rudkin S , Cerejo R , Tayal A . Imaging of acute ischemic stroke. Emerg Radiol. 2018; 25 (6 ):659–672.29980872
6. Vagal A , Wintermark M , Nael K . Automated CT perfusion imaging for acute ischemic stroke: Pearls and pitfalls for real-world use. Neurology. 2019;93 (20 ):888–898.31636160
7. Heit JJ , Wintermark M . Perfusion computed tomography for the evaluation of acute ischemic stroke: strengths and pitfalls. Stroke; J. Cerebral Circulation. 2016;47 (4 ): 1153–1158.
8. Rajpurkar P Lungren MP the current and future state of ai interpretation of medical images. New England J. Med. 2023;388 :1981–1990.37224199
9. Liu QC , Jia ZY , Zhao LB . Agreement and accuracy of ischemic core volume evaluated by three CT perfusion software packages in acute ischemic stroke. J. Stroke Cerebrovascular Diseases: Official J. National Stroke Association. 2021;30 (8 ), 105872.
10. Pérez-Pelegrí M , Biarnés C , Thió-Henestrosa S . Higher agreement in endovascular treatment decision-making than in parametric quantifications among automated CT perfusion software packages in acute ischemic stroke. J Xray Sci Technol. 2021;29 (5 ):823–834.34334443
11. Lu Q , Fu J , Lv K . Agreement of three CT perfusion software packages in patients with acute ischemic stroke: A comparison with RAPID.AI. Eur J Radiol. 2022;156 , 110500.36099834
12. Koopman MS , Berkhemer OA , Geuskens RREG. Comparison of three commonly used CT perfusion software packages in patients with acute ischemic stroke. J Neurointerv Surg. 2019 Dec;11 (12 ):1249–1256.31203208
13. Stanton RJ , Wang LL , Smith MS . Differences in automated perfusion software: do they matter clinically? Stroke: Vascular Int Neurol. 2022;2 (6 ), e000424.
14. Pisani L , Haussen DC , Mohammaden M . Comparison of CT perfusion software packages for thrombectomy eligibility. Ann Neurol. 2023. 10.1002/ana.26748.
15. Karamchandani RR , Helms AM , Satyanarayana S . Automated detection of intracranial large vessel occlusions using Viz.ai software: Experience in a large, integrated stroke network. Brain Behavior. 2023;13 (1 ):e2808.36457286
16. Antipova D , Eadie L , Macaden A . Diagnostic accuracy of clinical tools for assessment of acute stroke: a systematic review. BMC Emergency Med. 2019;19 (1 ):49.
17. Giavarina D Understanding bland Altman analysis. Biochem Med. 2015;25 (2 ): 141–151.
18. Konstas AA , Goldmakher GV , Lee TY . Theoretic basis and technical implementations of CT perfusion in acute ischemic stroke, part 1: Theoretic basis. AJNR. 2009;30 (4 ): 662–668.19270105
19. Sarraj A , Hassan AE , Abraham MG . Trial of endovascular thrombectomy for large ischemic strokes. NEJM. 2023;6 (14 ):1259–1271, 388.
20. Huo X , Ma G , Tong X . Trial of endovascular therapy for acute ischemic stroke with large infarct. NEJM. 2023;6 (14 ):1272–1283, 388.
21. Olthuis Susanne GH , Pirson FAV , Pinckaera FME , Endovascular treatment versus no endovascular treatment after 6–24h in patients with ischaemic stroke and collateral flow on CT angiography (MR CLEAN-LATE) in the Netherlands: a multicenter, open-label, blinded-endpoint, randomized, controlled, phase 3 trial. The Lancet. 2023;401 (10385 )P1371–1380.
22. Bendszus M , Fiehler J , Subtil F , Bonekamp S , Aamodt AH , Fuentes B , Gizewski ER , Hill MD , Krajina A , Pierot L , Simonsen CZ , Zeleňák K , Blauenfeldt RA , Cheng B , Denis A . Endovascular thrombectomy for acute ischaemic stroke with established large infarct: multicentre, open-label, randomised trial. Lancet. 2023;402 (10414 ): 1753–1763.37837989
23. AlMajali M , Dibas M , Ghannam M , Galecio Castillo M , Qudah AA , Khasiyev F , Vivanco Suarez J , Rodriguez Calienes A , Farooqui M , Shogren SL , AlMajali F , Yoo A , Samaniego E , Jovin T , Sarraj A , Does the ischemic core really matter? an updated systematic review and meta analysis of large core trials after TESLA, TENSION, and LASTE. stroke: vascular and interventional Neurology. 0 (0 ):e001243.
