
==== Front
Circulation
Circulation
CIR
Circulation
0009-7322
1524-4539
Lippincott Williams & Wilkins Hagerstown, MD

38881496
CIRCULATIONAHA2024068996
00004
10.1161/CIRCULATIONAHA.124.068996
3
10124
10204
Original Research Articles
Deep Learning for Echo Analysis, Tracking, and Evaluation of Mitral Regurgitation (DELINEATE-MR)
Long Aaron MS 12*
https://orcid.org/0000-0002-3227-4490
Haggerty Christopher M. PhD 24chris.m.haggerty@gmail.com
*
Finer Joshua MS 4jof2031@nyp.org

Hartzel Dustin BS 4ldw9007@nyp.org

https://orcid.org/0000-0003-3476-2882
Jing Linyuan PhD 4jinglinyuan2008@gmail.com

https://orcid.org/0000-0001-7197-2788
Keivani Azadeh PhD 4azastron@gmail.com

Kelsey Christopher BS 4dym9015@nyp.org

Rocha Daniel BA 4qxn9003@nyp.org

https://orcid.org/0000-0002-8741-979X
Ruhl Jeffrey MS 4hvx9001@nyp.org

vanMaanen David MS 4pjg9014@nyp.org

https://orcid.org/0000-0001-7816-7722
Metser Gil MD 1gm2886@cumc.columbia.edu

https://orcid.org/0000-0002-2354-6394
Duffy Eamon MD 1ed2947@cumc.columbia.edu

Mawson Thomas MD 1tm3183@cumc.columbia.edu

https://orcid.org/0000-0001-5400-5008
Maurer Mathew MD 1msm10@cumc.columbia.edu

https://orcid.org/0000-0003-2583-9278
Einstein Andrew J. MD, PhD 13ae2214@cumc.columbia.edu

https://orcid.org/0000-0002-1504-1169
Beecy Ashley MD 5asb9028@nyp.org

Kumaraiah Deepa MD 15DK2699@cumc.columbia.edu

Homma Shunichi MD 1sh23@cumc.columbia.edu

Liu Qi MD 1QL2341@CUMC.COLUMBIA.EDU

https://orcid.org/0000-0001-5075-989X
Agarwal Vratika MD 1va2374@cumc.columbia.edu

Lebehn Mark MD 1mal2327@cumc.columbia.edu

Leon Martin MD 16mleon@crf.org

https://orcid.org/0000-0003-1342-9191
Hahn Rebecca MD 1rth2@columbia.edu

https://orcid.org/0000-0002-9643-3024
Elias Pierre MD 124pae2115@cumc.columbia.edu
†
https://orcid.org/0000-0001-7284-3937
Poterucha Timothy J. MD 1†
Seymour, Paul, and Gloria Milstein Division of Cardiology, Department of Medicine, Columbia University Irving Medical Center/New York–Presbyterian Hospital, NY (A.L., G.M., E.D., T.M., M.M., A.J.E., D.K., S.H., Q.L., V.A., M. Lebehn, M. Leon, R.H., P.E., T.J.P.).
Departments of Biomedical Informatics (A.L., C.M.H., P.E.), Columbia University, New York, NY.
Radiology (A.J.E.), Columbia University, New York, NY.
Information Technology Data Science, New York–Presbyterian Hospital, NY (C.M.H., J.F., D.H., L.J., A.K., C.K., D.R., J.R., D.v.M., P.E.).
Division of Cardiology, Department of Medicine, Weill Cornell Medicine, New York, NY (A.B., D.K.).
Cardiovascular Research Foundation, New York, NY (M. Leon).
Correspondence to: Timothy Poterucha, MD, 177 Fort Washington Ave, Milstein 2, New York, NY 10032. Email tp2558@cumc.columbia.edu
17 6 2024
17 9 2024
150 12 911922
8 2 2024
24 5 2024
© 2024 The Authors.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Circulation is published on behalf of the American Heart Association, Inc., by Wolters Kluwer Health, Inc. This is an open access article under the terms of the Creative Commons Attribution Non-Commercial-NoDerivs License, which permits use, distribution, and reproduction in any medium, provided that the original work is properly cited, the use is noncommercial, and no modifications or adaptations are made.

BACKGROUND:

Artificial intelligence, particularly deep learning (DL), has immense potential to improve the interpretation of transthoracic echocardiography (TTE). Mitral regurgitation (MR) is the most common valvular heart disease and presents unique challenges for DL, including the integration of multiple video-level assessments into a final study-level classification.

METHODS:

A novel DL system was developed to intake complete TTEs, identify color MR Doppler videos, and determine MR severity on a 4-step ordinal scale (none/trace, mild, moderate, and severe) using the reading cardiologist as a reference standard. This DL system was tested in internal and external test sets with performance assessed by agreement with the reading cardiologist, weighted κ, and area under the receiver-operating characteristic curve for binary classification of both moderate or greater and severe MR. In addition to the primary 4-step model, a 6-step MR assessment model was studied with the addition of the intermediate MR classes of mild-moderate and moderate-severe with performance assessed by both exact agreement and ±1 step agreement with the clinical MR interpretation.

RESULTS:

A total of 61 689 TTEs were split into train (n=43 811), validation (n=8891), and internal test (n=8987) sets with an additional external test set of 8208 TTEs. The model had high performance in MR classification in internal (exact accuracy, 82%; κ=0.84; area under the receiver-operating characteristic curve, 0.98 for moderate or greater MR) and external test sets (exact accuracy, 79%; κ=0.80; area under the receiver-operating characteristic curve, 0.98 for moderate or greater MR). Most (63% internal and 66% external) misclassification disagreements were between none/trace and mild MR. MR classification accuracy was slightly higher using multiple TTE views (accuracy, 82%) than with only apical 4-chamber views (accuracy, 80%). In subset analyses, the model was accurate in the classification of both primary and secondary MR with slightly lower performance in cases of eccentric MR. In the analysis of the 6-step classification system, the exact accuracy was 80% and 76% with a ±1 step agreement of 99% and 98% in the internal and external test set, respectively.

CONCLUSIONS:

This end-to-end DL system can intake entire echocardiogram studies to accurately classify MR severity and may be useful in helping clinicians refine MR assessments.

artificial intelligence
deep learning
echocardiography
mitral valve insufficiency
OPEN-ACCESSTRUE
CMECME
SDCT
==== Body
pmcClinical Perspective

What Is New?

Mitral regurgitation assessment by echocardiography is subject to modest reliability, and deep learning analysis may aid in assessment.

The deep learning system presented here intakes entire transthoracic echocardiogram studies and yields an accurate classification of mitral regurgitation severity.

What Are the Clinical Implications?

These types of deep learning systems may be useful to aid clinicians at the time of echocardiogram interpretation or as quality improvement tools to identify studies that may have been misinterpreted.

The generalizability of this artificial intelligence–based algorithm to assess the function of other valves or left ventricular function should be tested in additional studies.

Editorial, see p 934

Mitral regurgitation (MR) is the most common valvular heart disease, with a clinically significant MR prevalence of 1.7% in the adult US population affecting at least 24 million people worldwide.1,2 This prevalence increases over the life span, peaking at 9.3% in adults >75 years of age. More than half of all valvular heart disease is undiagnosed.3 Patients with MR have increased mortality and worse quality of life, and deaths attributable to MR have been increasing during the past decade.4 An increasing degree of MR is associated in a stepwise fashion with increased mortality in moderate and severe MR over matched controls.5

The cornerstone of MR assessment remains the transthoracic echocardiogram (TTE) because of its ubiquity and advantages in cost and patient experience compared with other techniques.6 The accurate classification of MR is a 2-step process by which a trained sonographer or physician acquires diagnostic images followed by a cardiologist’s interpretation of these images. Doppler assessments enable visualization of the MR jet and quantification of MR severity through a variety of methods, including the proximal isovelocity surface area (PISA) method, quantitative Doppler, and 3-dimensional vena contracta area measurement.7 However, MR assessment by TTE is subject to significant interobserver variability in measurement of the MR jet area, vena contracta, and estimated regurgitant orifice area by PISA, leading to variable classifications of nonsevere compared with severe MR.8 This may be a result of the 3-dimensional nature of the regurgitant jet, as well as the temporal variability and effect of loading conditions that may occur even within the same study. As a result, there is a significant unmet need for reliable MR assessment techniques.9

Artificial intelligence (AI) is a general term for computing techniques that seek to simulate logical, often human-like thinking. Deep learning (DL), a subset of AI in which multiple calculation layers are used, has been particularly useful in tasks related to interpreting images and medical information.10 When trained on large, feature-rich data sets in focused problems, the accuracy of DL neural networks can outstrip even expert human observers. Since the publication of the first major DL analyses of echocardiography in 2018, algorithms have been developed that can classify echocardiogram views, quantify the left ventricular (LV) ejection fraction, determine strain, detect specific disease states, and predict outcomes.11–14 Segmentation technologies may reduce the time needed for sonographers and cardiologists to make measurements, including LV wall thickness, atrial size, and LV ejection fraction.15–19 These technologies may yield improvements in intraobserver and interobserver variability, a common issue in echocardiography.20 However, DL analyses of lesions assessed by color Doppler such as MR have remained largely unexplored in the literature and represent a more technically challenging DL task than previous milestones. Successful MR evaluation requires both view and valve classification, must account for variable frame rates per clip altering input size, and suffers from sometimes unreliable ground truth labels given the interobserver variability and the complex nature of the regurgitant jet. In a study with severe MR, perhaps only a few frames from a few clips show severe MR, and integration of results across multiple views is needed for accurate diagnosis. Thus, most of the input data for a case of severe MR look identical to those of lesser disease, increasing the difficulty in learning the difference.

In this study, we aimed to: (1) develop a DL system that can intake complete echocardiogram videos and accurately classify study-level MR severity, (2) determine whether integration of multiple echocardiographic views will lead to improved performance compared with single-view models, (3) validate this model in an independent outside data set, and (4) conduct an in-depth analysis of differences in cardiologist and DL model MR assessments (Figure 1). If integrated into echocardiography reading systems, an accurate DL system based on these principles could serve as a second check and potentially increase the accuracy of echocardiogram interpretation.

Figure 1. Deep learning classification of MR from complete echocardiograms. In this study, an end-to-end deep learning system was trained to classify the severity of mitral regurgitation (MR). The system intakes an entire transthoracic echocardiogram (TTE) study, identifies color Doppler clips that assess the mitral valve, makes a clip-level MR classification, and combines those clip-level classifications into a study-level MR classification on a 4-step scale of none/trace, mild, moderate, and severe. This system demonstrated high accuracy in internal and external test sets with exact agreement between the deep learning model of 79% to 82% and weighted κ coefficients of 0.80 to 0.84. There was strong binary classification of moderate or greater MR with area under the receiver-operating characteristic curve (AUROC) of 0.98 for the detection of moderate or greater MR. Furthermore, the primary deep learning system that integrates multiple TTE views had superior performance compared with an MR classification model using only the 4-chamber view. In the 1% of cases in the internal test set with the most significant disagreement between the TTE report and the deep learning MR classification, an expert panel review indicated that the interpreting cardiologist had misclassified the MR severity in 29% of those cases. These findings indicate that this deep learning system may be useful for improving and monitoring MR classification by echocardiography. AI indicates artificial intelligence.

METHODS

Patients and End-Point Definition

The “internal” development population of the study comprised adult patients undergoing a complete TTE at Columbia University Irving Medical Center, New York–Presbyterian Allen Hospital, or an outpatient cardiology Columbia Doctors clinic site from 2015 to 2023. These TTEs were performed with Philips (Amsterdam, Netherlands) ie33 and Epiq TTE machines with TTE data accessed using the Syngo Dynamics system (Siemens Healthineers, Erlangen, Germany). An external test set was curated from patients undergoing complete TTEs at Weill Cornell Medical Center from 2019 to 2021. These patients had TTEs performed with Philips TTE machines with analysis using the Philips Xcelera system. Patients with congenital heart disease, LV assist devices, or surgical or transcatheter mitral valve repair or replacement were excluded. TTEs with severely limited quality, those containing <50 total clips or <5 clips containing color Doppler images of the mitral valve, and intraprocedural echocardiograms were excluded. Notably, neither the internal nor external data sets were selectively enriched for disease and instead reflect the natural proportion of MR seen in these clinical practices.

MR assessments were abstracted from clinical echocardiogram reports using structured fields and natural language processing of free text fields. This MR determination was initially abstracted by a 6-step ordinal scale (none/trace MR, mild MR, mild-moderate MR, moderate MR, moderate-severe MR, and severe MR) in clinical use at both institutions. These MR determinations were made by attending echocardiographers according to American Society of Echocardiography guidelines, which can include a combination of visual assessment, quantitative Doppler, and PISA assessments.7 A total of >200 TTE reports were reviewed by investigators to confirm accurate MR tokenization.

To maintain consistency with institutional practices, the 6-step scale was used for model training (described later); however, in accordance with national guideline recommendations, we converted predictions to a 4-step system (none/trace, mild, moderate, and severe MR) for model testing in our primary analyses. This conversion was performed by mapping intermediate grades of MR to the next higher class (mild-moderate to moderate and moderate-severe to severe). As a secondary analysis, model testing with respect to the 6-step classification system was also performed, with performance detailed in the Supplemental Material.

This study was conducted with approval from the Columbia University Irving Medical Center and Weill Cornell Medicine institutional review boards. Summary data describing DL model performance in the test sets are provided in the article. Because of patient privacy concerns, echocardiogram images and reports cannot be shared as part of this document.

Model Selection and Development

TTEs were split on a per-patient level into train (60%), validation, (20%), and test (20%) sets. Multiple TTEs were used from each patient in the train set to maximize available training data with analysis restricted to only the single most recent echocardiogram per unique patient in the validation and test sets. The final model was selected according to performance in the validation set and then tested in the held-out test set. Model development and testing were performed in a Health Insurance Portability and Accountability Act–compliant, secure Microsoft Azure DL environment.

The resulting DL system, Deep Learning for Echo Analysis, Tracking, and Evaluation of Mitral Regurgitation (DELINEATE-MR), is designed to intake complete TTE studies, identify relevant color Doppler clips that evaluate the mitral valve, and yield a study-level MR determination. Color Doppler clips were extracted with a heuristic algorithm that detects the presence of the image attributes associated with color Doppler imaging, particularly the color bar present in the upper right corner of each color Doppler image. These color Doppler clips were passed through a 2-headed model that determined the view (parasternal long axis, parasternal short axis, apical 4 chamber, apical 5 chamber, apical 2 chamber, apical 3 chamber, and other) and the presence of the mitral valve in the color Doppler view. The accuracy of this model is described in Figure S1. Color Doppler clips that assessed for MR were passed to the MR classification model.

The MR classification model is designed to mimic the approach of a cardiologist when approaching the visual classification of MR through: (1) classifying MR severity on each single video, (2) ranking of videos with the highest degree of MR across a study, and (3) combining the most severe video-level MR assessments to yield a final study-level prediction of MR. First, the model intakes the TTE color Doppler MR clips and extracts spatial features from each still frame in combination with temporal features (such as motion and flow patterns) across the video sequence. This spatiotemporal convolutional neural network is similar in structure to models shown to be effective for LV ejection fraction determination.12 Notably, this model design does not rely on or yield formal segmentations of an MR jet area. Instead, this approach mirrors the human approach to video interpretation in which spatial patterns are integrated across time, potentially yielding a more nuanced and comprehensive assessment than what a traditional Doppler jet segmentation may be able to achieve. To generate models that are robust to baseline shift in the Doppler aliasing velocity and to yield an opportunity for the model to integrate an assessment of the PISA shell, images with baseline shift were included in both the training and test data. Using a supervised learning approach with study-level MR classification applied to each video, the model was trained to classify each individual video according to the 6-step ordinal MR scale of none/trace, mild, mild-moderate, moderate, moderate-severe, and severe with predictions rolled up to match the 4-step scale recommended by guidelines. These model outputs used the rank-consistent ordinal regression framework.21 This model design results in clip-level MR classification. We further refined the video-level prediction by taking advantage of the clinical observation that the amount of MR visualized in each color Doppler clip can vary widely across views and that the clip showing the most MR across the entire study corresponds to the true MR severity. For instance, a study demonstrating severe MR may have only a few clips that clearly show severe MR, with others showing a less severe degree of MR. Our training method takes advantage of this by using a stochastic TopK process.22 In short, each batch of training data contained up to 10 clips per study; the “k” (in this case, k=3) clips with most severe MR for a given study, as determined by the model, were retained for model optimization (ie, back-propagation) for that batch. This procedure was repeated for each batch of data throughout the training process. Of note, all relevant clips for a given study were included in the test sets rather than randomly subsetting to 10. Finally, the study-level MR classification was determined from an approach that took the 3 clips with the most MR and pooled the model outputs to yield a study-level MR classification for inference in the test sets.

Comparison of Single-View and Multiple-View Models

A series of experiments were conducted to compare MR models that use multiple views (the primary model in the present study) with MR models that use only apical 4-chamber views. The model to detect the severity of MR using only the 4-chamber view was developed from the same data sets and procedure as the multiple-view model, except for the additional restriction to inputting only apical 4-chamber images with color Doppler assessment of the mitral valve for model training and testing. The performance of this 4-chamber view model was compared with that of the multiple-view model with the primary comparison being the exact agreement and weighted κ agreement with the clinician read.

Statistical Analysis

The data were described with descriptive statistics, including mean and SD, median and 95% CI, and frequencies as appropriate. Model performance was assessed in the test cohorts on the 4-step ordinal scale using a variety of methods. A weighted κ was calculated with the 4×4 contingency table comparing the model results with the cardiologist’s clinical MR diagnosis with quadratic weighting to more significantly penalize model classifications that were highly discordant from the cardiologist’s assessments. The weighted κ is traditionally interpreted as: 0.0 to 0.20 as slight agreement, 0.21 to 0.40 as fair agreement, 0.41 to 0.60 as moderate agreement, 0.61 to 0.80 as substantial agreement, and 0.81 to 1.00 as almost perfect agreement.23,24 Given the importance of accurate MR classification and recommendation by some authors that it may overestimate the level of agreement in medical testing,25 a weighted κ>0.60 was interpretated as substantial correlation with deferral of the almost perfect agreement interpretation. The proportion of studies with agreement between the model prediction and the cardiologist’s interpretation was assessed (eg, model classifies mild MR in a study in which the cardiologist classifies mild MR). The area under the receiver-operating curve and the area under the precision-recall (positive-predictive value–sensitivity) curve were performed both for the detection of severe MR (ie, binary classification of severe versus nonsevere) and for detection of moderate or severe MR (ie, moderate or severe versus none/trace or mild). The F1 score was calculated as the harmonic mean of precision (positive predictive value) and recall (sensitivity) for both severe MR and moderate or severe MR. In addition to the primary analyses using the 4-step scale, the rate of exact and ±1 step agreement on the 6-step scale (with the addition of mild-moderate and moderate-severe MR classes) between the model output and cardiologist interpretation was determined.

Metrics were computed by bootstrapping 1000 random subsets of the data, and 95% CIs were computed. All testing was performed with Python version 3.8.

Subset Analyses and Comparison With Quantitative Data

To explore the performance of the system in subsets of MR, the structured and free text data in the TTE reports of the internal test set were abstracted to ascertain the cause of the MR as primary MR or secondary MR. Second, the performance of the model in cases of primary MR caused by mitral valve prolapse or myxomatous mitral valve disease was analyzed and compared with patients without mitral valve prolapse or myxomatous mitral valve disease. Third, the performance of the model in assessment of eccentric MR was compared with patients without eccentric MR. The performance of the model in these subsets was analyzed as the rate of agreement between the DL model and the interpreting cardiologist in each of these subsets. In addition, an adjusted analysis was performed matching the MR severity in each of these subsets to control for differing disease prevalence levels.

The correlation of DL model results with MR quantification metrics of regurgitant volume and estimated regurgitant orifice area by the PISA method and regurgitant volume and fraction by the quantitative Doppler method was studied in the internal test set with the Kendall coefficient of rank correlation. Further details on subset and quantitation analyses are provided in the Supplemental Material.

Analysis of Discordant Cases

An analysis was conducted of cases with the highest degree of discordance between the AI model and cardiologist classification of MR. For this analysis, the 6-step ordinal scale was used to enable more granular assessments. Those cases with a significant discordance (≥2 steps on the 6-point ordinal scale) between the cardiologist and model determination of MR were identified (eg, the cardiologist interpreted mild-moderate MR and the model makes a classification of moderate-severe MR). To fully characterize this discordance, the primary TTE data and the model-identified MR clips were manually reviewed for these cases by a level 3–trained, National Board of Echocardiography–certified attending echocardiographer (T.J.P.) while blinded to both the clinical read and model prediction of MR. For cases in which this overread disagreed with the initial cardiologist’s interpretation, the TTE was additionally reviewed in a blinded fashion by 2 additional level 3–trained, board-certified attending echocardiographers (2 of Q.L., V.A., or M. Lebehn). The 3 blinded interpretations were used to make a consensus MR determination to compare against both the model and the interpreting cardiologist’s assessment of MR using a majority voting method. For cases in which there were 3 different MR assessments, the middle determination was chosen as the consensus MR value (eg, mild, mild-moderate, and moderate votes yield a consensus of mild-moderate MR).

This analysis aimed to establish concordance among the DL model, clinical interpretation, and assessment by the expert panel. We classified all studies into 1 of 4 categories: model success–complete (when the DL model and panel read agreed), model success–±1 (when the DL model interpretation was within 1 step of the panel read), model failure–clinical (when the DL model was off by ≥2 steps from the panel read but correctly analyzed MR images), and model failure–preprocessing (when the inaccurate prediction of the model was due to preprocessing errors with classification predictions made on non-MR views).

RESULTS

Model Development and Validation

A total of 61 689 TTEs in 44 628 unique patients were included in the internal development data set. These TTEs were split into train (n=43 811 in 26 750 patients), validation (n=8991 in 8991 patients), and test (n=8987 in 8987 patients) sets, with only one TTE for any given patient included in the validation and test sets (Figure 2). Characteristics of patients are described on a TTE level in Table 1. In the test set, mean age was 65.1±16.8 years, and 4-grade MR severity was: none-trace (n=5895; 65.6%), mild (n=1878; 20.9%), moderate (n=1034; 11.5%), and severe (n=180; 2.0%).

Table 1. Patient Characteristics

Figure 2. Patient flow. This figure demonstrates the internal and external data sets used to train, validate, and test the mitral regurgitation deep learning (DL) model. LVAD indicates left ventricular assist device; MV, mitral valve; and TTE, transthoracic echocardiography.

In the internal (Columbia) test set, exact agreement on MR severity was achieved in 7411 studies (82%) on the 4-step ordinal scale (Figure 2; Table 2). The weighted κ was 0.84, indicating a substantial level of agreement between the DL model and the clinician read. Of note, most (n=990; 63%) cases of disagreement between the DL model and the cardiologist were in cases in which none/trace MR was classified as mild MR or vice versa. The area under the receiver-operating characteristic curve, area under the precision-recall curve, and F1 scores were 0.98, 0.90, and 0.80, respectively, for the detection of at least moderate MR and 0.99, 0.75, and 0.71, respectively, for the detection of severe MR (Figure 3). These metrics were determined in the context of 14.9% prevalence of moderate or greater MR and 2.0% prevalence of severe MR, respectively. In addition to the primary 4-step MR analysis, MR classification on the 6-step system demonstrated an exact accuracy of 80% and a ±1 step accuracy of 99% (Supplemental Material).

Table 2. Model Performance

Figure 3. Model confusion matrices in internal and external test sets. This figure demonstrates the confusion matrix in the internal (A) and external (B) test sets comparing the cardiologist clinical interpretation (vertical axis) and the DL model prediction (horizontal axis) for mitral regurgitation on a 4-step ordinal scale. Exact agreement between the DL mode and cardiologist was present in 82% in the internal test set and 79% of the external test set. AI indicates artificial intelligence.

Figure 4. Receiver-operating characteristic and precision-recall curves for the detection of moderate or greater MR. These figures demonstrate the model performance in the detection of mitral regurgitation (MR) in the internal (black) and external (red) test sets. A, Receiver operator characteristic (ROC) curve for at least moderate MR and the ROC curve for severe MR (B). C and D, The precision-recall (positive-predictive value–sensitivity) curves for detection of at least moderate MR (C) and severe MR (D). AU-PRC indicates area under the precision-recall curve; and AUROC, area under the receiver operator characteristic curve.

In the external (Cornell) test set, patients were younger than those in the internal data set (60.6±17.8 years of age), and the prevalence of MR classes was similar (Table 1). The model generalized well to these data, with substantial agreement on MR classification (accuracy, 79%; weighted κ=0.80).

Subanalyses by MR Type and View

In subset analyses, the model demonstrated similar performance in the assessment of both primary (accuracy, 72%) and secondary (accuracy in adjusted cohort with matched prevalence, 73%) MR. The accuracy of the model in patients with mitral valve prolapse or myxomatous mitral valve disease was similar (72%) to that in patients without mitral valve prolapse or myxomatous mitral valve disease when adjusted for prevalence (72%). The rate of agreement of the DL model assessment of MR with the cardiologist was lower (61%) for eccentric MR than in patients without eccentric MR (unadjusted accuracy, 83%), a difference that persisted even when prevalence of noneccentric MR was matched to eccentric MR (adjusted accuracy, 68%). Further details on subset analyses are contained in the Supplemental Material.

We hypothesized that restricting the model to a single echocardiographic view (apical 4-chamber view) would reduce performance compared with the DELINEATE-MR performance using all available views. We therefore trained this 4-chamber-view–only model and found a drop in performance in both the internal and external test sets (Table 2).

Association With Quantitative Measures

The results of the model with quantitative metrics of estimated regurgitant orifice area by PISA, regurgitant volume by PISA, regurgitant volume by quantitative Doppler, and regurgitant fraction by quantitative Doppler were analyzed. The model-generated 6-step ordinal MR classification was significantly associated with each of these metrics by the Kendall coefficient (P<0.001 for all analyses; Supplemental Material).

Analysis of Discordant Cases

A total of 106 cases were identified in the internal test set in which there was a significant discordance of 2 or more steps on the 6-point ordinal scale between the cardiologist and model determination of MR (1.2% of the test set). A total of 104 cases had images available for review. Of these 104 cases, 30 (29%) were classified as model success–complete, with the DL model correctly classifying MR severity compared with the panel assessment. In these cases, the initial clinical MR interpretation was incorrect by panel adjudication, which included 4 cases of panel-determined severe MR and 3 cases of moderate-severe MR. An additional 46 (44%) were classified as model success–±1 with the panel assessment of MR severity in between the interpreting cardiologist and the DL model. Among the 28 cases (27%) in which the model failed, 6 (6%) were attributable to preprocessing errors, with the DL model making predictions on non-MR videos, with the remaining 22 (21%) classified as model failure–clinical with incorrect MR classification on appropriate images.

DISCUSSION

In summary, we present here the DELINEATE-MR system, which uses an end-to-end approach to intake an entire TTE study and yield a study-level classification of MR. The principal findings of this study are: (1) the DELINEATE-MR system can accurately classify MR using the entire TTE study, achieving ≈80% agreement compared with clinical interpretation; (2) the model generalized well to an external data set; (3) use of this full system integrating MR classification across multiple views has improved accuracy compared with a model using only the apical 4-chamber view; and (4) the application of this system to retrospective data sets identified cases with probable cardiologist misclassification of MR compared with a panel consensus read (Figure 1).

Enhancing Diagnostic Precision in Valvular Heart Disease

Echocardiography is the most common dedicated cardiac diagnostic imaging test in the world. As an ultrasound-based modality, the images from these studies are highly heterogeneous and constrained by the availability of imaging windows. A trained and experienced sonographer is essential in obtaining diagnostic images that adequately interrogate every structure of interest with particular attention to those that appear most significantly diseased in initial routine imaging views. The interpreting cardiologist must then integrate the ≈50 to 150 resulting clips in a comprehensive TTE into a series of diagnoses, many of which rely on visual qualitative assessment, using images scattered across the entirety of the TTE study. Given these challenges, it is unsurprising to note significant variability in cardiologist interpretation of TTE.8,26–28

Previous work in the analysis of echocardiography using DL has documented successes in reproducing standard measurements such as the determination of LV ejection fraction and detection of specific diseases such as aortic stenosis.12,19 These approaches start with the TTE study and apply a series of DL models to identify specific views (eg, the 4-chamber and 2-chamber views) that can then have a distinct model perform a segmentation- or disease-related task (eg, determination of the biplane LV ejection fraction). Assessment of valvular regurgitation, including MR, is a more challenging task.7 The pathology and mechanics of regurgitation can vary widely among cases, leading to differences in the view and orientation required to fully capture and characterize the degree of regurgitation.29,30 Hence, there may only be a few clips that fully assess the severity of the regurgitation. An effective DL system to classify valvular regurgitation must mirror the approach of an expert clinician by evaluating every relevant clip and flexibly integrating them into a single unifying diagnosis.

The present study was designed with these design principles in mind: using an end-to-end system to intake a complete TTE study, identify color Doppler clips, classify the views and valves imaged, determine the severity of MR on a clip-level classification, and integrate those clip-level predictions into a single study-level MR classification. The results demonstrate the accuracy and generalizability of this approach across 2 distinct cohorts representing diverse patients. This high level of accuracy is consistent across multiple metrics, including weighted κ, area under the receiver-operating characteristic curve to detect moderate or greater MR or severe MR, and exact agreement with the reading cardiologist. Moreover, the results of a rigorous, blinded panel overread suggest that model performance may in some cases identify cases of MR misclassification by the interpreting cardiologist. Model performance was high in both primary and secondary MR, although a signal was noted for decreased agreement with the interpreting cardiologist in cases of eccentric MR. Eccentric MR is known to be challenging for grading by echocardiography both in overall assessment and in accurate quantification.7,9,31 Differences noted in model performance in eccentric cases could be attributable to uncertainty in the underlying cardiologist interpretations in addition to truly worse model performance, and these cases may be particularly likely to benefit from quantitative transesophageal echocardiogram and cardiac magnetic resonance imaging for more definitive assessment.

Although additional prospective validation and testing of this DL system are needed, we envision this technology functioning as an adjunct to cardiologist interpretation within existing clinical workflows. When used prospectively during echocardiogram interpretation, this type of AI system could serve as an independent “second check” to cue the interpreting cardiologist in cases at risk for possible MR misclassification. In addition, this type of technology could be deployed on a laboratory level to audit and monitor the accuracy and reliability of MR determinations. If deployed in this manner, this model could identify patients who could benefit from overreading of their TTE for confirmation or revision of the MR severity. However, we would specifically caution against using DL technologies to replace clinicians. The insight of an experienced cardiologist who can combine image interpretation with nuanced clinical understanding is irreplaceable. In addition, such an approach ignores the critical importance of quality control and laboratory oversight not only in interpreting but also in performing diagnostic studies. Of note, although these models have good statistical performance, they remain imperfect, and these technologies are inadequate for independent use outside the supervision of an experienced cardiologist.

Charting the Course for Future Innovation

The potential of DL extends beyond the classification of MR to potentially revolutionize the interpretation of echocardiography. In the short term, the end-to-end approach described here can be optimized and used in a broader array of echocardiographic diagnoses. Although most easily adapted for the other valvular regurgitation states of aortic regurgitation, tricuspid regurgitation, and pulmonic regurgitation, this approach may also be helpful for LV and right ventricular assessments and specific disease diagnoses such as cardiac amyloidosis. The approach pioneered in this analysis allows the integration of multiple video images into comprehensive diagnoses. In the medium term, these technologies need to be tested in real-world clinical scenarios in clinical trials evaluating both accuracy and changes in efficiency. In addition, these studies can help us explore the human-machine interaction of how clinicians can engage with DL-assisted diagnosed tools. This is particularly relevant given that DL systems with good top-line accuracy can err in unacceptable ways such as yielding predictions based on incorrect views. By identifying and understanding these failure modes, we can develop strategies that integrate the best traits of both human and DL approaches. For instance, cases with significant discrepancies between the cardiologist’s initial interpretation and DL model MR classification may highlight studies that warrant further investigation.

Looking to the future, the implications of near-human or superhuman performance by DL medical imaging models are profound. These systems will challenge the conventional view of what it means to be an expert imager and cardiologist. The long-term optimistic vision is that these tools will be used to improve patient care, reduce missed diagnoses, and improve efficiency such that cardiac imaging can be used as a tool for improved population health. However, the risks in this future are substantial, including clinician overreliance, workforce changes, and inequitable deployment of technological improvements, leading to increases in disparities. Ensuring rigorous prospective clinical studies with accuracy-based end points, continued vigilant laboratory oversight, and dedicated programs to ensure equitable access to these technologies are needed to mitigate these risks.

Limitations

This work does have limitations. All studies were performed in the New York region in academic echocardiographic laboratories with Philips machines. Follow-up studies using a range of machines and practice types will be helpful for further assessment. Model training and assessment of model accuracy are reliant on cardiologist TTE interpretations, which are subject to technical limitations and modest reliability. This “label noise” may limit peak model performance and accuracy assessments compared with an alternative reference standard such as cardiac magnetic resonance imaging, quantitative transesophageal echocardiogram, or a panel TTE interpretation with extensive quantification. Pediatric patients and adult patients with complex congenital heart disease, LV assist devices, and mitral valve interventions were excluded from the study, and results are not generalizable to these populations. Although tested in subtypes of MR, this current model assesses only MR severity and does not assess for the cause and mechanism of MR. The model architecture studied here does not replicate the quantification (PISA; quantitative Doppler) that may have been used by the reading cardiologist to yield an MR determination. Although replicating quantitative metrics may be helpful in supporting disease classification, using an end-to-end approach as was done has been shown to be more accurate than focusing on intermediate steps in machine learning approaches as varied as self-driving car steering and speech recognition.32,33 Future work should focus on automating quantifications and further understanding whether DL-assisted quantification improves MR classification. This work does not address the critical role of the sonographer in obtaining diagnostic imaging. This study is retrospective and does not address the performance of the system in clinical use. A prospective clinical study is being planned to develop and study optimal methods of clinical deployment.

Conclusions

A DL end-to-end system (DELINEATE-MR) demonstrated strong, generalizable performance in the classification of MR using complete TTE studies from retrospective analysis. This approach has potential to improve MR diagnosis compared with human interpretation alone.

ARTICLE INFORMATION

Sources of Funding

This study was funded by grants from the American Heart Association (awards 933452 and 23SCISA1077494; https://doi.org/10.58275/AHA.23SCISA1077494.pc.gr.172160).

Disclosures

Dr Poterucha owns stock in Abbott Laboratories and Baxter International with research support provided to his institution from the Amyloidosis Foundation, American Heart Association (awards 933452 and 23SCISA1077494; https://doi.org/10.58275/AHA.23SCISA1077494.pc.gr.172160), Eidos Therapeutics, Pfizer, Edwards Lifesciences, and Janssen. Dr Hahn reports speaker fees from Abbott Structural, Baylis Medical, Edwards Lifesciences, Medtronic, Philips Healthcare, and Siemens Healthineers; she has institutional consulting contracts for which she receives no direct compensation with Abbott Structural, Anteris, Edwards Lifesciences, Medtronic and Novartis; she is chief scientific officer for the Echocardiography Core Laboratory at the Cardiovascular Research Foundation for multiple industry-sponsored valve trials, for which she receives no direct industry compensation. Dr Einstein reports receiving authorship fees from Wolters Kluwer Healthcare—UpToDate and serving on a scientific advisory board for Canon Medical Systems; his institution has grants/grants pending from Attralus, BridgeBio, Canon Medical Systems, GE HealthCare, Intellia Therapeutics, Ionis Pharmaceuticals, Neovasc, Pfizer, Roche Medical Systems, and W.L. Gore & Associates. The other authors report no conflicts.

Supplemental Material

Tables S1–S4

Figures S1–S3

Supplementary Material

Nonstandard Abbreviations and Acronyms

AI artificial intelligence

DELINEATE-MR Deep Learning for Echo Analysis, Tracking, and Evaluation of Mitral Regurgitation

DL deep learning

LV left ventricular

MR mitral regurgitation

PISA proximal isovelocity surface area

TTE transthoracic echocardiogram

* A. Long and C.M. Haggerty contributed equally.

† P. Elias and T.J. Poterucha contributed equally.

Supplemental Material, the podcast, and transcript are available with this article at https://www.ahajournals.org/doi/suppl/10.1161/CIRCULATIONAHA.124.068996.

Continuing medical education (CME) credit is available for this article. Go to http://cme.ahajournals.org to take the quiz.

For Sources of Funding and Disclosures, see page 921.

Circulation is available at www.ahajournals.org/journal/circ
==== Refs
REFERENCES

1. Nkomo VT Gardin JM Skelton TN Gottdiener JS Scott CG Enriquez-Sarano M . Burden of valvular heart diseases: a population-based study. Lancet. 2006;368 :1005–1011. doi: 10.1016/S0140-6736(06)69208-8 16980116
2. Coffey S Roberts-Thomson R Brown A Carapetis J Chen M Enriquez-Sarano M Zuhlke L Prendergast BD . Global epidemiology of valvular heart disease. Nat Rev Cardiol. 2021;18 :853–864. doi: 10.1038/s41569-021-00570-z 34172950
3. d’Arcy JL Coffey S Loudon MA Kennedy A Pearson-Stuttard J Birks J Frangou E Farmer AJ Mant D Wilson J . Large-scale community echocardiographic screening reveals a major burden of undiagnosed valvular heart disease in older people: the OxVALVE Population Cohort Study. Eur Heart J. 2016;37 :3515–3522. doi: 10.1093/eurheartj/ehw229 27354049
4. Parcha V Patel N Kalra R Suri SS Arora G Arora P . Mortality due to mitral regurgitation among adults in the United States: 1999-2018. Mayo Clin Proc. 2020;95 :2633–2643. doi: 10.1016/j.mayocp.2020.08.039 33276836
5. Dziadzko V Clavel MA Dziadzko M Medina-Inojosa JR Michelena H Maalouf J Nkomo V Thapa P Enriquez-Sarano M . Outcome and undertreatment of mitral regurgitation: a community cohort study. Lancet. 2018;391 :960–969. doi: 10.1016/S0140-6736(18)30473-2 29536860
6. Otto CM Nishimura RA Bonow RO Carabello BA Erwin JP 3rd Gentile F Jneid H Krieger EV Mack M McLeod C . 2020 ACC/AHA guideline for the management of patients with valvular heart disease: a report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. Circulation. 2021;143 :e72–e227. doi: 10.1161/CIR.0000000000000923 33332150
7. Zoghbi WA Adams D Bonow RO Enriquez-Sarano M Foster E Grayburn PA Hahn RT Han Y Hung J Lang RM . Recommendations for noninvasive evaluation of native valvular regurgitation: a report from the American Society of Echocardiography developed in collaboration with the Society for Cardiovascular Magnetic Resonance. J Am Soc Echocardiogr. 2017;30 :303–371. doi: 10.1016/j.echo.2017.01.007 28314623
8. Biner S Rafique A Rafii F Tolstrup K Noorani O Shiota T Gurudevan S Siegel RJ . Reproducibility of proximal isovelocity surface area, vena contracta, and regurgitant jet area for assessment of mitral regurgitation severity. JACC Cardiovasc Imaging. 2010;3 :235–243. doi: 10.1016/j.jcmg.2009.09.029 20223419
9. Grayburn PA Weissman NJ Zamorano JL . Quantitation of mitral regurgitation. Circulation. 2012;126 :2005–2017. doi: 10.1161/CIRCULATIONAHA.112.121590 23071176
10. Dey D Slomka PJ Leeson P Comaniciu D Shrestha S Sengupta PP Marwick TH . Artificial intelligence in cardiovascular imaging: JACC state-of-the-art review. J Am Coll Cardiol. 2019;73 :1317–1335. doi: 10.1016/j.jacc.2018.12.054 30898208
11. Zhang J Gajjala S Agrawal P Tison GH Hallock LA Beussink-Nelson L Lassen MH Fan E Aras MA Jordan C . Fully automated echocardiogram interpretation in clinical practice. Circulation. 2018;138 :1623–1635. doi: 10.1161/CIRCULATIONAHA.118.034338 30354459
12. Ouyang D He B Ghorbani A Yuan N Ebinger J Langlotz CP Heidenreich PA Harrington RA Liang DH Ashley EA . Video-based AI for beat-to-beat assessment of cardiac function. Nature. 2020;580 :252–256. doi: 10.1038/s41586-020-2145-8 32269341
13. Salte IM Ostvik A Smistad E Melichova D Nguyen TM Karlsen S Brunvand H Haugaa KH Edvardsen T Lovstakken L . Artificial intelligence for automatic measurement of left ventricular strain in echocardiography. JACC Cardiovasc Imaging. 2021;14 :1918–1928. doi: 10.1016/j.jcmg.2021.04.018 34147442
14. Ulloa Cerna AE Jing L Good CW vanMaanen DP Raghunath S Suever JD Nevius CD Wehner GJ Hartzel DN Leader JB . Deep-learning-assisted analysis of echocardiographic videos improves predictions of all-cause mortality. Nat Biomed Eng. 2021;5 :546–554. doi: 10.1038/s41551-020-00667-9 33558735
15. Duffy G Cheng PP Yuan N He B Kwan AC Shun-Shin MJ Alexander KM Ebinger J Lungren MP Rader F . High-throughput precision phenotyping of left ventricular hypertrophy with cardiovascular deep learning. JAMA Cardiol. 2022;7 :386–395. doi: 10.1001/jamacardio.2021.6059 35195663
16. Lau ES Achille PD Kopparapu K Andrews CT Singh P Reeder C Al-Alusi M Khurshid S Haimovich JS Ellinor PT . Deep learning–enabled assessment of left heart structure and function predicts cardiovascular outcomes. J Am Coll Cardiol. 2023;82 :1936–1948. doi: 10.1016/j.jacc.2023.09.800 37940231
17. He B Kwan AC Cho JH Yuan N Pollick C Shiota T Ebinger J Bello NA Wei J Josan K . Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. 2023;616 :520–524. doi: 10.1038/s41586-023-05947-3 37020027
18. Holste G Oikonomou EK Mortazavi BJ Coppi A Faridi KF Miller EJ Forrest JK McNamara RL Ohno-Machado L Yuan N . Severe aortic stenosis detection by deep learning applied to echocardiography. Eur Heart J. 2023;44 :4592–4604. doi: 10.1093/eurheartj/ehad456 37611002
19. Krishna H Desai K Slostad B Bhayani S Arnold JH Ouwerkerk W Hummel Y Lam CSP Ezekowitz J Frost M . Fully automated artificial intelligence assessment of aortic stenosis by echocardiography. J Am Soc Echocardiogr. 2023;36 :769–777. doi: 10.1016/j.echo.2023.03.008 36958708
20. Yuan N Jain I Rattehalli N He B Pollick C Liang D Heidenreich P Zou J Cheng S Ouyang D . Systematic quantification of sources of variation in ejection fraction calculation using deep learning. JACC Cardiovasc Imaging. 2021;14 :2260–2262. doi: 10.1016/j.jcmg.2021.06.018 34274282
21. Shi X Cao W Raschka S . Deep neural networks for rank-consistent ordinal regression based on conditional probabilities. Pattern Analysis Applications. 2021;26 :941–955. doi: 10.1007/s10044-023-01181-9
22. Metwally A Agrawal D El Abbadi A . Efficient computation of frequent and Top-k elements in data streams. Database Theory ICDT 2005. 2005;3363 :398–412.
23. Cohen J . A coefficient of agreement for nominal scales. Educ Psychol Measurement. 1960;20 :37–46.
24. Cohen J . Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychol Bull. 1968;70 :213–220. doi: 10.1037/h0026256 19673146
25. McHugh ML . Interrater reliability: the kappa statistic. Biochem Med (Zagreb). 2012;22 :276–282. doi: 10.1016/j.jocd.2012.03.005 23092060
26. Moura LM Ramos SF Pinto FJ Barros IM Rocha-Goncalves F . Analysis of variability and reproducibility of echocardiography measurements in valvular aortic valve stenosis. Rev Port Cardiol. 2011;30 :25–33.21425741
27. Coisne A Aghezzaf S Edmé J-L Bernard A Ma I Bohbot Y Di Lena C Nicol M Lavie Badie Y Eyharts D . Reproducibility of reading echocardiographic parameters to assess severity of mitral regurgitation: insights from a French multicentre study. Arch Cardiovasc Dis. 2020;113 :599–606. doi: 10.1016/j.acvd.2020.02.004 32994143
28. Galema TW Geleijnse ML Yap SC van Domburg RT Biagini E Vletter WB Ten Cate FJ . Assessment of left ventricular ejection fraction after myocardial infarction using contrast echocardiography. Eur J Echocardiogr. 2008;9 :250–254. doi: 10.1016/j.euje.2007.03.025 17475569
29. Grayburn PA Thomas JD . Basic principles of the echocardiographic evaluation of mitral regurgitation. JACC Cardiovasc Imaging. 2021;14 :843–853. doi: 10.1016/j.jcmg.2020.06.049 33454273
30. Sabbagh AE Reddy YNV Nishimura RA . Mitral valve regurgitation in the contemporary era. JACC Cardiovasc Imaging. 2018;11 :628–643. doi: 10.1016/j.jcmg.2018.01.009 29622181
31. Zoghbi WA Enriquez-Sarano M Foster E Grayburn PA Kraft CD Levine RA Nihoyannopoulos P Otto CM Quinones MA Rakowski H ; American Society of Echocardiography. Recommendations for evaluation of the severity of native valvular regurgitation with two-dimensional and Doppler echocardiography. J Am Soc Echocardiogr. 2003;16 :777–802. doi: 10.1016/S0894-7317(03)00335-3 12835667
32. Bojarski M Del Testa D Dworakowski D Firner B Flepp B Goyal P Jackel LD Monfort M Muller U Zhang J . End to end learning for self-driving cars. arXiv. Preprint posted online April 25, 2016. arXiv:1604.07316.
33. Prabhavalkar R Hori T Sainath TN Schlüter R Watanabe S . End-to-end speech recognition: a survey. arXiv. Preprint posted online March 3, 2023. arXiv:2303.03329
