
==== Front
Circulation
Circulation
CIR
Circulation
0009-7322
1524-4539
Lippincott Williams & Wilkins Hagerstown, MD

39129623
CIRCULATIONAHA2024069047
00005
10.1161/CIRCULATIONAHA.124.069047
3
10094
10103
10198
Original Research Articles
High-Throughput Deep Learning Detection of Mitral Regurgitation
https://orcid.org/0000-0001-8895-1238
Vrudhula Amey MD 13
Duffy Grant BS 1Grant.Duffy@cshs.org

https://orcid.org/0000-0002-0431-1493
Vukadinovic Milos BS 14milosvuk@ucla.edu

https://orcid.org/0000-0002-7343-7583
Liang David MD, PhD 5david.ouyang@cshs.org

https://orcid.org/0000-0002-4977-036X
Cheng Susan MD, MMSc, MPH 1susan.cheng@cshs.org

https://orcid.org/0000-0002-3813-7518
Ouyang David MD 2
Department of Cardiology, Smidt Heart Institute (A.V., G.D., M.V., S.C.), Cedars-Sinai Medical Center, Los Angeles, CA.
Division of Artificial Intelligence in Medicine (D.O.), Cedars-Sinai Medical Center, Los Angeles, CA.
Icahn School of Medicine at Mt Sinai, New York, NY (A.V.).
Department of Bioengineering, University of California Los Angeles (M.V.).
Department of Medicine, Division of Cardiology, Stanford University, Palo Alto, CA (D.L.).
Correspondence to: David Ouyang, MD, Smidt Heart Institute, Cedars-Sinai Medical Center, 127 S San Vicente Blvd, AHSP Pavilion Suite A3100, Los Angeles, CA 90048. Email david.ouyang@cshs.org
12 8 2024
17 9 2024
150 12 923933
4 2 2024
27 6 2024
© 2024 The Authors.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Circulation is published on behalf of the American Heart Association, Inc., by Wolters Kluwer Health, Inc. This is an open access article under the terms of the Creative Commons Attribution Non-Commercial-NoDerivs License, which permits use, distribution, and reproduction in any medium, provided that the original work is properly cited, the use is noncommercial, and no modifications or adaptations are made.

BACKGROUND:

Diagnosis of mitral regurgitation (MR) requires careful evaluation by echocardiography with Doppler imaging. This study presents the development and validation of a fully automated deep learning pipeline for identifying apical 4-chamber view videos with color Doppler echocardiography and detecting clinically significant (moderate or severe) MR from transthoracic echocardiograms.

METHODS:

A total of 58 614 transthoracic echocardiograms (2 587 538 videos) from Cedars-Sinai Medical Center were used to develop and test an automated pipeline to identify apical 4-chamber view videos with color Doppler across the mitral valve and then assess MR severity. The model was tested internally on a test set of 1800 studies (80 833 videos) from Cedars-Sinai Medical Center and externally evaluated in a geographically distinct cohort of 915 studies (46 890 videos) from Stanford Healthcare.

RESULTS:

In the held-out Cedars-Sinai Medical Center test set, the view classifier demonstrated an area under the curve (AUC) of 0.998 (0.998–0.999) and correctly identified 3452 of 3539 echocardiography videos as having color Doppler information across the mitral valve (sensitivity of 0.975 [0.968–0.982] and specificity of 0.999 [0.999–0.999] compared with manually curated videos). In the external test cohort from Stanford Healthcare, the view classifier correctly identified 1051 of 1055 manually curated videos with color Doppler information across the mitral valve (sensitivity of 0.996 [0.990–1.000] and specificity of 0.999 [0.999–0.999]). In the Cedars-Sinai Medical Center test cohort, MR moderate or greater in severity was detected with an AUC of 0.916 (0.899–0.932) and severe MR was detected with an AUC of 0.934 (0.913–0.953). In the Stanford Healthcare test cohort, the model detected MR moderate or greater in severity with an AUC of 0.951 (0.924–0.973) and severe MR with an AUC of 0.969 (0.946–0.987).

CONCLUSIONS:

In this study, a novel automated pipeline for identifying clinically significant MR from full transthoracic echocardiography studies demonstrated excellent performance across large numbers of studies and across multiple institutions. Such an approach has the potential for automated screening and surveillance of MR.

artificial intelligence
deep learning
echocardiography
mitral valve insufficiency
OPEN-ACCESSTRUE
SDCT
==== Body
pmcClinical Perspective

What Is New?

We developed and validated an automated pipeline for detection of mitral regurgitation from transthoracic echocardiogram studies.

This pipeline automatically screens for appropriate apical 4-chamber view videos with color Doppler on the mitral valve and then assesses mitral regurgitation severity.

What Are the Clinical Implications?

The deep learning model proposed could aid in the screening of mitral regurgitation.

This deep learning model, and others like it, could facilitate review of institutional databases, or expand access for screening in low-resource settings.

Editorial, see p 934

Mitral regurgitation (MR) is one of the most common forms of valvular heart disease, affecting >4 million Americans.1–4 Often progressing insidiously and frequently underrecognized,4 both primary MR as well as secondary MR can be initially asymptomatic, but lead to worsening heart failure and death.2,5–9 There has been an increased focus on early MR diagnosis given advances in surgical and transcatheter treatment options.5,10,11 Echocardiography with color Doppler is the most common method of initial evaluation of MR, with a holistic assessment combining left atrial size, effective regurgitant orifice area, regurgitant fraction, regurgitant volume, as well as other key clinical factors to assess disease severity accurately.12,13 Although ultrasound technology has become more widely available, accurate assessment of MR still requires experienced expert image acquisition and evaluation.

Recent advances in machine learning offer opportunities to automate time-consuming steps in the interpretation of medical imaging. Artificial intelligence (AI) has the ability to precisely phenotype subtle cardiac physiology as well as identify imaging features of disease severity not recognized by clinicians.14–17 Deep learning has been applied to echocardiography to improve the precision of common measurements, such as left ventricular ejection fraction (LVEF)14 and wall thickness,16,18 as well as streamlining assessment of aortic stenosis,19 hypertrophic cardiomyopathy,16 and cardiac amyloidosis.16,20 With increased ultrasound availability, AI guidance has been developed for both image acquisition and interpretation.14,21 Considering the increasing prevalence of MR in an aging population with comorbid heart failure, AI could aid in MR screening and surveillance.22–26

In this study, we developed and evaluated performance of a deep learning pipeline in automating identification of MR from standard transthoracic echocardiogram studies. We hypothesized that a deep learning approach can identify and assess MR severity on color Doppler apical 4-chamber (A4C) view videos with high-throughput automation. The automated pipeline was evaluated in 2 geographically distinct cohorts (Figure 1). Such an approach can be used for serial surveillance and screening of MR with a focus on changes over time.

Figure 1. Computer Vision–Based Mitral Regurgitation Detection. An automated deep learning pipeline was trained to detect and stratify mitral regurgitation (MR) severity using large-scale data consisting of apical 4-chamber echocardiogram videos with color Doppler across the mitral valve (Cedars-Sinai Medical Center). The automated pipeline showed strong and consistent performance in test sets at Cedars-Sinai Medical Center and Stanford Healthcare. These results show that deep learning can accurately detect clinically significant MR using single-view transthoracic echocardiogram videos with Doppler information. Deep learning–based MR detection tools could serve as a part of point-of-care ultrasound screening as part of clinic visits or in resource-limited settings where imaging may be obtained by individuals with minimal training.

METHODS

Study Population and Data Source

Cedars-Sinai Medical Center Cohort

A total of 58 614 transthoracic echocardiograms from 38 461 patients receiving care at Cedars-Sinai Medical Center (CSMC) between October 11, 2011, and June 4, 2022, were used to train and evaluate the deep learning models. A total of 2 587 538 videos (an average of 44 videos per study after excluding still images) were initially sourced from DICOM (Digital Imaging and Communications in Medicine) files and underwent de-identification, view classification, and preprocessing into AVI (Audio Video Interleave) videos (as described in the Supplemental Material). A total of 354 117 videos were classified as A4C view videos using an automated view classifier and then manually curated to identify 34 714 videos with color Doppler across the mitral valve.27

After view selection, the CSMC cohort included 34 714 videos across 30 453 unique echocardiograms from 22 661 patients, and a subset enriched for moderate and severe MR was used for training. All examples of moderate and severe MR were used for training; mild MR and control videos were sampled to be approximately the same number as moderate MR videos to mitigate class imbalance when training the MR severity model. A total of 20 604 videos from 18 133 studies in the data set were randomly split on a patient level into train (80%), validation (10%), and test (10%) cohorts to train a deep neural network for MR severity classification (Figure 2). MR severity for each study was determined on the basis of the clinical echocardiogram report determined in a high-volume echocardiography laboratory in accordance with American Society of Echocardiography guidelines.28 When MR was characterized as an intermediate category (“trace to mild,” “mild to moderate,” or “moderate to severe”), videos were placed in the more severe categories. For example, “mild to moderate MR” was considered moderate MR, and “moderate to severe MR” was considered severe MR for model training. Both primary and secondary MR were included. Studies with concomitant mitral stenosis, prosthetic valves, and reduced as well as normal LVEF were also included in both training and validation data sets. For subgroup analysis of characteristics not frequently present in the test cohort, additional validation cohorts enriched for previous mitral valve intervention, concomitant moderate or severe mitral stenosis, cases of eccentric MR, and cases with clinician quantification were obtained and used for additional testing. Patients from the MR severity model training and validation cohort were excluded from all test cohorts.

Figure 2. Cedars-Sinai Medical Center and Stanford Data Set Isolation. A total of 34 714 color Doppler apical 4-chamber (A4C) videos were isolated from a larger set of videos from Cedars-Sinai Medical Center. A view classifier was trained and used to isolate A4C mitral Doppler videos from 915 studies containing 1055 suitable videos from Stanford Healthcare. The mitral regurgitation (MR) classification model was then benchmarked on an internal test set from Cedars-Sinai Medical Center and an external test set from Stanford Healthcare.

Stanford Healthcare Cohort

The pipeline was evaluated on 915 studies (containing a total of 46 890 videos) from the Stanford Healthcare (SHC) high-volume academic echocardiography laboratory. The automated view classification pipeline was compared with manual curation of videos within those studies to evaluate specificity. All videos identified by the view classifier were used for downstream MR severity model validation. Model output was compared with MR severity determined by expert cardiologists from the clinical reports. This study was approved by the institutional review boards at CSMC and SHC. The need for informed consent was waived because the study involved secondary analysis of existing data.

AI Model Training

Deep learning models were trained using the PyTorch Lightning deep learning framework. When patients had multiple echocardiograms, each video was considered an independent example during training, with care not to have patient overlap across training, validation, and test cohorts. Video-based convolutional neural networks (R2+1D) were used for view classification and MR severity assessment.29 This model architecture was previously used for other echocardiography tasks and shown to be effective.18 The models were initialized with random weights and trained using a binary cross entropy loss function for up to 100 epochs, using an ADAM optimizer, an initial learning rate of 1e-2, and a batch size of 24 on 2 NVIDIA RTX 3090 graphics processing units. Early stopping was performed on the basis of the validation loss.

The view classifier was trained using the 34 714 manually curated videos with color Doppler across the mitral valve as cases and 49 263 other A4C videos as controls. Controls were a combination of videos that did not have color Doppler or had color Doppler window not focused on the mitral valve (videos focused on the tricuspid valve, intra-atrial septum, or ventricular septum). The MR severity model was trained on 6206 videos without MR, 6128 videos with mild MR, 6174 videos with moderate MR, and 2042 videos with severe MR. This process is summarized in Figure 2.

Statistical Analysis

Model performance was evaluated using area under the receiver operating characteristic curve and confusion matrices. F1 score, recall (sensitivity), positive predictive value, and negative predictive value (NPV) were also evaluated for both greater than moderate MR and severe MR. During external validation, the view classifier and MR classifier were evaluated serially as an automated pipeline. Statistical analysis was performed in Python (version 3.8.0) and R (version 4.2.2). Confidence intervals were computed by bootstrapping with 10 000 samples. Reporting of study results is consistent with guidelines put forth by CONSORT-AI (Consolidated Standards of Reporting Trials–Artificial Intelligence; Supplemental Material).30,31

Subgroup analysis was conducted to assess model performance in patients with different pathogeneses of MR, different ranges of LVEF, study characteristics, associated comorbidities, and other clinical characteristics. Primary MR was identified by study reports with findings predominantly highlighting mitral valve prolapse, rheumatic changes of the mitral valve, flail leaflet, or eccentric MR jet. Secondary MR was defined as moderately or severely reduced LVEF and no identifiers of primary MR. All other videos were considered to be of unclear or mixed MR pathogenesis. Clinical characteristics were obtained from the electronic health record or associated echocardiography report. Echocardiogram study quality was determined by clinicians and extracted from the clinical report. Studies in which clinicians commented on technical difficulty, with poor study quality, or in which ≥1 major cardiac structures (eg, left ventricle, right ventricle, pulmonary artery) were not well visualized were classified as technically difficult.

Model Explainability

The key imaging features identified by the MR severity model were evaluated using saliency mapping generated using the integrated gradients method.32 This method generated a heatmap for every frame of the video, summarized as a final 2-dimensional heatmap generated by using the maximum value along the temporal axis for each pixel location in the video. Pixels brighter in intensity and closer to yellow were more salient to model predictions; darker pixels were less important to the model’s final prediction. When assessing videos with no MR, heatmaps were obtained by taking the maximum of saliency maps for the moderate and severe class output neurons for each pixel location (Figure 4). Correlation between model predictions and quantitative metrics of MR severity was examined in cases with moderate or severe MR showing model prediction severity correlated with more extreme metrics (Table S6).

Serial Monitoring of MR Severity

The performance of the MR severity model in monitoring patients longitudinally and detecting changes in MR across serial studies was evaluated. All patients in the test cohort with >1 echocardiogram with a mitral A4C Doppler video were included in the analysis. This cohort consisted of 3538 echocardiograms across 768 patients, with patients having anywhere from 2 to 25 serial echocardiogram studies. Predicted MR severity across serial studies was compared with cardiologist-determined MR severity across serial studies. Sankey plots showcasing MR severity across time in cardiologist-determined severity and model-predicted severity across serial studies were compared (Figure 5).

Data and Code Availability

The data set of videos and reports used to train EchoNet-MR is not publicly available because of its potentially identifiable nature; it is available for use with approval by the CSMC institutional review board. Our code and model weights are available at https://github.com/echonet/MR.

RESULTS

Study Population

A total of 58 614 transthoracic echocardiograms from 38 461 patients were used to train the deep learning pipeline. From 2 587 538 initial videos, a total of 354 117 videos were identified as A4C view and subsequently manually curated to identify videos that had color Doppler across the mitral valve. The manually curated color Doppler videos were used to train a view-classification model and linked with clinician reports to train the MR severity model. Patient characteristics are presented in Tables 1 and 2 and are representative of the general CSMC and SHC patient populations that received echocardiograms. The data were split on patient level for training and validation and had similar patient age, LVEF, left atrial volume index, and proportions of male sex, coronary artery disease, and atrial fibrillation.

Table 1. View Classifier Demographic Characteristics

Table 2. Mitral Regurgitation Cohort Demographic Characteristics

View Classifier Performance Across 2 Institutions

On a test set of 3109 transthoracic echocardiograms (132 767 videos) from CSMC not seen during model training, the view classifier identified 3452 videos (97.5% of manually identified cases) with an AUC of 0.998 (0.998–0.999). When the Youden Index was used as a threshold, videos with color Doppler across the mitral valve were identified with a sensitivity of 0.975 (0.968–0.982) and specificity of 0.999 (0.999–0.999). To evaluate generalization of the view classification model at a geographically distinct site, we evaluated its performance on 915 studies from SHC. The view classifier isolated 1091 videos from a total of 46 890 videos, and manual review identified 1055 videos with color Doppler across the mitral valve. The view classifier correctly identified 1051 (99.6%) of manually curated videos, with 4 videos not found by the AI pipeline and 40 false positives. This corresponds to a sensitivity of 0.996 (0.990–1.000) and specificity of 0.999 (0.999–0.999).

MR Severity Performance Across 2 Institutions

The MR severity model showed strong performance in distinguishing MR severity and identifying clinically significant MR (Figure 3). In the internal CSMC test set not used during model training, the model demonstrated an AUC of 0.916 (0.899–0.932) in detecting at least moderate MR and an AUC of 0.934 (0.913–0.953) for severe MR. The AI model had an NPV of 0.954 (0.940–0.967) for severe MR and an NPV of 0.863 (0.835–0.890) for at least moderate MR. Further information on MR model performance is presented in Table 3. The MR severity model performance was similar across institutions. In the SHC cohort, the model identified severe MR with an AUC of 0.969 (0.946–0.987) and moderate or severe MR (defined as at least moderate MR) with an AUC of 0.951 (0.924–0.973). In this cohort, the model had an NPV of 0.977 (0.962–0.990) for severe MR and an NPV of 0.986 (0.974–0.995) for moderate or greater MR. The potential for serial surveillance and evaluation of change over time is shown in the CSMC studies among patients with ≥4 echocardiogram studies, showing progression of MR over time in a select population identified by the AI model similar to rate of progression identified by cardiologists (Figure 5).

Figure 3. Model Performance Across Severity and Institution. A, Receiver operating characteristic curves for detection of severe or moderate or greater mitral regurgitation (MR) at Cedars-Sinai Medical Center (CSMC) and Stanford Healthcare (SHC). Moderate includes moderate, moderate to severe, and severe MR. B and C, MR classification on test set videos from CSMC and SHC, respectively. Confusion matrix color map values were scaled on the basis of the proportion of actual disease cases in each class that were predicted in each possible disease category. This was done to allow for relative comparison of model performance across disease classes (none, mild, moderate, and severe) given class imbalance. AUC indicates area under the curve.

Figure 4. Saliency Map Visualization for Mitral Regurgitation Classification Models. Echocardiogram videos with severe mitral regurgitation (MR) from Cedars-Sinai Medical Center (CSMC; top left) and Stanford Healthcare (SHC; bottom left) are shown on the left; videos with no MR from CSMC (top right) and SHC (bottom right) are shown on the right. Saliency maps were computed using the integrated gradients method. A final 2-dimensional heatmap was generated by using the maximum value along the temporal axis for each pixel location in the video. Brighter and closer to yellow pixels were more salient to model predictions; darker pixels were less important to the final prediction of the model. Severe MR was assessed by using the activation function for severe disease to generate a heatmap. When assessing controls, heatmaps were generated by stacking heatmaps for severe and moderate classes and taking the maximum between the 2 at each pixel location.

Figure 5. Tracking Mitral Regurgitation Severity Across Serial Studies. Sankey diagram of patients with ≥4 echocardiogram studies with severity of mitral regurgitation (MR) tracked over time as evaluated by clinicians and by artificial intelligence model. Progression and change demonstrated by flow diagrams moving across MR severity.

Table 3. Model Performance

Subgroup Analysis

The MR severity model showed strong performance across test set subgroups (Table 4). Moderate or more severe primary MR was detected with an AUC of 0.913 (0.877–0.943), and severe primary MR was detected with an AUC of 0.982 (0.970–0.992). Moderate or more severe secondary MR was detected with an AUC of 0.952 (0.930–0.970), and severe secondary MR was detected with an AUC of 0.975 (0.956–0.990). Similar performance was observed for moderate or more severe MR of unclear or mixed pathogenesis (AUC, 0.903 [0.881–0.923]) and severe MR of unclear or mixed pathogenesis (AUC, 0.939 [0.909–0.965]). The model performed well across the range of LVEF, preexisting atrial fibrillation, mitral annular calcification, left atrial dilation, previous mitral valve intervention (surgical repair or surgical replacement), or concomitant mitral stenosis or aortic valve disease. AUCs for moderate or severe MR ranged from 0.850 (0.807–0.973) to 0.899 (0.816–0.964) in these groups; AUCs for severe MR ranged from 0.868 (0.808–0.920) to 0.914 (0.871–0.949). AUCs were similar irrespective of study quality and acquisition date as well as eccentric MR jets.

Table 4. Mitral Regurgitation Model Subset Analysis

Comparison With Cardiac Magnetic Resonance Imaging

We identified an additional held-out test population of 1085 videos from 831 transthoracic echocardiograms across 569 CSMC patients who underwent cardiac magnetic resonance imaging (MRI) and echocardiography within 180 days. The MRI assessment of MR severity was then used to evaluate our deep learning model. Cardiac MRI was performed in patients with varying degrees of MR, and studies where patients had moderate or severe MR on the basis of echocardiographic imaging comprised 146 studies (17.5%) in the cohort (Table S2). Of these echocardiograms with moderate or severe MR, 23.2% exhibited primary MR, 33.1% had secondary MR, and 43.7% had MR of unclear pathogenesis, showing no clear bias for a given pathogenesis (Table S3). In this cohort, our AI model had an AUC of 0.805 (0.717–0.884) in detecting at least moderate MR and an AUC of 0.828 (0.688–0.940) for severe MR (Table S4). When comparing the cardiologist’s echocardiography-based assessment of MR severity with an MRI-based assessment, cardiologists had an AUC of 0.833 (0.586–0.937) in detecting at least moderate MR and an AUC of 0.791 (0.753–0.902) for severe MR (Table S5). This difference in AUC between AI and cardiologist models when compared with MRI was not significant for at least moderate MR (DeLong test; P=0.295) or severe MR (DeLong test; P=0.556).

Model Interpretation

Saliency maps for our model demonstrate that the model focuses on the clinically relevant imaging features of MR. Saliency maps from integrated gradients were used to identify regions of interest in each video contributing the most to detection of MR severity (Figure 4).32 These interpretability techniques demonstrated localization of the activation signal in the color Doppler window and primarily highlighting the MR jet, indicating that the model used appropriate, physiologic features of MR to make predictions. Frame-by-frame saliency visualizations are shown in Videos S1 through S4. When comparing of model predictions with quantitative MR metrics, predicted severe MR cases had more extreme measurements than moderate MR cases across all quantitative metrics (Table S6).

DISCUSSION

We developed and validated an automated pipeline for detection and assessment of MR severity with echocardiography. From a full transthoracic echocardiogram, the algorithm automatically identifies appropriate A4C videos with color Doppler on the mitral valve and then assesses MR severity. For both severe MR as well as at least moderate MR, the model demonstrated strong performance (>0.916 AUC and >0.863 NPV). This automated workflow worked in unselected external validation studies without preselection or exclusion of other concomitant comorbidities and performed similarly for both primary and secondary MR.

Our algorithm learns features of MR that generalize across variability in imaging practices in 2 geographically distinct sites. Many previous echocardiography AI models have focused on standard 2-dimensional black-and-white B-mode images and videos.14,16,19,33 Meanwhile, our study focuses on the AI assessment of color Doppler videos and used a video-based model for the incorporation of rich temporal and Doppler information, both crucial for accurate MR assessment. Incorporation of color Doppler information greatly expands the opportunities for AI in echocardiography, particularly in valve disease. Previous works have applied deep learning to color Doppler echocardiography videos to assess valvular heart disease and rheumatic heart disease.34,35 Similar to image-based B-mode models mentioned previously, these approaches rely on image-based deep learning models for frame selection, localization of the left atrium, and jet segmentation to yield a probability of MR from either A4C video frames or a combination of A4C and parasternal long-axis videos. Meanwhile, the current approach relies on a state-of-the-art video-based architecture to directly detect and stratify MR from A4C videos, which are automatically isolated from entire echocardiogram studies. This novel approach allows for the use of information across multiple frames while reducing the need for multiple deep learning models for intermediate steps (eg, frame selection or jet segmentation). The current work also includes more representative training data (including both mild and moderate cases as well as severe MR cases during training) and used more training data (>10 000 videos of MR used for training in our model compared with 777 and 282 training cases, respectively, in previous work).

For comprehensive expert clinical evaluation of MR, a variety of views, including color Doppler, as well as Doppler metrics and calculations, are synthesized together for a holistic assessment of MR severity. In this work, our current AI model focuses on color Doppler videos in the A4C view, which might be insufficient to visualize eccentric or off-axis MR jets. Despite this limitation, our AI algorithm results in concordant interpretations with the comprehensive clinical approach, suggesting there is overlapping information or ancillary information such as left atrial size or mitral valve motion that can assist an AI model in evaluating MR severity even when there is a limited MR jet. In clinical practice, cardiologists often integrate conflicting calculations of MR severity, and a provider’s clinical gestalt for a particular severity of MR might correlate with hidden findings in the A4C view as well as jet area on color Doppler. Regardless, this clinical context suggests that future work synthesizing information across multiple echocardiographic views will improve the performance of future MR models.

The current work builds upon previous work in the space of echocardiography and AI. Several recent works have reported strides in computer vision and echocardiography, including automated view classification,36,37 phenotyping of left ventricular hypertrophy,16 assessment of left ventricular systolic function,14 aortic stenosis risk stratification, and detection of complex congenital heart defects.38 Previous work in machine learning applied to MR has primarily focused on structured data and non-deep learning approaches. The combination of our algorithm with previously published work using AI to guide novices in acquiring imaging could increase access to screening of MR.21,39 Future work could focus on automatically quantifying measures such as valve leaflet thickness, effective regurgitant orifice area, regurgitant volume, and fraction.

We introduce a model to screen for and stratify MR severity from transthoracic echocardiogram videos. To do so, we provide a workflow for isolating mitral valve color Doppler videos and automation of MR severity assessment. The models were evaluated to have good performance in internal and external test cohorts. The use of such a deep learning pipeline, with a high AUC, NPV, and generalizability across sites, could aid in the preliminary assessment of MR, facilitate review of institutional databases, or expand access for screening in primary care or low-resource settings.

Study Limitations

Although promising, the current work has limitations. Echocardiographic assessment of MR depends on appropriate images being obtained, with different views potentially maximizing the visualized MR jet. This algorithm would not overcome incomplete input information and insufficient images that would result in underestimation of MR. The model was trained on clinical assessment of MR, which did not always have comprehensive quantitative assessment (Table S7), although model output correlated with quantitative assessments when available. When compared with MRI assessment of MR severity, our AI model was reduced in performance compared with the comparison with echocardiographic assessment of MR severity. In parallel, the cardiologists’ assessment of echocardiographic MR severity was no better in concordance with MRI findings. Given the recognized discordance between clinician assessment of MR severity across modalities40 and the fact that our AI model was trained on labels from cardiologists’ echocardiographic assessments, the modest reduction in performance when compared across modalities should be expected.

ARTICLE INFORMATION

Acknowledgments

Dr Vrudhula acknowledges support from the Sarnoff Cardiovascular Research Award. Dr Cheng acknowledges support from the Erika J Glazer Family Foundation.

Sources of Funding

At the time of this work, Dr Vrudhula was a research fellow supported by the Sarnoff Cardiovascular Research Award. Dr Cheng acknowledges support from the Erika J. Glazer Family Foundation, NIH R01-HL131532 and NIH R01-HL142983.Dr Ouyang reports research grants NIH R00-HL157421, NIH R01-HL173526 and NIH R01-HL173487.

Disclosures

Dr Cheng reports consulting fees from UCB and Viz.ai. Dr Ouyang reports research support from AstraZeneca Alexion, as well as consulting fees from EchoIQ, Ultromics, Pfizer, and InVision. The other authors report no conflicts of interest.

Supplemental Material

Expanded Methods

Tables S1–S7

Figures S1–S2

Videos S1–S4

Supplemental Video Legends

CONSORT AI Checklist

Supplementary Material

Nonstandard Abbreviations and Acronyms

A4C apical 4-chamber

AI artificial intelligence

AUC area under the curve

AVI Audio Video Interleave

CONSORT-AI Consolidated Standards of Reporting Trials–Artificial Intelligence

CSMC Cedars-Sinai Medical Center

DICOM Digital Imaging and Communications in Medicine

LVEF left ventricular ejection fraction

MR mitral regurgitation

MRI magnetic resonance imaging

NPV negative predictive value

SHC Stanford Healthcare

Supplemental Material, the podcast, and transcript are available with this article at https://www.ahajournals.org/doi/suppl/10.1161/CIRCULATIONAHA.124.069047.

For Sources of Funding and Disclosures, see page 932.

Circulation is available at www.ahajournals.org/journal/circ
==== Refs
REFERENCES

1. Yadgir S Johnson CO Aboyans V Adebayo OM Adedoyin RA Afarideh M Alahdab F Alashi A Alipour V Arabloo J ; Global Burden of Disease Study 2017 Nonrheumatic Valve Disease Collaborators. Global, regional, and national burden of calcific aortic valve and degenerative mitral valve diseases, 1990–2017. Circulation. 2020;141 :1670–1680. doi: 10.1161/CIRCULATIONAHA.119.043391 32223336
2. Messika-Zeitoun D Candolfi P Vahanian A Chan V Burwash IG Philippon J-F Toussaint J-M Verta P Feldman TE Iung B . Dismal outcomes and high societal burden of mitral valve regurgitation in France in the recent era: a nationwide perspective. J Am Heart Assoc. 2020;9 :e016086. doi: 10.1161/JAHA.120.016086 32696692
3. Harb SC Griffin BP . Mitral valve disease: a comprehensive review. Curr Cardiol Rep. 2017;19 :73. doi: 10.1007/s11886-017-0883-5 28688022
4. Dziadzko V Clavel M-A Dziadzko M Medina-Inojosa JR Michelena H Maalouf J Nkomo V Thapa P Enriquez-Sarano M . Outcome and undertreatment of mitral regurgitation: a community cohort study. Lancet. 2018;391 :960–969. doi: 10.1016/S0140-6736(18)30473-2 29536860
5. Enriquez-Sarano M Akins CW Vahanian A . Mitral regurgitation. Lancet. 2009;373 :1382–1394. doi: 10.1016/S0140-6736(09)60692-9 19356795
6. Otto CM Verrier ED . Mitral regurgitation: what is best for my patient? N Engl J Med. 2011;364 :1462–1463. doi: 10.1056/NEJMe1102013 21463152
7. Del Forno B De Bonis M Agricola E Melillo F Schiavi D Castiglioni A Montorfano M Alfieri O . Mitral valve regurgitation: a disease with a wide spectrum of therapeutic options. Nat Rev Cardiol. 2020;17 :807–827. doi: 10.1038/s41569-020-0395-7 32601465
8. Ocher R May M Labin J Shah J Horwich T Watson KE Yang EH Marcella A . Mitral regurgitation in female patients: sex differences and disparities. Catheter Cardiovasc Interv. 2023;2 :101032. doi: 10.1016/j.jscai.2023.101032
9. Simpson TF Kumar K Samhan A Khan O Khan K Strehler K Fishbein S Wagner L Sotelo M Chadderdon S . Clinical predictors of mortality in patients with moderate to severe mitral regurgitation. Am J Med. 2022;135 :380–385.e3. doi: 10.1016/j.amjmed.2021.09.004 34648779
10. Tribouilloy CM Enriquez-Sarano M Schaff HV Orszulak TA Bailey KR Tajik AJ Frye RL . Impact of preoperative symptoms on survival after surgical correction of organic mitral regurgitation: rationale for optimizing surgical indications. Circulation. 1999;99 :400–405. doi: 10.1161/01.cir.99.3.400 9918527
11. David TE Ivanov J Armstrong S Rakowski H . Late outcomes of mitral valve repair for floppy valves: implications for asymptomatic patients. J Thorac Cardiovasc Surg. 2003;125 :1143–1152. doi: 10.1067/mtc.2003.406 12771888
12. Fadel BM Bakarman H Dahdouh Z Di Salvo G Mohty D . Spectral Doppler interrogation of mitral regurgitation: spot diagnosis. Echocardiography. 2015;32 :1179–1183. doi: 10.1111/echo.12891 25611451
13. Hagendorff A Knebel F Helfen A Stöbe S Haghi D Ruf T Lavall D Knierim J Altiok E Brandt R . Echocardiographic assessment of mitral regurgitation: discussion of practical and methodologic aspects of severity quantification to improve diagnostic conclusiveness. Clin Res Cardiol. 2021;110 :1704–1733. doi: 10.1007/s00392-021-01841-y 33839933
14. Ouyang D He B Ghorbani A Yuan N Ebinger J Langlotz CP Heidenreich PA Harrington RA Liang DH Ashley EA . Video-based AI for beat-to-beat assessment of cardiac function. Nature. 2020;580 :252–256. doi: 10.1038/s41586-020-2145-8 32269341
15. Elias P Poterucha TJ Rajaram V Moller LM Rodriguez V Bhave S Hahn RT Tison G Abreau SA Barrios J . Deep learning electrocardiographic analysis for detection of left-sided valvular heart disease. J Am Coll Cardiol. 2022;80 :613–626. doi: 10.1016/j.jacc.2022.05.029 35926935
16. Duffy G Cheng PP Yuan N He B Kwan AC Shun-Shin MJ Alexander KM Ebinger J Lungren MP Rader F . High-throughput precision phenotyping of left ventricular hypertrophy with cardiovascular deep learning. JAMA Cardiol. 2022;7 :386–395. doi: 10.1001/jamacardio.2021.6059 35195663
17. He B Kwan AC Cho JH Yuan N Pollick C Shiota T Ebinger J Bello NA Wei J Josan K . Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. 2023;616 :520–524. doi: 10.1038/s41586-023-05947-3 37020027
18. Soto JT Weston Hughes J Sanchez PA Perez M Ouyang D Ashley EA . Multimodal deep learning enhances diagnostic precision in left ventricular hypertrophy. Eur Heart J Digit Health. 2022;3 :380–389. doi: 10.1093/ehjdh/ztac033 36712167
19. Holste G Oikonomou EK Mortazavi BJ Coppi A Faridi KF Miller EJ Forrest JK McNamara RL Ohno-Machado L Yuan N . Severe aortic stenosis detection by deep learning applied to echocardiography. Eur Heart J. 2023;44 :4592–4604. doi: 10.1093/eurheartj/ehad456 37611002
20. Vrudhula A Stern L Cheng PC Ricchiuto P Daluwatte C Witteles R Patel J Ouyang D . Impact of case and control selection on training artificial intelligence screening of cardiac amyloidosis. JACC Adv. 2024;100998 :100998. doi: 10.1016/j.jacadv.2024.100998
21. Narang A Bae R Hong H Thomas Y Surette S Cadieu C Chaudhry A Martin RP McCarthy PM Rubenson DS . Utility of a deep-learning algorithm to guide novices to acquire echocardiograms for limited diagnostic use. JAMA Cardiol. 2021;6 :624–632. doi: 10.1001/jamacardio.2021.0185 33599681
22. Nkomo VT Gardin JM Skelton TN Gottdiener JS Scott CG Enriquez-Sarano M . Burden of valvular heart diseases: a population-based study. Lancet. 2006;368 :1005–1011. doi: 10.1016/S0140-6736(06)69208-8 16980116
23. Martin A-C Bories M-C Tence N Baudinaud P Pechmajou L Puscas T Marijon E Achouh P Karam N . Epidemiology, pathophysiology, and management of native atrioventricular valve regurgitation in heart failure patients. Front Cardiovasc Med. 2021;8 :713658. doi: 10.3389/fcvm.2021.713658 34760937
24. Avierinos J-F Tribouilloy C Grigioni F Suri R Barbieri A Michelena HI Ionico T Rusinaru D Ansaldi S Habib G ; Mitral Regurgitation International Database (MIDA) Investigators. Impact of ageing on presentation and outcome of mitral regurgitation due to flail leaflet: a multicentre international study. Eur Heart J. 2013;34 :2600–2609. doi: 10.1093/eurheartj/eht250 23853072
25. Grave C Tribouilloy C Tuppin P Weill A Gabet A Juillière Y Cinaud A Olié V . Fourteen-year temporal trends in patients hospitalized for mitral regurgitation: the increasing burden of mitral valve prolapse in men. J Clin Med Res. 2022;11 :3289. doi: 10.3390/jcm11123289
26. Mitchell E Walker R . Global ageing: successes, challenges and opportunities. Br J Hosp Med (Lond). 2020;81 :1–9. doi: 10.12968/hmed.2019.0377
27. Zhang J Gajjala S Agrawal P Tison GH Hallock LA Beussink-Nelson L Lassen MH Fan E Aras MA Jordan C . Fully automated echocardiogram interpretation in clinical practice. Circulation. 2018;138 :1623–1635. doi: 10.1161/CIRCULATIONAHA.118.034338 30354459
28. Zoghbi WA Adams D Bonow RO Enriquez-Sarano M Foster E Grayburn PA Hahn RT Han Y Hung J Lang RM . Recommendations for noninvasive evaluation of native valvular regurgitation. J Am Soc Echocardiogr. 2017;30 :303–371. doi: 10.1016/j.echo.2017.01.007 28314623
29. Tran D Wang H Torresani L Ray J LeCun Y Paluri M . A closer look at spatiotemporal convolutions for action recognition. Paper presented at: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2018; Salt Lake City, UT. doi: 10.1109/CVPR.2018.00675
30. Selvaraju RR Cogswell M Das A Vedantam R Parikh D Batra D . Grad-CAM: visual explanations from deep networks via gradient-based localization. Int J Comput Vis. 2020;128 :336–359. doi: 10.1007/s11263-019-01228-7
31. Saito T Rehmsmeier M . The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10 :e0118432. doi: 10.1371/journal.pone.0118432 25738806
32. Sundararajan M Taly A Yan Q . Axiomatic attribution for deep networks. Proceedings of the 34th International Conference on Machine Learning; 2017; Sydney, Australia. doi: 10.48550/arXiv.1703.01365
33. Yuan N Jain I Rattehalli N He B Pollick C Liang D Heidenreich P Zou J Cheng S Ouyang D . Systematic quantification of sources of variation in ejection fraction calculation using deep learning. JACC Cardiovasc Imaging. 2021;14 :2260–2262. doi: 10.1016/j.jcmg.2021.06.018 34274282
34. Yang F Chen X Lin X Chen X Wang W Liu B Li Y Pu H Zhang L Huang D . Automated analysis of Doppler echocardiographic videos as a screening tool for valvular heart diseases. JACC Cardiovasc Imaging. 2022;15 :551–563. doi: 10.1016/j.jcmg.2021.08.015 34801459
35. Brown K Roshanitabrizi P Rwebembera J Okello E Beaton A Linguraru MG Sable CA . Using artificial intelligence for rheumatic heart disease detection by echocardiography: Focus on mitral regurgitation. J Am Heart Assoc. 2024;13 :e031257. doi: 10.1161/JAHA.123.031257 38226515
36. Madani A Arnaout R Mofrad M Arnaout R . Fast and accurate view classification of echocardiograms using deep learning. NPJ Digit Med. 2018;1 :6. doi: 10.1038/s41746-017-0013-1 30828647
37. Steffner KR Christensen M Gill G Bowdish M Rhee J Kumaresan A He B Zou J Ouyang D . Deep learning for transesophageal echocardiography view classification. Sci Rep. 2024;14 :11. doi: 10.1038/s41598-023-50735-8 38167849
38. Arnaout R Curran L Zhao Y Levine JC Chinn E Moon-Grady AJ . An ensemble of neural networks provides expert-level prenatal detection of complex congenital heart disease. Nat Med. 2021;27 :882–891. doi: 10.1038/s41591-021-01342-5 33990806
39. Chiu I-M Lin C-HR Yau F-FF Cheng F-J Pan H-Y Lin X-H Cheng C-Y . Use of a deep-learning algorithm to guide novices in performing focused assessment with sonography in trauma. JAMA Netw Open. 2023;6 :e235102. doi: 10.1001/jamanetworkopen.2023.5102 36976564
40. Uretsky S Gillam L Lang R Chaudhry FA Argulian E Supariwala A Gurram S Jain K Subero M Jang JJ . Discordance between echocardiography and MRI in the assessment of mitral regurgitation severity: a prospective multicenter trial. J Am Coll Cardiol. 2015;65 :1078–1088. doi: 10.1016/j.jacc.2014.12.047 25790878
