
==== Front
9918590689006676
52479
Glob Pediatr
Glob Pediatr
Global pediatrics
2667-0097

39301448
10.1016/j.gpeds.2024.100220
nihpa2022313
Article
Leveraging machine learning to study how temperament scores predict pre-term birth status
Seamon Erich a
Mattera Jennifer.A. b
Keim Sarah.A. c
Leerkes Esther.M. d
Rennels Jennifer. L. e
Kayl Andrea.J. e
Kulhanek Kirsty.M. e
Narvaez Darcia f
Sanborn Sarah.M. g
Grandits Jennifer.B. g
Schetter Christine Dunkel h
Coussons-Read Mary i
Tarullo Amanda.R. j
Schoppe-Sullivan Sarah.J. k
Thomason Moriah.E. l
Braungart-Rieker Julie.M. m
Lumeng Julie. C. n
Lenze Shannon.N. o
Christian Lisa M. p
Saxbe Darby.E. q
Stroud Laura.R. r
Rodriguez Christina.M. s
Anzman-Frasca Stephanie t
Gartstein Maria.A. b*
a University of Idaho Department of Design and Environments, 875 Perimeter Drive MS 2481, Moscow, Idaho 83844-2481, United States
b Washington State University, Department of Psychology, P.O. Box 644820, Pullman WA 99164-4820, United States
c Nationwide Children’s Hospital & The Ohio State University, Center for Biobehavioral Health, Abigail Wexner Research Institute 700 Children’s Drive, Columbus OH 43205, United States
d University of North Carolina Greensboro, P.O. Box 26170, Greensboro NC 27402-6170, United States
e University of Nevada, Las Vegas, 4505 S. Maryland Way, Las Vegas, NV 89154, United States
f University of Notre Dame, 390 Corbett, Notre Dame IN 46556, United States
g Clemson University, College of Behavioral, Social and Health Sciences, 116 Edwards Hall, Clemson South Carolina 29634, United States
h University of California, Los Angeles, Department of Psychology, 1285 Franz Hall, Box 951563, Los Angeles CA 90095, United States
i University of Colorado-Colorado Springs Psychology Department, Columbine Hall, 1420 Austin Bluffs Pkwy, Colorado Springs CO 80918, United States
j Boston University, Department of Psychological & Brain Sciences 64 Cummington Mall, Room 149 Boston, Massachusetts 02215, United States
k The Ohio State University, 243 Psychology Building, 1835 Neil Ave, Columbus OH, 43210, United States
l New York University, Langone One Park Ave, New York, NY 10016, United States
m Colorado State University, Human Development and Family Studies, College of Health and Human Sciences, 1570 Campus Delivery, Fort Collins, CO, 80523-1501, United States
n University of Michigan Medical School, Division of Developmental and Behavioral Pediatrics, 1600 Huron Parkway, Building 520, Ann Arbor, Michigan, 48109, United States
o Washington University School of Medicine Institute for Public Health, 660 S. Euclid, MSC 8217-0094-02, St. Louis MO 63110, United States
p The Ohio State University Wexner Medical Center, 460 Medical Center Drive, Columbus, OH 43210, United States
q University of Southern California, 3616 Trousdale Parkway, AHF 108, Los Angeles, CA 90089-0376, United States
r Department of Psychiatry and Human Behavior Warren Alpert Medical School, Brown University, Coro West, Suite 309, 164 Summit Avenue, Providence, RI 02906, United States
s Old Dominion University, 115 Hampton Blvd, Norfolk, VA 23529, United States
t University at Buffalo Jacobs School of Medicine and Biomedical Sciences Division of Behavioral Medicine, G56 Farber Hall, 3435 Main Street, Buffalo New York 14214, United States
* Corresponding author. gartstma@wsu.edu (Maria.A. Gartstein).
13 9 2024
9 2024
22 7 2024
19 9 2024
9 100220https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Background:

Preterm birth (birth at <37 completed weeks gestation) is a significant public heatlh concern worldwide. Important health, and developmental consequences of preterm birth include altered temperament development, with greater dysregulation and distress proneness.

Aims:

The present study leveraged advanced quantitative techniques, namely machine learning approaches, to discern the contribution of narrowly defined and broadband temperament dimensions to birth status classification (full-term vs. preterm). Along with contributing to the literature addressing temperament of infants born preterm, the present study serves as a methodological demonstration of these innovative statistical techniques.

Study design:

This study represents a metanalysis conducted with multiple samples (N = 19) including preterm (n = 201) children and (n = 402) born at term, with data combined across investigations to perform classification analyses.

Subjects:

Participants included infants born preterm and term-born comparison children, either matched on chronological age or age adjusted for prematurity.

Outcome measures:

Infant Behavior Questionnaire-Revised Very Short Form (IBQ-R VSF) was completed by mothers, with factor and item-level data considered herein.

Results and conclusions:

Accuracy estimates were generally similar regardless of the comparison groups. Results indicated a slightly higher accuracy and efficiency for IBQR-VSF item-based models vs. factor-level models. Divergent patterns of feature importance (i.e., the extent to which a factor/item contributed to classification) were observed for the two comparison groups (chronological age vs. adjusted age) using factor-level scores; however, itemized models indicated that the two most critical items were associated with effortful control and negative emotionality regardless of comparison group.

Preterm birth
Infancy
Temperament
Quantitative methodology
==== Body
pmcIntroduction

Preterm birth is considered to be birth at <37 completed weeks of gestation1 Among the roughly 4 million new births each year in the United States, approximately 12 % of infants are born preterm,2 and worldwide the annual rate of preterm birth is 10.6% (14.84 million).3 The earlier an infant is born, the more underdeveloped or medically compromised the child is likely to be, with preterm status conferring risk via biological (e.g., immature neurobehavioral profile) and contextual pathways. For example, neonatal intensive care unit hospitalizations involve a variety of stressful experiences, such as pain exposure and separation from parents/caregivers. It is thus perhaps not surprising that many preterm infants experience regulation difficulties, and prematurity is associated with increased negative emotionality and behavior problems.4 A recent meta-analysis5 demonstrated a pattern of higher Activity level, lower Attentional Focusing and Attention Span/Persistence for infants born preterm relative to their full-term counterparts, concluding these infants exhibit less regulated temperament than those born full term.

Multiple temperament theories or frameworks have been proposed, and Rothbart’s psychobiological model is generally viewed as the most widely accepted at this time. According to this approach, temperament represents a set of constitutionally based individual differences in reactivity and self-regulation, with “constitutional” referring to the relatively enduring biological make-up of the individual, influenced by heredity, maturation, and experience.6 Within this temperament framework, reactivity encompasses arousability of emotional, motor, and attentional responses, typically assessed by threshold, latency, intensity, time to peak intensity, and reaction time recovery. Self-regulation refers to processes that can serve to modulate reactivity, such as soothability and inhibitory control. Structurally, temperament has often been delineated into three overarching factors in infancy: Negative Emotionality, Positive Affectivity/Surgency, and Regulatory Capacity/Orienting,7 with parent-report serving as the primary means of assessing infant temperament. Parent-report measures provide descriptors of child temperament across time and situations, not just a “snapshot” of reactivity and/or regulation gleaned from brief observations. Their widespread use is also a function of ease of administration, scoring, and accessibility. Importantly, this utility of parent-report has been further enhanced by more recent additions of brief formats, such as the Infant Behavior Questionnaire-Revised Very Short Form (IBQ-R VSF) included in this investigation.8

In this secondary analysis, we leveraged IBQ-R VSF data collected across 19 laboratories (N = 603) to further investigate differences between infants born preterm and their full-term counterparts, addressing yet unanswered questions. Given previously identified significant differences among these groups with respect to temperament development,5 algorithmic modeling techniques were used to discern the extent to which the IBQ-R VSF factor and item scores (referred to as features) accurately classified participating children as preterm (n = 201) or full-term (n = 402). This study addresses an important gap in research, being the first to determine the ability of temperament attributes to reliably differentiate preterm vs. full-term status, quantifying the extent to which early reactivity and regulation provide the features necessary for accurate classification. Importantly, two different comparison full-term groups were considered – infants matched on chronological age (in weeks) and those matched based on age adjusted for prematurity. This effort provides a new direction for research addressing the development of infants born preterm, serving as a methodological demonstration of machine learning applications not yet utilized in these areas of scientific inquiry. This meta-analytic data driven effort is the first to rely on advanced machine learning techniques using temperament features to classify infants into born preterm and full-term groups, rather than compare temperament of children who were born pre- and full-term. Implementation of multiple machine learning techniques provides for robust group classification/prediction and feature importance (relative contribution of predictors) which can accommodate a large number of predictor variables. This cross-laboratory effort also overcomes prior limitations associated with small samples that were not representative, producing results with superior generalizability.

Material and methods

Participants in all studies providing data for the present meta-analytic effort completed the IBQR-VSF.8 The IBQR-VSF includes 37 items (Supplemental Tables S1–S3) which have been shown to form three broadband scales including Positive Affectivity/Surgency (13 items; e.g., “How often during the week did your baby move quickly toward new objects”), Negative Affectivity (12 items; e.g., “When tired, how often did your baby show distress”), and Effortful Control (12 items; e.g., “Play with one toy for 5 to 10 min”). For each item, caregivers rate the extent to which their child engaged in the target behavior during the past 7 days on a scale from 1 (Never) to 7 (Always). In case certain infant responses were not observed in the past week (e.g., encountering an unfamiliar adult), not applicable is a response option. In prior research, the IBQR-VSF has demonstrated adequate internal consistency, test-retest reliability, and inter-rater agreement.8,9

Data sets (N = 19) were acquired by emailing researchers who requested the IBQ-R or published research using the instrument between 2006 and 2019. Contributors were asked to provide item level data from the IBQ-R as well as infant age, sex, and race. Participants included infants born preterm as well as term-born comparison children, either matched based on chronological age or age adjusted for prematurity. For all participants, the IBQ-R scores considered herein were based on mother-report. See Supplemental Table S4 for a brief description of the samples.

Descriptive statistics across different term status groups were computed first (Supplemental Table S5). We then constructed a model framework to assess the utility of IBQR-VSF temperament factors and individual items with respect to term status classifications, resulting in a total of 4 models. Specifically, we considered 2 temperament score types (factor and item-level) and 2 comparison groups, based on: (1) chronological age; and (2) prematurity adjusted age, matched within 2 weeks. Classification of term status groups based on both factor scores and individual items allowed us to determine if accuracy is improved by a more comprehensive representation of infant temperament or through a data reduction/composite creation resulting in factor scores (Supplemental Figure S1).

Established machine learning techniques, methodologically rigorous and shown to provide reliable/reproducible results, were used in this study.10,11 Specifically, for all models, we used repeated 10-fold cross-validation partitioning with random assignment: a training dataset including 70 % of the sample, and 30 % reserved as a hold-out dataset (testing) to evaluate the predictive utility of the trained models. A total of 11 different algorithms were considered for each model type, including: (1) linear discriminant analysis; (2) generalized linear modeling; (3) support vector machines; (4) K-nearest neighbor; (5) naïve Bayes; (6) classification and regression trees; (7) C5.0 classification; (8) bootstrapped aggregated trees; (9) ensembled decision trees (Random Forest12,13;); (10) gradient boosting; and (11) multi-class adaptive boosting (AdaBoost). These algorithms were chosen based on their applicability and widespread use in the classification modeling literature,12,13 and to achieve the most robust and replicable results discernable across multiple modeling techniques. The aforementioned models were then compared to discern the most effective classification of infant term status with temperament features based on misclassification rates, Cohen’s kappa coefficients, and sensitivity and specificity via the area under the curve (AUC) from Receiver Operator Curves (ROC), considered as indicators of model accuracy.

Misclassification provides a simplistic posterior assessment of model classification based on contingency tables and is often used for initial classification and model accuracy evaluation. Accuracy indicators, reported herein, represent the inverse of misclassification rates. Cohen’s kappa coefficient assesses reliability of categorization, which incorporates chance agreement, is normalized, and can range from −1 to 1. Kappa values will typically be lower than overall misclassification indictors, as it represents a more conservative estimate given its assessment of accuracy compared to random assignment. The area under a ROC curve (AUC) is a third metric used to evaluate the accuracy of binary classifiers, which encapsulates both Type I and Type II errors.14 However, ROC-AUC is limited insofar as it does not take predicted probability values and goodness-of-fit of evaluated models into account. While all three indicators (misclassification rates, Kappa coefficients, AUC) provide unique assessments of classification accuracy, overall misclassification rate (or, inversely, accuracy) is the most broadly used metric for classification evaluation.15 For all of the model classification indices, higher values (i.e., closer to 1) can be considered indicative of more optimal performance. Feature importance was also considered at factor and item level, enabling us to discern which specific temperament dimensions were most prominent contributors to classification.

Results

Accuracy estimates were generally similar regardless of comparison group. Model results (model 1 - chronological age, itemized; model 2 – adjusted age itemized; model 3 – chronological age factorized; model 4 – adjusted age factorized) indicated a slightly higher accuracy for IBQR-VSF item-based models vs. factor-level models, with a roughly 10 % increase in accuracy/R2 estimates (model 1 maximum R2 =0.644; model 2 maximum R2 =0.632; model 3 maximum R2 =0.552; model 4 maximum R2 =0.581). Area under the curve (AUC) scores, which provide an assessment of model efficiency, also indicated superior model performance for item-based models (model 1 maximum AUC =0.646; model 2 maximum AUC =0.623; model 3 maximum AUC =0.564; model 4 maximum AUC =0.642) (Supplemental Figures S14 - S17). Itemized vs. factorized AUC scores were similar for three of the four model approaches, with itemized chronological age vs. pre-terms being the only model with notable model performance differentiation (Fig. 1). Similarly, Kappa coefficients suggested greater accuracy for itemized models in comparison to factorized (Supplemental figures 1,3,5, and 7). When examining feature importance rankings (mean decrease in impurity) for both chronological age and adjusted age models, distress proneness/negative affectivity was a prominent contributor to model performance at the item level. Effortful control and positive affectivity/surgency were associated with minimal model importance (<10 %) for preterm vs. adjusted age, while preterm vs. chronological age indicated more secondary influences from effortful control (Fig. 2). Divergent patterns of feature importance were observed for the two comparison groups (chronological age vs. adjusted age) using factor-level scores (Supplemental figures 9 and 10). Overall, itemized model feature importance analyses provided a more nuanced picture of meaningful contributions to term status classification. While there were several surgency-associated itemized variables that emerged as important, the top two most critical items were associated with effortful control and negative affect, for both preterm vs. chronological age as well as preterm vs. adjusted age (Fig. 2).

Discussion

Consistent with existing studies demonstrating significant differences in temperament between infants born pre-term and their full-term counterparts,4,5 the present study leveraging machine learning techniques demonstrated that temperament-based models provided the basis for classification into term status groups (i.e., pre-term vs. full-term, matched by both chronological age and adjusted age). Prediction accuracy was generally similar across the two comparison groups, yielding comparable estimates. In this study, both factor-level and item-level features were considered to ascertain whether or not frequently undertaken data reduction (i.e., forming factor or composite scores) is optimal with respect to classification. Alternatively, we considered the possibility that items possess classification utility not captured by their average/factorized representations. Itemized models outperformed those based on factorized scores, especially in the context of pre-term vs. chronologically matched comparison classification.

Feature importance analyses are of particular interest in terms of making connections to the existing literature, and providing recommendations for future directions. These analyses were conducted at the factor and item level, with somewhat different patterns of results emerging for these variable types and depending on the comparison sample examined. Specifically, factor level results indicated that Positive Affectivity and Effortful Control were associated with greater importance in group differentiation, although the primary contribution varied depending on the comparison group: Effortful Control was more critical in the comparison to chronologically matched peers, whereas Positive affectivity emerged as most significant when comparing pre-term infants and an age-adjusted comparison group. At the item level, effortful control and negative affect scores were most influential in classification for both comparisons performed (preterm vs. chronological age; preterm vs. adjusted age).

Supplementary Material

1

Acknowledgments

The following funding sources contributed to data collection only: Sarah A. Keim, Nationwide Children’s Hospital & The Ohio State University

Data used for the purposes of this study was collected with grant support from the U.S. Health Resources and Services Administration [R40MC28316], the March of Dimes [12-FY14-171], the Allen Foundation, Cures Within Reach, the National Center for Advancing Translational Sciences/National Institutes of Health [UL1TR001070], Eunice Kennedy Shriver National Institute of Child Health and Human Development [R01HD100493], and internal funds from The Research Institute at Nationwide Children’s Hospital.

Esther M. Leerkes, University of North Carolina Greensboro

Data collected and utilized for this publication was funded by the Eunice Kennedy Shriver National Institute for Child Health and Human Development [R01HD058578].

Jennifer L. Rennels, University of Nevada, Las Vegas

This work was funded by a National Science Foundation Award [BCS-1148049].

Christine Dunkel Schetter and Mary Coussons-Read, University of California, Los Angeles

Research reported in this publication was supported by the Eunice Kennedy Shriver National Institute for Child Health and Human Development [R01HD073491] and the National Institute of Drug Abuse Grant [F31DA051181].

Sarah J. Schoppe-Sullivan, The Ohio State University

The New Parents Project was funded by the National Science Foundation [CAREER 0746548, Schoppe-Sullivan], with additional support from the Eunice Kennedy Shriver National Institute of Child Health and Human Development [NICHD; 1K01HD056238], and The Ohio State University’s Institute for Population Research [NICHD; P2CHD058484].

Moriah E. Thomason, New York University

Data used for the purposes in this publication was collected with grant support from the National Institute of Drug Abuse Grant [NIDA U01DA055338], National Institute of Mental Health [NIMH; R01MH122447, R01MH126468] and additional support from the National Institute of Environmental Health Sciences [NIEHS; R01ES032294].

Julie Lumeng, University of Michigan Medical School:Research reported in this publication was supported by the Eunice Kennedy Shriver National Institute for Child Health and Human Development (HD084163).

Julie M. Braungart-Rieker, Colorado State University

This study was supported by a grant from the Eunice Kennedy Shriver National Institute of Child Health and Human Development [NICHD; 5R03HD39802].

Shannon N. Lenze, Washington University School of Medicine Institute for Public Health

This data collection was supported by a grant from the National Institute of Mental Health [NIMH; K23090245].

Lisa M. Christian, The Ohio State University Wexner Medical Center

This data collection was supported by a grant from the Eunice Kennedy Shriver National Institute of Child Health and Human Development [NICHD; R21HD067670].

Laura R. Stroud, The Miriam Hospital & Brown University Alpert Medical School

Research was supported by grants from the National Institute on Drug Abuse [NIDA; 1R01DA031188, 5R01DA044504, 1R01DA056787], the National Institute of Environmental Health Sciences [NIEHS; U24ES028507] and the National Institute of General Medical Sciences [NIGMS; 1P20GM139767].

Stephanie Anzman-Frascat, University at Buffalo

The INSIGHT Study providing data for the present manuscript is supported by an award from the National Institute of Diabetes and Digestive and Kidney Diseases [NIDDK; R01DK088244

Fig. 1. a) R2 model comparisons for chronological age groups – itemized vs. factorized variables; b) area under the curve (AUC) model comparisons for chronological age groups – itemized vs. factorized variables; c) R2 model comparisons for age adjusted groups – itemized vs. factorized variables; d) area under the curve (AUC) model comparisons for age adjusted groups – itemized vs. factorized variables.

Fig. 2. Random forest itemized variable feature importance, relative influence in classification, for: a) infant born pre-term vs. the chronological age control group; b) infant born preterm vs. the adjusted age control group.

CRediT authorship contribution statement

Erich Seamon: Writing – original draft, Visualization, Software, Methodology, Formal analysis, Conceptualization. Jennifer.A. Mattera: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Sarah.A. Keim: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Esther.M. Leerkes: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Jennifer.L. Rennels: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Andrea.J. Kayl: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Kirsty.M. Kulhanek: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Darcia Narvaez: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Sarah.M. Sanborn: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Jennifer.B. Grandits: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Christine Dunkel Schetter: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Mary Coussons-Read: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Amanda.R. Tarullo: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Sarah.J. Schoppe-Sullivan: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Moriah.E. Thomason: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Julie.M. Braungart-Rieker: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Julie.C. Lumeng: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Shannon.N. Lenze: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Lisa M. Christian: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Darby.E. Saxbe: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Laura.R. Stroud: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Christina.M. Rodriguez: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Stephanie Anzman-Frasca: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation. Maria.A. Gartstein: Writing – review & editing, Resources, Investigation, Funding acquisition, Data curation.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Supplementary materials

Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.gpeds.2024.100220.
==== Refs
References

1. World Health Organization, 2012.
2. Martin Hamilton , Ventura Osterman , & Matthews , 2013.
3. Chawanpaiboon , 2019.
4. Hwang AW , Soong WT , Liao HF . Influences of biological risk at birth and temperament on development at toddler and preschool ages. Child Care Health Dev. 2009;35 (6 ):817–825.19702642
5. Cassiano RGM , Provenzi L , Linhares MBM , Gaspardo CM , Montirosso R . Does preterm birth affect child temperament? A meta-analytic study. Infant Behav Dev. 2020;58 , 101417. 10.1016/j.infbeh.2019.101417.31927307
6. Gartstein MA , Putnam SP , Aaron E , Rothbart M . Temperament and personality. editor. In: Maltzman S , ed. Oxford Handbook of Treatment Processes and Outcomes in Counseling Psychology. New York, NY: Oxford University Press; 2016:11–41.
7. Gartstein MA , Rothbart MK . Studying infant temperament via the revised infant behavior questionnaire. Infant Behav Dev. 2003;26 :64–86.
8. Putnam SP , Helbig AL , Gartstein MA , Rothbart MK , Leerkes E . Development and assessment of short and very short forms of the infant behavior questionnaire – revised. J Pers Assess. 2014;96 :445–458.24206185
9. Leerkes EM , Su J , Reboussin BA , Daniel SS , Payne CC , Grzywacz JG . Establishing the measurement invariance of the very short form of the infant behavior questionnaire revised for mothers who vary on race and poverty status. J Pers Assess. 2017;99 (1 ): 94–103. 10.1080/00223891.2016.1185612.27292626
10. Prasad AM , Iverson LR , Liaw A . Newer classification and regression tree techniques: bagging and random forests for ecological prediction. Ecosystems. 2006;9 (2 ):181–199.
11. Kotsiantis SB , Zaharakis ID , Pintelas PE . Supervised machine learning : a review of classification techniques general issues of supervised learning algorithms. Inform. 2007;31 :249–268.
12. Breiman L Bagging predictors. Mach Learn. 1996;24 (2 ):123–140.
13. Ho TK . Random decision forest. In: Proceedings of the 3rd International Conference on Document Analysis and Recognition. Montreal; 1995:278–282.
14. Ben-David S , Blitzer J , Crammer K , Kulesza A , Pereira F , Vaughan JW . A theory of learning from different domains. Mach Learn. 2010;79 (1–2 ):151–175.
15. Hanley JA , McNeil BJ . The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143 (1 ):29–36.7063747
