
==== Front
Perm J
tpj
tpj
The Permanente Journal
1552-5767
1552-5775
The Permanente Press

39219312
10.7812/TPP/23.139
TPJ-23-139
Original Research
Protocol for Designing a Model to Predict the Likelihood of Psychosis From Electronic Health Records Using Natural Language Processing and Machine Learning
http://orcid.org/0000-0001-9040-3724
Stavers-Sosa Icelini MD 1 2
http://orcid.org/0000-0002-2797-5102
Cronkite David J MS 3
Gerstley Lawrence D PhD 4
Kelley Ann MHA 3
Kiel Linda MA 3
http://orcid.org/0000-0002-6303-7308
Kline-Simon Andrea H MS 4
Marafino Ben J PhD 4
Ramaprasan Arvind MS 3
Carrell David S PhD 3
http://orcid.org/0000-0001-8654-1825
Hirschtritt Matthew E MD, MPH 1 2 4
1 Department of Psychiatry, Kaiser Permanente Oakland Medical Center, Oakland, CA, USA
2 Department of Psychiatry and Behavioral Sciences, University of California, San Francisco, San Francisco, CA, USA
3 Kaiser Permanente Washington Health Research Institute, Seattle, WA, USA
4 Division of Research, Kaiser Permanente Northern California, Oakland, CA, USA
Icelini Stavers-Sosa, MD icelini.x.stavers-sosa@kp.org
2024
02 9 2024
28 3 2336
© 2024 The Authors.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ Published by The Permanente Federation LLC under the terms of the CC BY-NC-ND 4.0 license https://creativecommons.org/licenses/by-nc-nd/4.0/.

Abstract

Introduction

Rapid identification of individuals developing a psychotic spectrum disorder (PSD) is crucial because untreated psychosis is associated with poor outcomes and decreased treatment response. Lack of recognition of early psychotic symptoms often delays diagnosis, further worsening these outcomes.

Methods

The proposed study is a cross-sectional, retrospective analysis of electronic health record data including clinician documentation and patient–clinician secure messages for patients aged 15–29 years with ≥ 1 primary care encounter between 2017 and 2019 within 2 Kaiser Permanente regions. Patients with new-onset PSD will be distinguished from those without a diagnosis if they have ≥ 1 PSD diagnosis within 12 months following the primary care encounter. The prediction model will be trained using a trisourced natural language processing feature extraction design and validated both within each region separately and in a modified combined sample.

Discussion

This proposed model leverages the strengths of the large volume of patient-specific data from an integrated electronic health record with natural language processing to identify patients at elevated chance of developing a PSD. This project carries the potential to reduce the duration of untreated psychosis and thereby improve long-term patient outcomes.
==== Body
pmcIntroduction

The incidence and prevalence of psychotic spectrum disorder (PSD) are understudied due to PSD’s nonspecific diagnostic category, compared to schizophrenia, for example. Psychotic disorders are estimated at a 2.5% to 3% lifetime prevalence in the adult population, representing a major public health concern.1 PSDs, most notably schizophrenia, are associated with substantial personal and societal burden.2 A shorter duration of untreated psychosis (DUP) correlates with a significantly higher remission rate over the course of the illness, such that DUP of 4 weeks predicts < 20% more severe symptoms at follow-up relative to a DUP of 1 week.3 DUP is an important prognostic factor; therefore, addressing psychotic symptoms sooner leads to improved outcomes pertaining to remission rate and preservation of brain volume.3–11

Over the last several years, various research teams have designed and tested automated prediction systems using electronic health record (EHR) data and patient-generated data such as speech and language to predict which individuals will develop psychosis among other conditions.12–20 However, delayed treatment remains a major challenge. To the authors’ knowledge, none of these prior efforts have used a machine learning approach applied to only EHR data to develop a psychosis risk prediction tool. The timely identification and treatment of early psychosis is crucial for improving patient outcomes. Such a risk-prediction tool could provide real-time risk assessment during the encounter to flag patients at potential risk for PSD and who may need additional assessment and intervention.

The rate at which patients with psychotic symptoms present to primary care is largely unknown. Primary care practitioners play a crucial role in the early identification and management of psychosis, as they often serve as the initial point of contact for individuals experiencing mental health concerns. Studies estimate that approximately 35% of first contacts prior to psychosis were to a primary care physician.21 Additionally, a shorter DUP is associated with more family doctor visits prior to diagnosis.22,23 Family doctors are therefore a vital part of the psychosis care pathway.

The proposed study details the methodology for developing a risk-prediction model to identify patients at elevated chance of developing a PSD via natural language processing (NLP) and machine learning by analyzing EHR data from an integrated health system.

Methods

Setting

Kaiser Permanente Northern California and Kaiser Permanente Washington are private, not-for-profit integrated health systems and are part of the nation’s largest Health Plan; Kaiser Permanente Northern California has 4.3 million members and Kaiser Permanente Washington has 640,000 members. Members receive coverage through employers, government programs (eg, Medicaid and Medicare), and individual plans. The demographic characteristics of Kaiser Permanente members resemble those of the general population within each of the Kaiser Permanente regions.24

Study design and cohort construction

The cohort included encounter-level data for patients potentially at risk for PSD where PSD was defined as having an F2* international classification of diseases (ICD-10-CM) diagnosis code.25 Screen-eligible encounters (SEEs) included all primary care encounters for eligible patients made between January 1, 2017, and December 31, 2019, where primary care was defined as encounters in any of the following departments: family practice, internal medicine, pediatrics, preventive medicine, primary care, and geriatrics. Eligible patients could have multiple SEEs during the study period (Figure 1). The index date was defined as the date of the SEE.

Figure 1: Study design: possible scenarios. NLP = natural language processing; PSD = psychotic spectrum disorder; SEE, screen-eligible encounter.

Patients were eligible for inclusion if they: 1) were age 15 to 29 years; 2) maintained 2 years of continuous enrollment prior to the SEE; and 3) did not have a history of PSD prior to the SEE (eg, there were no PSD diagnoses in the patient’s EHR prior to the SEE). PSD outcomes included any PSD diagnoses made in primary care, inpatient care, urgent care, emergency department, addiction medicine, or psychiatry departments within 12 months after the SEE. Patient demographics including age, sex, race, and ethnicity as well as mental health diagnoses in the 2 years prior through 1 year post each SEE were extracted from the EHR (Figure 2). Mental health diagnoses were categorized following the clinical classifications software–refined categories. All encounter notes made by physicians and patient messages sent to practitioners within each Health Plan during the 2 years prior to the SEE were also extracted from the EHR. For modeling the authors excluded any individual with no prior index dates or if the index date coincided with the index diagnoses.

Figure 2: Cohort construction. Rule assigment for model to identify cases and avoid noncases. PSD = psychotic spectrum disorder; Dx = diagnosis.

The final sample included 2,647,224 SEEs from 627,437 unique patients across both regions. This includes 5309 and 2,641,915 encounters from 1848 case encounters that included patients who received a PSD diagnosis during the 1 year post the SEE, and 625,589 control encounters that included patients who did not receive a PSD diagnosis during the study, respectively (Figure 3; Tables 1, 2, and 3).

Figure 3: Study sample. PSD = psychotic spectrum disorder; SEE = screen-eligible encounters.

Table 1: Demographic characteristics of patients included in model development and validation

Characteristic	Kaiser Permanente Northern California (n = 586,885)	Kaiser Permanente Washington (n = 40,552)	Total sample (n = 627,437)	
	Cases
(n = 1730)	Controls
(n = 585,155)	Total	Cases
(n = 118)	Controls
(n = 40,434)	Total	Cases
(n = 1848)	Controls
(n = 625,589)	Total	
Age at first SEE date	 	 	 	 	 	 	 	 	 	
Mean (SD)	20.9 (4.0)	21.1 (4.7)	21.1 (4.7)	20.5
(4.0)	21.1
(4.6)	21.1
(4.6)	20.9^	21.1^	21.1^	
Median (IQR)	20.0 (6.0)	21.0 (8.0)	21.0
(8.0)	20.0
(7.0)	21.0
(8.0)	21.0
(8.0)	n/a^	n/a^	n/a^	
Sex (no., %)	 	 	 	 	 	 	 	 	 	
Male	982
(56.8)	272,761
(46.6)	273,743
(46.6)	63
(53.4)	17,459
(43.2)	17,522
(43.2)	1045
(56.5)	290,220
(46.4)	291,265
(46.4)	
Female	746
(43.1)	312,305
(53.4)	313,051
(53.3)	55
(46.6)	22,974
(56.8)	23,029 (56.8)	801
(43.3)	335,279
(53.6)	336,080
(53.6)	
Other/unknown	2
(0.1)	89
( < 0.1)	91
( < 0.1)	0
(0)	1
( < 0.1)	1 ( < 0.1)	2
(0.1)	90
( < 0.1)	92
(0.0)	
Race/ethnicity (no., %)	 	 	 	 	 	 	 	 	 	
White	662
(38.3)	222,201
(38.0)	222,863
(38.0)	74
(62.7)	24,673
(61.2)	24,747 (61.0)	736 (39.8)	246,874
(39.5)	247,610
(39.5)	
Asian/Native Hawaiian/Pacific Islander	249
(14.4)	120,731
(20.6)	120,980
(20.6)	10
(8.5)	6060 (15.0)	6070
(15.0)	259 (14.0)	126,791
(20.3)	127,050
(20.2)	
Black	328
(19.0)	52,551 (9.0)	52,879
(9.0)	19
(16.1)	3886
(9.6)	3905
(9.6)	347 (18.8)	56,437
(9.0)	56,784 (9.1)	
Latino/Hispanic	432
(25.0)	162,695
(27.8)	163,127
(27.8)	6
(5.1)	2888
(7.1)	2894
(7.1)	438 (23.7)	165,583
(26.5)	166,021
(26.5)	
American Indian/Alaska Native	16
(0.9)	3689
(0.6)	3705
(0.6)	2
(1.7)	758
(1.9)	760
(1.9)	18
(1.0)	4456
(0.7)	4465 (0.7)	
a ^ Due to data sharing/privacy limitations patient-level data could not be combined across sites; therefore only total means are reported.

IQR, interquartile range; PSD, psychotic spectrum disorder; SD, standard deviation; SEE, screen-eligible encounter (primary care encounters between January 1, 2017, and December 31, 2019, among eligible patients).

Table 2: Non–psychotic spectrum disorder mental health and substance use diagnoses among patients included in model development and validation

Diagnostic group
(no., %)	Kaiser Permanente Northern California (n = 586,885)	Kaiser Permanente Washington (n = 40,552)	Total sample
(n = 627,437)	
	Cases(n = 1730)	Controls(n = 585,155)	Total	Cases
(n = 118)	Controls
(n = 40,434)	Total	Cases
(n = 1848)	Controls(n = 625,589)	Total	
Psychiatric conditions a										
Anxiety	1247
(72.1)	130,970
(22.4)	132,217
(22.5)	102
(86.4)	13,269
(32.8)	13,371
(33.0)	1349
(73.0)	144,239
(23.1)	145,588
(23.2)	
Mood	1398
(80.8)	93,288
(15.9)	94,686
(16.1)	100
(84.7)	11,375
(28.1)	11,475
(28.3)	1498
(81.1)	104,663
(16.7)	106,161
(16.9)	
Suicidal ideation/attempt/intentional self-harm	708
(40.9)	12,660
(2.2)	13,368
(2.3)	54
(45.8)	2040
(5.0)	2094
(5.2)	762
(41.2)	14,700
(2.3)	15,462
(2.5)	
Trauma- and stressor-related	753
(43.5)	77,465
(13.2)	78,218
(13.3)	59
(50.0)	6342
(15.7)	6401
(15.8)	812
(43.9)	83,807
(13.4)	84,619
(13.5)	
Neuro-developmental	560
(32.4)	48,854
(8.3)	49,414
(8.4)	53
(44.9)	5232
(12.9)	5285
(13.0)	613
(33.2)	54,086
(8.6)	54,699
(8.7)	
Personality	264
(15.3)	8814
(1.5)	9078
(1.5)	20
(16.9)	550
(1.4)	570
(1.4)	284
(15.4)	9364
(1.5)	9648
(1.5)	
Feeding and eating	101
(5.8)	7844
(1.3)	7945
(1.4)	8
(6.8)	703
(1.7)	711
(1.8)	109
(5.9)	8547
(1.4)	8656
(1.4)	
Other/unspecified	283
(16.4)	21,695
(3.7)	21,978
(3.7)	33
(28.0)	2248
(5.6)	2281
(5.6)	316
(17.1)	23,943
(3.8)	24,259
(3.9)	
Substance use conditions a										
Tobacco	489
(28.3)	31,423
(5.4)	31,912
(5.4)	31
(26.3)	2132
(5.3)	2163
(5.3)	520
(28.1)	33,555
(5.4)	34,075
(5.4)	
Alcohol	348
(20.1)	13,597
(2.3)	13,945
(2.4)	24
(20.3)	1405
(3.5)	1429
(3.5)	372
(20.1)	15,002
(2.4)	15,374
(2.5)	
Stimulant	293
(16.9)	4614
(0.8)	4907
(0.8)	21
(17.8)	436
(1.1)	457
(1.1)	314
(17.0)	5050
(0.8)	5364
(0.9)	
a Non-PSD mental health and substance use diagnoses were assessed during 1 year before and 1 year after the screen-eligible visit.

PSD = psychotic spectrum disorder;

Table 3: Characteristics of incident psychotic spectrum disorder cases

Characteristic	Kaiser Permanente Northern California
(n = 1730)	Kaiser Permanente Washington
(n = 118)	Total (n = 1848)	
Age at incident PSD diagnosis, y				
 Mean (SD) a	21.4 (4.0)	21.1 (4.0)	21.4	
 Median (IQR)	21.0 (6.0)	21.0 (6.0)	n/a	
Incident PSD diagnosis (no., %)				
 Schizophrenia (F20.0, F20.1, F20.2, F20.3, F20.5, F20.89, F20.9)	100 (5.8)	8 (6.8)	108 (5.8)	
 Delusional disorder (F22)	235 (13.6)	36 (30.5)	271 (14.7)	
 Brief psychotic disorder (F23)	163 (9.4)	10 (8.5)	173 (9.4)	
 Schizotypal disorder (F21)	17 (1.0)	0 (0.0)	17 (0.9)	
 Schizophreniform disorder (F20.81)	10 (0.6)	1 (0.8)	11 (0.6)	
 Schizoaffective disorder (F25.0, F25.1, F25.8, F25.9)	108 (6.2)	11 (9.3)	119 (6.4)	
 Unspecified psychosis (F29)	1067 (61.7)	51 (43.2)	1118 (60.5)	
 Other psychotic disorder (F24, F28)	30 (1.7)	1 (0.8)	31 (1.7)	
SEEs prior to incident PSD diagnosis (no., %) b				
 1	689 (39.8)	33 (28.0)	722 (39.1)	
 2	373 (21.6)	24 (20.3)	397 (21.5)	
 3	238 (13.8)	17 (14.4)	255 (13.8)	
 4	132 (7.6)	8 (6.8)	140 (7.6)	
 5	98 (5.7)	9 (7.6)	107 (5.8)	
 6	58 (3.4)	8 (6.8)	66 (3.6)	
 ≥ 7	142 (8.2)	19 (16.1)	161 (8.7)	
Duration between first SEE date and incident PSD diagnosis date, d				
 Mean (SD) a	210.1 (114.2)	203.7 (115.2)	209.7	
 Median (IQR)	233.0 (204.0)	211.0 (204.0)	n/a	
Note: Only PSD diagnoses made in primary care, ED, urgent care, hospital, mental health, and chemical dependency departments were included.

a Due to data sharing/privacy limitations patient-level data could not be combined across sites; therefore only total means are reported.

b Other encounters included laboratory-only, nonacute institutional stays, and other encounters.

ED, emergency department; IQR, interquartile range; PSD, psychotic spectrum disorder; SD, standard deviation; SEE, screen-eligible encounter (primary care encounters between January 1, 2017, and December 31, 2019, among eligible patients).

NLP feature engineering

The application of machine learning to EHR data has shown promise in identifying various medical conditions such as anaphylaxis,17,26 suicide,18 rheumatoid arthritis, and coronary artery disease.27 Utilizing text mining and NLP, these methods extract relevant information from EHRs.

The authors’ approach builds upon previous methods outlined in reports by Yu,26–28 Carrell,17 and Simon.18 The development of algorithms to identify specific medical conditions from EHRs involves 2 main approaches: manual construction, which relies on human expertise to select relevant features, and statistical or machine learning methods, which algorithmically infer relevant features from the data. The latter approach aims to automate the process of feature selection, reducing the need for labor-intensive manual curation, and identifying relevant features that would otherwise escape manual selection. A method called automated feature extraction for phenotyping (AFEP) automatically identifies features from publicly available medical knowledge sources using NLP and selects informative features for phenotype classification with data-driven screening using EHR data. AFEP aims to create accurate phenotyping algorithms without the need for manual feature selection. AFEP involves several steps, including concept collection from public knowledge resources, note parsing using NLP, data-driven concept screening to remove uninformative concepts, and model training using penalized logistic regression. The AFEP method holds promise for improving accuracy and efficiency in EHR-based clinical studies.

To explore an innovative approach in their methods the authors incorporated less commonly used feature engineering, combining unigrams and bigrams in addition to the keyword extraction from structured data representations (features) of unstructured text (clinical notes and messages). This information is typically represented as binary indicators (ie, presence/absence of any relevant concept mention) and/or counts (ie, the number of distinct mentions in a patient’s chart, optionally normalized by the quantity of chart text processed).

The team will use 3 complementary approaches to generate a large set of NLP features potentially relevant to predicting PSD risk, which include 1) AFEP, 2) unigrams and bigrams extracted from clinical notes, and 3) clinical expertise–based terms.

AFEP requires the identification and mining of publicly available clinical knowledge base articles on the topic of interest for relevant Unified Medical Language System concepts (eg, “psychotic disorders”) under the assumption that relevant keywords would appear in these sources. Each concept has a concept unique identifier (CUI; eg, “psychotic disorders” is represented by the CUI C0033975) and is associated with a number of synonymous terms and phrases (eg, “psychotic disorders” includes “psychosis,” “psychoses,” “psychotic disorder,” etc). Yu et al26–28 recommended using publicly available knowledge sources (ie, Wikipedia, Medscape, and Merck Manuals) but noted a possible limitation in that these sources may not contain the concepts of interest or concepts that are actually used within the clinical record. To address this concern, the field of knowledge sources was expanded to include UpToDate, MayoClinic.org, MedlinePlus.gov, Wikipedia, National Alliance on Mental Illness, and National Institute of Mental Health online publications when extracting concepts related to psychosis, psychosis risk, and first-episode psychosis. Additionally, the authors incorporated peer-reviewed papers that describe the most current and validated screening tools to identify individuals as ultra-high-risk for developing psychosis such as the Comprehensive Assessment of At-Risk Mental States29 and the Structured Interview for Psychosis-Risk Syndromes.30,31

Rather than requiring concepts to appear in at least half of the sources as defined by AFEP, the authors adopted a more inclusive approach and included all identified terms. These were then reviewed by clinical domain experts to exclude unrelated concepts for their final list of 1293 terms. MetaMapLite32 will be used to extract concepts from patient notes and secure messages using the 2022 Unified Medical Language System Level 0 + 4 + 9 dataset, and the output CUIs will be restricted to only those selected by the clinical domain experts.

The second set of features will come from unigrams and bigrams extracted directly from the clinical notes. Unigrams include all words (eg, “psychotic”) and bigrams include all possible 2-word strings constructed from neighboring pairs of words (eg, from the string “consider psychotic disorder” one can extract the bigrams “consider psychotic” and “psychotic disorder”). In processing these notes, stop words such as “the” or “and,” numbers, punctuation, and other symbols are removed.

The third NLP feature set will include clinical expertise–based terms believed to be useful for predicting PSD risk. A domain expert generated a list of terms derived from questions contained in validated screening tools to detect psychosis-risk states.33–40

The “raw” data produced by all 3 approaches were used to operationalize the NLP features (binary and count) for each term and concept, which were then used as predictor variables during model development (Figure 4).

Figure 4: Natural language processing feature extraction. CUI = concept unique identifier.

Machine learning model development and evaluation

Together with the NLP features detailed above, the team will also use variables extracted from the EHR, including diagnosis codes and patient characteristics, to develop and validate machine learning models to predict incident PSD diagnosis in the 12 months following the SEE. Two machine learning algorithms, gradient boosting and elastic net regression, will be used to develop a pair of prediction models, and their validation performance will then be compared.

Both gradient boosting and regularized regression (here, the elastic net), and gradient boosting in particular, represent the gold standard for prediction modeling on tabular data, ie, patient-level data without inherent higher-order structure unlike as in image or text data. Studies41 have shown that gradient boosting consistently outperforms even deep-learning algorithms on tabular data. Moreover, the team also opted to use regularized regression due to its ability to build parsimonious (and thus potentially more interpretable) models while still yielding good predictive performance.

The authors plan to develop a hybrid-data model using data from both regions present in the training and test datasets, rather than either set being derived from only 1 region, as will be the case during external validation. To ensure that no data containing protected health information are shared across sites, the team will undertake a stepwise model development process. Initially, the unigram and bigram variables will be excluded from the training dataset, retaining only the rich NLP variables based on the expert-curated list of terms and AFEP-generated CUIs. Kaiser Permanente Northern California will be the main site training the model on all data from both sites; Kaiser Permanente Washington will share this data from its cohort at the encounter level without sharing patient-level characteristics with Kaiser Permanente Northern California.

Finally, the authors will train Kaiser Permanente Northern California models on Kaiser Permanente Northern California data, including the rich NLP variables listed above combined with unigrams and bigrams. The team will identify variables selected by the models—that is, variables having a nonzero coefficient (for elastic net regression) or those present in the component trees (for gradient boosting). Simultaneously, this process repeats at Kaiser Permanente Washington with Kaiser Permanente Washington data. Next, Kaiser Permanente Washington will strip protected health information, if any, from all unigram and bigram variables selected by the Kaiser Permanente Washington–trained model before sharing those variables with Kaiser Permanente Northern California. Only a set of selected predictor variables is exchanged with the other region. Once Kaiser Permanente Northern California is aware which unigram and bigram variables from Kaiser Permanente Washington need to be added to its own set of unigrams and bigrams (ie, those selected by Kaiser Permanente Northern California and Kaiser Permanente Washington models), then Kaiser Permanente Northern California trains a new hybrid model using this set of variables on data for both sites.

For both algorithms, the team will perform 10-fold cross-validation within a training set (consisting of 80% of the data, drawn at random) to tune the hyperparameters for each algorithm (Figure 5). Then, both the trained gradient-boosted tree and elastic net regression models will be validated in the test set (consisting of the remaining 20% of the data) to estimate performance metrics. The team will also perform external validation by training a model using either exclusively Kaiser Permanente Northern California or Kaiser Permanente Washington data, then validating them at the other site. They will repeat this process using both sites as training sets.

Figure 5: Study flow.

The purpose of cross-validation is to tune the hyperparameters for each machine learning algorithm,42 thus guiding selection of the final fitted model. These hyperparameters control how each algorithm fits the data43; for example, for the elastic net, the key hyperparameters (alpha and lambda) control the overall strength of regularization as well as the weights placed on the L 1 and L 2 penalties.44 Gradient boosting requires a larger set of hyperparameters to be tuned, which places constraints on how the weak learners are built.45 If the model were trained on the same data used to select the hyperparameters (rather than as within a cross-validation step over the training data), this would potentially lead to overly optimistic estimates of out-of-sample performance. Performing cross-validation within the training set guards against this bias and thus is a recommended practice.46

The primary performance metrics of interest are the areas under the receiver operating characteristic47 and the precision-recall curves.48 In addition, the team will compute the sensitivity, specificity, positive predictive value, and negative predictive value of the risk score over the top 10%, 5%, 1%, and 0.5% of scores. Finally, to characterize potential impacts on clinician and system-wide workload, they also will compute the expected number of alerts for a range of plausible risk cutoffs, as well as the number needed to evaluate (ie, 1/positive predictive value) at each cutoff.

Sensitivity analyses

The model will also be trained and tested on 3 increasingly stringent PSD definitions: 1 PSD diagnosis during the 12 months post SEE, ≥ 2 PSD diagnoses, and 1 schizophrenia/schizoaffective disorder diagnosis (ie, specifically schizophrenia or schizoaffective disorder, not an unspecified psychotic disorder). Additionally, the authors plan to train the model to predict other (non-PSD) incident mental health diagnoses as a sensitivity analysis to determine whether the model is also able to accurately detect signals of these non-PSD diagnoses. The team will also compare the most influential features in the PSD models to those in the non-PSD models, allowing them to understand which terms occurring in notes are driving these predictions. The authors hypothesize that incident psychosis will have signals distinct from those of other incident mental health diagnoses, which would be important to confirm given the high degree of overlap between psychiatric diagnosis categories.

These analyses could assist in creating a graded alert system based on the estimated 12-month PSD risk at the SEE. Rather than providing a binary alert (at high risk/not at high risk), such a graded system could alert immediately (ie, at the SEE) for patients more likely to experience PSD, enabling opportunities to intervene. For patients at moderate and low (though still elevated) chance of PSD, increased monitoring without alerting (and thus without immediate intervention) may be appropriate. For patients in the moderate-risk category, whether to ultimately alert or not could be predicated on a rule such as the presence of a new rule-out diagnosis of PSD in the EHR.

This study was reviewed and approved by the Kaiser Permanente Northern California Institutional Review Board.

Discussion

The strengths of the proposed method lie in its passive data collection approach, utilizing EHR data that include all data from routine clinical care. By doing so, this method can identify patients who may not otherwise complete patient-reported assessment tools. Additionally, the proposed technique is automated, generating an initial risk score without the need for extensive clinical review. This method combines both free-text and structured data, enabling a more comprehensive analysis of clinical information compared to relying solely on structured data. Additionally, the authors believe that the patient-generated text collected through secure messages to practitioners adds nuance to the clinician-generated text. Specific use of words or expressions are part of the analyzed data, reflecting a patient-derived element. This is a unique strength of this model. This inclusive approach has the potential to enhance the sensitivity of psychosis-specific screening, improving the identification of patients who would benefit from early intervention. Moreover, this study will generate psychometric properties associated with different definitions of psychotic disorder onset. This information will empower health care systems to choose a definition that aligns with their specific clinical needs.

Another strength of the proposed method is the combined data from 2 different Kaiser Permanente regions. The authors found enough differences in the distribution of race and gender between Kaiser Permanente Northern California and Kaiser Permanente Washington to consider their combined data to be complementary. Moreover, care processes also likely substantially differ across regions, which is reflected in part by the number of encounters available to the models. Future development of the final model will be strengthened by external validation at non–Kaiser Permanente sites or with more broad-based data.

The authors anticipate that one feasible next step is to perform a prospective validation, where the model is ported to Kaiser Permanente HealthConnect (Kaiser Permanente’s EHR system) and runs in “silent mode” without being linked to an intervention. In this prospective validation, the team would track model predictions and compare them to actual outcomes to monitor model performance over time, including signs of any model drift or decrements in predictive performance, and to understand in which subpopulations or settings the model may perform better or worse and adjust accordingly. Depending on the results of this proposed model, the team may consider adding additional features, including encounter and secure-messaging frequency, which may be related to the intensity of the patient’s symptoms and/or relative level of impairment.

The attractiveness of this model is how it can be built on with evolving performance. In this first attempt, the authors are looking to build the most parsimonious version given the enormous analytic load, and fully acknowledge that it is not complete and may not include everything that may be contributory to psychosis risk. Although the authors acknowledge the possibility that their prediction models may not be able to accurately predict new-onset psychosis, such a finding may have multiple causes, among them potential heterogeneity in the data, which may be difficult to model. In that case, the model may not be powerful enough to distinguish signals of new-onset psychosis from background noise. The authors anticipate several possible solutions, including performing feature engineering (creating new, more expressive sets of patient features [predictors] based on the data) as well as model stacking or ensembling. The latter techniques fit multiple models to the same outcome and combine their predictions—a method that has been shown to reliably improve predictive performance over using a single model alone. In addition, the authors may also consider restricting data to subpopulations and/or settings (eg, certain types of visits) among which there may be less heterogeneity. Although this potential solution would limit the applicability of the model, it may also result in a more accurate model that will be tailored to the subgroup(s) of interest, and the findings at this stage (eg, observing that the model performs better in a more restricted setting compared to another) could be fed back into the overall model development process (eg, resulting in more precisely engineered features) to create a more accurate model designed to work across all subpopulations and care settings.

The interpretation and implementation of the proposed model require careful consideration of several factors. Firstly, the application of NLP in this technique is specific to the health systems in which it was developed. To ensure its validity and generalizability, the authors will validate the model by splitting data between training, testing, and validation sets. They will also test the extent to which the model can generalize to data from different Kaiser Permanente regions. However, it is important to recognize that other health systems may need to adapt the model to their specific patient populations. Moreover, even for a model developed and validated wholly within a Kaiser Permanente region, continuous reassessment of the NLP algorithm is necessary to account for potential changes in clinical documentation practices and other care processes over time.

Even within a large, integrated, and diverse health care system like Kaiser Permanente,24 patients who are less inclined to seek clinical care due to various factors such as transportation limitations, financial distress, or distrust of the medical system may be underrepresented in the data used to train the model. As a result, these systematic biases may lead to inaccurate PSD risk scores for these individuals even if they ultimately do have a SEE yielding a risk score. Considering these factors is crucial to ensure an accurate interpretation and effective implementation of the proposed model in clinical practice.

Future implementation of such a model will require careful planning and consideration of various factors to ensure its effectiveness and integration into clinical practice. Given the major stigma associated with mental health diagnoses,49 particularly psychotic disorders,50 it is crucial to approach the deployment of this model in the EHR system with caution. One important consideration in building predictive models is to anticipate and mitigate biases such as the type described above. The final stages of the model development will undergo evaluation for biases using a recently reported approach in preparation for deployment.51

During and following the implementation process, it will be crucial to elicit feedback from system leaders, primary care clinicians, mental health clinicians, and patients with recent onset of psychotic disorders and their family members. Conducting qualitative, semistructured interviews with these stakeholders can provide valuable insights into current diagnostic methods, comfort levels and skills in diagnosing psychotic disorders, and perceptions of automated screening techniques including those based on machine learning models. Stakeholder-informed ethical frameworks to guide implementation of prediction models derived from EHR are underway.52

After developing and testing the model, a controlled pilot should be conducted where regular feedback from clinicians will be actively solicited and used to modify the application of the model in the EHR. This iterative process ensures that the model aligns with clinical workflows and addresses the needs of health care practitioners. Synthesizing the results from feedback interviews will inform design and workflow considerations. Questions such as the frequency of the model running among eligible patients, conveying the relative chance of psychosis to clinicians, responsibility for responding to alerts, communication of alerts to patients, and conducting further psychosis-specific screening should be addressed. These considerations will guide the development of an effective workflow for implementing the model.

Incorporating a figure or graphic to illustrate the implementation process can enhance clarity and understanding. For example, a theoretical design could involve applying the model to all eligible patients every 30 days. Primary care clinicians would receive reminders in the EHR during clinical encounters with patients whose risk scores exceed a predetermined threshold. This prompts them to invite the patient to complete a brief psychosis screening tool, with the results prepopulating a clinician-to-clinician message. If the score indicates a high chance of a psychotic disorder, the mental health department can then reach out to the patient to schedule an appointment or provide appropriate follow-up care.

It is crucial to acknowledge that the implementation method, like the model itself, should be adjusted to accommodate the specific requirements of each clinical system. For instance, in more fragmented health care systems, primary care clinicians may need to directly communicate the elevated psychosis risk to the patient and their family, interpret the results, and make appropriate referrals to mental health practitioners. By considering these implementation factors, the model can be effectively integrated into clinical practice, ultimately improving the early detection and treatment of psychosis.

Conclusion

This proposed model leverages the strengths of the large volume of patient-specific data from an integrated EHR together with NLP and machine learning methods to identify patients at elevated chance of developing a PSD. This project carries the potential to reduce the DUP, and thereby to improve long-term patient outcomes.

Author Contributions: Icelini Stavers-Sosa, MD, was responsible for the conceptualization and writing of the original draft as well as reviewing and editing subsequent drafts of the manuscript. David J Cronkite, MS, Lawrence D Gerstley, PhD, Linda Kiel, MA, Andrea H Kline-Simon, MS, Ben J Marafino, PhD, Arvind Ramaprasan, MS, and David S Carrell, PhD, contributed to the data curation, methodology, and reviewing and editing of the manuscript. Andrea Kline-Simon and Arvind Ramaprasan additionally contributed to the data curation and formal analysis. Ann Kelley, MHA, contributed to the validation and methodology. Linda Kiel contributed to the project administration. Matthew E Hirschtritt, MD, MPH, contributed to the conceptualization, methodology, funding acquisition, supervision, and reviewing and editing of the manuscript. All authors have given final approval for the article.

Conflicts of Interest: None declared

Funding: Funding for this study was provided by the Kaiser Permanente Garfield Memorial Fund (principal investigator: Hirschtritt).

Data-Sharing Statement: Underlying data in this study are not publicly available.
==== Refs
References

1. Chang WC , Wong CSM , Chen EYH , et al. Lifetime prevalence and correlates of schizophrenia-spectrum, affective, and other non-affective psychotic disorders in the Chinese adult population. Schizophr Bull. 2017;43 (6 ):1280–1290. 10.1093/schbul/sbx056 28586480
2. Kotzeva A , Mittal D , Desai S , Judge D , Samanta K . Socioeconomic burden of schizophrenia: A targeted literature review of types of costs and associated drivers across 10 countries. J Med Econ. 2023;26 (1 ):70–83. 10.1080/13696998.2022.2157596 36503357
3. Howes OD , Whitehurst T , Shatalina E , et al. The clinical significance of duration of untreated psychosis: An umbrella review and random-effects meta-analysis. World Psychiatry. 2021;20 (1 ):75–95. 10.1002/wps.20822 33432766
4. Fraguas D , Del Rey-Mejías A , Moreno C , et al. Duration of untreated psychosis predicts functional and clinical outcome in children and adolescents with first-episode psychosis: A 2-year longitudinal study. Schizophr Res. 2014;152 (1 ):130–138. 10.1016/j.schres.2013.11.018 24332406
5. Breitborde NJK , Bell EK , Dawley D , et al. The Early Psychosis Intervention Center (EPICENTER): Development and six-month outcomes of an American first-episode psychosis clinical service. BMC Psychiatry. 2015;15 . 10.1186/s12888-015-0650-3
6. Morrison AP , Pyle M , Maughan D , et al. Antipsychotic medication versus psychological intervention versus a combination of both in adolescents with first-episode psychosis (MAPS): A multicentre, three-arm, randomised controlled pilot and feasibility study. Lancet Psychiatry. 2020;7 (9 ):788–800. 10.1016/S2215-0366(20)30248-0 32649925
7. Thomson A , Griffiths H , Fisher R , McCabe R , Abbott-Smith S , Schwannauer M . Treatment outcomes and associations in an adolescent-specific early intervention for psychosis service. Early Interv Psychiatry. 2019;13 (3 ):707–714. 10.1111/eip.12778 30690896
8. Bhullar G , Norman RMG , Klar N , Anderson KK . Untreated illness and recovery in clients of an early psychosis intervention program: A 10-year prospective cohort study. Soc Psychiatry Psychiatr Epidemiol. 2018;53 (2 ):171–182. 10.1007/s00127-017-1464-z 29188310
9. Tang JY-M , Chang W-C , Hui CL-M , et al. Prospective relationship between duration of untreated psychosis and 13-year clinical outcome: A first-episode psychosis study. Schizophr Res. 2014;153 (1–3 ):1–8. 10.1016/j.schres.2014.01.022 24529612
10. Ružić Baršić A , Rubeša G , Mance D , Miletić D , Gudelj L , Antulov R . Influence of psychotic episodes on grey matter volume changes in patients with schizophrenia. Psychiatr Danub. 2020;32 (3–4 ):359–366. 10.24869/psyd.2020.359 33370733
11. Balamurugan VP , Chew QH , Sim K . Relationship between volume reductions of hippocampal subfields and thalamus and duration of untreated psychosis in schizophrenia. Asian J Psychiatr. 2022;71 :103082. 10.1016/j.ajp.2022.103082 35299141
12. Cohen DJ , Keller SR , Hayes GR , Dorr DA , Ash JS , Sittig DF . Integrating patient-generated health data into clinical care settings or clinical decision-making: Lessons learned from project HealthDesign. JMIR Hum Factors. 2016;3 (2 ):e26. 10.2196/humanfactors.5919 27760726
13. Figueroa-Barra A , Del Aguila D , Cerda M , et al. Automatic language analysis identifies and predicts schizophrenia in first-episode of psychosis. Schizophrenia (Heidelb). 2022;8 (1 ). 10.1038/s41537-022-00259-3
14. Corcoran CM , Carrillo F , Fernández-Slezak D , et al. Prediction of psychosis across protocols and risk cohorts using automated language analysis. World Psychiatry. 2018;17 (1 ):67–75. 10.1002/wps.20491 29352548
15. Bedi G , Carrillo F , Cecchi GA , et al. Automated analysis of free speech predicts psychosis onset in high-risk youths. NPJ Schizophr. 2015;1 :15030. 10.1038/npjschz.2015.30 27336038
16. Mackinley M , Limongi R , Silva AM , et al. More than words: Speech production in first-episode psychosis predicts later social and vocational functioning. Front Psychiatry. 2023;14 . 10.3389/fpsyt.2023.1144281
17. Carrell DS , Gruber S , Floyd JS , et al. Improving methods of identifying anaphylaxis for medical product safety surveillance using natural language processing and machine learning. Am J Epidemiol. 2023;192 (2 ):283–295. 10.1093/aje/kwac182 36331289
18. Simon GE , Johnson E , Lawrence JM , et al. Predicting suicide attempts and suicide deaths following outpatient visits using electronic health records. AJP. 2018;175 (10 ):951–960. 10.1176/appi.ajp.2018.17101167
19. Irving J , Patel R , Oliver D , et al. Using natural language processing on electronic health records to enhance detection and prediction of psychosis risk. Schizophr Bull. 2021;47 (2 ):405–414. 10.1093/schbul/sbaa126 33025017
20. Sullivan SA , Kounali D , Morris R , et al. Developing and internally validating a prognostic model (p risk) to improve the prediction of psychosis in a primary care population using electronic health records: The MAPPED study. Schizophr Res. 2022;246 :241–249. 10.1016/j.schres.2022.06.031 35843156
21. Kennedy L , Johnson KA , Cheng J , Woodberry KA . A public health perspective on screening for psychosis within general practice clinics. Front Psychiatry. 2019;10 . 10.3389/fpsyt.2019.01025
22. Skeate A , Jackson C , Wood MB , Jones C . Duration of untreated psychosis and pathways to care in first-episode psychosis: investigation of help-seeking behaviour in primary care. Br J Psychiatry. 2002;181 (S43 ):s73–s77. 10.1192/bjp.181.43.s73
23. Addington J , Van Mastrigt S , Hutchinson J , Addington D . Pathways to care: Help seeking behaviour in first episode psychosis. Acta Psychiatr Scand. 2002;106 (5 ):358–364. 10.1034/j.1600-0447.2002.02004.x 12366470
24. Davis AC , Voelkel JL , Remmers CL , Adams JL , McGlynn EA . Comparing Kaiser Permanente members to the general population: Implications for generalizability of research. Perm J. 2023;27 (2 ):87–98. 10.7812/TPP/22.172 37170584
25. Clinical Classifications Software (CCS) for ICD-10-PCS (beta version) [Agency for Healthcare Research and Quality]. Accessed 17 April 2024. https://hcup-us.ahrq.gov/toolssoftware/ccs10/ccs10.jsp
26. Yu W , Zheng C , Xie F , et al. The use of natural language processing to identify vaccine-related anaphylaxis at five health care systems in the Vaccine Safety Datalink. Pharmacoepidemiol Drug Saf. 2020;29 (2 ):182–188. 10.1002/pds.4919 31797475
27. Yu S , Liao KP , Shaw SY , et al. Toward high-throughput phenotyping: Unbiased automated feature extraction and selection from knowledge sources. J Am Med Inform Assoc. 2015;22 (5 ):993–1000. 10.1093/jamia/ocv034 25929596
28. Yu S , Chakrabortty A , Liao KP , et al. Surrogate-assisted feature extraction for high-throughput phenotyping. J Am Med Inform Assoc. 2017;24 (e1 ):e143–e149. 10.1093/jamia/ocw135 27632993
29. Yung AR , Yuen HP , McGorry PD , et al. Mapping the onset of psychosis: The comprehensive assessment of at-risk mental states. Aust N Z J Psychiatry. 2005;39 (11–12 ):964–971. 10.1080/j.1440-1614.2005.01714.x 16343296
30. Miller TJ , McGlashan TH , Woods SW , et al. Symptom assessment in schizophrenic prodromal states. Psychiatr Q. 1999;70 (4 ):273–287. 10.1023/a:1022034115078 10587984
31. Woods SW , Parker S , Kerr MJ , et al. Development of the PSYCHS: Positive symptoms and diagnostic criteria for the CAARMS harmonized with the SIPS. medRxiv. 2023. 10.1101/2023.04.29.23289226
32. Demner-Fushman D , Rogers WJ , Aronson AR . MetaMap Lite: An evaluation of a new Java implementation of MetaMap. J Am Med Inform Assoc. 2017;24 (4 ):841–844. 10.1093/jamia/ocw177 28130331
33. Dean K , Murray RM . Environmental risk factors for psychosis. Dialogues Clin Neurosci. 2005;7 (1 ):69–80. 10.31887/DCNS.2005.7.1/kdean 16060597
34. Laurens KR , Luo L , Matheson SL , et al. Common or distinct pathways to psychosis? A systematic review of evidence from prospective studies for developmental risk factors and antecedents of the schizophrenia spectrum disorders and affective psychoses. BMC Psychiatry. 2015;15 . 10.1186/s12888-015-0562-2
35. Fusar-Poli P , Borgwardt S , Bechdolf A , et al. The psychosis high-risk state: A comprehensive state-of-the-art review. JAMA Psychiatry. 2013;70 (1 ):107–120. 10.1001/jamapsychiatry.2013.269 23165428
36. Wang P , Yan C-D , Dong X-J , et al. Identification and predictive analysis for participants at ultra-high risk of psychosis: A comparison of three psychometric diagnostic interviews. World J Clin Cases. 2022;10 (8 ):2420–2428. 10.12998/wjcc.v10.i8.2420 35434048
37. Oliver D , Arribas M , Radua J , et al. Prognostic accuracy and clinical utility of psychometric instruments for individuals at clinical high-risk of psychosis: A systematic review and meta-analysis. Mol Psychiatry. 2022;27 (9 ):3670–3678. 10.1038/s41380-022-01611-w 35665763
38. Vollmer-Larsen A , Handest P , Parnas J . Reliability of measuring anomalous experience: The Bonn Scale for the assessment of basic symptoms. Psychopathology. 2007;40 (5 ):345–348. 10.1159/000106311 17657133
39. Fusar-Poli P , Hobson R , Raduelli M , Balottin U . Reliability and validity of the comprehensive assessment of the at risk mental state, Italian version (CAARMS-I). Cur Pharm Des. 2012;18 (4 ):386–391. 10.2174/138161212799316118
40. Ciarleglio AJ , Brucato G , Masucci MD , et al. A predictive model for conversion to psychosis in clinical high-risk patients. Psychol Med. 2019;49 (7 ):1128–1137. 10.1017/S003329171800171X 29950184
41. Shwartz-Ziv R , Armon A . Tabular data: Deep learning is not all you need. 2021. Accessed 24 November 2021. https://arxiv.org/pdf/2106.03253
42. Berrar D. Cross-validation. In: Ranganathan S , Gribskov M , Nakai K , Schönbach C , eds. Encyclopedia of Bioinformatics and Computational Biology. Academic Press; 2019 :542-545. 10.1016/j.schres.2022.06.031
43. Hastie T , Tibshirani R , Friedman J . The Elements of Statistical Learning. 2nd ed. New York: Springer; 2009. 10.1007/978-0-387-84858-7
44. Zou H , Hastie T . Regularization and variable selection via the elastic net. J R Stat Soc Series B Stat Methodol. 2005;67 (2 ):301–320. 10.1111/j.1467-9868.2005.00503.x
45. de Lacy N , Ramshaw MJ , Kutz JN . Integrated evolutionary learning: An artificial intelligence approach to joint learning of features and hyperparameters for optimized, explainable machine learning. Front Artif Intell. 2022;5 . 10.3389/frai.2022.832530
46. Wainer J , Cawley G . Nested cross-validation when selecting classifiers is overzealous for most practical applications. Expert Syst Appl. 2021;182 :115222. 10.1016/j.eswa.2021.115222
47. Bradley AP . The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recognition. 1997;30 (7 ):1145–1159. 10.1016/S0031-3203(96)00142-2
48. Saito T , Rehmsmeier M . The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE. 2015;10 (3 ). 10.1371/journal.pone.0118432
49. Thornicroft G . Shunned: Discrimination Against People with Mental Illness. Oxford: Oxford University Press; 2006.
50. Kular A , Perry BI , Brown L , et al. Stigma and access to care in first-episode psychosis. Early Interv Psychiatry. 2019;13 (5 ):1208–1213. 10.1111/eip.12756 30411522
51. Wang HE , Landers M , Adams R , et al. A bias evaluation checklist for predictive models and its pilot application for 30-day hospital readmission models. J Am Med Inform Assoc. 2022;29 (8 ):1323–1333. 10.1093/jamia/ocac065 35579328
52. Yarborough BJH , Stumbo SP , Schneider J , Richards JE , Hooker SA , Rossom R . Clinical implementation of suicide risk prediction models in healthcare: A qualitative study. BMC Psychiatry. 2022;22 (1 ). 10.1186/s12888-022-04400-5
