
==== Front
HGG Adv
HGG Adv
Human Genetics and Genomics Advances
2666-2477
Elsevier

S2666-2477(24)00081-2
10.1016/j.xhgg.2024.100341
100341
Article
Estimating prevalence of rare genetic disease diagnoses using electronic health records in a children’s hospital
Herr Kate 1
Lu Peixin 1
Diamreyan Kessi 2
Xu Huan 3
Mendonca Eneida 12
Weaver K. Nicole 234
Chen Jing jing.chen2@cchmc.org
125∗
1 Division of Biomedical Informatics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229, USA
2 University of Cincinnati College of Medicine, Cincinnati, OH 45229, USA
3 Division of Human Genetics, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229, USA
4 Heart Institute, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH 45229, USA
∗ Corresponding author jing.chen2@cchmc.org
5 Lead contact

14 8 2024
10 10 2024
14 8 2024
5 4 10034120 2 2024
9 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Summary

Rare genetic diseases (RGDs) affect a significant number of individuals, particularly in pediatric populations. This study investigates the efficacy of identifying RGD diagnoses through electronic health records (EHRs) and natural language processing (NLP) tools, and analyzes the prevalence of identified RGDs for potential underdiagnosis at Cincinnati Children’s Hospital Medical Center (CCHMC). EHR data from 659,139 pediatric patients at CCHMC were utilized. Diagnoses corresponding to RGDs in Orphanet were identified using rule-based and machine learning-based NLP methods. Manual evaluation assessed the precision of the NLP strategies, with 100 diagnosis descriptions reviewed for each method. The rule-based method achieved a precision of 97.5% (95% CI: 91.5%, 99.4%), while the machine-learning-based method had a precision of 73.5% (95% CI: 63.6%, 81.6%). A manual chart review of 70 randomly selected patients with RGD diagnoses confirmed the diagnoses in 90.3% (95% CI: 82.0%, 95.2%) of cases. A total of 37,326 pediatric patients were identified with 977 RGD diagnoses based on the rule-based method, resulting in a prevalence of 5.66% in this population. While a majority of the disorders showed a higher prevalence at CCHMC compared with Orphanet, some diseases, such as 1p36 deletion syndrome, indicated potential underdiagnosis. Analyses further uncovered disparities in RGD prevalence and age of diagnosis across gender and racial groups. This study demonstrates the utility of employing EHR data with NLP tools to systematically investigate RGD diagnoses in large cohorts. The identified disparities underscore the need for enhanced approaches to guarantee timely and accurate diagnosis and management of pediatric RGDs.

This study leverages electronic health records and natural language processing to analyze rare genetic disease diagnoses in pediatric patients. Findings reveal a 5.66% prevalence of such diagnoses with disparities across gender and racial groups. The research highlights the need for improved diagnostic strategies and expanded genetic testing.

Keywords

rare genetic diseases
natural language processing
bioinformatics
Orphanet
electronic health record
genetic testing
==== Body
pmcIntroduction

Rare genetic diseases (RGDs) pose significant challenges in healthcare. Although they individually affect a small percentage of the population, collectively, they can affect a substantial number of people. Particularly concerning is the fact that about 70% of these diseases emerge during early childhood.1,2 Due to the complicated phenotypes and limited awareness, patients with RGDs often face a prolonged and arduous journey from the onset of their diseases to the time of diagnosis. It is estimated that on average, a patient remains undiagnosed or misdiagnosed for 4 to 5 years before a correct diagnosis can be obtained.3

To evaluate the effectiveness of using electronic health records (EHRs) for identifying diagnoses corresponding to RGDs, our study aims to estimate the prevalence of RGD diagnoses among pediatric patients based on EHRs at Cincinnati Children’s Hospital Medical Center (CCHMC). Through this analysis, our study aims to gain insights into RGDS that are potentially underdiagnosed in pediatric patients. Although traditionally used for billing purposes, EHRs are increasingly being repurposed for clinical research on RGDs. For instance, phenotypic signatures from EHRs were shown to aid the identification of patients for genetic testing.4 Our previous research also demonstrated the potential of EHR texts in detecting undiagnosed cases of Noonan syndrome (MIM: 163950).5 However, most existing studies emphasize phenotypic information from EHRs,6,7,8 with limited focus on the systematic analysis of RGD diagnostic data. This is likely due to the nuanced nature of EHR data. While unstructured data offer rich clinical information, the large-scale recognition of RGD diagnoses is extremely difficult as it requires sophisticated natural language processing (NLP) in combination with laborious manual curation. On the other hand, structured EHR data, such as ICD-10 codes, often lack the granularity required for RGD diagnoses.9,10,11

Recognizing this challenge, we chose to utilize diagnosis description texts based on the Intelligent Medical Objects (IMO) terminology. According to IMO,12 its terminology covers 310,000 clinical concepts, which is 17 times and 3 times more comprehensive than ICD-10-UK and SNOMED CT, respectively. This offers more detailed diagnostic information compared with conventional diagnosis codes while being structured for efficient large-scale analysis. In addition, we chose to map these diagnosis texts to ORPHA codes from the Orphanet Rare Disease Ontology (ORDO),13 which is a descriptive ontology curated specifically for rare diseases by Orphanet (www.orpha.net). The ontology structure of ORDO, with synonyms and relations among ORPHA codes, enhances its compatibility with NLP tasks.

Although recent developments in NLP have seen an advancement in machine learning (ML) approaches,14 especially deep learning models, they rely on large sets of training samples and accurate annotations, both of which are still limited for RGD.15 Two recent studies16,17 investigated the potential of large language models (LLMs) for recognizing names of RGDs. While their results suggest that LLMs achieve performance close to or sometimes surpassing conventional rule-based methods, both studies fine-tuned and reported results based on the RareDis corpus,18 which consists of standard text from a rare disease database. This standard text differs significantly from the diagnosis descriptions used in our study. Given the lack of suitable training data for the diagnosis descriptions, we opted to compare a rule-based method with an ML-based method for mapping RGD diagnoses to ORPHA codes. In the following sections, we introduce the data of the study, the development and evaluation of our NLP pipelines, and the findings related to RGD diagnoses.

Material and methods

Data source

The de-identified structured EHRs were obtained from the Cincinnati Children’s Hospital Medical Center TriNetX database on June 2, 2023. The study included all patients born on or after January 1, 2005, ensuring a comprehensive and relevant dataset. The raw data were in a tabular format, with columns corresponding to information related to patient diagnosis and demographics, and rows representing individual diagnoses in the EHRs. Each row included a diagnosis code, a diagnosis description text based on IMO, and the patient’s age (in days) at the time of diagnosis. Demographic information included the patient’s year of birth, self-reported gender, and racial identity, as shown in Figure 1. Examples of these data can be found in Table S1.Figure 1 Two NLP methods (NOBLE Coder and ScispaCy) used to map diagnosis descriptions to RGDs

Both methods were applied separately and tested to map RGD diagnoses back to de-identified patient diagnosis descriptions in EHRs.

Disease annotations were downloaded from Orphanet knowledge base on June 16, 2023. The annotations included rare disease alignments, epidemiology, natural history, classifications, and genes associated with rare diseases. All data, initially in XML files, were downloaded and converted to CSV files for subsequent analysis.

This study protocol was reviewed and approved by the Cincinnati Children’s Hospital Institutional Review Board (protocol no. 2023-0484).

Diagnosis mapping using NLP

A list of all the unique diagnosis descriptions was extracted from the raw data. Two separate NLP approaches, NOBLE Coder19 and ScispaCy,20 were used to map these diagnosis descriptions to ORPHA codes representing RGDs. The workflow is shown in Figure 1. Both methods were applied and evaluated for the performance in this study, as they represent two categories of NLP, with the former a rule-based approach and the latter a pre-trained ML method.

NOBLE Coder is an open-source rule-based concept recognition algorithm implemented in Java. It allows configurations to map input text to terms in a selected ontology. NOBLE Coder was downloaded from the GitHub repository and configured to link each diagnosis description to an ORPHA code defined in the ORDO ontology. The ORDO ontology OWL file was downloaded from BioPortal on June 8, 2023. The default settings of NOBLE Coder were used to identify the most specific matching ORPHA code for each diagnosis description.

SpaCy is a Python-based NLP library for processing large volumes of text and has multiple trainable pipeline components. ScispaCy is an open-source repository for Spacy with custom models trained on large quantities of biomedical text. Spacy (v.3.4.4) and ScispaCy (v.0.5.2) were downloaded in Python (v.3.8.0) with v.0.5.1 of the “en_core_sci_sm” model from ScispaCy. We used ScispaCy’s pre-trained entity linker (v.2.5.0) to map diagnosis descriptions to concepts in the United Medical Language System (UMLS),21 a comprehensive meta thesaurus for biomedical terminology. Terms without definitions were included to ensure mappings could still be made to rare diseases without a definition in UMLS. ScispaCy assigns a “score” to mappings describing how well the extracted entity matches the mapped UMLS term. The default score threshold of 0.7 was kept for terms with definitions, and the score threshold for terms without definitions was set to 0.9 to allow mapping to more terms without definitions. The maximum number of linkings to UMLS terms for each entity was set to 3 to limit the number of mappings per diagnosis description. UMLS codes were converted to ORPHA codes based on the mapping file, downloaded on June 16, 2023 from orphadata. This mapping file contained mappings between 4,744 ORPHA codes corresponding to RGDs and 5,301 UMLS codes.

A comparative analysis was conducted on three distinct strategies: NOBLE Coder based, ScispaCy based, and a combination of both. In the combined approach, NOBLE Coder’s results were prioritized, with ScispaCy being used only when NOBLE Coder failed to map a diagnosis text to an ORPHA code. The analysis was exclusively focused on diagnoses corresponding to RGDs, following the definitions outlined in the previous study.22 Specifically, a diagnosis needed to be mapped to a “disorder” in the Orphanet classification, excluding “group of disorders” and “disorders subtype,” and the corresponding ORPHA code had to be found in the “classification of rare genetic diseases” in Orphanet. For RGD prevalence information, only regions from “United States” and “World” in Orphanet were considered.

Post-processing of mapped data

The mapped data, obtained from the aforementioned NLP processes, underwent additional refinement to eliminate conspicuous false positives. Any mappings based solely on acronyms were systematically excluded from both the NOBLE Coder and ScispaCy results. In addition, text entries referring to family history, such as “Family history of,” “No family history,” “Family historic risk,” “FH:” were excluded as they were determined as false positives. Similar action was taken for text entries referring to screening procedures, indicated by terms such as “Screening for” and “Screening examination.”

Mappings for all RGDs by NOBLE Coder that are present in our cohort and have a listed Orphanet prevalence of less than 1 in a million were manually reviewed. Given our cohort’s size of under one million patients, such mappings were examined for potential mapping errors or false positives by an undergraduate student in a pre-medicine program.

Evaluation of RGD diagnosis mappings

The performance of mapping between diagnosis descriptions to RGDs was evaluated for individual diagnosis descriptions as well as for individual patients. For the description-level evaluation, a random set of 100 diagnosis description-to-ORPHA code mappings from each of three strategies were evaluated manually by two annotators (K.H. and K.D.). If the two annotators could not reach an agreement the mapping was marked as partially correct with a score of 0.5. Precision, defined as totalscoreofcorrectmappingstotalnumberofidentifiedmappings, was calculated for the three mapping strategies.

For the patient-level evaluation, a manual chart review was performed. A random set of 70 patients who received diagnosis mapped to RGDs was selected based on NOBLE Coder mappings. We selected patients based on NOBLE Coder mappings because NLP based on NOBLE Coder alone achieved the best performance in our evaluation (see results). The chart review was performed to confirm the presence of RGD diagnoses identified from the de-identified data. The review process was conducted in EPIC, the EHR system used at CCHMC. All relevant parts of the patient records were reviewed. During the review, the primary focus was on locating the RGD diagnoses within the "Problem List." When the diagnoses were not found in this section, the reviewer examined the "Encounters" and "Notes" sections to identify the presence of the diagnoses and determine the earliest recorded time of diagnosis. The chart review was conducted by an undergraduate biomedical engineering student in a pre-medicine program, who was trained in navigating and interpreting electronic health records (K.H.). The student worked under the direct supervision of a clinician specialized in genetic disorders during the review process (K.N.W.). The numbers of random diagnosis descriptions and patients reviewed were primarily limited by the study timeline. As shown in the results, these random sets were representative of the data in the cohort, as no obvious bias was observed in disease types or patient demographic characteristics.

From the de-identified EHR data, the date of the first occurrence of the RGD diagnosis in each patient was selected as the time of diagnosis of the RGD for the patient. In this chart review, we evaluated the accuracy of the age at diagnosis (in days) in the de-identified data by comparison with those documented in EPIC.

Identification of pediatric patients with genetic testing

A list of 147 current procedural terminology (CPT) codes corresponding to genetic testing was compiled from a previous study23 with four additional codes identified in CCHMC TrinetX (Table S3). Patients with these CPT codes in their EHRs were classified as having undergone at least one genetic test.

Statistical analyses

Prevalence was calculated for each RGD based on the number of patients with a specific ORPHA code out of the total number of patients included in the cohort. Comparisons of prevalence between male and female patients, as well as white and black patients, were made for each RGD using Fisher’s exact test. Age at diagnosis was compared between male and female patients, and white and black patients using Wilcoxon rank-sum test. The Benjamini-Hochberg correction was applied to p values to account for multiple hypothesis testing. Corrected p value <0.05 was considered as significant. Confidence intervals for precision estimated from manual reviews were calculated using Wilson score interval.24 All statistical analyses were performed in R and Python.

Results

EHR data and basic demographics of the cohort

The cohort contained 659,139 unique patients born on or after January 1, 2005, from the CCHMC EHR database. The basic demographic information of the cohort is shown in Table 1. White and black are the predominant racial groups, with other racial groups each encompassing less than 5% of the cohort. There are slightly more males than females in the cohort, and the average birth year is 2012.86, translating to an average age of approximately 10.56 years as of June 2, 2023. For these patients, a total of 26,801 unique ICD-10 codes and 142,122 unique diagnosis descriptions were obtained in the raw data table. A total of 30,045,595 diagnoses existed for these patients, averaging about 46 diagnoses for each patient.Table 1 Demographic information of the study cohort

	Non-RGD (N = 621,813)	RGD (N = 37,326)	Total (N = 659,139)	p value	
Race				<0.001	
 White	427,448 (68.74%)	27,123 (72.67%)	454,571 (68.96%)		
 Black or African American	88,495 (14.23%)	4,824 (12.92%)	93,319 (14.16%)		
 Unknown	44,896 (7.22%)	1,869 (5.01%)	46,765 (7.09%)		
 Identify with two or more races	20,254 (3.26%)	1,206 (3.23%)	21,460 (3.26%)		
 Other	20,049 (3.22%)	986 (2.64%)	21,035 (3.19%)		
 Asian	16,894 (2.72%)	948 (2.54%)	17,842 (2.71%)		
 Middle eastern	1,808 (0.29%)	255 (0.68%)	2,063 (0.31%)		
 Native Hawaiian and other pacific islander	1,122 (0.18%)	58 (0.16%)	1,180 (0.18%)		
 American Indian and Alaska native	847 (0.14%)	57 (0.15%)	904 (0.14%)		
Gender				<0.001	
 Male	322,582 (51.88%)	20,478 (54.86%)	343,060 (52.05%)		
 Female	298,961 (48.08%)	16,844 (45.13%)	315,805 (47.91%)		
 Unknown	270 (0.04%)	4 (0.01%)	274 (0.04%)		
Birth year				<0.001	
 Mean (SD)	2,012.83 (5.06)	2,013.37 (5.09)	2,012.86 (5.07)		
Genetic testing				<0.001	
 Genetic tested	19,849 (3.19%)	9,513 (25.49%)	29,362 (4.45%)		

Diagnosis mapping using NLP

The NLP pipelines utilizing NOBLE Coder, ScispaCy, or a combination of the two methods were applied to the diagnosis descriptions to map them to ORPHA codes corresponding to RGDs. As described in the methods, mappings referring to family history and screening were excluded. In addition, 381 diagnosis descriptions mapped by NOBLE coder to 296 RGDs with a prevalence of less than 1 in a million people were reviewed for potential false positives. This manual review confirmed that the mappings associated with eight RGDs (listed in Table S2) were false positives and were excluded from subsequent analysis. The number of diagnoses mapped to RGDs, the number of unique RGDs identified, and the number of patients with RGDs by each of the methods are listed in Table 2.Table 2 NLP validation and results for three methods of mapping diagnosis descriptions

	NOBLE Coder	ScispaCy	NOBLE Coder + ScispaCy	
NLP results for rare genetic disorders	
	
No. of unique diagnosis texts mapped to RGDs	3,092	3,872	4,828	
No. of unique RGDs identified from patients	977	699	1,036	
No. of patients with diagnoses mapped to RGD	37,326	34,485	42,100	
	
NLP validation results	
	
Precision at diagnosis level (%) (n = 100)	97.50 (91.5, 99.4)	73.5 (63.6, 81.6)	78.5 (68.9, 85.8)	
Precision at patient level (%) (n = 93)	90.32 (82.0, 95.2)	–	–	
Accuracy of age at diagnosis (%) (n = 93)	51.61 (41.1, 62.0)	–	–	

Orphanet also provides mappings between ICD-10 and ORPHA codes, which are classified as Exact, BTNT (broader term to narrower term), NTBT (narrower term to broader term), and ND (unable to decide). When considering all mapping classes, 5,527 RGD ORPHA codes were mapped to ICD-10 codes in our cohort. However, only 884 ICD-10 codes were used to describe these ORPHA codes, and a total of 149,222 patients in the cohort (a prevalence of 22.6%) had at least one of these 884 ICD-10 codes. This result suggests that direct mappings from ICD-10 to ORPHA codes are not only non-specific but also contain substantial false positives that may arise when a narrower ORPHA code is mapped to a broader ICD-10 code. When considering only the exact mappings between ICD-10 and ORPHA codes, a mere 154 ICD-10 codes were mapped to 150 ORPHA codes in 29,777 patients in our cohort. Compared with the results of text-based mapping using the NOBLE Coder-based NLP pipeline, this represents a striking 88% reduction in RGD diagnoses and a 20% reduction in the number of patients identified. Moreover, there is a higher likelihood of false positives in these 29,777 patients, as many ICD-10 codes are mapped to multiple ORPHA codes with only one of the mappings as exact. For instance, H90.5 is mapped to four ORPHA codes (998, 87884, 330029, and 363396); however, only the mapping to 87884 (non-syndromic genetic deafness) is exact. In contrast, the diagnosis descriptions from IMO terminology are often more accurate than ICD-10 codes. For example, in our EHR data, diagnosis texts such as “Aarskog syndrome,” “Cornelia de Lange syndrome,” “Noonan syndrome,” “Russell’s syndrome,” and many more syndromes were all noted as Q87.19. This ICD-10 code, however, is not used at all in Orphanet. Instead, a different ICD-10 code, Q87.1 is used, which is mapped to more than 60 RGD ORPHA codes.

To evaluate the performance of the mappings at description level, a random set of 100 unique diagnosis descriptions that were mapped to RGDs were selected for each of the three NLP pipelines. Manual reviews were conducted by two annotators and the precision of each method was calculated. As can be seen in Table 2, using NOBLE Coder alone resulted in a precision of 97.5% with a 95% confidence interval (91.5%, 99.4%), while inclusion of ScispaCy would significantly reduce the precision to 78.5% with a 95% confidence interval (68.9%, 85.8%). In the 100 disease descriptions identified by NOBLE Coder, the two annotators agreed on 97 mappings as correct and 2 as incorrect. They disagreed on only 1 mapping. In the 100 disease descriptions identified by ScispaCy, the two annotators agreed on 73 mappings as correct and 26 as incorrect. They disagreed on 1 mapping. Due to its superior performance, our subsequent analyses utilized the NOBLE Coder-based results.

NOBLE Coder identified 89 unique ORPHA codes in the 100 random descriptions. According to the classification information from Orphanet, these codes cover all 23 disease classes by Orphanet (some diseases could be classified to multiple types, e.g., Ehlers-Danlos syndrome is classified as a rare skin disease and a rare systemic and rheumatological disease).

To evaluate the performance of the mappings at the patient level, we conducted a manual chart review for 70 randomly selected patients, who, based on NOBLE Coder pipeline, had received at least one RGD diagnosis. A total of 93 diagnosis descriptions were mapped to RGDs for these patients as some patients carried multiple RGD diagnoses. For instance, one patient diagnosed with Down syndrome also had other RGD diagnoses, including atrial septal defect, congenital laryngomalacia, congenital central alveolar hypoventilation syndrome, and Hirschsprung disease. Out of the 93 diagnoses of RGDs, 84 (90.32%) were present in their clinical records in EPIC. When reviewing the age at diagnosis in days, it is found that 61 (65.59%) RGD diagnoses had a discrepancy less than 6 months against EPIC, and 48 (51.61%) of them were the correct time of the initial diagnosis in EPIC. At least 11 (11.82%) of the RGD diagnoses appeared to be obtained prior to visiting CCHMC, although this information was not always available in EPIC.

We examined the disease classifications of the 70 patients for their RGD diagnoses. Their diagnoses were mapped to 50 unique ORPHA codes, covering all 23 disease classes recognized by Orphanet. Furthermore, we compared the demographic information of the 70 randomly selected patients against the remaining 37,256 patients with RGD diagnoses. As shown in Table S10, no obvious bias was observed in race, gender, and age in this random set.

Prevalence of RGD diagnoses

A total of 5.66% of the cohort, or 37,326 patients, received at least one diagnosis mapped to RGDs. The demographic information of patients with RGD diagnoses is listed in Table 1. Indicated by p values <0.001, there are significant differences in race, gender, and birth year between the patients with and without RGD diagnoses. For instance, white patients are more likely to have RGD diagnoses. White patients make up 69% of patients without RGD diagnoses and 73% of patients with RGD diagnoses. Similarly, male patients are more likely to have an RGD diagnosis, as 52% of patients without RGD diagnoses are male while 55% patients with RGD diagnoses are male. The birth year for patients with RGD diagnoses is significantly larger, i.e., these patients are younger compared with patients without RGD diagnoses.

A total of 977 unique ORPHA codes were identified among these RGD diagnoses. The prevalence of the RGD diagnoses, in accordance with Orphanet, was classified into 6 bins: “<1/1,000,000”, “1–9/1,000,000”, “1–9/100,000”, “6–9/10,000”, “1–5/10,000”, and “>1/1,000”. They were also assigned to bins numbered 1 to 6, respectively, as their ordinal categories. The number of RGDs in each prevalence bin is shown in Figure 2A. Note that the bin of “<1/1,000,000” is not present in CCHMC due to our cohort having fewer than 1,000,000 patients.Figure 2 Prevalence of RGD diagnoses

Histogram of prevalence of RGD diagnosis in CCHMC (A), and heatmap showing numbers of RGD diagnoses found in each prevalence bin in CCHMC vs. Orphanet (B).

Out of the 977 RGDs, 552 (56.50%) did not have a worldwide or US point prevalence or prevalence at birth estimate available in Orphanet. For the remaining 425 RGDs, the prevalence in CCHMC is compared with Orphanet and plotted in a heatmap in Figure 2B. Among them, 334, 65, and 26 RGDs have a prevalence level higher, equal, and lower in CCHMC, respectively.

Table 3 lists the top RGDs for which prevalence at CCHMC notably deviates from that in Orphanet. Nine RGDs have a prevalence at least 2 levels higher than Orphanet with at least 75 patients found in CCHMC. Both Wiskott-Aldrich syndrome (MIM: 301000) and PHACE syndrome (MIM: 606519) are estimated to affect very few people (1–9 out of a million and less than 1 out of a million, respectively) based on Orphanet but have a substantially higher prevalence at CCHMC (1–5 out of 10,000 people). Mappings for both disorders appear to be correct based on their respective diagnosis descriptions. Conversely, another 11 RGDs with a prevalence of at least 1/10,000 in Orphanet show a prevalence at least one level lower at CCHMC. Hemoglobin C disease (MIM: 141900), partial chromosome Y deletion (MIM: 400042), and hereditary persistence of fetal hemoglobin-sickle cell disease syndrome (MIM: 141749) are relatively prevalent in Orphanet’s data, yet their prevalence is substantially lower at CCHMC.Table 3 Top RGDs that have a prevalence in CCHMC different from Orphanet

ORPHA Code	RGD name	Age of onset	Prevalence region	Orphanet prevalence type	Orphanet prevalence bin	CCHMC prevalence bin	Difference	
Top 9 RGDs that have a prevalence higher in CCHMC than Orphanet	
	
42775	PHACE syndrome	antenatal; infancy; neonatal	worldwide	point prevalence	<1/1,000,000	1–5/10,000	3	
906	Wiskott-Aldrich syndrome	infancy; neonatal	United States	prevalence at birth	1-9/1,000,000	1–5/10,000	2	
805	tuberous sclerosis complex	all ages	worldwide	point prevalence	1–9/100,000	6–9/10,000	2	
98896	Duchenne muscular dystrophy	childhood	worldwide	point prevalence	1–9/100,000	6–9/10,000	2	
870	Down syndrome	antenatal; neonatal	worldwide	point prevalence	1–5/10,000	>1/1,000	2	
388	Hirschsprung disease	childhood; infancy; neonatal	worldwide	point prevalence	1–5/10,000	>1/1,000	2	
903	Von Willebrand disease	all ages	worldwide	point prevalence	1–5/10,000	>1/1,000	2	
636	neurofibromatosis type 1	infancy; neonatal	worldwide	prevalence at birth	1–5/10,000	>1/1,000	2	
846	α-thalassemia	all ages	United States	prevalence at birth	1–5/10,000	>1/1,000	2	
	
Top 11 RGDs that have a prevalence lower in CCHMC than Orphanet	
	
2132	hemoglobin C disease	all ages	United States	point prevalence	>1/1,000	1–5/10,000	−2	
1646	partial chromosome Y deletion	adult	worldwide	point prevalence	1–5/10,000	1–9/1,000,000	−2	
251380	hereditary persistence of fetal hemoglobin-sickle cell disease syndrome	all ages	United States	prevalence at birth	1–5/10,000	1–9/1,000,000	−2	
791	retinitis pigmentosa	adolescent; adult; childhood	worldwide	point prevalence	1–5/10,000	1–9/100,000	−1	
3378	trisomy 13	antenatal; neonatal	United States	prevalence at birth	1–5/10,000	1–9/100,000	−1	
288	hereditary elliptocytosis	all ages	worldwide	point prevalence	1–5/10,000	1–9/100,000	−1	
1606	1p36 deletion syndrome	antenatal; neonatal	United States	prevalence at birth	1–5/10,000	1–9/100,000	−1	
43	X-linked adrenoleukodystrophy	adolescent; adult; childhood; elderly	United States	prevalence at birth	1–5/10,000	1–9/100,000	−1	
101081	Charcot-Marie-Tooth disease type 1A	childhood	worldwide	point prevalence	1–5/10,000	1–9/100,000	−1	
214	cystinuria	all ages	worldwide	Point prevalence	1–5/10,000	1–9/100,000	−1	

Besides the RGDs which showed different prevalence in CCHMC compared with Orphanet, we found 26 disorders that had a prevalence estimate greater than or equal to 1–9/100,000 people in Orphanet but did not have any patients identified in CCHMC. Given their relatively high prevalence yet absence in our cohort, we manually searched for these disease names in the diagnosis descriptions to identify potential false negatives missed by the NLP tool. Manual inspection suggested most of these were attributed to their late onset and were therefore missing in our pediatric cohort. Eleven of these 26 disorders, however, do have an early onset, shown in Table 4. Diagnosis descriptions corresponding to proximal 16p11.2 microdeletion syndrome (MIM: 611913), Prader-Willi syndrome (MIM: 176270), and hemophilia A (MIM: 306700) were present in EHRs but not correctly mapped to RGDs by NOBLE Coder, therefore were considered as false negatives. Specifically, “Prader-Willi syndrome” was recognized as “Prader-Willi-like syndrome,” and “hemophilia A” was recognized as “hemophilia.” Chromosome 16p11.2 microdeletion syndrome was not recognized at all. The rest of the eight RGDs were not found in the diagnosis descriptions in the raw data.Table 4 Eleven early-onset RGDs with prevalence of at least 1/100,000 in Orphanet that were not identified in CCHMC

ORPHA Code	RGD name	Age of onset	Prevalence region	Orphanet prevalence type	Orphanet prevalence bin	
261197	proximal 16p11.2 microdeletion syndrome∗	childhood	United States	point prevalence	1–5/10,000	
141132	oculo-auriculo-vertebral spectrum	antenatal; neonatal	worldwide	point prevalence	1–9/100,000	
87503	Mal de Meleda	childhood; infancy; neonatal	worldwide	point prevalence	1–9/100,000	
739	Prader-Willi syndrome∗	antenatal; neonatal	United States	point prevalence	1–9/100,000	
98878	hemophilia A∗	childhood; infancy; neonatal	United States	point prevalence	1–9/100,000	
3306	inverted duplicated chromosome 15 syndrome	neonatal	worldwide	prevalence at birth	1–9/100,000	
95702	X-linked adrenal hypoplasia congenita	childhood; infancy	worldwide	prevalence at birth	1–9/100,000	
2157	histidinemia	infancy; neonatal	United States	prevalence at birth	1–9/100,000	
158	systemic primary carnitine deficiency	infancy; neonatal	United States	prevalence at birth	1–9/100,000	
2655	thanatophoric dysplasia	antenatal; neonatal	United States	prevalence at birth	1–9/100,000	
54	X-linked recessive ocular albinism	infancy; neonatal	United States	prevalence at birth	1–9/100,000	

Comparison of prevalence of RGD diagnoses across gender and race

Prevalence of 41 RGD diagnoses were significantly different between male and female patients in CCHMC (listed in Table S4). Of these, 16 had a known X-linked inheritance pattern obtained from the Orphanet “natural history.” Most of these RGDs have been previously documented to have different prevalence in males and females.

Another set of 43 RGD diagnoses showed significant prevalence differences between white and black patients (listed in Table S5). Some of these RGDs were known to have different prevalence in the racial groups. For instance, α-thalassemia (MIM: 604131), sickle cell anemia (MIM: 603903), and hemoglobin C disease were predominantly associated with black patients. While many of the remaining RGDs were expected to affect both races equally, they showed notable differences in CCHMC. For instance, the diagnoses of Down syndrome (MIM: 190685), fragile X syndrome (FXS) (MIM: 300624), neurofibromatosis type 1 (NF1) (MIM:162200), 22q11.2 deletion syndrome (22q11.2DS) (MIM: 188400), Duane retraction syndrome (MIM: 126800), and tuberous sclerosis complex (MIM: 191100) were all considerably less frequent in black patients.

Comparison of age of diagnosis across gender and race

Age of diagnosis was notably different between male and female patients for five RGDs (listed in Table S6). These five RGDs, Von Willebrand disease (MIM: 193400), hypermobile Ehlers-Danlos (MIM: 130020), prune belly syndrome (MIM: 100100), familial dysautonomia (MIM: 223900), and FXS were all diagnosed earlier in male patients. FXS and hypermobile Ehlers-Danlos syndrome in particular showed differences in both prevalence and age at diagnosis. Figure 3A shows the time until diagnosis for male and female patients with these two disorders.Figure 3 Age at diagnosis for RGDs based on a specific demographic

Age at diagnosis for fragile X syndrome and hypermobile Ehlers-Danlos syndrome in female and male patients (A), and age at diagnosis for 22q11.2 deletion syndrome and neurofibromatosis type 1 in black and white patients (B).

For racial discrepancies in age of diagnosis, seven RGDs showed significant difference between white and black patients (listed in Table S7). These disorders were all diagnosed at earlier ages in black patients. Intriguingly, despite NF1 and 22q11.2DS having lower prevalence in black patients, their diagnoses occurred much earlier when compared with white patients. As illustrated in Figure 3B, all black patients with 22q11.2DS are diagnosed before age 5, whereas white patients continued to receive diagnoses into their teenage years. For NF1, a third of black patients had been diagnosed by age 1 compared with only 14.12% of white patients who were diagnosed by this same age.

Prevalence of genetic testing

Out of the 659,139 patients in the cohort, only 29,362 (or 4.45%) had at least one CPT code corresponding to genetic testing recorded in their EHRs. As shown in Table 1, among the 37,326 patients with RGD diagnoses, 9,513 (25.49%) had at least one genetic testing CPT code. In contrast, only 3.19% of patients without an RGD diagnosis had undergone genetic testing, a rate nearly 8 times lower than that of patients with RGD diagnoses. The demographic information of patients with genetic testing is shown in Table S8. Similar to the results in Table 1, racial and gender disparities were observed: white patients were more likely to undergo genetic testing than black patients, and males were more frequently tested than females.

The frequency of genetic testing among patients also exhibited substantial variation based on their specific RGD diagnosis (Table S9). For conditions such as hemoglobin C disease and Noonan syndrome, over 70% patients underwent genetic testing. However, for Rolandic epilepsy (MIM: 117100) and Hirschsprung disease (MIM: 142623), only approximately 10% patients were tested.

Discussion

This study utilized NLP and EHR data to gain insights into the prevalence of RGD diagnoses, particularly within the pediatric population of CCHMC. The overall estimated prevalence of RGD diagnosis is 5.66%. While our NLP process achieved high precision, the estimate is based on only 977 unique RGD diagnoses, suggesting that the real prevalence in this population might be higher.

Comparison of RGD prevalence between CCHMC and Orphanet

A significant takeaway from our study is the comparison between the RGD prevalence in CCHMC and Orphanet. A majority, 334 out of 425 (or 78.59%), of the evaluated RGDs showed prevalence in CCHMC higher than Orphanet’s data. As shown in Table 3, certain RGD diagnoses demonstrated significantly higher prevalence at CCHMC than in Orphanet. A plausible explanation for this disparity is “referral bias.” CCHMC’s recognized expertise in certain disorders could lead to enrichment of patients with these diagnoses. For instance, nearly 70% of CCHMC patients diagnosed with Wiskott-Aldrich syndrome came from states outside of Ohio, based on their EHR home addresses. This is likely true for other diseases such as tuberous sclerosis complex, Duchenne muscular dystrophy (MIM: 310200), Down syndrome, and NF1.

Conversely, some RGD diagnoses showed a markedly lower prevalence at CCHMC compared with Orphanet. Conditions such as partial chromosome Y deletion, retinitis pigmentosa (MIM: 180100), X-linked adrenoleukodystrophy (MIM: 300100), cystinuria (MIM: 220100), and Mayer-Rokitansky-Küster-Hauser syndrome (MIM: 277000) may exhibit fewer diagnoses due to their onset generally occurring beyond the age range of the pediatric cohort in this study.25,26,27,28,29 In conditions such as hemoglobin C disease and hereditary elliptocytosis (MIM: 130600), the reduced frequency of diagnoses could likely be due to the generally milder phenotypes associated with these conditions.30,31 For trisomy 13, the reduced diagnosis rate could be attributed to births occurring outside CCHMC, with families potentially opting for hospice care, resulting in these patients never entering the CCHMC system.32 Interestingly, the apparent low prevalence of Charcot-Marie-Tooth disease type 1A (MIM: 118220) was mainly caused by differences in diagnosis terms used in EHRs. Our manual inspection discovered that many patients were diagnosed with “CMT (Charcot-Marie-Tooth disease)” or “Charcot-Marti-Tooth disease” but were not matched to Charcot-Marie-Tooth disease type 1A in our data. Including these patients in our analysis would significantly increase the reported prevalence of the disease at CCHMC. For 1p36 deletion syndrome, underdiagnosis at CCHMC might explain its lesser number of cases, especially if it presents in milder forms that do not compel patients to seek care.33 Further investigations are warranted to gain more insights into the discrepancy in expected and observed 1p36 deletion diagnoses.

Certain RGD diagnoses, even with a prevalence exceeding 1 in 100,000 in Orphanet, were absent from CCHMC dataset (shown in Table 4). There could be several explanations: (1) some conditions existed in the CCHMC EHRs but were not recognized by our NLP pipeline, such as proximal 16p11.2 microdeletion syndrome (MIM: 611913), Prader-Willi syndrome, and hemophilia A. Although NOBLE Coder mappings exhibit high precision, estimating their recall remains challenging. (2) Certain diseases might be present in the EHRs but recorded under alternative names, such as oculo-auriculo-vertebral spectrum (MIM: 164210), which is sometimes used interchangeably as Goldenhar syndrome or hemifacial microsomia. (3) Some conditions, such as thanatophoric dysplasia (MIM: 156830), could be lethal, so affected patients never reach CCHMC for diagnosis and treatment. (4) There might be conditions that are not included in the IMO terminology, thus absent in our EHR data. (5) It is possible that prevalence of certain diseases is too low that their absence in CCHMC cohort is simply due to statistical randomness.

Difference in prevalence and age at diagnosis across patient groups

While the NLP process showed a high precision in recognizing RGD diagnoses in EHR data, the prevalence estimates in our study may not be comparable with the general population. Similarly, the age at diagnosis obtained from our EHR data may not reflect true age at diagnosis, as shown to be only around 50% accurate in our chart review. However, when comparing the prevalence and age of diagnosis of RGDs among patients within the study, biases may reduce, especially for the diagnoses with large numbers of cases. Notably, our results highlighted disparities in RGD prevalence and age of diagnosis across gender and racial groups.

Many RGD diagnoses show prevalence differences between males and females, as listed in Table S4. As expected, many of them have an X-linked inheritance pattern. For instance, FXS being X-linked dominant, was found more prevalent and diagnosed earlier in males (Figure 3A), which was consistent with previous reports.34 Some disorders, although having an autosomal inheritance pattern, are well documented to have different symptoms or age of onset across genders. For instance, for hypermobile Ehlers-Danlos syndrome, although males have lower prevalence, they appear to be diagnosed earlier, in line with previous reports.35,36 For malacia disorders, such as congenital laryngomalacia (MIM: 150280) and congenital tracheomalacia, higher prevalence was observed in males, as reported previously.37 However, it is currently unclear if the genetic conditions causing malacia disorders tend to be more common in males and are affecting the rate of gender-specific prevalence of malacia disorders.

Similarly, many RGD diagnoses display prevalence disparities across racial groups, especially between black and white patients (shown in Table S5). Blood cell disorders, such as α-thalassemia, β-thalassemia (MIM: 613985), sickle cell anemia, and hemoglobin C disease are known to be more common in black individuals.38,39 However, disorders such as NF1, 22q11.2DS, Duane retraction syndrome, and tuberous sclerosis complex are expected to affect all races equally yet showed a significantly lower prevalence in black patients. Several reasons could explain this discrepancy. For instance, NF 1 patients frequently present with cafe-au-lait macules, freckling, and neurofibromas while hypomelanotic macules, angiofibroma, ungual fibromas, and shagreen patch are included among the major criteria for tuberous sclerosis complex (TSC) diagnosis.40,41 Recent research has suggested a lack of representation of darker skin tones at the topic level in medical textbooks.42 This lack of representation could play a factor in diagnosing NF1 and TSC. In the meantime, NF1 and 22q11.2DS are diagnosed significantly earlier in black patients (Figure 3B), suggesting underdiagnosis for black patients presenting milder phenotypes later in life.

These observed discrepancies raise essential questions about the factors contributing to the diagnostic disparities, be they genetic, environmental, or related to healthcare accessibility. For certain conditions, diagnosis timing varied noticeably across racial and gender groups, which could have significant implications for early interventions and treatment plans.

Prevalence of genetic testing

Genetic testing prevalence also provided interesting insights. Notably, patients diagnosed with RGDs were almost eight times more likely to undergo genetic testing than those without an RGD diagnosis. This result underscores the critical role of genetic tests in diagnosing and managing RGDs. At the same time, it is important to notice the overall prevalence of genetic testing, 4.45% in the cohort, similar to a previous study,23 is still relatively low. For example, the rate of testing for Rolandic epilepsy and childhood absence epilepsy (MIM: 600131) is about 10%. This is not surprising, given that the electroencephalogram remains the primary diagnostic tool for epilepsy. However, recent studies have shown that genetic testing can not only help with the diagnosis of epilepsy but also help guide the selection of treatment plans.43 Our results emphasize the need and potential to expand genetic testing, especially in the patient groups who tend to be underdiagnosed or tested less frequently.

Performance of NLP and de-identified EHRs for genetic disease diagnoses

Our results demonstrate the potential of integrating NLP with de-identified EHRs to effectively identify diagnosis of RGDs on a large scale. The remarkable precision of 97.5% achieved using the NOBLE Coder not only underscores the value of rule-based NLP approaches but also points toward the potential improvement needed for ML-based methods, which are limited by the lack of training data and accurate annotations. In fact, given the high precision of the rule-based NLP pipeline, the identified patients and their corresponding EHR data have the potential to become the training sets for refining ML-based methodologies in future studies.

Furthermore, our manual chart reviews have indicated a small yet substantial discrepancy between the de-identified EHRs used in the study and EPIC regarding RGD diagnoses. Such findings highlight the limitations of the current de-identified EHR datasets, and the imperative for creating and improving methods that can extract and standardize clinical information more reliably from EHRs.

Limitations of the study

This study provides valuable insights into the prevalence of RGD diagnoses using EHRs. However, several inherent limitations must be acknowledged. (1) Population specificity: the data for our study were extracted from a regional pediatric hospital. As such, our findings may not be representative to broader populations, including adults, or other regions with different healthcare access or genetic backgrounds. (2) Data constraints: the use of de-identified diagnosis descriptions imposes limitations in terms of both size and accuracy. The de-identification process can sometimes remove or alter crucial pieces of information, potentially impacting the results. In addition, as with any EHR-based study, our findings are confined by the quality and comprehensiveness of the recorded data. (3) NLP recall not evaluated: while the precision of the pipelines was evaluated, we did not evaluate their recall. As a result, the reported prevalence of RGD diagnoses might represent a conservative estimate, with the actual number of cases in the cohort possibly being higher. (4) Complications with CPT codes: relying on CPT codes to identify genetic testing presents its challenges. CPT codes can be complex, frequently updated, and may not capture all relevant genetic tests, particularly newer or less common ones. This could lead to an underestimation of the number of genetic tests performed. (5) Limitations of manual chart review: manual reviews of EHR data, while essential for validating automated approaches, are resource intensive and were, therefore, limited to a relatively small subset of diagnoses and patients. Future studies with larger sample sizes and more extensive manual reviews would be beneficial to further validate our findings and assess the generalizability of our methods. (6) Study cohort size: although all pediatric patients in CCHMC were included, the cohort size is still relatively small for many RGDs, resulting in inaccurate estimates for their prevalence. (7) Choice of NLP tools: our approach relies on two “generalist” concept recognition tools, which may be considered low-tech compared with more recent LLMs. While our choice of tools was guided by the specific requirements of the task and the available resources, it is unclear if using more advanced generative models would have produced better results. Future research should explore the performance of LLMs on EHR data specific to RGDs.

Given these limitations, it is necessary to conduct future studies to replicate our findings using data from different medical centers to produce potentially actionable steps. This endeavor is facilitated by the widespread adoption of the IMO terminology, which is used by "over 4,500 hospitals; 89% of physicians, nurses, and PAs in the US and 20 countries outside of the US."12 Moreover, it is worth noting that the landscape of the EHR systems and genetic testing is rapidly evolving. This study provides a snapshot based on current methodologies and available data, and future research might yield different insights as these fields continue to progress.

Conclusion

Our study demonstrates the vast potential of EHRs when combined with NLP tools in investigating RGDs. The disparities observed in RGD prevalence and age of diagnosis across different demographics reiterate the need for more advanced approaches and awareness in clinical care and research. Our findings also highlight the importance of expanding genetic testing, particularly in populations that are currently underdiagnosed or have limited access to these diagnostic tools. As we move forward, the integration of EHRs and NLP offers a promising avenue for enhanced patient care and more in-depth research into RGDs.

Data and code availability

The Orphanet datasets analyzed during the current study were downloaded from the Orphanet knowledge base. The CCHMC dataset is not publicly available due to privacy restrictions. The analyses performed during the current study did not use any custom code that is deemed central to the conclusions but may be made available to qualified researchers on reasonable request from the corresponding author.

Acknowledgments

This study was partially supported by the Center for Pediatric Genomics at Cincinnati Children's Hospital Medical Center. No additional external funding was received for this study.

Author contributions

K.H. played a pivotal role in data curation and formal analysis and was a significant contributor to the writing of the manuscript. P.L. was instrumental in developing the methodology and contributed to the formal analysis. K.D. contributed to data curation. H.X. provided substantial input to the formal analysis. E.M. made valuable contributions to the methodology and played a key role in reviewing the manuscript. N.W. was a major force in leading the clinical investigation and provided critical reviews of the manuscript. J.C. was central to the conceptualization of the manuscript and investigation of the study and was a significant contributor to the writing of the manuscript. All authors have read and provided their approval for the final version of the manuscript.

Declaration of interests

The authors declare no competing interests.

Declaration of generative AI and AI-assisted technologies in the writing process

During the preparation of this work the author(s) used ChatGPT in order to improve language and readability. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.

Web resources

GitHub: https://github.com/navd/nobletools

OMIM: http://www.omim.org.

ORPHA: http://www.orphadata.com/.

Orphanet: http://www.orpha.net.

Orphanet knowledge base: https://www.orphadata.com/orphanet-scientific-knowledge-files/.

ScispaCy: https://allenai.github.io/scispacy.

SpaCy: http://spacy.io.

UMLS: http://uts.nlm.nih.gov/uts/umls.

Supplemental information

Document S1. Tables S1, S2, S6–S8, and S10

Table S3. CPT codes for genetic testing

Table S4. RGD diagnoses with prevalence difference in male and female

Table S5. RGD diagnoses with prevalence difference in black and white

Table S9. Prevalence of genetic testing for RGDs with at least 100 patients

Document S2. Article plus supplemental information

Supplemental information can be found online at https://doi.org/10.1016/j.xhgg.2024.100341.
==== Refs
References

1 Hartin S.N. Means J.C. Alaimo J.T. Younger S.T. Expediting rare disease diagnosis: a call to bridge the gap between clinical and functional genomics Mol. Med. 26 2020 117 33238891
2 The Lancet Child Adolescent Rare diseases: clinical progress but societal stalemate Lancet Child Adolesc. Health 4 2020 251 32119839
3 Marwaha S. Knowles J.W. Ashley E.A. A guide for the diagnosis of rare and undiagnosed disease: beyond the exome Genome Med. 14 2022 23 35220969
4 Morley T.J. Han L. Castro V.M. Morra J. Perlis R.H. Cox N.J. Bastarache L. Ruderfer D.M. Phenotypic signatures in clinical data enable systematic identification of patients for genetic testing Nat. Med. 27 2021 1097 1104 34083811
5 Yang Z. Shikany A. Ni Y. Zhang G. Weaver K.N. Chen J. Using deep learning and electronic health records to detect Noonan syndrome in pediatric patients Genet. Med. 24 2022 2329 2337 36098741
6 Liu C. Ta C.N. Havrilla J.M. Nestor J.G. Spotnitz M.E. Geneslaw A.S. Hu Y. Chung W.K. Wang K. Weng C. OARD: Open annotations for rare diseases and their phenotypes based on real-world data Am. J. Hum. Genet. 109 2022 1591 1604 35998640
7 Bastarache L. Hughey J.J. Hebbring S. Marlo J. Zhao W. Ho W.T. Van Driest S.L. McGregor T.L. Mosley J.D. Wells Q.S. Phenotype risk scores identify patients with unrecognized Mendelian disease patterns Science 359 2018 1233 1239 29590070
8 Chen J. Xu H. Jegga A. Zhang K. White P.S. Zhang G. Novel phenotype-disease matching tool for rare genetic diseases Genet. Med. 21 2019 339 346 29895857
9 Strashny A. Alford J. Rappole C. Santo L. The National Hospital Care Survey Is a Unique Source of Data on Rare Diseases Value Health 25 2022 1814 1817
10 Aref L. Bastarache L. Hughey J.J. The phers R package: using phenotype risk scores based on electronic health records to study Mendelian disease and rare genetic variants Bioinformatics 38 2022 4972 4974 36083022
11 Fung K.W. Richesson R. Bodenreider O. Coverage of rare disease names in standard terminologies and implications for patients, providers, and research AMIA Annu. Symp. Proc. 2014 2014 564 572 25954361
12 Staff I. The clinical terminology behind the health IT curtain https://www.imohealth.com/ideas/article/the-clinical-terminology-behind-the-health-it-curtain/ 2021
13 Vasant D. Chanas L. Malone J. Hanauer M. Olry A. Jupp S. Robinson P.N. Parkinson H. Rath A. Ordo: an ontology connecting rare disease, epidemiology and genetic data. Proceedings of ISMB, 30 2014 1 4
14 Banda J.M. Seneviratne M. Hernandez-Boussard T. Shah N.H. Advances in Electronic Phenotyping: From Rule-Based Definitions to Machine Learning Models Annu. Rev. Biomed. Data Sci. 1 2018 53 68 31218278
15 Lee J. Liu C. Kim J. Chen Z. Sun Y. Rogers J.R. Chung W.K. Weng C. Deep learning for rare disease: A scoping review J. Biomed. Inf. 135 2022 104227
16 Hong N. Liu C. Gao J. Han L. Chang F. Gong M. Su L. State of the Art of Machine Learning-Enabled Clinical Decision Support in Intensive Care Units: Literature Review JMIR Med. Inform. 10 2022 e28781
17 Segura-Bedmar I. Camino-Perdones D. Guerrero-Aspizua S. Exploring deep learning methods for recognizing rare diseases and their clinical manifestations from texts BMC Bioinf. 23 2022 263
18 Martinez-deMiguel C. Segura-Bedmar I. Chacón-Solano E. Guerrero-Aspizua S. The RareDis corpus: A corpus annotated with rare diseases, their signs and symptoms J. Biomed. Inf. 125 2022 103961
19 Tseytlin E. Mitchell K. Legowski E. Corrigan J. Chavan G. Jacobson R.S. NOBLE - Flexible concept recognition for large-scale biomedical natural language processing BMC Bioinf. 17 2016 32
20 Neumann M. King D. Beltagy I. Ammar W. ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing 2019 Association for Computational Linguistics
21 Bodenreider O. The Unified Medical Language System (UMLS): integrating biomedical terminology Nucleic Acids Res. 32 2004 D267 D270 14681409
22 Nguengang Wakap S. Lambert D.M. Olry A. Rodwell C. Gueydan C. Lanneau V. Murphy D. Le Cam Y. Rath A. Estimating cumulative point prevalence of rare diseases: analysis of the Orphanet database Eur. J. Hum. Genet. 28 2020 165 173 31527858
23 Schroeder B.E. Gonzaludo N. Everson K. Than K.S. Sullivan J. Taft R.J. Belmont J.W. The diagnostic trajectory of infants and children with clinical features of genetic disease NPJ Genom. Med. 6 2021 98 34811359
24 Wilson E.B. Probable inference, the law of succession, and statistical inference J. Am. Stat. Assoc. 22 1927 209 212
25 Liu T. Song Y.X. Jiang Y.M. Early detection of Y chromosome microdeletions in infertile men is helpful to guide clinical reproductive treatments in southwest of China Medicine (Baltim.) 98 2019 e14350
26 Tsujikawa M. Wada Y. Sukegawa M. Sawa M. Gomi F. Nishida K. Tano Y. Age at onset curves of retinitis pigmentosa Arch. Ophthalmol. 126 2008 337 340 18332312
27 Sadiq S. Cil O. Cystinuria: An Overview of Diagnosis and Medical Management Turk. Arch. Pediatr. 57 2022 377 384 35822468
28 Herlin M.K. Petersen M.B. Brannstrom M. Mayer-Rokitansky-Kuster-Hauser (MRKH) syndrome: a comprehensive update Orphanet J. Rare Dis. 15 2020 214 32819397
29 Patel S. Gutowski N. The difficulty in diagnosing X linked adrenoleucodystrophy and the importance of identifying cerebral involvement BMJ Case Rep. 2015 2015 bcr2015209732
30 Jha, S.K. and S. Vaqar, Hereditary Elliptocytosis StatPearls. 2023: Treasure Island (FL) ineligible companies. Disclosure: Sarosh Vaqar Declares No Relevant Financial Relationships with Ineligible Companies.
31 Karna, B., S.K. Jha, and E. Al Zaabi, Hemoglobin C Disease, in StatPearls. 2023: Treasure Island (FL) ineligible companies. Disclosure: Suman Jha declares no relevant financial relationships with ineligible companies. Disclosure: Eiman Al Zaabi Declares No Relevant Financial Relationships with Ineligible Companies.
32 Cortezzo D.E. Tolusso L.K. Swarr D.T. Perinatal Outcomes of Fetuses and Infants Diagnosed with Trisomy 13 or Trisomy 18 J. Pediatr. 247 2022 116 123.e5 35452657
33 Nistico D. Guidolin F. Navarra C.O. Bobbo M. Magnolato A. D'Adamo A.P. Giorgio E. Pivetta B. Barbi E. Gasparini P. Dental anomalies as a possible clue of 1p36 deletion syndrome due to germline mosaicism: a case report BMC Pediatr. 20 2020 201 32386509
34 Bartholomay K.L. Lee C.H. Bruno J.L. Lightbody A.A. Reiss A.L. Closing the Gender Gap in Fragile X Syndrome: Review on Females with FXS and Preliminary Research Findings Brain Sci. 9 2019 11
35 Gocentas A. Jascaniniene N. Pasek M. Przybylski W. Matulyte E. Mieliauskaite D. Kwilecki K. Jaszczanin J. Prevalence of generalised joint hypermobility in school-aged children from east-central European region Folia Morphol. 75 2016 48 52
36 Demmler J.C. Atkinson M.D. Reinhold E.J. Choy E. Lyons R.A. Brophy S.T. Diagnosed prevalence of Ehlers-Danlos syndrome and hypermobility spectrum disorder in Wales, UK: a national electronic cohort study and case-control comparison BMJ Open 9 2019 e031365
37 Masters I.B. Chang A.B. Patterson L. Wainwright C. Buntain H. Dean B.W. Francis P.W. Series of laryngomalacia, tracheomalacia, and bronchomalacia disorders and their associations with other conditions in children Pediatr. Pulmonol. 34 2002 189 195 12203847
38 Pokhrel A. Olayemi A. Ogbonda S. Nair K. Wang J.C. Racial and ethnic differences in sickle cell disease within the United States: From demographics to outcomes Eur. J. Haematol. 110 2023 554 563 36710488
39 Beutler E. West C. Hematologic differences between African-Americans and whites: the roles of iron deficiency and alpha-thalassemia on hemoglobin levels and mean corpuscular volume Blood 106 2005 740 745 15790781
40 Pounders A.J. Rushing G.V. Mahida S. Nonyane B.A.S. Thomas E.A. Tameez R.S. Gipson T.T. Racial differences in the dermatological manifestations of tuberous sclerosis complex and the potential effects on diagnosis and care Ther. Adv. Respir. Dis. 3 2022 26330040221140125
41 Legius E. Messiaen L. Wolkenstein P. Pancza P. Avery R.A. Berman Y. Blakeley J. Babovic-Vuksanovic D. Cunha K.S. Ferner R. Revised diagnostic criteria for neurofibromatosis type 1 and Legius syndrome: an international consensus recommendation Genet. Med. 23 2021 1506 1513 34012067
42 Louie P. Wilkes R. Representations of race and skin tone in medical textbook imagery Soc. Sci. Med. 202 2018 38 42 29501717
43 Striano P. Minassian B.A. From Genetic Testing to Precision Medicine in Epilepsy Neurotherapeutics 17 2020 609 615 31981099
