==== Front PLOS Digit Health PLOS Digit Health plos PLOS Digital Health 2767-3170 Public Library of Science San Francisco, CA USA 10.1371/journal.pdig.0000269 PDIG-D-22-00278 Research Article Medicine and Health Sciences Medical Conditions Genetic Diseases Autosomal Recessive Diseases Gaucher's Disease Medicine and Health Sciences Clinical Genetics Genetic Diseases Autosomal Recessive Diseases Gaucher's Disease Medicine and Health Sciences Medical Conditions Metabolic Disorders Inherited Metabolic Disorders Gaucher's Disease Research and Analysis Methods Database and Informatics Methods Information Retrieval Medicine and Health Sciences Diagnostic Medicine Biology and Life Sciences Developmental Biology Morphogenesis Morphogenic Segmentation Medicine and Health Sciences Medical Conditions Genetic Diseases Fabry Disease Medicine and Health Sciences Clinical Genetics Genetic Diseases Fabry Disease Medicine and Health Sciences Clinical Medicine Signs and Symptoms Medicine and Health Sciences Pharmaceutics Drug Therapy Enzyme Replacement Therapy Medicine and Health Sciences Clinical Genetics X-Linked Traits Biology and Life Sciences Genetics Heredity Genetic Linkage Sex Linkage X-Linked Traits FindZebra online search delving into rare disease case reports using natural language processing FindZebra search into rare disease case reports Liévin Valentin Conceptualization Data curation Formal analysis Investigation Methodology Resources Software Validation Visualization Writing – original draft Writing – review & editing 1 2 Hansen Jonas Meinertz Data curation Investigation Methodology Resources Software Writing – review & editing 2 Lund Allan Validation Writing – review & editing 3 Elstein Deborah Validation Writing – review & editing 4 Matthiesen Mads Emil Conceptualization Funding acquisition Methodology Project administration Resources Supervision Writing – review & editing 2 Elomaa Kaisa Conceptualization Funding acquisition Investigation Project administration Resources Supervision Writing – original draft Writing – review & editing 5 Zarakowska Kaja Project administration Supervision Writing – review & editing 6 ¤ Himmelhan Iris Funding acquisition Investigation Project administration Resources Writing – review & editing 6 Botha Jaco Data curation Formal analysis Methodology Resources Writing – review & editing 6 Borgeskov Hanne Funding acquisition Project administration Supervision Writing – review & editing 7 https://orcid.org/0000-0002-1966-3205 Winther Ole Conceptualization Data curation Formal analysis Funding acquisition Investigation Methodology Project administration Resources Supervision Writing – original draft Writing – review & editing 1 2 8 9 * 1 DTU Compute, Technical University of Denmark, Lyngby, Denmark 2 FindZebra, Denmark 3 Centre Inherited Metabolic Diseases, Department of Clinical Genetics and Paediatrics, Copenhagen University Hospital, Rigshospitalet, Copenhagen Ø, Denmark 4 Independent consultant; Jerusalem, Israel 5 Takeda Oy, Helsinki, Finland 6 Takeda Pharmaceuticals International AG, Zürich, Switzerland 7 Department of Clinical Pharmacology, Aalborg University Hospital, Aalborg, Denmark 8 Department of Biology, University of Copenhagen, Copenhagen N, Denmark 9 Genomic Medicine, Copenhagen University Hospital, Rigshospitalet, Copenhagen Ø, Denmark Harrison Ewen M. Editor University of Edinburgh, UNITED KINGDOM I have read the journal’s policy and the authors of this manuscript have the following competing interests: KE, IH and JB are employed by Takeda and hold Takeda stocks/stock options. HB is a former employee in Takeda Pharma A/S, Denmark, holding a current position at Department of Clinical Pharmacology, Aalborg University Hospital in Denmark. KZ was employed by Takeda at the time of the study and holds Takeda stocks/stock options. OW, MM, VL and JH are employed by FindZebra, which received funding from Takeda for conducting the study. AL received reimbursement from FindZebra for clinical expertise in this study. AL reports also personal consultancy fees and travel grants from Takeda during the study, as well as grant support paid to his institution. DE received consultancy fees from Takeda for clinical expertise in this study. ¤ Current address: UCB Farchim SA, Bulle, Switzerland * E-mail: ole.winther@bio.ku.dk 29 6 2023 6 2023 2 6 e000026927 9 2022 3 5 2023 © 2023 Liévin et al 2023 Liévin et al https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Early diagnosis is crucial for well-being and life quality of the rare disease patient. Access to the most complete knowledge about diseases through intelligent user interfaces can play an important role in supporting the physician reaching the correct diagnosis. Case reports may offer information about heterogeneous phenotypes which often further complicate rare disease diagnosis. The rare disease search engine FindZebra.com is extended to also access case report abstracts extracted from PubMed for several diseases. A search index for each disease is built in Apache Solr adding age, sex and clinical features extracted using text segmentation to enhance the specificity of search. Clinical experts performed retrospective validation of the search engine, utilising real-world Outcomes Survey data on Gaucher and Fabry patients. Medical experts evaluated the search results as being clinically relevant for the Fabry patients and less clinically relevant for the Gaucher patients. The shortcomings for Gaucher patients mainly reflect a mismatch between the current understanding and treatment of the disease and how it is reported in PubMed, notably in the older case reports. In response to this observation, a filter for the publication date was added in the final version of the tool available from deep.findzebra.com/ with = gaucher, fabry, hae (Hereditary angioedema). Author summary Rare diseases affect a substantial part of the population. However, they are especially challenging to diagnose. Because of their rarity, physicians often ignore rare diseases in the differential diagnosis. When confronted with hard-to-diagnose patients, physicians often turn to online resources like Google or PubMed, which index both general disease information as well as case reports. Case reports are a unique asset in helping the diagnosis of rare diseases because they often present with a varied and complex phenotype, which might not appear in the general literature. Nonetheless, searching for patient-relevant case reports is challenging. A tool dedicated to searching case reports assisting diagnosis is still missing because general-purpose search engines, like PubMed search, are primarily set up for literature search and because advanced search tools like FindZebra do not handle case reports. In this study, we present a novel online search tool https://deep.findzebra.com/ dedicated to searching PubMed case reports based on a patient description (age, sex, symptoms, negative findings, etc.). Two medical experts evaluated the tool on forty challenging cases (twenty Fabry and twenty Gaucher). To our knowledge, this is the first specialized search tool for case reports that is built to assist diagnosis. This study provides a clear recipe for building and validating modern medical information retrieval systems to index and search complex and heterogeneous data. http://dx.doi.org/10.13039/100007723 Takeda Pharmaceuticals U.S.A. http://dx.doi.org/10.13039/501100009708 Novo Nordisk Fonden NNF20OC0062606 https://orcid.org/0000-0002-1966-3205 Winther Ole http://dx.doi.org/10.13039/100006785 Google Liévin Valentin The study was funded by Takeda. OW and VL are supported by the Novo Nordisk Foundation (NNF20OC0062606) and DeepMind/Google through their employment at UCph and DTU. Data AvailabilityThis project relies on two data sources: a collection of PubMed abstracts (https://huggingface.co/datasets/findzebra/case-reports) and clinical Takeda owned Outcome Survey data for Fabry (FOS) and Gaucher (GOS). De-identified records from 4484 Fabry and 1095 Gaucher patients served as the basis for selecting 20 Fabry and 20 Gaucher patients with atypical symptoms used in the expert validation. These 40 patient cases are available in S1 Data. Inquiries about the Takeda Outcome Survey can be addressed to GMA.Research@Takeda.com. Data Availability This project relies on two data sources: a collection of PubMed abstracts (https://huggingface.co/datasets/findzebra/case-reports) and clinical Takeda owned Outcome Survey data for Fabry (FOS) and Gaucher (GOS). De-identified records from 4484 Fabry and 1095 Gaucher patients served as the basis for selecting 20 Fabry and 20 Gaucher patients with atypical symptoms used in the expert validation. These 40 patient cases are available in S1 Data. Inquiries about the Takeda Outcome Survey can be addressed to GMA.Research@Takeda.com. ==== Body pmcIntroduction A disease is considered rare when it affects: in Europe 1:2000 and in the US about 1:1600 people. Currently, there are more than 6000 distinct rare diseases in the EU [1]. Around 80% of rare diseases are of genetic origin and, of those, 70% manifest already in childhood. Many rare diseases are chronic, progressive, and life-threatening. Early diagnosis may save lives, slow disease progression and/or prevent further irreversible organ damage, and ultimately improve the quality of life for these patients. A recent population-based telephone survey in Germany revealed a median duration of the diagnostic delay of 20 or more years for some rare lysosomal storage disorders (LSD) [2]. The diagnostic odyssey is complicated since the signs and symptoms may intuitively implicate more common pathologies, and most primary care professionals may have no experience with any of these disorders. Moreover, recognizing the disease trajectory is also complicated by variable ages of onset, the progressive natural history with change in manifestations with age, variable presentation of multiple and diverse organs and tissues, and the lack of awareness of specialised diagnostic markers. In today’s era of quick access to the internet and social media it is becoming equally true that patients find their diagnosis acting upon irrelevant information garnered from trawling these sources [3,4]. Two challenging rare diseases with complex phenotypes, a sufficient amount of patients available and good expert understanding are Fabry and Gaucher. Fabry disease is a rare X-linked inherited LSD [5]. Major morbidity of several organ systems often begins in childhood with gastrointestinal, dermatological, and often ocular signs; as patients age, cardiac and renal disease may rather rapidly culminate in end-stage organ failure; stroke and other cardiovascular events that are life-threatening are also common. Early intervention with disease-specific therapies, enzyme replacement therapies (ERT), or pharmacological chaperones (PC) is critical and universally recommended before the development of more devastating and irreversible renal, cardiac, and/or cerebrovascular signs [6]. Gaucher disease has a clinical spectrum from a perinatal lethal neuronopathic form (type 2), to a form that has variable neurological and visceral signs (types 3a, 3b, and 3c), to a chronic, non-neuronopathic form (type 1), where some patients are truly asymptomatic through-out their normative life-spans but others suffer from mild to severe visceral and haematological signs that appear variously anywhere from childhood to old age [7]. All these patients can benefit from early administration of several disease-specific treatment options to some extent: ERT and SRT will improve visceral and haematological signs as well as well-being; the SRTs and PCs can partly impact the neurological trajectory in the neuronopathic forms as well as the encroaching signs of Parkinson disease and other Lewy Body Dementia symptoms in type 1 patients [8,9]. Given the diagnostic delay, there are only a few published algorithms to assist in earlier diagnosis for either disease and none of them have been deployed into clinical use [10–13]. PubMed is the most complete information source on medical scientific knowledge. It comes with good information retrieval capabilities but is not designed for aiding diagnosis. This is demonstrated when benchmarking on medical cases against dedicated rare disease tools like FindZebra [14,15]. FindZebra, a search engine that has been popular within the medical community for the last past ten years, indexes a collection of curated disease articles from sources such as OrphaNet, OMIM and Wikipedia. The articles mostly describe generic disease phenotypes and therefore FindZebra only offers limited coverage of the more exceptional disease phenotypes, which are described in the specialised medical literature. A tool that applies advanced information retrieval techniques to searching the specialised literature has to our knowledge so far been missing. The case reports registered on PubMed are great candidates for extending the coverage of FindZebra. Over two million case reports are registered, and each one of them covers a unique medical case. Case reports could be used to improve the clinical management of today’s patients, for instance by tailoring the treatment to the patient profile, or by supporting the healthcare providers in their disease management choices [16,17]. However, case reports show higher variability in style and quality than the curated articles gathered from OrphaNet, OMIM, Wikipedia, etc. Searching case reports is thus a more complex task and care must be taken to prioritize the information that is essential to recognize disease phenotypes in both the case reports and the search queries (age, sex, symptoms, genetics, negative findings, etc.). The proposed tool relies on two main components (Fig 1) built on publicly available tools (PubMedBERT [18] and Solr [19]): a deep learning model that allows transforming unstructured text documents into structured patient profiles a full-text search engine (Solr) that allows searching for similar patients across the profile dimensions. 10.1371/journal.pdig.0000269.g001 Fig 1 Converting unstructured PubMed abstract into a structured search index using 1) text segmentation and 2) a Solr instance with composite fields corresponding to the text segmentation categories. In this study, we evaluate the tool based on real Fabry and Gaucher patients to identify issues particular to the task that will be essential for rolling the tool out to all rare diseases. In summary, the main contributions of the study are: 1) a recipe for building and validating an information retrieval system for heterogeneous medical information and 2) the search tool made available at deep.findzebra.com. Material and methods We discuss how the novel search tool integrates with the existing FindZebra.com. We present the collection of case report abstracts indexed by our system before detailing the clinical data used to evaluate the tool. We conclude by describing the development of the search engine (segmentation of the abstracts and search ranking algorithm setup) and last the setup of the expert validation. FindZebra workflow FindZebra.com allows searching across a collection of curated medical articles. For canonical disease profiles, this step is sufficient to find information that is relevant to the patient. For rare phenotypes, this new tool allows “diving in” the large pool of case reports within a particular disease to retrieve case reports that match the patient profile (Fig 2). 10.1371/journal.pdig.0000269.g002 Fig 2 Workflow for retrieving documentation relevant to atypical patients. The patient presents with two typical findings where one leads to identification of Disease A. The case report search (the contribution of this paper) leads to a case report with the same atypical symptom combination. Data PubMed case reports We collected case reports from 803 PubMed articles for Fabry disease and 883 for Gaucher disease. For each article, we retained only the abstract, which in most cases summarises information about the case at hand. We detail the data collection process in S1 Text. Clinical data—Fabry and Gaucher Outcomes Surveys We based the study on the real-world data from long-term observational Fabry and Gaucher Outcomes Surveys, FOS and GOS respectively. FOS and GOS aim at improving the clinical management of patients (see the S1 Text for further details). This was a non-interventional study, limited to the use of readily available data. It did not involve recontacting patients, and the informed consents, captured for the original FOS and GOS, allowed the use of their data for the validation purposes of this study. Data was anonymized by removing all information that could potentially identify a patient; a new randomization number was assigned to the Patient ID, all other potential identifiers, such as country, site name and date of birth were removed, as well as all other dates, e.g. visit dates and dates of laboratory assessments. Records from 4484 Fabry patients and 1095 Gaucher patients were collected. Each record features demographic information (age, sex and mutation when available), a list of signs and symptoms and a quantitative evaluation of the relevant organs (Fabry: eGFR, LVMI; Gaucher: liver size, spleen size, haemoglobin value and platelet count). We selected two anonymized patients, called patient F for Fabry and patient G for Gaucher. Their records are presented in Table 1. Throughout the text, we use patients F and G to showcase the data processing and evaluation steps. 10.1371/journal.pdig.0000269.t001 Table 1 Examples of survey data for Fabry and Gaucher patients. Fabry patient F Gaucher patient G Diagnosis Fabry Diagnosis Gaucher Age 66 Age 43 Sex male Sex male CKD stage 4 Haemoglobin 154 g/l eGFR 76·02 ml/min/1·73m2 Platelet count 87 109/l LVMI 109·51 g/m**2·7 Liver size 4·60 multiples of normal ·· ·· Spleen size 0·85 multiples of normal Symptoms sign angiokeratomas, sign haemorrhoids, sign lv hypertrophy, symptoms vertigo, sign arrhythmia, haematuria, sign stroke, tumours, heart failure Symptoms lipid profile-low ldl, jaw-big osteolytic lesion, elevated ast, no hepatosplenomegaly Converting survey entries to search queries We converted the survey data (tabular format) into full-text queries and numerical features into the corresponding signs using reference tables [20,21]. For instance, a low platelet count value was translated as “thrombocytopenia” whereas a normal value is converted as “no thrombocytopenia”. The resulting textual features are combined into a comma-separated list of terms, see examples in Table 2. 10.1371/journal.pdig.0000269.t002 Table 2 Example of generated queries. Query Patient F male, elderly, 66-year-old, sign angiokeratomas, sign haemorrhoids, sign lv hypertrophy, sign arrhythmia, sign stroke, symptoms vertigo, haematuria, tumours, heart failure, severe chronic kidney disease Patient G 43-year-old, male, adult, thrombocytopenia, lipid profile-low ldl, jaw-big osteolytic lesion, elevated ast, no splenomegaly, no hepatosplenomegaly, normal haemoglobin level, no hepatomegaly We designed a segmentation model that transforms raw text into structured representations by extracting non-overlapping spans of text. Each span—or segment—is labelled using eight categories, which we summarise in Table 3. Each category is selected to represent a particular clinical feature that might be useful for diagnosis. 10.1371/journal.pdig.0000269.t003 Table 3 Segmentation categories. Each category represents a dimension of the patient profile. The categories are used to index the case reports and parse the queries. Category Examples 1 Sex male, female, her, his 2 Age young adult, 23-year-old, infant 3 Ethnicity Ashkenazi Jewish, African American 4 Diagnosis, Signs and Symptoms Fabry disease, Gaucher disease, Morbus Fabry, Chronic renal failure, recurrent posterior stroke-like symptoms, mild retardation, Gaucher cells, zebra bodies found in kidney biopsy, Creatinine level (200 micromol/L), anaemia, abnormal blood count 5 Medications and interventions Enzyme replacement therapy, splenectomy 6 Genetics c.427G>A (p.A143T) variant, rare mutation in the GBA gene 7 Negative findings Covid-19 negative, no history of diabetes, normal blood count 8 Family history History of early strokes in the family Text segmentation The model builds upon a domain-specific masked language model [22], PubMedBERT [18], which follows the same architecture as the popular BERT model [23]. BERT allows computing contextual language representations, which we augmented with a conditional random field likelihood to improve the local coherence of the segments [24]. We randomly selected and labelled 100 abstracts for each disease. Each of the 200 abstracts were manually labeled into text segments. We chose to label spans of text such that each span of text encapsulates a single clinical feature completely. This results in segments of text that might overlap multiple text entities (see examples in Table 3). 20 documents were set aside for testing and the remaining documents were used for training (160) and validation (20). We detail the fine-tuning process in S1 Text. Our implementation relies on popular machine learning libraries [25–27]. We used the same model for both parsing the user queries and indexing the abstracts. To make the model robust to both the abstracts and the comma-separated queries, we augmented the training data by swapping abstracts with pseudo-queries for half the training iterations. Pseudo-queries were obtained by concatenating N ~ Poisson(λ = 5) segments extracted from the replaced abstract. Search engine The search engine is built on a composite Solr index for each disease separately. The ranking function is a weighted combination of the BM25 scores computed across each segmentation category. We detail the configuration and design of the index in S1 Text. We analysed the corpora based on the extracted profiles. We used the SciSpaCy library to link symptoms to the Unified Medical Language System (UMLS) entities [28,29]. We report the frequency of symptoms as well as a summary of the demographic data in S1 Text. Validation protocol The case report search engine is specifically designed for the use cases where the diagnosis is established or suspected but a deeper understanding of the phenotype is needed. Therefore, we focused on the subset of patients with atypical symptoms, which we defined in this study as the symptoms occurring in less than 10% of the population. Step 1—Preliminary non-expert validation During development, we inspected the quality of retrieval based on simple tests. For a selection of patients, we tested if the retrieved articles corresponded to the age, sex, mutations (if any) and the domain of symptoms (e.g., skeletal, psychiatric involvement, …). Step 2—Expert validation Two rare disease experts (DE and AL) evaluated the relevance of the search results given for the 20 patients for each disease. The anonymized data for the 40 patients is available in S1 Data. We selected patients with atypical phenotypes and diverse disease profiles. The experts were asked to evaluate the clinical relevance of each of the top three returned articles using a scale from one to five and using a text field. For each retrieved document, we report the maximum grade given by the two experts. For each patient, we report the precision for the top three results (P@3) based on the maximum grade and for multiple relevance thresholds. In S1 Text, we provide further details about the patient selection, the rating scale, the evaluation interface and experts’ agreement. Step 3—Population and corpus level analysis To gain a better understanding of the diseases, their differences and how it affects retrieval, we analysed how the search engine maps the population of patients to the corpus of PubMed articles. We built a bipartite graph, using patients and articles as nodes, and created edges if an article was retrieved as top three. We used the resulting network to study the relationship between patients and PubMed articles. Ethics statement This was a retrospective, non-interventional study, limited to the use of readily available patient data in Takeda-owned Fabry and Gaucher Outcomes surveys (FOS and GOS, respectively). The written informed consents obtained from the patients participating in the original FOS and GOS, allowed the use of their data for the validation purposes of this study which did not involve recontacting patients. Therefore, this study was not a subject for Ethics Committee approval. Data was furthermore anonymized by removing all information that could potentially identify a patient; a new randomization number was assigned to the Patient ID, all other potential identifiers, such as country, site name and date of birth were removed, as well as all other dates, e.g. visit dates and dates of laboratory assessments. Results This section begins with a quantitative and qualitative evaluation of the final segmentation model. It continues with a description of the population of the selected patients. We then review the case report search tool in three acts: i) we display the case of two patients, ii) we present the expert review and iii) we illustrate how the search tools maps the cohort of patients to the PubMed corpus. As presented in S1 Text, the two diseases exhibit different profiles of symptoms, which supports the need to adapt the ranking function to each disease. Text segmentation The final model scored 0·75 F1 score on the test set (0·76 F1 score on the validation set). The model appeared to be robust to a wide diversity of case reports and user queries. In Fig 3, we present a Fabry case report segmented using the final model. In Table 4, we present an example of a segmented query. In S1 Text, we present three additional labelled abstracts: one for Fabry, one for Gaucher and one out-of-domain example (COVID-19, see Supplement III). 10.1371/journal.pdig.0000269.g003 Fig 3 Segmentation example of the article (test set): “Two cases of Fabry’s disease: A hemizygote with a point mutation in the alpha-galactosidase A gene and his relative [30]”. Each colour corresponds to one of eight segmentation categories listed in Table 3. 10.1371/journal.pdig.0000269.t004 Table 4 Example of segmented query. Fabry patient F Gaucher patient G age 66 43 sex male male symptoms sign angiokeratomas, sign haemorrhoids, sign lv hypertrophy, sign arrhythmia, sign stroke, symptoms vertigo, haematuria, tumours, heart failure, severe chronic kidney disease lipid profile-low ldl, jaw-big osteolytic lesion, elevated ast Negative findings - no splenomegaly, no hepatosplenomegaly, normal haemoglobin level, no hepatomegaly Clinical data and queries We summarise the demographic features as well as the distribution of symptoms in S1 Text. We found that the populations of patients from the survey and from the PubMed articles follow similar demographics. Furthermore, we found that symptoms stated in the records were often discussed in the PubMed corpora. The retrieval workflow (Fig 2) has two steps. For completeness, we also evaluate step 1 (FindZebra search) for all patients in the two Surveys. The correct diagnosis appeared in the top ten search results for 68·4% of the Fabry patients, whereas this was the case for only 21·7% of the Gaucher patients (see S1 Text for further details). Subsets of typical and atypical patients The medical expert validation of rare disease search (step 2 in Fig 2) focuses on the group of patients with atypical symptoms. Using the criteria defined in the previous section, we labelled 56% of the Fabry patients and 64% of the Gaucher patients as atypical. In Table 5, we report the number of patients for each group, the proportion of patients treated for Fabry or Gaucher and the mean number of symptoms recorded for each patient. We found that atypical patients have on average twice the number of symptoms and were more likely to be treated than the typical ones, indicating a more serious form of the disease. 10.1371/journal.pdig.0000269.t005 Table 5 Statistics for the groups of typical and atypical patients. Fabry Gaucher Group all atypical typical all atypical typical Patients 4484 2491 1993 1079 685 394 Treated (%) 56 66 39 82 84 79 Mean number of symptoms 6·6 9·8 3·1 5·6 6·6 3·6 Expert validation We found considerable disparities in the evaluation of the two diseases. Whereas retrieval was judged to be effective when applied to the Fabry patients, articles were more often judged as irrelevant for the Gaucher patients. We report the precision in Table 6 and display the distribution of maximum grades given to each article in Fig 4. 10.1371/journal.pdig.0000269.t006 Table 6 Precision given 3 retrieved articles per patient (20 patients). Each document is labelled as relevant if the maximum rating given by the two experts (AL and DE) is greater or equal to the threshold. The rating scale is described in S1 Text. Threshold Fabry (P@3) Gaucher (P@3) Max. rating ≥ 2 98·3% 35·0% Max. rating ≥ 3 88·3% 13·3% Max. rating ≥ 4 51·7% 8·3% Max. rating ≥ 5 15·0% 6·7% 10.1371/journal.pdig.0000269.g004 Fig 4 Distribution of scores assigned to each document for each disease (Fabry disease left and Gaucher disease right). For 20 patients per disease, we retrieve the top-3 abstracts. For each article, we use the maximum among the two expert ratings as evaluation score. In the case of the Fabry patients, a minority of articles were judged irrelevant (11·7% of patients were assigned with a maximum rating lower than three) and 51·7% of the search results were graded with a maximum rating of at least four. The Gaucher patients were more difficult to match with relevant case reports, as 65·0% of the articles were rated one. Only 8·3% of the articles received at least one grade above three. The experts’ comments revealed six failure patterns listed in Table 7. The most common cause of failure (2 for Fabry, 18 for Gaucher) was attributed to retrieving articles that presented a radically different clinical picture, despite sharing a few symptoms and/or demographic features with the referenced patient. The second most prevalent cause of failure was associated with returning abstracts that were no longer considered valid by the medical experts. Other identified causes were diagnosis mismatch (failure pattern #3), missing data about the reference patient (#4), symptom mismatch (failure pattern #5), and age mismatch (failure pattern #6). 10.1371/journal.pdig.0000269.t007 Table 7 Failure patterns. Number of identified failure patterns for each disease. # Pattern # Fabry # Gaucher 1 The article presents a rare or very specific clinical profile or treatment that is irrelevant to the patient despite significant lexical overlap and/or similar symptoms (“buzz words”) 2 18 2 The article is outdated, its content is no longer clinically relevant 2 11 3 The case presented in the article presents similar clinical features, but the diagnosis was dissimilar (similar features, different causes) 2 6 4 The article studies a specific characteristic (mutation, ethnicity, family history), that is unknown for the reference patient 3 3 5 The case and the patients don’t share similar symptoms, the case might have been matched solely based on the age and sex of the reference patient 1 4 6 Age mismatch (infant vs. adult) 1 4 Case study In S1 Text, we illustrate the whole retrieval and validation process based on patients F and G. For each patient, we present the raw data, the segmented queries, and the retrieved abstracts associated with their corresponding expert ratings and comments. Population and corpora To illustrate how the search scoring function maps a population of patients to the PubMed case reports, we sampled a subset of 500 patients for each disease to exclude the effect of the population size on the analysis. We created a patient-article network for each disease using the top three retrieved articles and the subset of patients. Both networks are visualised in Fig 5. We provide an analysis of the networks in S1 Text. 10.1371/journal.pdig.0000269.g005 Fig 5 Visualization of the patient-article networks for Gaucher (left) and Fabry (right). Nodes represent patients and articles; edges are drawn if an article is retrieved as the top three for a given patient. Each colour corresponds to a cluster of patients and articles. This visualization shows how the population of patients maps to the corpus of articles. The Fabry network shows a higher degree of clustering than the Gaucher Network. We recorded the number of retrieved articles for each disease as well as the mean number of patients connected to each article. We found significant differences between the two diseases: Fabry articles were connected to 7·6 patients on average (with a total of 295 retrieved articles) whereas Gaucher articles were connected to 3·9 patients per article (with a total of 386 articles). Discussion We have built a search engine that allows searching case reports based on patient features that are automatically extracted from the query and the indexed reports using deep learning. The tool allows case reports based on multiple features (sex, age, gender, mutation, symptoms, etc), which performs robustly thanks to the simple BM25-based design. Nonetheless, we tested more than semantic overall: we evaluated if the top three search results were clinically relevant according to medical experts. The articles were more often judged as clinically relevant for Fabry patients than for Gaucher patients, and we felt, looking at our results, that this could be partly explained by the dichotomy in explanations of the clinical manifestations and by the divergence in disease-specific management options of these two diseases. Fabry disease is an X-linked disorder which implies that the male patients are generally more severely affected and at an earlier age than females plus there is also the impact of the various mutations that may be predictive of a specific phenotypic expression in both genders. On the other hand, Gaucher disease is usually divided into genotypes, and each genotype might lead to radically different disease trajectories (e.g., lethal, severe neuropathic genotype versus mild, non-neuropathic genotype). Therefore, matching Gaucher patients was more challenging, because the greater diversity of profiles made it easier to miss, and because age and sex were not as informative as in Fabry (Gaucher is not X-linked). This highlights the limitations of our ranking function. In some cases, we found the ranking function to be misaligned with the expert judgement, as it placed too little weight on sex (failure pattern #6), or on symptoms (failure pattern #5). In other cases, failure was attributed to BM25, as it only allows handling symptoms independently, thus failing to grasp the whole clinical picture. This might lead to placing too much weight on a few rare terms (“buzz words”), which is linked to failure pattern #1. Furthermore, we found that many retrieved articles were outdated, especially in the case of the Gaucher disease (failure pattern #1), this is explained by the recent changes in the clinical management of the two diseases. In Fabry disease, the triad of end-stage organ failure of renal, cardiac, and cerebrovascular events in patients may be partly prevented, but they are still irreversible once established. The ramifications on the other organ systems remain poorly controlled, which may also impact quality of life and longevity, so that patients today face many of the same challenges as those of decades past. However, disease management in Gaucher disease has been transformed. Several new modalities of therapeutics can now assure normative function by reversing visceral signs and symptoms in the non-neuropathic patients. Furthermore, the disease phenotypes have evolved due to a tendency to diagnose earlier and due to longer survival. Ultimately, this was seen in our study which underscored the explosion of recent developments for the several types of Gaucher, so that older case reports were of limited value for patients being seen today. This uncovers a broader problem that not all case reports contain valid information. Whereas validating the content of case reports is challenging, recency can be easily controlled. For the released version of the search tool, we added the possibility of filtering on publication date. The validation protocol was designed to mimic the real-world use of our tool, but a discrepancy remains between the evaluation setup and the real-world usage. First, we used all the recorded data for each patient, whereas in a clinical context, the healthcare professionals rely only on the subset of the features that might be relevant in that particular clinical setting. Second, the generated queries only contained information about age, sex, mutation, symptoms and negative findings. Our tool handles additional profile dimensions such as medications, family history and ethnicity (which is a critical factor in the management of Fabry and Gaucher diseases). Third, whereas information systems are traditionally evaluated using the top ten results, we restricted the evaluation to the top three results. Scrolling past the top three results might be required in the more difficult cases such as the ones observed in the Gaucher evaluation. Lastly, we acknowledge the challenge and limitations of data anonymization especially in the rare disease space. However, as the FOS and GOS patient data collected for the purpose of this study is global and consists of relatively high numbers of patients, we consider anonymization sufficient. Conclusion We have built a search engine specialised for case reports and submitted it to a thorough expert validation process using real-world clinical Outcomes Surveys data. To the best of our knowledge, our tool pioneers the task of indexing and retrieving cases reports with the aim of aiding diagnosis of rare diseases. It allows browsing large quantities of case reports natural language and clinical descriptions using a structured ranking function (age, sex, mutation, symptom, etc…). Our approach details a general approach that can be used to make the clinical literature more readily useful for healthcare practitioners. Based on real-world rare diseases information, we found that retrieved articles were often clinically relevant for the Fabry patients, whereas articles retrieved for the Gaucher patients had less clinical value. Further analysis of the expert comments and the patient-abstract network allowed us to identify shortcomings associated with our method. Our study highlights the gap that remains between modern search technologies and clinical practice. Although this study was restricted to the Fabry and Gaucher diseases, we will now focus on scaling the process to all rare diseases recorded at FindZebra.com. We learned from the expert evaluation and will use their feedback to improve our tool, by adding a temporal filter to the query field. It is hoped that our tool will be used in the field to help the healthcare professionals to improve the clinical management of the many patients who suffer from rare diseases. Our tool will remain publicly available. Supporting information S1 Text The docx file contains the supplementary information referenced in the main text. (DOCX) Click here for additional data file. S1 Data This zip file contains the 40 patient cases in JSON format (20 Fabry and 20 Gaucher) used for the evaluation. (ZIP) Click here for additional data file. 10.1371/journal.pdig.0000269.r001 Decision Letter 0 Pant Pai Nitika Section Editor Harrison Ewen M. Academic Editor © 2023 Pant Pai, Harrison 2023 Pant Pai, Harrison https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Submission Version0 26 Dec 2022 PDIG-D-22-00278 FindZebra online search delving into rare disease case reports using natural language processing PLOS Digital Health Dear Dr. Winther, Thank you for submitting your manuscript to PLOS Digital Health. After careful consideration, we feel that it has merit but does not fully meet PLOS Digital Health's publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. ============================== EDITOR: We would like you to respond to the comments made by reviewers. Please proof read your manuscript before submission. ============================== Please submit your revised manuscript within 60 days Feb 24 2023 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at digitalhealth@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pdig/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. Please include the following items when submitting your revised manuscript: * A rebuttal letter that responds to each point raised by the editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. * A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. * An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. We look forward to receiving your revised manuscript. Kind regards, Ewen M. Harrison, PhD FRCS Academic Editor PLOS Digital Health Journal Requirements: 1. Please amend your detailed Financial Disclosure statement. This is published with the article. It must therefore be completed in full sentences and contain the exact wording you wish to be published. a. State the initials, alongside each funding source, of each author to receive each grant. b. State what role the funders took in the study. If the funders had no role in your study, please state: “The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.” c. If any authors received a salary from any of your funders, please state which authors and which funders. If you did not receive any funding for this study, please simply state: “The authors received no specific funding for this work.” 2. We ask that a manuscript source file is provided at Revision. Please upload your manuscript file as a .doc, .docx, .rtf or .tex. 3. Please provide separate figure files in .tif or .eps format only and remove any figures embedded in your manuscript file. Please also ensure that all files are under our size limit of 10MB. For more information about figure files please see our guidelines: https://journals.plos.org/globalpublichealth/s/figures https://journals.plos.org/globalpublichealth/s/figures#loc-file-requirement 4. In the online submission form, you indicated that "This project relies on two data sources: a collection of PubMed abstracts and clinical Takeda owned Outcome Survey data for Fabry (FOS) and Gaucher (GOS). De-identified records from 4484 Fabry and 1095 Gaucher patients were used for validation purposes as described in the Material and Methods, and some de-identified raw data is released in the Results. Neither the entire raw Outcomes Survey data nor related documents will be made publicly available. The list of the PubMed articles used in this study will be made available on request.". All PLOS journals now require all data underlying the findings described in their manuscript to be freely available to other researchers, either 1. In a public repository, 2. Within the manuscript itself, or 3. Uploaded as supplementary information. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If your data cannot be made publicly available for ethical or legal reasons (e.g., public availability would compromise patient privacy), please explain your reasons by return email and your exemption request will be escalated to the editor for approval. Your exemption request will be handled independently and will not hold up the peer review process, but will need to be resolved should your manuscript be accepted for publication. One of the Editorial team will then be in touch if there are any issues. Additional Editor Comments (if provided): [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Does this manuscript meet PLOS Digital Health’s publication criteria? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented. Reviewer #1: Yes Reviewer #2: Partly -------------------- 2. Has the statistical analysis been performed appropriately and rigorously? Reviewer #1: N/A Reviewer #2: N/A -------------------- 3. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)? The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified. Reviewer #1: Yes Reviewer #2: Yes -------------------- 4. Is the manuscript presented in an intelligible fashion and written in standard English? PLOS Digital Health does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here. Reviewer #1: Yes Reviewer #2: Yes -------------------- 5. Review Comments to the Author Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters) Reviewer #1: Thank you for submitting this interesting work. The authors describe the evaluation of extension of the search engine FindZebra with PubMed search for case reports based on patient characteristics. To do this, they conducted an extensive validation process by medical experts. The validation of the approach as performed against the Fabry and Gaucher disease. Feedback from the experts was highlighted and will be part of future research. The study contributes to current research in NLP and rare diseases and is of high importance. The approach presented is a useful combination of established knowledge sources or useful extension of an established tool. MAJOR ISSUES (1) Focus of the article should be sharpened: Should the study describe the evaluation process or how the tool works, or both? (2) Introduction: (2a) The combined use of already existing models should be added when mentioning the two components for the tool (p. 3). Otherwise, the impression is created that the components were completely newly developed by the authors. (2b) The goal of the study is stated as the evaluation of the tool (p. 3). However, the title and the methods also describe the functional area of the extension. Therefore, the research objective (and possibly also the title) should be modified accordingly. (3) Material and methods: (3a) When explaining the workflow, only the goal is named and not the method of functioning. A more detailed explanation – especially of Figure 2 – is necessary here (p. 3) (4) Results: (4a) The description for "Search index" should focus more on actual results and less on the procedure (p. 7). (4b) It is not always clear which results were obtained by which methods; especially with respect to the validation process, it would be interesting to know which results resulted from which step. Here, a reference to the individual steps would be important. (An example of this is the evaluation of the previous FindZebra tool (p. 8)). MINOR ISSUES (5) Introduction: (5a) Transition from explanations of diseases to explanation of PubMed could seem more natural by adding the research gap again (p. 2). (5b) Claim that a tool is missing is not substantiated, so perhaps add "to our knowledge" (p. 3). (5c) Incomplete sentence: “Case reports could be used to improve the clinical management of today’s patients, for instance by tailoring the treatment to the patient profile, or by supporting the healthcare providers in their [16,17]” (p. 3). (5d) Figure 1 is referenced when the two components are mentioned, but the two components are not explicitly shown in the figure. Here, you could highlight the individual components separately (p. 3). (6) Material and Methods: (6a) The introduction to the chapter "Material and methods" does not exactly match the following subchapters. The title and sequences should be harmonized here (p. 3). (6b) Is this really de facto anonymization? Or can individual patients be identified by the combination of sex, age and mutation (p. 4)? (6c) Explanation of text segmentation should be under the heading "Text Segmentation" and not be content of "Data" (p. 5). (7) Results: (7a) The introduction to the chapter "Results" does not exactly match the following subchapters. The title and sequences should be harmonized here (p. 7). (7b) Revision of figures and tables necessary: (7b.i) Figure 4: Name / describe units (p. 10). (7b.ii) Labeling of table 7 differs from other labeling (p. 9). (7b.iii) No table 6 included (p. 8-9). (7b.iv) Two figures are named “Figure 4” (p. 9 and 11). (8) Conclusion: (8a) Improve the conclusion by describing the novelty of the approach in more detail and by stating scientific and technical implications for the entire research field (p. 13). OTHER POINTS (9) Great unified and beautiful presentation of illustrations. (10) Great detailed discussion of the results. Reviewer #2: The manuscript describes a new search functionality for rare disease case reports of FindZebra.com service. Two rare diseases of Gaucher and Fabry were used as the exemplar study diseases. Natural language processing models/tools were utilised to support the modelling/indexing of case reports and the matching between user queries and case reports. Particularly, user queries were generated from 'real' patient cases including age/gender and clinical features like symptoms and phenotypes. Evaluation protocols and metrics were proposed to validate the utilities of the service in supporting clinical decision making for patients with those diseases, seemingly in scenarios of both diagnosis and treatments. Overall, from clinical utility point of view, this would potentially be a very valuable work and much needed service for supporting rare disease diagnosis, treatment and managements. However, technically - from information retrieval and NLP point of view, the work requires further developments and clarifications to make it publishable - in other words, making substantial contribution to the field and useful for the community. 1. It is not clear how NLP models were used and developed. Named entity recognition was mentioned only in the abstract. In the main text and supplementary it was called segmentations. The two might be totally different NLP tasks. The use of terminology aside, there is no information how the NER was done. PubmedBERT was mentioned to be the language model for fine-tuning the segmentation task (assuming the NER for clinical features like symptoms etc). However, there was no mention where the ground truth of NER came from. 2. The ranking algorithm is key in the methodology of an information retrieval system. But it seems very limited contribution was proposed to that aspect - the default BM25 was used instead. 3. The key evaluation result of precision@3 seems not very good. There were 5 grades as detailed in the supplementary. From the descriptions, only grade 4 & 5 can be assessed to be relevant. From table 7, the 'real' useful results for the two diseases were 52% and 8% for the two diseases respectively. This seems a bit low for a search service according to these numbers. There was no baseline provided to compare their service to. Also, there was no ablation evaluation on different techniques. Therefore, it is not clear how difficult the task was and it is hard to justify how good their service was. 4. From table 8 (analysis of the failure patterns), first, clearly the domain experts were evaluating the service from facilitating diagnosis/treatment points of view. For example, item 2 was "The article is outdated", which is clearly not a direct IR/NLP problem. However, the method/models proposed by the authors seem not dealing with these requirements directly. This leads to a general question - have the authors properly defined their technical tasks based on the intended use of the service? 5. A side and puzzled finding from table 8 was the big difference of the last two rows between the two diseases. There seems to be no obvious reasons why the second disease was harder for the same (relatively simpler) IR tasks. 6. I struggled to understand the relevance of "Population and Corpora" section. First, it seems not directly relevant to information retrieval and NLP tasks. Second, the networks/graphs were generated from top 3 articles of the service. However, the performance of the current service as it is seems not good enough to support the generation of a reasonably good network for this kind of analysis. Presentation issues: - There are no justifications on why the two rare diseases. Do they present two exemplar (distinct) IR/NLP challenges? - The abstract does not provide clear descriptions on what the main validation purposes are. - There are many abbreviations which should be expanded when they were first used. - There are grammar and incomplete sentences at various places in the manuscript. -------------------- 6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files. Do you want your identity to be public for this peer review? If you choose “no”, your identity will remain anonymous but your review may still be made public. For information about this choice, including consent withdrawal, please see our Privacy Policy. Reviewer #1: No Reviewer #2: Yes: Honghan Wu -------------------- [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step. 10.1371/journal.pdig.0000269.r002 Author response to Decision Letter 0 Submission Version1 3 Mar 2023 Attachment Submitted filename: Response to reviewers.pdf Click here for additional data file. 10.1371/journal.pdig.0000269.r003 Decision Letter 1 Pant Pai Nitika Section Editor Harrison Ewen M. Academic Editor © 2023 Pant Pai, Harrison 2023 Pant Pai, Harrison https://creativecommons.org/licenses/by/4.0/ This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Submission Version1 3 May 2023 FindZebra online search delving into rare disease case reports using natural language processing PDIG-D-22-00278R1 Dear Prof. Winther, We are pleased to inform you that your manuscript 'FindZebra online search delving into rare disease case reports using natural language processing' has been provisionally accepted for publication in PLOS Digital Health. Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow-up email from a member of our team.  Please note that your manuscript will not be scheduled for publication until you have made the required changes, so a swift response is appreciated. IMPORTANT: The editorial review process is now complete. PLOS will only permit corrections to spelling, formatting or significant scientific errors from this point onwards. Requests for major changes, or any which affect the scientific understanding of your work, will cause delays to the publication date of your manuscript. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they'll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact digitalhealth@plos.org. Thank you again for supporting Open Access publishing; we are looking forward to publishing your work in PLOS Digital Health. Best regards, Sarah Mayo Staff Admin PLOS Digital Health *********************************************************** Reviewer Comments (if any, and for reference): ==== Refs References 1 Rare diseases. (n.d.). Retrieved April 4, 2022, from https://ec.europa.eu/health/non-communicable-diseases/steering-group/rare-diseases_en. 2 Mengel E , Gaedeke J , Gothe H , Krupka S , Lachmann A , Reinke J , et al . The patient journey of patients with Fabry disease, Gaucher disease and Mucopolysaccharidosis type II: A German-wide telephone survey. PLoS One. 2020;15 : e0244279. doi: 10.1371/journal.pone.0244279 eCollection 2020. .33382737 3 Nicholl H , Tracey C , Begley T , King C , Lynch AM . Internet Use by Parents of Children With Rare Conditions: Findings From a Study on Parents’ Web Information Needs. J Med Internet Res. 2017; 19 : e51. doi: 10.2196/jmir.5834 28246072 4 Wake Forest Baptist Medical Center. "Internet can be valuable tool for people with undiagnosed rare disorders." ScienceDaily 2019 Aug 7. . 5 Kok K , Zwiers KC , Boot RG , Overkleeft HS , Aerts JMFG , Artola M . Fabry Disease: Molecular Basis, Pathophysiology, Diagnostics and Potential Therapeutic Directions. Biomolecules. 2021;11 : 271. doi: 10.3390/biom11020271 .33673160 6 Hughes DA , Aguiar P , Lidove O , Nicholls K , Nowak A , Thomas M , et al . Do clinical guidelines facilitate or impede drivers of treatment in Fabry disease? Orphanet Journal of Rare Diseases. 2022;17 : 42. doi: 10.1186/s13023-022-02181-4 35135579 7 Zimran A, Elstein D. Lipid storage diseases. In: K. Kaushansky, M, Lichtman, J Prchal, M.M. Levi, O. Press, L. Burns, M. Caligiuri (Eds.), Williams Hematology, 9th edition; New York: McGraw-Hill, Chapter 72 (2016). 8 Revel-Vilk S , Szer J , Mehta A , Zimran A . How we manage Gaucher Disease in the era of choices. Br J Haematol. 2018;182 : 467–480. doi: 10.1111/bjh.15402 29808905 9 Mehta A , Kuter DJ , Salek SS , Belmatoug N , Bembi B , Bright J , et al . Presenting signs and patient co-variables in Gaucher disease: outcome of the Gaucher Earlier Diagnosis Consensus (GED-C) Delphi initiative [published correction appears in Intern Med J. 2019 Aug;49(8 ):1059]. Intern Med J. 2019;49 : 578–591.31387147 10 Mehta A , Rivero-Arias O , Abdelwahab M , Campbell S , McMillan A , Rolfe MJ , et al . Scoring system to facilitate diagnosis of Gaucher disease. Intern Med J. 2020; 50 : 1538–1546. doi: 10.1111/imj.14942 .33174353 11 Savolainen MJ , Karlsson A , Rohkimainen S ,Toppila I , Lassenius MI , Vaca Falconi C , et al . The Gaucher earlier diagnosis consensus point-scoring system (GED-C PSS): Evaluation of a prototype in Finnish Gaucher disease patients and feasibility of screening retrospective electronic health record data for the recognition of potential undiagnosed patients in Finland. Molecular Genetics and Metabolism Reports. 2021;21 : 100725. doi: 10.1016/j.ymgmr.2021.100725 33604241 12 Jefferies JL , Spencer AK , Lau HA , Nelson MW , Giuliano JD , Zabinski JW , et al . A new approach to identifying patients with elevated risk for Fabry disease using a machine learning algorithm. Orphanet J Rare Dis. 2021 20;16 : 518. doi: 10.1186/s13023-021-02150-3 .34930374 13 Andrade-Campos MM , de Frutos LL , Cebolla JJ , Serrano-Gonzalo I , Medrano-Engay B , Roca-Espiau M , et al . Identification of risk features for complication in Gaucher’s disease patients: a machine learning anal NNF20OC0062606ysis of the Spanish registry of Gaucher disease. Orphanet J Rare Dis. 2020;15 : 256. doi: 10.1186/s13023-020-01520-7.32962737 14 Dragusin R , Petcu P , Lioma C , Larsen B , Jørgensen HL , Cox IJ , et al . FindZebra: a search engine for rare diseases. Int J Med Inform. 2013;82 : 528–538. doi: 10.1016/j.ijmedinf.2013.01.005 Epub 2013 Feb 23. .23462700 15 Svenstrup D , Jørgensen HL , Winther O . Rare disease diagnosis: A review of web search, social media and large-scale data-mining approaches. Rare Diseases. 2015;3 :1. doi: 10.1080/21675511.2015.1083145 26442199 16 Kawamoto K , Houlihan CA , Balas EA , Lobach DF . Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success. BMJ. 2005;330 : 765. doi: 10.1136/bmj.38398.500764.8F Epub 2005 Mar 14. 15767266 17 Garg AX , Adhikari NKJ , McDonald H , Rosas-Arellano MP , Devereaux PJ , Beyene J , et al . Effects of computerized clinical decision support systems on practitioner performance and patient outcomes: a systematic review. JAMA. 2005;293 : 1223–1238. doi: 10.1001/jama.293.10.1223 .15755945 18 Gu Y , Tinn R , Cheng H , Lucas M , Usuyama N , Liu X , et al . Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans. Comput. Healthcare. 2021;3 :1, Article 2 (January 2022), 23 pages. doi: 10.1145/3458754 19 Sparck JK , Walker S , Robertson SE . "A probabilistic model of information retrieval: development and comparative experiments: Part 2". Information processing & management 36.6 (2000 ): 809–840. 20 Zimran A , Elstein D , Gonzalez DE , Lukina EA , Qin Y , Dinh Q , et al . Treatment-naïve Gaucher disease patients achieve therapeutic goals and normalization with velaglucerase alfa by 4 years in phase 3 trials. Blood Cells Mol Dis. 2018;68 : 153–159. doi: 10.1016/j.bcmd.2016.10.007 Epub 2016 Oct 21. .27839979 21 Kampmann C , Linhart A , Baehner F , Palecek T , Wiethoff CM , Miebach E , et al . Onset and progression of the Anderson-Fabry disease related cardiomyopathy. Int J Cardiol. 2008;130 : 367–373. doi: 10.1016/j.ijcard.2008.03.007 18572264 22 Lee J , Yoon W , Kim S , Kim D , Kim S , So CH , et al . BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Bioinformatics. 2020;36 : 1234–1240, doi: 10.1093/bioinformatics/btz682 31501885 23 Devlin J, Chang M-W, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2019;1 (Long and Short Papers): pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics. 24 Lafferty JD, McCallum A, Pereira FCN. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. In Proceedings of the Eighteenth International Conference on Machine Learning (ICML ’01). Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 282–289. 25 Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics. 26 Falcon W. "Pytorch lightning" GitHub. Note: https://github.com/PyTorchLightning/pytorch-lightning 3 (2019): 6. 27 Liaw R, Liang E, Nishihara R, Moritz P, Gonzalez JE Stoica I., 2018. Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118. 28 Neumann M, King D, Beltagy I, Ammar W. 2019. ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 319–327, Florence, Italy. Association for Computational Linguistics. 29 Bodenreider O. The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Res. 2004 Jan 1; 32 (Database issue ): D267–70. doi: 10.1093/nar/gkh061 14681409 30 Inaoki M , Otsuki N , Ishise S , Ueda Y , Sakuraba H . Two cases of Fabry’s disease: a hemizygote with a point mutation in the alpha-galactosidase A gene and his relative. J Dermatol. 1992;19 : 481–486. doi: 10.1111/j.1346-8138.1992.tb03266.x 1328341