
==== Front
Comput Struct Biotechnol J
Comput Struct Biotechnol J
Computational and Structural Biotechnology Journal
2001-0370
Research Network of Computational and Structural Biotechnology

S2001-0370(24)00275-7
10.1016/j.csbj.2024.08.016
Software/Web server Article
BioTextQuest v2.0: An evolved tool for biomedical literature mining and concept discovery☆
Theodosiou Theodosios a
Vrettos Konstantinos a
Baltsavia Ismini a
Baltoumas Fotis b
Papanikolaou Nikolas a
Antonakis Andreas Ν. a
Mossialos Dimitrios c
Ouzounis Christos A. d
Promponas Vasilis J. e
Karaglani Makrina f
Chatzaki Ekaterini f
Brandau Sven g
Pavlopoulos Georgios A. b
Andreakos Evangelos h
Iliopoulos Ioannis ioannis@uoc.gr
a⁎
a Division of Basic Sciences, University of Crete Medical School, Heraklion 71110, Greece
b Institute for Fundamental Biomedical Research, BSRC "Alexander Fleming", Vari, Athens 16672, Greece
c Department of Biochemistry and Biotechnology, University of Thessaly, 41500 Larissa, Greece
d Biological Computation & Computational Biology Group, AIIA Lab, School of Informatics, Aristotle University of Thessalonica, 57001 Thessalonica, Greece
e Bioinformatics Research Laboratory, Department of Biological Sciences, University of Cyprus, Nicosia 1678, Cyprus
f Medical School, Democritus University of Thrace, 68100 Alexandroupolis, Greece
g Experimental and Translational Research, Department of Otorhinolaryngology, University Hospital Essen, Essen, Germany
h Center for Immunology and Transplantation, Biomedical Research Foundation Academy of Athens, Athens, Greece
⁎ Corresponding author. ioannis@uoc.gr
21 8 2024
12 2024
21 8 2024
23 32473253
19 4 2024
5 8 2024
15 8 2024
© 2024 Published by Elsevier B.V. on behalf of Research Network of Computational and Structural Biotechnology.
2024

https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
The process of navigating through the landscape of biomedical literature and performing searches or combining them with bioinformatics analyses can be daunting, considering the exponential growth of scientific corpora and the plethora of tools designed to mine PubMed(®) and related repositories. Herein, we present BioTextQuest v2.0, a tool for biomedical literature mining. BioTextQuest v2.0 is an open-source online web portal for document clustering based on sets of selected biomedical terms, offering efficient management of information derived from PubMed abstracts. Employing established machine learning algorithms, the tool facilitates document clustering while allowing users to customize the analysis by selecting terms of interest. BioTextQuest v2.0 streamlines the process of uncovering valuable insights from biomedical research articles, serving as an agent that connects the identification of key terms like genes/proteins, diseases, chemicals, Gene Ontology (GO) terms, functions, and others through named entity recognition, and their application in biological research. Instead of manually sifting through articles, researchers can enter their PubMed-like query and receive extracted information in two user-friendly formats, tables and word clouds, simplifying the comprehension of key findings. The latest update of BioTextQuest leverages the EXTRACT named entity recognition tagger, enhancing its ability to pinpoint various biological entities within text. BioTextQuest v2.0 acts as a research assistant, significantly reducing the time and effort required for researchers to identify and present relevant information from the biomedical literature.

Keywords

Biomedical literature mining
Concept discovery
==== Body
pmc1 Introduction

The exponential growth of biomedical literature and the proliferation of repositories housing biological data present formidable challenges for researchers. PubMed® alone contains over 35 million MEDLINE entries, underscoring the sheer volume of information contained in publications from peer-reviewed journals. Furthermore, starting from January 2023 and onwards, the service also includes selected preprints from sources such as bioRxiv, medRxiv, and Research Square, further increasing its content (link: https://www.ncbi.nlm.nih.gov/pmc/about/nihpreprints/). At the same time, data from publications is frequently coupled with annotations derived from genomic, proteomic, and clinical repositories, presenting new opportunities and challenges. Overall, the volume of biomedical information demands meticulous tracking of advancements within specific domains and necessitates extracting actionable knowledge from scientific texts. This involves tasks such as identifying, visualizing, and analyzing connections between various biological entities and discovering novel concepts.

In the face of this data deluge, researchers increasingly rely on computational tools to sift through this information overload. Natural language processing methods, and especially text mining, have emerged as a vital tool to tackle the complexities inherent in biomedical research. Various applications have been developed to address diverse challenges, ranging from concept discovery in PubMed abstracts and OMIM records to identifying associations between entries in various databases.

Examples include BioTextQuest+ [1], which facilitates concept discovery in Pubmed abstracts and OMIM records, DrugQuest [2] for identifying associations between drugs and other entries of the DrugBank database, PolySearch [3] for identifying gene, mutation, and disease associations and DISEASES [4] for extracting disease-gene associations from biomedical abstracts. Tools like Darling [5] and OnTheFly [6] mine disease-related databases and perform named entity recognition in various texts and databases, respectively. UniProt Related Documents (UniReD) [7] assists wet lab biologists in their quest to find novel counterparts in a protein interaction network generated by text mining using UniProt records. CoPub is a text mining system for microarray data analysis but also for general exploration of biomedical literature [8]. Other tools, like EXTRACT [9], NETME [10], GNormPlus [11], and SciLite [12] detect associations between biomedical terms and construct knowledge networks incorporating term associations within the literature. PREGO [13] links environments to biological processes and organisms. PubAnnotation [14] is an open, Agile text mining framework assisting researchers for annotation purposes. In contrast, Medline Ranker [15] ranks abstracts from Medline, according to a training set of abstracts or a MeSH term. LipiDisease [16] focuses on identifying associations between lipids and diseases in biomedical literature data. PESCADOR [17] extracts a network of gene and protein interactions from a set of PubMed abstracts selected by a user and PubTator [18] offers automated annotations generated by cutting-edge text mining systems for genes/proteins, genetic variants, diseases, chemicals, species, and cell lines.

Databases such as STRING [19], STITCH [20]), DisGeNET [21], and CancerMine [22] periodically mine MEDLINE to computationally identify novel protein-protein, protein-chemical compound, and gene-disease associations to produce interaction networks, as well as enrich their existing evidence with additional annotations. Similarly, web servers offering NLP functionalities such as GePI[23]. Functional enrichment analysis tools such as aGOtool [24], SciMiner [25], ToppGene [26] and Flame [27] utilize text-mining annotations alongside other sources to annotate gene and protein lists. Finally, text-mining resources and datasets have been used to derive training sets for deep learning methods aimed at the functional characterization of unannotated proteins. Notable examples include ProtNLM which has been trained by text mining the annotation fields of UniProtKB [28] and Pfam [29] entries, and KV-PLM [30], trained using the S2ORC corpus of English academic research publications [31].

All the tools and resources mentioned above, along with numerous others, aid researchers in uncovering new knowledge within extensive lisis of biomedical articles, as well as in organizing and analyzing the vast amount of associated information.

Here, we present BioTextQuest v2.0 – an updated version of a previously published web tool BiotextQuest (+) – a tool for biomedical literature analysis and concept discovery. BioTextQuest v2.0 is an open-source online web portal for document clustering and efficient management of the overload of information in PubMed abstracts. It performs automated knowledge extraction and bridges the gap between named entity recognition and bioinformatics-related data sources. It takes as input PubMed queries and the output is presented in tabular and word cloud formats. In its current form, it utilizes EXTRACT, a tool that identifies genes/proteins, chemical compounds, organisms, environments, tissues, diseases, phenotypes, and Gene Ontology terms mentioned in Biomedical Literature. Substantial improvements of this significantly evolved version include: i) Use of Extract tool (BioTextQuest 2.0 recognizes more biomedical entities). ii) A user option to choose the type of biomedical entities to cluster. iii) Analysis of up to 10,000 PubMed abstracts (instead of 5000 in the previous version) iv) Implementation of two additional, state-of-the-art machine learning clustering algorithms. v) A more user-friendly interface, based also on suggestions from users, vi) Option to download the results in publication quality files/images and vii) Implementation of the Docker technology.

By enabling the user to choose the preferable biomedical entities to perform the clustering, users are assisted in their quest to identify all kinds of associations between these entities e.g. protein-disease associations, and thus extracting known or novel biomarkers. With all the enhancements, BioTextQuest v2.0 empowers researchers to navigate and extract valuable insights from the vast landscape of biomedical literature more effectively.

2 Methods

BioTextQuest version 2 is an easy-to-use web application with a responsive web interface. It is implemented in R [32] using the Shiny package1 to build the web application. The Shiny package enables the creation of a user-friendly web interface aimed at enhancing the user experience. It provides a simple web interface with just a query field and graphical representation of clusters related to biological terms. To ensure high availability and scalability to internet-level usage without limitations on concurrent users of our Shiny app, we implemented the Shinyproxy open-source framework, which is built on JVM technology and Docker. The analysis process starts with entering a query into the designated field for the PubMed database. Users also have the option to customize analysis parameters through the advanced settings button. (Fig. 1).Fig. 1 The main page of BiotextQuest v.2 query and the Advanced parameters that users may define.

Fig. 1

The main difference from the previous version of BiotextQuest is that it uses extracted terms from the EXTRACT tagger [9]. EXTRACT is a text-mining-assisted interactive annotation app of biomedical named entities and ontology terms. It tags significant biomedical terms from PubMed abstracts, belonging to different categories, like gene ontology terms, diseases, genes, proteins, etc. The tagged terms are precomputed and stored locally in a MySQL database. BiotextQuest v.2 uses the tagged terms of each PubMed abstract of a query and clusters abstracts into subjects according to their similarity based on the extracted terms.

2.1 Query system

BioTextQuest v.2 currently queries PubMed. The PubMed database is locally stored in MongoDB. The query field allows input of any valid PubMed query supporting all features offered by the search mechanism of PubMed, such as field tags, Boolean operators, or grouping parentheses. BioTextQuest v.2 uses the easyPubMed R package to post a query directly to PubMed and obtain the PubMed identifiers of matching entries. In a subsequent step and depending on user-defined parameters, the platform uses these identifiers to retrieve the appropriate combination of abstracts and their associated EXTRACT tagged terms from the local database, thus maximizing the speed of information retrieval. The PubMed retrieval is performed in a few seconds. BioTextQuest v.2 users may also specify the number of documents to be retrieved/processed (with a maximum of 10,000 articles/entries as default). In cases of queries returning more than 10,000 abstracts, the 10,000 most recent are retained for analysis. The user can also choose specific categories of tagged terms, for example only genes/proteins to include in the analysis of abstracts.

2.2 Document similarities

We represent each abstract with a binary vector indicating the presence or absence of the EXTRACT tagged terms found in the text collection. Similarity metrics available in BioTextQuest v.2 are the Cosine similarity, Jaccard, and Euclidean distance [33] that we have been deemed as best to obtain optimal results.

2.3 Document clustering

Based on our experience from the previous version of BiotextQuest, we incorporated four different clustering algorithms, namely K-means [34], MCL [35], Louvain [36] and Top2vec [37] that can be seen when pressing the Advanced parameters button at the web interface to cluster documents based on their EXTRACT tagged terms. The methods take as input the similarity matrices described in the document similarities section. In order not to overwhelm non-expert users, we chose to hide by default the relevant options from the main interface; K-means is selected as the default clustering algorithm with 3 clusters and the cosine similarity metric. Some may consider that the plethora of integrated algorithms can become confusing; however, it is a useful feature for experts, as each algorithm takes a different angle on how to cluster data. An important feature of BioTextQuest v.2 is that it hosts two classes of clustering algorithms, depending on whether the number of resulting clusters is required as input. For example, MCL automatically detects the number of clusters formed by the data (depending on the choice of inflation parameter), whereas K-Means requires that the number of clusters k is known beforehand. Empirically, the MCL and k-means clustering algorithms are among the fastest.

2.4 Visualization of results

The visualization of the results is mainly based on a tag cloud for each cluster. Each cluster’s tag cloud contains by default the 50 most frequent EXTRACT tagged terms based on the abstracts grouped in that cluster. The terms are colored based on the category they belong to, for example, green is used for gene ontology terms, etc. The user can choose to include in each cluster more or less than 50 terms. Furthermore, the user can download the word clouds in publication quality figures as a PDF file. It is also noteworthy that the terms in each word cloud are links to the relevant external databases that contain information about the term, for example, GO terms are linked to the gene ontology database to retrieve all the relevant information. Users can also choose not to show tagged terms belonging to a specific category by deselecting it from the checkboxes at the right panel of the interface (Fig. 2).Fig. 2 Results presented as a word cloud.

Fig. 2

Two additional tabs of results are available to assist users in delving deeper into the information generated by the analysis. The "Biomedical Terms" tab displays a comprehensive list of all the terms tagged with EXTRACT that have been retrieved, allowing users to sort them by frequency. Users can also search for specific terms within this tab (Fig. 3).Fig. 3 The list of the tagged terms extracted from the abstracts and their frequency.

Fig. 3

The second tab, named "Clustered Documents," provides an overview of the results, including the total number of clusters and the number of articles within each cluster. Additionally, this tab features a table that presents details about the abstracts utilized in the analysis, such as the PMID, abstract content, title, publication year, and the corresponding cluster assignment. Users have the option to filter and export the table based on their specific requirements. (Fig. 4).Fig. 4 Information about the clustered documents.

Fig. 4

3 Results and discussion

The use of BiotextQuest v.2 involves two primary steps. First, users commence their analysis by selecting the "Start" tab. This page hosts a text field that allows input of any valid PubMed query, mirroring the syntax of PubMed's query system. Users also have the option to adjust analysis parameters via the "Advanced" button. The analysis is initiated by clicking the "Run Analysis" button, which triggers a progress bar display. Upon completion of the analysis, users are automatically directed to the "Results" tab. The initial page of results, labeled "Clustered documents" presents information such as the number of articles, their respective clusters, article titles, etc. By navigating to the "Wordclouds" tab, users gain access to interactive word clouds that showcase the EXTRACT tagged terms within each cluster. Finally, there is a third page of results called “Biological terms” containing in a sortable and filterable table all the EXTRACT terms utilized in the analysis of the PubMed articles.

Substantial improvements of this significantly evolved version include:

•Use of EXTRACT named entity recognition tool (BioTextQuest 2.0 recognizes more biomedical entities).

•The user has the option to choose the type of biomedical entities to cluster.

•Analysis of up to 10,000 PubMed abstracts (instead of 5000 in the previous version).

•Implementation of two additional and well-established clustering algorithms.

•A more user-friendly interface, based also on suggestions from users.

•Option to download the results in publication quality files/images.

•Implementation of Docker technology.

The usefulness of BioTextQuest v2.0 has been shown previously elsewhere [38]. To assess the representation of the myeloid-derived suppressor cell (MDSC) concept and the tumor-associated neutrophils (TAN) concept in the literature, we employed the following text mining methodology using a preliminary version of Biotextquest version 2.0. We identified genes, proteins, molecular functions, pathways, and biological processes in PubMed abstracts. We used the following pattern to form our queries: “A and B and C,” where A was “neutrophil” or “MDSC,” B was always “cancer” and C was either “progression, suppression, mechanism, human, mice, maturation, clinic, or pathway.” We analyzed sixteen queries in total and the advanced parameters were set as “use terms from the database: extract, choose extract entities: all, similarity matrix: cosine, clustering algorithm: K-means, number of clusters: 5.”. Using the frequency of terms returned, as well as the word clouds of each cluster, we observed a clear focus of mentioning aspects of granulopoiesis and neutrophil differentiation (central term “G-CSF”) in TAN-related research papers.

Key terms retrieved from MDSC literature reflected the central function of MDSC, which is their capacity to limit the activity of lymphocytes, in particular T cells (central words “CD4″ and “CD8″). The text mining also revealed the most prominent cell biological mechanism in the context of MDSC research, which appears to be linked to the transcription factor STAT3. This term appeared as a central word in connection with the search terms “progression”, “suppression”, “mechanism”, “pathway” and “clinic”. Our MDSC-TAN use case also revealed shared “key” words between MDSC and TAN, such as IL-6, which seems to be seen as relevant in both research areas.

Another use case involves RICTOR, a fundamental subunit of the mTORC2. The mechanistic target of rapamycin (mTOR) signaling pathway involves two distinct complexes, mTOR complex 1 (mTORC1) and 2 (mTORC2), which coordinate numerous vital cellular processes.

We used the following query to retrieve related abstracts from PubMed ” Rictor ”. This process retrieved 1143 abstracts and we analyze the articles using the default parameters. The result was three clusters and we set the number of terms in the wordcloud to be 100 (Fig. 1 supplementary data).

The mTORC2-RICTOR regulates numerous crucial cellular processes as indicated by the terms apoptosis, autophagy, cell growth, etc. (Wordcloud 2) [39], [40]. Genomic alterations in the mTOR pathway components are frequently detected in cancers, subsequently modifying pathway activity is implicated in oncogenesis of different tumor types [41]. There is growing evidence reporting that RICTOR is aberrantly regulated across numerous cancer types. Wordcloud 1 contains several types of cancers and more terms related to cancer than in the other two other Wordclouds, for example by the terms non-small cell lung cancer, ovarian cancer, melanoma, etc) and is associated with tumorigenesis and poor prognosis (Wordcloud 1) [42]. Moreover, different studies have underlined the correlation between RICTOR expression and glycolysis and other metabolic functions such as glycolysis and lipogenesis as indicated by the terms glycolysis, insulin receptor and fatty acid oxidation (Wordcloud 3) [43]. It must be noted that although some terms from the aforementioned examples are included in two clusters/wordclouds, their frequency is higher in the cluster they are more associated (table 1 supplementary data).

We then test the performance of BiotextQuest v.2 when compared to its predecessor BioTextQuest(+)[1] using the same case study. We performed four different MeSH-term–based queries and retrieved all PubMed articles that mention a specific phase of the human cell cycle by excluding all others (phases M, G1, S, G2). We used limitations in publication date in order to retrieve the same articles as in the previous test. The process returned 8 genes for M-phase, 7 for G1-Phase, 8 for S-phase and 11 for G2-phase, in the case of BioTextQuest(+) whereas using the updated version of BiotextQuest v.2 we were able to retrieve 60 genes for M-phase, 117 for G1-Phase, 138 for S-phase and 99 for G2-phase (see table 2 in supplementary data). This result clearly shows that BiotextQuest v.2 significantly outperforms the previous version.

In the future, BioTextQuest is expected to benefit from Large Language Models (LLMs) in text mining for biomedical literature. These advanced AI models could process vast amounts of scientific text with exceptional precision, revealing intricate relationships between genes, diseases, and drug targets. This capability will enable researchers to uncover hidden connections, generate novel hypotheses and accelerate the discovery process. The integration of LLMs into text mining has the potential to revolutionize biomedical research, leading to breakthroughs that could significantly advance human health. Moreover, by incorporating new types of biomedical text association studies, such as Publication-Wide Association Studies (PWAS), into BioTextQuest could open up new avenues for further understanding complex biological relationships. [44].

Even in its current form, BioTextQuest v2.0 proves to be a highly useful tool for analyzing extensive collections of articles. It supports biomedical researchers in uncovering new concepts and relationships between biological entities, including the discovery of biomarkers.

Summarizing, our tool can help researchers in prioritizing topics and focus on classical state-of-the art and systematic reviews on biomedical subjects in an unbiased manner. Since cited literature and chosen topics in such reviews are often and obviously subjective and mainly driven by the authors’ individual knowledge, BioTextQuest v2.0 can help review authors identify central terms and themes, providing a balance between individual study selection and systematic literature survey.

Source code and Docker image availability

The source code is available at: https://github.com/theodos/biotextquest_v2.

The docker image is available at: https://hub.docker.com/r/mpaltsai/biotextquest_4.2.2/tags.

Funding

This research has been supported by the project “ELIXIR-GR: Managing and Analysing Life Sciences Data” (MIS: 5002780 ), co-financed by Greece and the European Union—10.13039/501100008530 European Regional Development Fund . This work was supported by 10.13039/501100000921 COST (European Cooperation in Science and Technology) Action 20117 - Converting molecular profiles of myeloid cells into biomarkers for inflammation and cancer (Mye-InfoBank). This work has been supported by a research grant under the Horizon 2020 programme of the European Commission (TO_AITION, no. 848146 ). This work was also supported by 10.13039/501100013209 Hellenic Foundation for Research and Innovation (H.F.R.I) under the call ‘Greece 2.0 - Basic Research Financing Action (Horizontal support of all Sciences), Sub-action II’, Grant ID: 16718-PRPFOR ; ‘Greece 2.0 - National Recovery and Resilience Plan’, Grant ID: TAEDR-0539180 .

Author statement

We would like to submit the revised version of the article entitled: “BioTextQuest v2.0: an evolved tool for Biomedical Literature mining and concept discovery.” by Theodosiou et. al. Below you can find a point-by-point description with answers to the reviewers’ comments. We would like to thank you in advance for considering our article for publication and we are looking forward to hearing from you.

CRediT authorship contribution statement

Evangelos Andreakos: Writing – review & editing, Validation. Sven Brandau: Writing – review & editing, Validation. Theodosios Theodosiou: Writing – review & editing, Writing – original draft, Software, Methodology. Ekaterini Chatzaki: Writing – review & editing, Validation. Konstantinos Vrettos: Software. Christos A. Ouzounis: Writing – review & editing, Validation. Dimitrios Mossialos: Validation. Makrina Karaglani: Validation. Vasilis J. Promponas: Writing – review & editing, Validation. Fotis Baltoumas: Software. Ioannis Iliopoulos: Writing – review & editing, Writing – original draft, Validation, Supervision, Methodology, Funding acquisition, Formal analysis, Conceptualization. Ismini Baltsavia: Software. Andreas N. Antonakis: Validation. Nikolas Papanikolaou: Validation. Georgios A. Pavlopoulos: Writing – review & editing, Writing – original draft, Validation.

Declaration of Competing Interest

The authors declare no conflict of interest.

Appendix A Supplementary material

Supplementary material

.

Acknowledgments

We would like to thank Anastasios Gkountakos for his assistance in the “Rictor” use case.

☆ BioTextQuest v2.0 is available at: http://bioinformatics.med.uoc.gr/shinyapps/app/biotextquest

Appendix A Supplementary data associated with this article can be found in the online version at doi:10.1016/j.csbj.2024.08.016.

1 https://shiny.posit.co
==== Refs
References

1 Papanikolaou N. BioTextQuest(+): a knowledge integration platform for literature mining and concept discovery Bioinformatics vol. 30 22 2014 3249 3256 10.1093/bioinformatics/btu524 25100685
2 Papanikolaou N. Pavlopoulos G.A. Theodosiou T. Vizirianakis I.S. Iliopoulos I. DrugQuest - a text mining workflow for drug association discovery BMC Bioinforma vol. 17 Suppl 5 2016 182 10.1186/s12859-016-1041-6
3 Cheng D. Knox C. Young N. Stothard P. Damaraju S. Wishart D.S. PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites Nucleic Acids Res vol. 36 Web Server 2008 W399 W405 10.1093/nar/gkn296 18487273
4 Pletscher-Frankild S. Pallejà A. Tsafou K. Binder J.X. Jensen L.J. DISEASES: text mining and data integration of disease-gene associations Methods vol. 74 2015 83 89 10.1016/j.ymeth.2014.11.020 25484339
5 Karatzas E. Darling: a web application for detecting disease-related biomedical entity associations with literature mining Biomolecules vol. 12 4 2022 10.3390/biom12040520
6 Baltoumas F.A. OnTheFly2.0: a text-mining web application for automated biomedical entity recognition, document annotation, network and functional enrichment analysis NAR Genom Bioinform vol. 3 4 2021 lqab090 10.1093/nargab/lqab090
7 Theodosiou T. UniProt-Related Documents (UniReD): assisting wet lab biologists in their quest on finding novel counterparts in a protein network NAR Genom Bioinform vol. 2 1 2020 lqaa005 10.1093/nargab/lqaa005
8 Fleuren W.W.M. CoPub update: CoPub 5.0 a text mining system to answer biological questions Nucleic Acids Res vol. 39 Web Server issue 2011 W450 4 10.1093/nar/gkr310 21622961
9 Pafilis E. EXTRACT: interactive extraction of environment metadata and term suggestion for metagenomic sample annotation Database (Oxf) vol. 2016 2016 10.1093/database/baw005
10 Muscolino A. NETME: on-the-fly knowledge network construction from biomedical literature Appl Netw Sci vol. 7 1 2022 1 10.1007/s41109-021-00435-x
11 Wei C.-H. Kao H.-Y. Lu Z. GNormPlus: an integrative approach for tagging genes, gene families, and protein domains Biomed Res Int vol. 2015 2015 1 7 10.1155/2015/918710
12 Venkatesan A. SciLite: a platform for displaying text-mined annotations as a means to link research articles with biological data Wellcome Open Res vol. 1 2017 25 10.12688/wellcomeopenres.10210.2 28948232
13 Zafeiropoulos H. Paragkamian S. Ninidakis S. Pavlopoulos G.A. Jensen L.J. Pafilis E. PREGO: a literature and data-mining resource to associate microorganisms, biological processes, and environment types Microorganisms vol. 10 2 2022 293 10.3390/microorganisms10020293 35208748
14 Kim J.-D. Wang Y. Fujiwara T. Okuda S. Callahan T.J. Cohen K.B. Open Agile text mining for bioinformatics: the PubAnnotation ecosystem Bioinformatics vol. 35 21 2019 4372 4380 10.1093/bioinformatics/btz227 30937439
15 Fontaine J.-F. Barbosa-Silva A. Schaefer M. Huska M.R. Muro E.M. Andrade-Navarro M.A. MedlineRanker: flexible ranking of biomedical literature Nucleic Acids Res vol. 37 suppl_2 2009 W141 W146 10.1093/nar/gkp353 19429696
16 More P. Bindila L. Wild P. Andrade-Navarro M. Fontaine J.-F. LipiDisease: associate lipids to diseases using literature mining Bioinformatics vol. 37 21 2021 3981 3982 10.1093/bioinformatics/btab559 34358314
17 Barbosa-Silva A. Fontaine J.-F. Donnard E.R. Stussi F. Ortega J.M. Andrade-Navarro M.A. PESCADOR, a web-based tool to assist text-mining of biointeractions extracted from PubMed queries BMC Bioinforma vol. 12 1 2011 435 10.1186/1471-2105-12-435
18 Wei C.-H. Allot A. Leaman R. Lu Z. PubTator central: automated concept annotation for biomedical full text articles Nucleic Acids Res vol. 47 W1 2019 W587 W593 10.1093/nar/gkz389 31114887
19 Szklarczyk D. The STRING database in 2021: customizable protein-protein networks, and functional characterization of user-uploaded gene/measurement sets Nucleic Acids Res vol. 49 D1 2021 D605 D612 10.1093/nar/gkaa1074 33237311
20 Szklarczyk D. Santos A. von Mering C. Jensen L.J. Bork P. Kuhn M. STITCH 5: augmenting protein–chemical interaction networks with tissue and affinity data Nucleic Acids Res vol. 44 D1 2016 D380 D384 10.1093/nar/gkv1277 26590256
21 Piñero J. The DisGeNET knowledge platform for disease genomics: 2019 update Nucleic Acids Res 2019 10.1093/nar/gkz1021
22 Lever J. Zhao E.Y. Grewal J. Jones M.R. Jones S.J.M. CancerMine: a literature-mined resource for drivers, oncogenes and tumor suppressors in cancer Nat Methods vol. 16 6 2019 505 507 10.1038/s41592-019-0422-y 31110280
23 Faessler E. Hahn U. Schäuble S. GEPI: large-scale text mining, customized retrieval and flexible filtering of gene/protein interactions Nucleic Acids Res vol. 51 W1 2023 W237 W242 10.1093/nar/gkad445 37224532
24 Schölz C. Lyon D. Refsgaard J.C. Jensen L.J. Choudhary C. Weinert B.T. Avoiding abundance bias in the functional annotation of posttranslationally modified proteins Nat Methods vol. 12 11 2015 1003 1004 10.1038/nmeth.3621 26513550
25 Hur J. Schuyler A.D. States D.J. Feldman E.L. SciMiner: web-based literature mining tool for target identification and functional enrichment analysis Bioinformatics vol. 25 6 Mar. 2009 838 840 10.1093/bioinformatics/btp049 19188191
26 Chen J. Bardes E.E. Aronow B.J. Jegga A.G. ToppGene Suite for gene list enrichment analysis and candidate gene prioritization Nucleic Acids Res vol. 37 Web Server 2009 W305 W311 10.1093/nar/gkp427 19465376
27 Karatzas E. Flame (v2.0): advanced integration and interpretation of functional enrichment results from multiple sources Bioinformatics vol. 39 8 2023 10.1093/bioinformatics/btad490
28 Bateman A. UniProt: the Universal Protein Knowledgebase in 2023 Nucleic Acids Res vol. 51 D1 2023 D523 D531 10.1093/nar/gkac1052 36408920
29 Mistry J. Pfam: The protein families database in 2021 Nucleic Acids Res vol. 49 D1 2021 D412 D419 10.1093/nar/gkaa913 33125078
30 Zeng Z. Yao Y. Liu Z. Sun M. A deep-learning system bridging molecule structure and biomedical text with comprehension comparable to human professionals Nat Commun vol. 13 1 2022 862 10.1038/s41467-022-28494-3 35165275
31 K. Lo, L.L. Wang, M. Neumann, R. Kinney, and D. Weld, “S2ORC: The Semantic Scholar Open Research Corpus,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Stroudsburg, PA, USA: Association for Computational Linguistics, 2020, pp. 4969–4983. 10.18653/v1/2020.acl-main.447.
32 R Core Team, “R: A Language and Environment for Statistical Computing,” 2022, Vienna, Austria: https://www.R-project.org/.
33 Singh Lehal Manpreet Comparison of Cosine, Euclidean Distance and Jaccard Distance Int J Sci Res Sci, Eng Technol(IJSRSET) vol. 3 8 2017 1376 1381
34 Lloyd S. Least squares quantization in PCM IEEE Trans Inf Theory vol. 28 2 1982 129 137
35 Van Dongen S. Graph clustering via a discrete uncoupling process SIAM J Matrix Anal Appl vol. 30 1 2008 121 141 10.1137/040608635
36 Blondel V.D. Guillaume J.-L. Lambiotte R. Lefebvre E. Fast unfolding of communities in large networks J Stat Mech: Theory Exp vol. 2008 10 Oct. 2008 P10008 10.1088/1742-5468/2008/10/P10008
37 D. Angelov, “Top2Vec: Distributed Representations of Topics,” ArXiv, vol. abs/2008.09470, 2020.
38 Antuamwine B.B. N1 versus N2 and PMN-MDSC: a critical appraisal of current concepts on tumor-associated neutrophils and new directions for human oncology Immunol Rev vol. 314 1 2023 250 279 10.1111/imr.13176 36504274
39 Lee G. Chung J. Discrete functions of rictor and raptor in cell growth regulation in Drosophila Biochem Biophys Res Commun vol. 357 4 2007 1154 1159 10.1016/j.bbrc.2007.04.086 17462592
40 Ballesteros‐Álvarez J. Andersen J.K. mTORC2: The other mTOR in autophagy regulation Aging Cell vol. 20 8 Aug. 2021 10.1111/acel.13431
41 Saxton R.A. Sabatini D.M. mTOR signaling in growth, metabolism, and disease Cell vol. 168 6 2017 960 976 10.1016/j.cell.2017.02.004 28283069
42 Gkountakos A. Unmasking the impact of Rictor in cancer: novel insights of mTORC2 complex Carcinogenesis vol. 39 8 2018 971 980 10.1093/carcin/bgy086 29955840
43 Kocalis H.E. Rictor/mTORC2 facilitates central regulation of energy and glucose homeostasis Mol Metab vol. 3 4 2014 394 407 10.1016/j.molmet.2014.01.014 24944899
44 Narganes-Carlón David Crowther Daniel J. Pearson Ewan R. A publication-wide association study (PWAS), historical language models to prioritise novel therapeutic drug targets Sci Rep vol. 13 1 2023 8366 10.1038/s41598-023-35597-4 37225853
