
==== Front
Ophthalmol Ther
Ophthalmol Ther
Ophthalmology and Therapy
2193-8245
2193-6528
Springer Healthcare Cheshire

39180701
1018
10.1007/s40123-024-01018-6
Review
Utilizing Large Language Models in Ophthalmology: The Current Landscape and Challenges
Chotcomwongse Peranut 1
Ruamviboonsuk Paisan 1
http://orcid.org/0000-0002-3724-2391
Grzybowski Andrzej ae.grzybowski@gmail.com

23
1 https://ror.org/0238gtq84 grid.415633.6 0000 0004 0637 1304 Vitreoretina Unit, Department of Ophthalmology, Rajavithi Hospital, Rungsit University, Bangkok, Thailand
2 grid.412607.6 0000 0001 2149 6795 University of Warmia and Mazury, Olsztyn, Poland
3 https://ror.org/01pmj6109 Institute for Research in Ophthalmology, Foundation for Ophthalmology Development, 61-553 Poznan, Poland
24 8 2024
24 8 2024
10 2024
13 10 25432558
29 5 2024
1 8 2024
© The Author(s) 2024
2024
https://creativecommons.org/licenses/by-nc/4.0/ Open Access This article is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License, which permits any non-commercial use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc/4.0/.
A large language model (LLM) is an artificial intelligence (AI) model that uses natural language processing (NLP) to understand, interpret, and generate human-like language responses from unstructured text input. Its real-time response capabilities and eloquent dialogue enhance the interactive user experience in human–AI communication like never before. By gathering several sources on the internet, LLM chatbots can interact and respond to a wide range of queries, including problem solving, text summarization, and creating informative notes. Since ophthalmology is one of the medical fields integrating image analysis, telemedicine, AI, and other technologies, LLMs are likely to play an important role in eye care in the near future. This review summarizes the performance and potential applicability of LLMs in ophthalmology according to currently available publications.

Keywords

Large language model
Ophthalmology
ChatGPT
Bard
Copilot
Telemedicine
Artificial intelligence
issue-copyright-statement© Springer Healthcare Ltd., part of Springer Nature 2024
==== Body
pmcKey Summary Points

Since the advent of the large language model (LLM), its ability to generate human-like prompt response has been widely recognized, with a huge impact in various medical fields including ophthalmology.	
This study aims to review and summarize current publications regarding clinical applications, advantages, and limitations of using LLM chatbots in an ophthalmology context.	
LLM is an emerging technology that can potentially bring about changes in medical practice, including ophthalmology. Its applicability involves clinical, educational, and research applications. Although the current performance is promising, the real-world implementation remains controversial and open for further studies.	

Introduction

A large language model (LLM) is an artificial intelligence (AI) model that uses natural language processing (NLP) to understand, interpret, and generate human-like language responses from unstructured text input. Its real-time response capabilities and eloquent dialogue enhance the interactive user experience in human–AI communication like never before. By gathering information from several sources on the internet, the LLM chatbot can interact and respond to a wide range of queries, including problem solving, text summarization, and creating informative notes [1] (Fig. 1).Fig. 1 Diagram demonstrating how a large language model (LLM) chatbot generates output from the various available sources on the internet

Among the many LLMs, ChatGPT (Open AI, San Francisco, CA, USA) has gained immense popularity and has become widely established in recent years. In the initial version in 2022, Generative Pre-trained Transformer (GPT)-3.5 was integrated to back-end the chatbot. The later version, GPT-4, was launched in 2023 [2]. GPT-4 incorporates added features with improved problem-solving ability and allows multimodal input analysis, such as image analysis [3]. Besides ChatGPT, other generative AI chatbots such as Bard (Google, Menlo Park, CA, USA), and Microsoft Copilot (Microsoft, Redmond, WA, USA) are also freely available with different LLM backbones [2].

With regard to ophthalmology, the number of patients is growing as the aging population continues to increase. AI, deep learning, and other novel modes of care using technology can augment patient care. As this medical specialty deals with vast amounts of images, implementing algorithms to detect diseases could support ophthalmologists and enhance treatment workflows not only in the clinic but also through telemedicine. Previous studies have demonstrated the good performance of deep learning in detecting diabetic retinopathy, glaucoma, age-related macular degeneration, ocular surface diseases, and other visual impairment conditions. LLMs, combined with these technologies, could potentially facilitate workflow and improve the patient experience [4].

Methods

This is a comprehensive review of the publicly available articles on PubMed regarding LLMs and ophthalmology. We conducted a literature search in PubMed in February 2024 using the search terms “Large language model” OR “Generative artificial intelligence” OR “ChatGPT” OR “Bard” OR “Bing” OR “Ophthalmology.” This article is based on previously conducted studies and does not contain any new studies with human participants or animals performed by any of the authors. Ethics committee approval was not required for this study because no human subjects were involved.

LLM and ChatGPT Development

The dataset, model architecture, and number of parameters used to develop the GPT model are shown in Table 1. The first version, GPT-1, was launched in 2018 using semi-supervised training on the BookCorpus dataset with 11,308 books containing 1 × 109 words along with a fine-tuning process. In later years, GPT-2 was released with a 10-fold larger dataset from WebText, containing over 8 million documents [2]. Table 1 Evolution of GPT and what lies behind the development [2, 6]

	GPT-1	GPT-2	GPT-3	GPT-4	
Dataset	1 dataset: BookCorpus (11,380 books, 1 × 109 words)	1 dataset: Web text (40 GB of data, 8 million documents)	5 datasets: CommonCrawl, WebText2, Books1, Books2, Wikipedia (45 TB of data)	~13 trillion tokens, both text-based and code-based data (fine-tuning data from ScaleAI and internally)	
Model architecture	12 layers with 12 attention heads in each self-attention layer	48 layers with 1600-dimensional vectors for word embedding	96 layers with 96 attention heads	120 layers	
Parameters	115 million	1.5 billion	175 billion	~1.8 trillion	
GB gigabyte, GPT Generative Pre-trained Transformer, TB terabyte

In 2020, GPT-3 was established, with 100 times the number of parameters in GPT-2. Common Crawl, WebText2, Books1, Books2, and Wikipedia datasets were integrated into the training dataset, demonstrating higher performance. GPT-3.5 was later developed with human-generated text-based input–output pairs to fine-tune GPT-3. Reinforcement learning from human feedback creates human ranking of GPT-3.5-generated outputs, facilitating self-reinforcement learning based on human feedback [5]. In November 2022, ChatGPT was released, allowing users to experience the chatbot user interface under the GPT-3.5 model at no cost.

The following year, in March 2023, OpenAI launched a new version of ChatGPT embedded with the GPT-4 model, which is only available for paid chatbot users [2]. Information on model training and architecture was recently released. It was trained on approximately 13 trillion tokens, including text-based and code-based data, with fine-tuning of data from ScaleAI. OpenAI utilizes 16 experts within their model, at a cost of around $63 million [6]. Because of the much larger training dataset, GPT-4 outperforms the previous GPT-3.5 in problem-solving abilities and broader knowledge base [7]. For example, GPT-4 showed better performance in taking an ophthalmology board exam, and the results were comparable to a human response [8].

In addition to ChatGPT, other models were launched shortly afterward. Bard AI is a model developed by Google AI using Pathways Language Model 2 (PaLM 2) as a backbone. In contrast to ChatGPT, which is proficient in generating articles, coding, or scripts, Bard AI excels in accurate text summarization. Furthermore, Bard AI can access the internet in real time and is able to provide up-to-date information on current topics [9].

Microsoft Copilot has recently changed its name from Bing Chat. It uses GPT-4 as a fundamental language model in the underlying structure. It is designed to better conform to Microsoft architecture and is able to use the internet for up-to-date information. Furthermore, it offers users a free version, while GPT-4 is only available in the paid version of ChatGPT [2]. However, some functions are still preserved only for GPT-4.

As all the current LLMs are proficient in text-based analysis and response, Ophtha-LLaMA2 is a novel LLM aimed at the ophthalmology field. The advantage of Ophtha-LLaMA2 is that it exhibits satisfying accuracy and efficacy in making diagnoses using imaging analysis. In contrast to ChatGPT, which uses publicly available sources on the internet, Ophtha-LLaMA2 integrates various imaging modalities, including optical coherence tomography (OCT), color fundus photography (CFP), and ocular surface analysis (OSA), in the dataset for training [10]. However, its real-world performance has yet to be explored [18, 50, 52].

LLM Utility in Ophthalmology

Patient Perspective and Engagement

To use LLM chatbots in the field of ophthalmology, patients can manually ask questions related to eye disease, current treatments of choice, and available research. The ability to generate dependable medical information through LLMs can enhance user–AI communication with human-like context and real-time response [2].

Disease Information and Recommendations

Many studies have focused on recommendations regarding ophthalmic diseases. Most of them have reported quality suggestions for various kinds of diseases, such as diabetic retinopathy [11], myopia [12], ophthalmic plastic reconstruction [13, 14], surgical retinal diseases [15], keratitis [16], neuro-ophthalmology [17], uveitis [18], and postoperative care [19]. The concept and main content are generally acceptable across all subspecialties. This might be a significant benefit to patients’ education. The more data that are included for LLM training in the newer-generation models, such as GPT-4, the more accurate the response is [12]. However, some authors reported that while many relevant responses were provided, incomplete and inaccurate statements were also found in the generated recommendations. In some cases, harmful information was given—for example, ChatGPT suggested that patients should remove the conjunctiva for vernal keratoconjunctivitis treatment. Furthermore, the response did not include the potential serious side effects of topical steroids, which is considered unacceptable [20].

Diagnostic Support

Because LLM chatbots are only a few mouse clicks away, we cannot deny their usefulness for patients using them to initially gather information about the symptoms they are experiencing before consulting a real physician. Some authors report the use of ChatGPT for making diagnoses based on given symptoms [21–28]. While GPT-3.5 showed promise in the ability to make a correct diagnosis, GPT-4 showed superior performance in terms of diagnostic accuracy. GPT-4 was also able to provide a diagnosis faster than a specialist [22]. However, the efficiency is potentially inferior in rare diseases. For example, GPT-3 and GPT-4 showed only 60% and 66% diagnostic success rates, respectively, in the field of uveitis [26, 27]. GPT-4 even performed more poorly than family physicians and junior ophthalmologists in detecting rare eye diseases [25]. However, GPT-4 and Bard were effectively used as a triage tool for ophthalmic symptoms, with response correctness of 80–98% [29, 30]. In more common diseases such as glaucoma, ChatGPT exhibited 72.2% correctness in provisional diagnosis for primary and secondary glaucoma. Its performance was similar to or better than that of three senior ophthalmology residents in that study [22].

A recent study by Huang et al. demonstrated that GPT-4 had comparable diagnostic and treatment accuracy in both glaucoma retina specialties against fellowship-trained ophthalmologists. Although GPT-3.5 was found to have significant inaccuracies, the novel GPT-4 provides promising ability and may contribute to future applications [31]. Furthermore, the ability to significantly reduce workload and shorten the amount of time surpasses that of any previous AI models including earlier LLM. This could usher in groundbreaking changes in ophthalmology and other medical specialties [32].

This supports the use of LLM chatbots to facilitate user–AI interaction with knowledge in ophthalmology. Moreover, users can easily access the chatbot, which can function like a reliable personal assistant in medicine. Although the context given by the LLM chatbot is potentially trustworthy, knowledge accuracy and its real-world performance must be determined for proper validation before it can be integrated into real-world workflow. Moreover, the LLM chatbot may provide misleading information and cause harm to innocent patients. This raises ethical concerns that must be addressed before incorporating this technology into a large-scale health care system [2].

Telemedicine Integration and Implementation

There is a new visual aid tool that embeds an LLM to assist low-vision patients. Meta launched a new generation of smart glasses in September 2023. With the integration of GPT-4, smart glasses can create new functions that can read text, respond to specific questions, or summarize paragraphs [33]. This positive indication suggests that we can anticipate further implementation in the near future.

Although the novel LLM chatbot can understand image-based tasks, only a few studies have evaluated its performance. A study by Mihalache et al. reported that ChatGPT performed better in text-based questions than image-based questions (82% vs. 65%) [34]. Another study by Mihalache et al. demonstrated approximately 50% correct diagnoses in retinal cases requiring multimodal input [35]. In terms of ophthalmic images, its accuracy is not yet well established for real-world implementation.

Health Care Provider Perspective

LLMs can be used for a broad range of educational resources in ophthalmology. They represent an exciting opportunity for scientists to improve their research. LLM chatbots not only can suggest related articles but can also summarize and provide the key context for each article. This could create work-time efficiency for article reviews. LLM chatbots can also accurately generate references with the appropriate citation format for the article. Moreover, they can analyze text and provide grammatical suggestions for improving or rephrasing sentences to match the proper meaning. This can help create an LLM chatbot-like writing style [36].

Research Assistance

As of January 2024, many publishers have established authorship policies for the use of any LLM chatbot for writing manuscripts. Most publishers do not accept an LLM chatbot as a co-author; however, the use of LLM chatbots must be clarified in the research methodology. In addition, the author is responsible for the integrity of the content generated by these models and should thoroughly review it before submission. Furthermore, the editorial offices recommend caution on the part of authors and readers when LLM chatbots are used [37]. Some examples of authorship policies from well-known journals in ophthalmology are exhibited in Table 2. Table 2 List of authorship policies among different publishers (accessed January 18, 2024)

Journal	Authorship policy	Links	
JAMA Ophthalmology (JAMA)	The submission and publication of content created by artificial intelligence, language models, machine learning, or similar technologies is discouraged, unless part of formal research design or methods, and is not permitted without clear description of the content that was created and the name of the model or tool, version and extension numbers, and manufacturer. Authors must take responsibility for the integrity of the content generated by these models and tools	https://jamanetwork.com/journals/jamaophthalmology/pages/instructions-for-authors	
Eye (Nature)	Large language models (LLMs), such as ChatGPT, do not currently satisfy our authorship criteria. Notably an attribution of authorship carries with it accountability for the work, which cannot be effectively applied to LLMs. Use of an LLM should be properly documented in the Methods section (and if a Methods section is not available, in a suitable alternative part) of the manuscript	https://www.nature.com/nature-portfolio/editorial-policies/ai	
American Journal of Ophthalmology (Elsevier)	Where authors use generative artificial intelligence (AI) and AI-assisted technologies in the writing process, authors should only use these technologies to improve readability and language. Applying the technology should be done with human oversight and control, and authors should carefully review and edit the result, as AI can generate authoritative-sounding output that can be incorrect, incomplete or biased. AI and AI-assisted technologies should not be listed as an author or co-author, or be cited as an author. Authorship implies responsibilities and tasks that can only be attributed to and performed by humans, as outlined in Elsevier’s AI policy for authors	https://www.sciencedirect.com/journal/american-journal-of-ophthalmology/publish/guide-for-authors	

Medical Education

The performance with regard to ophthalmologic knowledge generated by GPT-3 and GPT-3.5 was found to be limited compared with that of humans [22, 26, 28, 38, 39]. In 2023, GPT-4, the latest available version of ChatGPT, provided significant improvements in data accuracy in ophthalmologic examinations including the American Board of Ophthalmology, European Board of Ophthalmology (EBO), OphthoQuestions, the American Academy of Ophthalmology's Basic and Clinical Science Course (BCSC), and the Fellowship of the Royal College of Ophthalmologists (FRCOphth) examinations [8, 39–44]. In addition to standard English, GPT-4 can achieve excellent EBO results in French, with 91% correctness [40]. This may mark a promising milestone toward the integration of the LLM chatbot into medical education and training systems.

Administrative Efficiency

Apart from knowledge of diseases, administrative tasks such as writing letters, rescheduling appointments, and refilling medication prescriptions are undeniably burdensome tasks for many health care providers. Documentation and administrative tasks are what LLMs specialized in. LLM applications can alleviate tedium and improve workflow efficiency by accelerating data synthesis and optimizing language on demand [2, 45].

For example, GPT-4 can draft cataract surgery operative notes in only a few seconds, with minimal refinements required [19, 46]. Moreover, not only can common operative documents be generated with GPT-4, but it can also handle complicated cases such as posterior capsule rupture layout remarks. However, some detailed information regarding operations is highly specific to its own specialties. Significant tuning according to the specific case is necessary, and the implementation of the LLMs should be used with caution and human verification [47].

The performance of LLMs in extracting International Classification of Diseases (ICD) code was demonstrated in a study conducted by Ong et al. GPT-3.5 showed 70% correctness in a true-positive result with at least one correct ICD code for retinal diseases in an outpatient department. This suggests the potential to alleviate physician burden by prompt provision of diagnosis codes [48].

GPT can also generate an ophthalmic discharge summary with specific medications, follow-up instructions, consultation time, and location in less than 20 s [46]. The ability to promptly conclude loads of medical information combined with manual review could help reduce workflow bottlenecks in the clinic. Compared to real doctors, LLMs outperform in terms of quality. LLM is able to identify data in the ophthalmic note, summarize piles of documents, and generate conclusions in a short time [49]. LLM could also be deployed in electronic health records (EHR) systems to provide key details for the consultation [4]. A list of recent articles regarding the use of LLMs in the ophthalmology field is presented in Table 3. Table 3 List of the current available publications regarding LLM in ophthalmology and their performance

Study authors	LLM model	Assessment	Performance	
Yu et al. [11]	BERT and RoBERTa	Identification of diabetic retinopathy-related clinical concepts from clinical narrative	This study demonstrated the efficiency of transformer-based NLP models for clinical concept extraction and relation extraction.	
Lim et al. [12]	GPT-3.5, GPT-4, Bard	Response to myopic queries	GPT-4 delivered accurate and comprehensive responses to myopia-related queries.	
Ali et al. [13]	GPT	Response to lacrimal drainage disorder queries	ChatGPT demonstrated average performance in the context of lacrimal drainage disorders.	
Al-Sharif et al. [14]	GPT-3.5, Bard	Response to oculoplastic queries	The rate of comprehensive responses was better in ChatGPT (71.4%) than Bard (53.1%). ChatGPT showed more empathy (48.9%) than BARD (13.2%).	
Momenaei et al. [15]	GPT-4	Response to surgical retina questions	GPT-4 responded correctly to 84% of multiple-choice practice questions on OphthoQuestions.	
Masalkhi et al. [16]	GPT-4	Providing educational support for ocular infectious diseases	GPT-4 provided a variety of useful information regarding ocular infectious disease instruction and recommendations.	
Waisberg et al. [17]	GPT-4	Providing medical information for neuro-ophthalmic medication	GPT-4 can potentially benefit neuro-ophthalmic education in the context of providing disease recommendations, generating management plans, summarizing new studies, and helping learners understand complex neuro-ophthalmic visual phenomena.	
Tan Yip Ming et al. [18]	GPT, Bard	Assessing role of LLM in uveitis	ChatGPT can provide accurate diagnosis, with 60–90% correctness.	
Waisberg et al. [19]	GPT-4	Generating detailed operative note for general ophthalmic surgery	GPT-4 has the potential to generate a detailed cataract surgery operative note.	
Rasmussen et al. [20]	GPT	Providing educational support for vernal keratoconjunctivitis (VKC)	ChatGPT often provided relevant responses to general patient and parent queries on VKC. However, it sometimes provided inaccurate and potentially dangerous statements.	
Balas et al. [21]	GPT-3	Providing ophthalmic diagnosis	ChatGPT identified the correct diagnosis in 9/10 cases, while having the correct diagnosis listed in all 10/10 of its lists of differentials.	
Delsoz et al. [22]	GPT-3	Providing glaucoma diagnosis	ChatGPT was correct in 8 out of 11 (72.7%) cases, and three ophthalmology residents were correct in 6, 8, and 8 (54.5–72.7%) cases, respectively.	
Shemer et al. [23]	GPT-3.5	Providing ophthalmic diagnosis	ChatGPT achieved a significantly lower accurate diagnosis rate (54%) than residents (75%) and attendings (71%).	
Liu et al. [24]	GPT-3.5	Providing retinal vascular disease diagnosis using fluorescein angiography reports in English and Chinese	ChatGPT could serve as a medical assistant to provide ophthalmic diagnosis using fluorescein angiography reports under non-English clinical environments.	
Hu et al. [25]	GPT-4	Providing diagnosis for rare eye diseases	GPT-4 could only output several possible diseases generally without correct diagnosis in the simulated patient scenarios.	
Rojas-Carabali et al. [26]	GPT-3.5, GPT-4	Providing educational support for uveitis diseases	Ophthalmologists achieved 76–100% success; AI attained 72%.	
Rojas-Carabali et al. [27]	GPT-4	Providing educational support for uveitis diseases	Uveitis experts accurately diagnosed all cases (100%), while ChatGPT achieved a diagnostic. success rate of 66%.	
Delsoz et al. [28]	GPT-3.5, GPT-4	Providing corneal disease diagnosis	The provisional diagnostic accuracy based on GPT-4.0 was 85% (17 correct out of 20 cases) while the accuracy of GPT-3.5 was 60% (12 correct cases out of 20).	
Tsui et al. [29]	GPT	Providing ophthalmic symptom triage recommendations	Eight out of ten sets of responses from ChatGPT received a grade of both precise and suitable.	
Lyons et al. [30]	GPT-4, Bing Chat	Providing ophthalmic symptom triage recommendations	The physician respondents, ChatGPT, Bing Chat, and WebMD listed the appropriate diagnosis among the top three suggestions in 42 (95%), 41 (93%), 34 (77%), and 8 (33%) cases, respectively.	
Huang et al. [31]	GPT-4	Assessing performance on glaucoma and retina questions	ChatGPT significantly improved both the efficiency and the quality of writing of review articles for scientists. However, it is important to keep in mind the limitations of ChatGPT’s capabilities for writing review articles.	
Waisberg et al. [33]	GPT	Enhancing smart glasses technology with LLM integration	Recent update included GPT integration with smart glasses, allowing users to ask glasses-specific questions, such as summarizing text or reading only vegan items from a menu.	
Antaki et al. [38]	GPT-3.5	Assessing performance on BCSC and OphthoQuestions	ChatGPT demonstrated accuracy of 55.8% on the BCSC exam and 42.7% on the OphthoQuestions.	
Mihalache et al. [39]	GPT	Assessing performance on OphthoQuestions	ChatGPT correctly answered 58 of 125 questions (46%) on the OphthoQuestions.	
Panthier et al. [40]	GPT-4	Assessing performance on FEBO in French	ChatGPT achieved 91.2% correctness on the French language EBO examination.	
Antaki et al. [41]	GPT-4	Assessing performance on BCSC and OphthoQuestions	GPT-4.0 achieved 75.8% correctness on the BCSC exam and 70.0% on the OphthoQuestions.	
Fowler et al. [42]	GPT-4, Bard	Assessing performance on FRCOphth practice questions	GPT-4 correctly answered 42 out of the 49 questions (85.7% correctness) and exceeded the exam’s passing threshold. Meanwhile, Bard answered 22 out of the 49 questions correctly, with 44.9% correctness.	
Lin et al. [43]	GPT-3.5, GPT-4	Assessing performance on ophthalmic written exam	GPT-3.5, GPT-4, and human users scored 63.1%, 76.9%, and 72.6%, respectively, and exceeded the 2022 American Board of Ophthalmology written exam’s passing threshold.	
Mihalache et al. [44]	GPT-4	Assessing performance on OphthoQuestions	Of the 125 text-based multiple-choice questions, 105 (84%) were answered correctly by GPT-4.0 on the OphthoQuestions.	
Singh et al. [46]	GPT	Providing discharge summaries and ophthalmic operative notes	The performance of ChatGPT for generating ophthalmic discharge summaries and operative notes was encouraging.	
Waisberg et al. [47]	GPT-4	Providing discharge ophthalmic operative notes with complications	GPT-4 produced a detailed operative note with all the required components of a good ophthalmic note.	
Ong et al. [48]	GPT	Generating ICD codes using text-based information from a retina clinic	ChatGPT’s ability could potentially alleviate physician burden in ICD coding by generating a selection of codes.	
AI artificial intelligence, BCSC Basic and Clinical Science Course, EBO European Board of Ophthalmology, FEBO Fellow of European Board of Ophthalmology, GPT Generative Pre-trained Transformer, ICD International Classification of Diseases, LLM large language model, NLP natural language processing

Although LLM allows the generation of administrative work, its capability cannot totally replace human abilities. The idea of LLM-assisted documentation should be utilized under supervision and tailored to individual patients. We anticipate more refined results for the LLM-generated document as it continues to learn from a more comprehensive EHR [4]. Moreover, comprehensive standards for LLM-generated summaries are still needed. Regardless of current U.S. Food and Drug Administration (FDA) regulations, LLMs should be tested to quantify both benefits and clinical harms prior to real-world implementation. Without FDA regulatory safeguards in place, the use of current LLMs for summarizing clinical data should be approached with caution [50].

Challenges and Limitations

The results of this review show that response correctness is arguably a point of discussion needed prior to clinical deployment. The performance of LLM chatbots has been found satisfactory for general data but less proficient in specific specialty knowledge [26, 27, 51]. Furthermore, the available data used for model development may not be up to date, as ChatGPT was trained on data prior to 2021. Aside from Google Bard and Microsoft Copilot, most LLM chatbots, including ChatGPT, are unable to access the internet in real time, which could cause data incoherence across platforms. For example, Microsoft Copilot was able to include pegcetacoplan, approved by the US FDA in February 2023, in one of the treatment plans, while ChatGPT did not mention this recent medical therapy [2]. Faricimab and brolucizumab might not be included as potential anti-vascular endothelial growth factor (VEGF) treatments, as this information was introduced more recently. This example demonstrates the incoherence of data across platforms.

LLMs generate varying results and occasionally lack repeatability. Despite running the identical medical information at the same time, the LLM generated summaries with differences due to random variability, since there are many “right” ways to summarize information. The output from LLMs is constantly changing due to their probabilistic processes. Variation in LLM-generated summaries could lead to inconsistencies in clinician decision-making [50].

While the ability for creativity in LLM chatbots is an attractive feature, a concern arises if they exploit this ability to generate nonexistent false scientific research [52]. Writing an article can be a tiresome process, which may lead to the misuse of LLM chatbots, violating ethics standards. This is potentially harmful not only to patients but also in the scientific community. There are tools that can detect LLM model use, such as OpenAI's text classifier, GPTZero, and Copyleaks, as such use creates an ethical issue regarding authorship disclosure of the use of LLM chatbots [53]. Plagiarism is another concern when using an LLM chatbot without reviewing it. The generated text may resemble text from other works, such as published articles or online sources. Hence, LLM chatbots should be considered a supplement for writing scientific articles rather than a replacement [36].

Observations show that LLM chatbots occasionally generate responses that align with the user’s incorrect or vague assumptions. This phenomenon is called “falsehood mimicry” and is observed when the user input lacks clarity or accuracy [2, 54, 55]. “Sycophancy bias” describes the bias behavior of LLMs that tend to generate responses that align with user expectations. Without reviewing for a confirmation bias, the LLM could increase diagnostic error [50]. Moreover, the generated response is grammatically polished and may be more convincing than that from a human. The LLM may produce writing that looks perfect but contains false information. Hence, all data should be carefully examined, as users may be inclined to agree with responses that conform to their beliefs. Furthermore, LLM chatbots tend to report data with a bias from the existing data used for training. Hence, it is important to build LLMs that have doubt concerns and acknowledge uncertainty rather than copying resources with preexisting incorrect information or biases [56].

The current concern regarding data privacy is a complex issue, especially when software needs to be trained on the EHR. To generate query responses, LLMs may need to access patients’ information, including medical history and ocular images. These data are considered confidential and require patient consent and acceptance. This issue is highly sensitive when the information is inadvertently disclosed. Providing for the security of the ophthalmic images would be necessary before incorporating the LLM into a real-world workflow to ensure that patient data will be kept safely encrypted and that patients will be promptly notified of any data breach [4].

In an era with a huge amount of information on the internet, LLMs can gather both correct relevant and false data regardless of references. Most LLMs are trained with little regard to source variability. Therefore, they could treat articles published in credible journals and random websites as equally reliable. Hence, comparing LLM-generated results with other reputable sources identified in Google searches or standard clinical guidelines could support decision-making and avoid pitfalls. Discordant results should be a concern, and should demand pursuit of further clarification [57]. Small errors with important clinical influence could induce faulty decision-making [50].

As real-world medical decision-making is a highly intricate process, it is crucial that ethical, technical, and legal implications are addressed to ensure the safety of LLM use in medical fields. The LLM should be considered a tool for augmenting the capabilities of health care professionals, not replacing them [58].

Conclusion

The LLM is an emerging technology that can potentially bring about changes in medical practice, including ophthalmology. Its applicability involves clinical, educational, and research applications. Although current performance is promising, real-world implementation remains controversial and open for further studies. As LLM is also continuously evolving, we can expect the future performance to improve the quality of eye health care.

Author Contributions

Peranut Chotcomwongse: concept and design, drafting and finalizing manuscript. Paisan Ruamviboonsuk: concept and design, drafting manuscript. Andrzej Grzybowski: concept and design, finalizing manuscript.

Funding

No funding or sponsorship was received for this study or publication of this article.

Declarations

Conflict of Interest

Peranut Chotcomwongse has nothing to disclose. Paisan Ruamviboonsuk has nothing to disclose. Andrzej Grzybowski is an Editorial Board member of Ophthalmology and Therapy. Andrzej Grzybowski was not involved in the selection of peer reviewers for the manuscript nor any of the subsequent editorial decisions.

Ethical Approval

This article is based on previously conducted studies and does not contain any new studies with human participants or animals performed by any of the authors. Ethics committee approval was not required for this study because of no involvement of human subjects.
==== Refs
References

1. Jin K Yuan L Wu H Grzybowski A Ye J Exploring large language model for next generation of artificial intelligence in ophthalmology Front Med 2023 10 1291404 10.3389/fmed.2023.1291404
Jin K, Yuan L, Wu H, Grzybowski A, Ye J. Exploring large language model for next generation of artificial intelligence in ophthalmology. Front Med. 2023;10:1291404.
2. Tan TF Thirunavukarasu AJ Campbell JP Keane PA Pasquale LR Abramoff MD Generative artificial intelligence through ChatGPT and other large language models in ophthalmology: clinical applications and challenges Ophthalmol Sci 2023 10.1016/j.xops.2023.100394 38313399
Tan TF, Thirunavukarasu AJ, Campbell JP, Keane PA, Pasquale LR, Abramoff MD, et al. Generative artificial intelligence through ChatGPT and other large language models in ophthalmology: clinical applications and challenges. Ophthalmol Sci. 2023. 10.1016/j.xops.2023.100394.38313399
3. Waisberg E Ong J Masalkhi M Zaman N Sarker P Lee AG GPT-4 and medical image analysis: strengths, weaknesses and future directions J Med Artif Intell 2023 6 29 29 10.21037/jmai-23-94
Waisberg E, Ong J, Masalkhi M, Zaman N, Sarker P, Lee AG, et al. GPT-4 and medical image analysis: strengths, weaknesses and future directions. J Med Artif Intell. 2023;6:29–29. 10.21037/jmai-23-94.
4. Betzler BK Chen H Cheng CY Lee CS Ning G Song SJ Large language models and their impact in ophthalmology Lancet 2023 5 e917 e924
Betzler BK, Chen H, Cheng CY, Lee CS, Ning G, Song SJ, et al. Large language models and their impact in ophthalmology. Lancet. 2023;5:e917–24.
5. Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, et al. Training language models to follow instructions with human feedback 2022 [Internet]. arXiv:2203.02155.
6. Wong G. GPT-4 architecture, infrastructure, training dataset, costs, vision, MoE. 2023 [Internet].
7. Waisberg E Ong J Masalkhi M Kamran SA Zaman N Sarker P GPT-4: a new era of artificial intelligence in medicine Ir J Med Sci 2023 192 3197 3200 10.1007/s11845-023-03377-8 37076707
Waisberg E, Ong J, Masalkhi M, Kamran SA, Zaman N, Sarker P, et al. GPT-4: a new era of artificial intelligence in medicine. Ir J Med Sci. 2023;192:3197–200.37076707
8. Cai LZ Shaheen A Jin A Fukui R Yi JS Yannuzzi N Performance of generative large language models on ophthalmology board-style questions Am J Ophthalmol 2023 254 141 149 10.1016/j.ajo.2023.05.024 37339728
Cai LZ, Shaheen A, Jin A, Fukui R, Yi JS, Yannuzzi N, et al. Performance of generative large language models on ophthalmology board-style questions. Am J Ophthalmol. 2023;254:141–9. 10.1016/j.ajo.2023.05.024.37339728
9. Waisberg E Ong J Masalkhi M Zaman N Sarker P Lee AG Google’s AI chatbot “Bard”: a side-by-side comparison with ChatGPT and its utilization in ophthalmology Eye (Basingstoke). 2023 38 4 642 645
Waisberg E, Ong J, Masalkhi M, Zaman N, Sarker P, Lee AG, et al. Google’s AI chatbot “Bard”: a side-by-side comparison with ChatGPT and its utilization in ophthalmology. Eye (Basingstoke). 2023;38(4):642–5.
10. Zhao H, Ling Q, Pan Y, Zhong T, Hu J-Y, Yao J, et al. Ophtha-LLaMA2: a large language model for ophthalmology. 2023 [Internet]. 2023. arXiv:2312.04906.
11. Yu Z Yang X Sweeting GL Ma Y Stolte SE Fang R Identify diabetic retinopathy-related clinical concepts and their attributes using transformer-based natural language processing methods BMC Med Inform Decis Mak 2022 10.1186/s12911-022-01996-2 36457119
Yu Z, Yang X, Sweeting GL, Ma Y, Stolte SE, Fang R, et al. Identify diabetic retinopathy-related clinical concepts and their attributes using transformer-based natural language processing methods. BMC Med Inform Decis Mak. 2022. 10.1186/s12911-022-01996-2.36457119
12. Lim ZW, Pushpanathan K, Min S, Yew E, Lai Y, Sun C-H, et al. Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard. 2023 [Internet]. www.thelancet.com.
13. Ali MJ ChatGPT and lacrimal drainage disorders: performance and scope of improvement Ophthalmic Plast Reconstr Surg 2023 39 3 221 225 10.1097/IOP.0000000000002418 37166289
Ali MJ. ChatGPT and lacrimal drainage disorders: performance and scope of improvement. Ophthalmic Plast Reconstr Surg. 2023;39(3):221–5. 10.1097/IOP.0000000000002418.37166289
14. Al-Sharif EM Penteado RC Dib El Jalbout N Topilow NJ Shoji MK Kikkawa DO Evaluating the accuracy of ChatGPT and Google BARD in fielding oculoplastic patient queries: a comparative study on artificial versus human intelligence Ophthalmic Plast Reconstr Surg 2024 10.1097/IOP.0000000000002567 38722772
Al-Sharif EM, Penteado RC, Dib El Jalbout N, Topilow NJ, Shoji MK, Kikkawa DO, et al. Evaluating the accuracy of ChatGPT and Google BARD in fielding oculoplastic patient queries: a comparative study on artificial versus human intelligence. Ophthalmic Plast Reconstr Surg. 2024. 10.1097/IOP.0000000000002567.38722772
15. Momenaei B Wakabayashi T Shahlaee A Durrani AF Pandit SA Wang K Appropriateness and readability of ChatGPT-4-generated responses for surgical treatment of retinal diseases Ophthalmol Retina 2023 7 10 862 868 10.1016/j.oret.2023.05.022 37277096
Momenaei B, Wakabayashi T, Shahlaee A, Durrani AF, Pandit SA, Wang K, et al. Appropriateness and readability of ChatGPT-4-generated responses for surgical treatment of retinal diseases. Ophthalmol Retina. 2023;7(10):862–8. 10.1016/j.oret.2023.05.022.37277096
16. Masalkhi M Ong J Waisberg E Zaman N Sarker P Lee AG ChatGPT to document ocular infectious diseases Eye (Basingstoke) 2023 38 5 826 828
Masalkhi M, Ong J, Waisberg E, Zaman N, Sarker P, Lee AG, et al. ChatGPT to document ocular infectious diseases. Eye (Basingstoke). 2023;38(5):826–8.
17. Waisberg E Ong J Masalkhi M Lee AG Large language model (LLM)-driven chatbots for neuro-ophthalmic medical education Eye (Basingstoke). 2023 38 4 639 641
Waisberg E, Ong J, Masalkhi M, Lee AG. Large language model (LLM)-driven chatbots for neuro-ophthalmic medical education. Eye (Basingstoke). 2023;38(4):639–41.
18. Tan Yip Ming C Rojas-Carabali W Cifuentes-González C Agrawal R Thorne JE Tugal-Tutkun I The potential role of large language models in uveitis care: perspectives after ChatGPT and Bard Launch Ocul Immunol Inflamm 2023 10.1080/09273948.2023.2242462 37562028
Tan Yip Ming C, Rojas-Carabali W, Cifuentes-González C, Agrawal R, Thorne JE, Tugal-Tutkun I, et al. The potential role of large language models in uveitis care: perspectives after ChatGPT and Bard Launch. Ocul Immunol Inflamm. 2023. 10.1080/09273948.2023.224246237562028
19. Waisberg E Ong J Masalkhi M Kamran SA Zaman N Sarker P GPT-4 and ophthalmology operative notes Ann Biomed Eng 2023 51 2353 2355 10.1007/s10439-023-03263-5 37266720
Waisberg E, Ong J, Masalkhi M, Kamran SA, Zaman N, Sarker P, et al. GPT-4 and ophthalmology operative notes. Ann Biomed Eng. 2023;51:2353–5.37266720
20. Rasmussen MLR Larsen AC Subhi Y Potapenko I Artificial intelligence-based ChatGPT chatbot responses for patient and parent questions on vernal keratoconjunctivitis Graefe’s Arch Clin Exp Ophthalmol. 2023 261 10 3041 3043 10.1007/s00417-023-06078-1 37129631
Rasmussen MLR, Larsen AC, Subhi Y, Potapenko I. Artificial intelligence-based ChatGPT chatbot responses for patient and parent questions on vernal keratoconjunctivitis. Graefe’s Arch Clin Exp Ophthalmol. 2023;261(10):3041–3. 10.1007/s00417-023-06078-1.37129631
21. Balas M Ing EB Conversational AI models for ophthalmic diagnosis: comparison of ChatGPT and the Isabel pro differential diagnosis generator JFO Open Ophthalmol 2023 1 100005 10.1016/j.jfop.2023.100005
Balas M, Ing EB. Conversational AI models for ophthalmic diagnosis: comparison of ChatGPT and the Isabel pro differential diagnosis generator. JFO Open Ophthalmol. 2023;1: 100005. 10.1016/j.jfop.2023.100005.
22. Delsoz M Raja H Madadi Y Tang AA Wirostko BM Kahook MY The use of ChatGPT to assist in diagnosing glaucoma based on clinical case reports Ophthalmol Ther 2023 12 6 3121 3132 10.1007/s40123-023-00805-x 37707707
Delsoz M, Raja H, Madadi Y, Tang AA, Wirostko BM, Kahook MY, et al. The use of ChatGPT to assist in diagnosing glaucoma based on clinical case reports. Ophthalmol Ther. 2023;12(6):3121–32. 10.1007/s40123-023-00805-x.37707707
23. Shemer A Cohen M Altarescu A Atar-Vardi M Hecht I Dubinsky-Pertzov B Diagnostic capabilities of ChatGPT in ophthalmology Graefe’s Arch Clin Exp Ophthalmol 2024 10.1007/s00417-023-06363-z
Shemer A, Cohen M, Altarescu A, Atar-Vardi M, Hecht I, Dubinsky-Pertzov B, et al. Diagnostic capabilities of ChatGPT in ophthalmology. Graefe’s Arch Clin Exp Ophthalmol. 2024. 10.1007/s00417-023-06363-z.
24. Liu X, Wu J, Shao A, Shen W, Ye P, Wang Y, et al. Transforming retinal vascular disease classification: a comprehensive analysis of ChatGPT’s performance and inference abilities on non-english clinical environment. medRxiv [preprint]. 2023. 10.1101/2023.06.28.23291931.
25. Hu X Ran AR Nguyen TX Szeto S Yam JC Chan CKM What can GPT-4 do for diagnosing rare eye diseases? A pilot study Ophthalmol Ther 2023 12 6 3395 3402 10.1007/s40123-023-00789-8 37656399
Hu X, Ran AR, Nguyen TX, Szeto S, Yam JC, Chan CKM, et al. What can GPT-4 do for diagnosing rare eye diseases? A pilot study. Ophthalmol Ther. 2023;12(6):3395–402. 10.1007/s40123-023-00789-8.37656399
26. Rojas-Carabali W Cifuentes-González C Wei X Putera I Sen A Thng ZX Evaluating the diagnostic accuracy and management recommendations of ChatGPT in uveitis Ocul Immunol Inflamm 2023 10.1080/09273948.2023.2253471 38133945
Rojas-Carabali W, Cifuentes-González C, Wei X, Putera I, Sen A, Thng ZX, et al. Evaluating the diagnostic accuracy and management recommendations of ChatGPT in uveitis. Ocul Immunol Inflamm. 2023. 10.1080/09273948.2023.2253471.38133945
27. Rojas-Carabali W Sen A Agarwal A Tan G Cheung CY Rousselot A Chatbots vs. human experts: evaluating diagnostic performance of chatbots in uveitis and the perspectives on AI adoption in ophthalmology Ocul Immunol Inflamm 2023 10.1080/09273948.2023.2266730 38133945
Rojas-Carabali W, Sen A, Agarwal A, Tan G, Cheung CY, Rousselot A, et al. Chatbots vs. human experts: evaluating diagnostic performance of chatbots in uveitis and the perspectives on AI adoption in ophthalmology. Ocul Immunol Inflamm. 2023. 10.1080/09273948.2023.2266730.38133945
28. Delsoz M, Madadi Y, Munir WM, Tamm B, Mehravaran S, Soleimani M, et al. Performance of ChatGPT in diagnosis of corneal eye diseases. medRxiv [preprint]. 2023. 10.1101/2023.08.25.23294635.
29. Tsui JC Wong MB Kim BJ Maguire AM Scoles D VanderBeek BL Appropriateness of ophthalmic symptoms triage by a popular online artificial intelligence Chatbot Eye (Basingstoke) 2023 37 17 3692 3693 10.1038/s41433-023-02556-2
Tsui JC, Wong MB, Kim BJ, Maguire AM, Scoles D, VanderBeek BL, et al. Appropriateness of ophthalmic symptoms triage by a popular online artificial intelligence Chatbot. Eye (Basingstoke). 2023;37(17):3692–3. 10.1038/s41433-023-02556-2.
30. Lyons RJ Arepalli SR Fromal O Choi JD Jain N Artificial intelligence Chatbot performance in triage of ophthalmic conditions Can J Ophthalmol 2023 10.1101/2023.06.11.23291247 37572695
Lyons RJ, Arepalli SR, Fromal O, Choi JD, Jain N. Artificial intelligence Chatbot performance in triage of ophthalmic conditions. Can J Ophthalmol. 2023. 10.1101/2023.06.11.23291247.37572695
31. Huang AS Hirabayashi K Barna L Parikh D Pasquale LR Assessment of a large language model’s responses to questions and cases about glaucoma and retina management JAMA Ophthalmol 2024 10.1001/jamaophthalmol.2023.6917 39235822
Huang AS, Hirabayashi K, Barna L, Parikh D, Pasquale LR. Assessment of a large language model’s responses to questions and cases about glaucoma and retina management. JAMA Ophthalmol. 2024. 10.1001/jamaophthalmol.2023.6917.39235822
32. Young BK Zhao PY Large language models and the shoreline of ophthalmology JAMA Ophthalmol. 2024 142 4 375 376 10.1001/jamaophthalmol.2023.6937 38386327
Young BK, Zhao PY. Large language models and the shoreline of ophthalmology. JAMA Ophthalmol. 2024;142(4):375–6.38386327
33. Waisberg E Ong J Masalkhi M Zaman N Sarker P Lee AG Meta smart glasses—large language models and the future for assistive glasses for individuals with vision impairments Eye (Basingstoke). 2023 38 6 1036 1038
Waisberg E, Ong J, Masalkhi M, Zaman N, Sarker P, Lee AG, et al. Meta smart glasses—large language models and the future for assistive glasses for individuals with vision impairments. Eye (Basingstoke). 2023;38(6):1036–8.
34. Mihalache A Huang RS Popovic MM Patil NS Pandya BU Shor R Accuracy of an artificial intelligence Chatbot’s interpretation of clinical ophthalmic images JAMA Ophthalmol 2024 142 4 321 326 10.1001/jamaophthalmol.2024.0017 38421670
Mihalache A, Huang RS, Popovic MM, Patil NS, Pandya BU, Shor R, et al. Accuracy of an artificial intelligence Chatbot’s interpretation of clinical ophthalmic images. JAMA Ophthalmol. 2024;142(4):321–6. 10.1001/jamaophthalmol.2024.0017.38421670
35. Mihalache A Huang RS Mikhail D Popovic MM Shor R Pereira A Interpretation of clinical retinal images using an artificial intelligence Chatbot Ophthalmol Sci. 2024 10.1016/j.xops.2024.100556 39139542
Mihalache A, Huang RS, Mikhail D, Popovic MM, Shor R, Pereira A, et al. Interpretation of clinical retinal images using an artificial intelligence Chatbot. Ophthalmol Sci. 2024. 10.1016/j.xops.2024.100556.39139542
36. Huang J, Tan M. The role of ChatGPT in scientific communication: writing better scientific review articles. Am J Cancer Res. 2023; 13 [Internet]. www.ajcr.us/.
37. Park JY Could ChatGPT help you to write your next scientific paper?: concerns on research ethics related to usage of artificial intelligence tools J Korean Assoc Oral Maxillofac Surg 2023 49 105 106 10.5125/jkaoms.2023.49.3.105 37394928
Park JY. Could ChatGPT help you to write your next scientific paper?: concerns on research ethics related to usage of artificial intelligence tools. J Korean Assoc Oral Maxillofac Surg. 2023;49:105–6.37394928
38. Antaki F Touma S Milad D El-Khoury J Duval R Evaluating the performance of ChatGPT in ophthalmology: an analysis of its successes and shortcomings Ophthalmol Sci 2023 10.1101/2023.01.22.23284882 37334036
Antaki F, Touma S, Milad D, El-Khoury J, Duval R. Evaluating the performance of ChatGPT in ophthalmology: an analysis of its successes and shortcomings. Ophthalmol Sci. 2023. 10.1101/2023.01.22.23284882.37334036
39. Mihalache A Popovic MM Muni RH Performance of an artificial intelligence chatbot in ophthalmic knowledge assessment JAMA Ophthalmol 2023 141 6 589 597 10.1001/jamaophthalmol.2023.1144 37103928
Mihalache A, Popovic MM, Muni RH. Performance of an artificial intelligence chatbot in ophthalmic knowledge assessment. JAMA Ophthalmol. 2023;141(6):589–97. 10.1001/jamaophthalmol.2023.1144.37103928
40. Panthier C Gatinel D Success of ChatGPT, an AI language model, in taking the French language version of the European Board of Ophthalmology examination: a novel approach to medical knowledge assessment J Fr Ophtalmol 2023 46 7 706 711 10.1016/j.jfo.2023.05.006 37537126
Panthier C, Gatinel D. Success of ChatGPT, an AI language model, in taking the French language version of the European Board of Ophthalmology examination: a novel approach to medical knowledge assessment. J Fr Ophtalmol. 2023;46(7):706–11. 10.1016/j.jfo.2023.05.006.37537126
41. Antaki F Milad D Chia MA Giguère C-É Touma S El-Khoury J Capabilities of GPT-4 in ophthalmology: an analysis of model entropy and progress towards human-level medical question answering Br J Ophthalmol 2023 10.1136/bjo-2023-324438 37923374
Antaki F, Milad D, Chia MA, Giguère C-É, Touma S, El-Khoury J, et al. Capabilities of GPT-4 in ophthalmology: an analysis of model entropy and progress towards human-level medical question answering. Br J Ophthalmol. 2023. 10.1136/bjo-2023-324438.37923374
42. Fowler T Pullen S Birkett L Performance of ChatGPT and Bard on the official part 1 FRCOphth practice questions Br J Ophthalmol 2023 10.1136/bjo-2023-324091 37932006
Fowler T, Pullen S, Birkett L. Performance of ChatGPT and Bard on the official part 1 FRCOphth practice questions. Br J Ophthalmol. 2023. 10.1136/bjo-2023-324091.37932006
43. Lin JC Younessi DN Kurapati SS Tang OY Scott IU Comparison of GPT-3.5, GPT-4, and human user performance on a practice ophthalmology written examination Eye (Basingstoke). 2023 37 17 3694 3695 10.1038/s41433-023-02564-2
Lin JC, Younessi DN, Kurapati SS, Tang OY, Scott IU. Comparison of GPT-3.5, GPT-4, and human user performance on a practice ophthalmology written examination. Eye (Basingstoke). 2023;37(17):3694–5. 10.1038/s41433-023-02564-2.
44. Mihalache A Huang RS Popovic MM Muni RH Performance of an upgraded artificial intelligence Chatbot for ophthalmic knowledge assessment JAMA Ophthalmol 2023 141 8 796 798 10.1001/jamaophthalmol.2023.2710 37410447
Mihalache A, Huang RS, Popovic MM, Muni RH. Performance of an upgraded artificial intelligence Chatbot for ophthalmic knowledge assessment. JAMA Ophthalmol. 2023;141(8):796–8. 10.1001/jamaophthalmol.2023.2710.37410447
45. Kleinig O Gao C Kovoor JG Gupta AK Bacchi S Chan WO How to use large language models in ophthalmology: from prompt engineering to protecting confidentiality Eye (Basingstoke). 2023 38 4 649 653
Kleinig O, Gao C, Kovoor JG, Gupta AK, Bacchi S, Chan WO. How to use large language models in ophthalmology: from prompt engineering to protecting confidentiality. Eye (Basingstoke). 2023;38(4):649–53.
46. Singh S Djalilian A Ali MJ ChatGPT and ophthalmology: exploring its potential with discharge summaries and operative notes Semin Ophthalmol 2023 38 5 503 507 10.1080/08820538.2023.2209166 37133418
Singh S, Djalilian A, Ali MJ. ChatGPT and ophthalmology: exploring its potential with discharge summaries and operative notes. Semin Ophthalmol. 2023;38(5):503–7. 10.1080/08820538.2023.2209166.37133418
47. Waisberg E Ong J Masalkhi M Zaman N Sarker P Lee AG GPT-4 to document ophthalmic post-operative complications Eye (Basingstoke). 2023 38 3 414 415
Waisberg E, Ong J, Masalkhi M, Zaman N, Sarker P, Lee AG, et al. GPT-4 to document ophthalmic post-operative complications. Eye (Basingstoke). 2023;38(3):414–5.
48. Ong J Kedia N Harihar S Vupparaboina SC Singh SR Venkatesh R Applying large language model artificial intelligence for retina International Classification of Diseases (ICD) coding J Med Artif Intell. 2023 10.21037/jmai-23-106
Ong J, Kedia N, Harihar S, Vupparaboina SC, Singh SR, Venkatesh R, et al. Applying large language model artificial intelligence for retina International Classification of Diseases (ICD) coding. J Med Artif Intell. 2023. 10.21037/jmai-23-106.
49. Wang SY Huang J Hwang H Hu W Tao S Hernandez-Boussard T Leveraging weak supervision to perform named entity recognition in electronic health records progress notes to identify the ophthalmology exam Int J Med Inform 2022 10.1016/j.ijmedinf.2022.104864 36577203
Wang SY, Huang J, Hwang H, Hu W, Tao S, Hernandez-Boussard T. Leveraging weak supervision to perform named entity recognition in electronic health records progress notes to identify the ophthalmology exam. Int J Med Inform. 2022. 10.1016/j.ijmedinf.2022.104864.36577203
50. Goodman KE Yi PH Morgan DJ AI-generated clinical summaries require more than accuracy JAMA 2022 10.1001/jama.2024.0555 36318140
Goodman KE, Yi PH, Morgan DJ. AI-generated clinical summaries require more than accuracy. JAMA. 2024. 10.1001/jama.2024.055536318140
51. Hu W Wang SY Predicting glaucoma progression requiring surgery using clinical free-text notes and transfer learning with transformers Transl Vis Sci Technol. 2022 10.1167/tvst.11.3.37 36264650
Hu W, Wang SY. Predicting glaucoma progression requiring surgery using clinical free-text notes and transfer learning with transformers. Transl Vis Sci Technol. 2022. 10.1167/tvst.11.3.37.36264650
52. Elali FR Rachid LN AI-generated research paper fabrication and plagiarism in the scientific community Patterns 2023 10.1016/j.patter.2023.100706 36960451
Elali FR, Rachid LN. AI-generated research paper fabrication and plagiarism in the scientific community. Patterns. 2023. 10.1016/j.patter.2023.10070636960451
53. Hosseini M Resnik DB Holmes K The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts Res Ethics 2023 19 4 449 465 10.1177/17470161231180449
Hosseini M, Resnik DB, Holmes K. The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Res Ethics. 2023;19(4):449–65. 10.1177/17470161231180449.
54. Ji Z Lee N Frieske R Yu T Su D Xu Y Survey of hallucination in natural language generation ACM Comput Surv 2022 10.1145/3571730
Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2022. 10.1145/3571730.
55. Alkaissi H McFarlane SI Artificial hallucinations in ChatGPT: implications in scientific writing Cureus 2023 10.7759/cureus.35179 37605676
Alkaissi H, McFarlane SI. Artificial hallucinations in ChatGPT: implications in scientific writing. Cureus. 2023. 10.7759/cureus.35179.37605676
56. Thirunavukarasu AJ Large language models will not replace healthcare professionals: curbing popular fears and hype J R Soc Med 2023 116 181 182 10.1177/01410768231173123 37199678
Thirunavukarasu AJ. Large language models will not replace healthcare professionals: curbing popular fears and hype. J R Soc Med. 2023;116:181–2.37199678
57. Mello MM Guha N ChatGPT and physicians’ malpractice risk JAMA Health Forum. 2023 4 e231938 10.1001/jamahealthforum.2023.1938 37200013
Mello MM, Guha N. ChatGPT and physicians’ malpractice risk. JAMA Health Forum. 2023;4: e231938.37200013
58. Meskó B The impact of multimodal large language models on health care’s future J Med Internet Res 2023 25 e52865 10.2196/52865 37917126
Meskó B. The impact of multimodal large language models on health care’s future. J Med Internet Res. 2023;25: e52865.37917126
