
==== Front
Explor Res Clin Soc Pharm
Explor Res Clin Soc Pharm
Exploratory Research in Clinical and Social Pharmacy
2667-2766
Elsevier

S2667-2766(24)00089-1
10.1016/j.rcsop.2024.100492
100492
Article
Unlocking the potential of advanced large language models in medication review and reconciliation: A proof-of-concept investigation
Sridharan Kannan skannandr@gmail.com
a⁎
Sivaramakrishnan Gowri b
a Department of Pharmacology & Therapeutics, College of Medicine & Medical Sciences, Arabian Gulf University, Manama, Bahrain
b Speciality Dental Residency Program, Primary Health Care Centers, Manama, Bahrain
⁎ Corresponding author. skannandr@gmail.com
17 8 2024
9 2024
17 8 2024
15 10049217 5 2024
8 8 2024
13 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Background

Medication review and reconciliation is essential for optimizing drug therapy and minimizing medication errors. Large language models (LLMs) have been recently shown to possess a lot of potential applications in healthcare field due to their abilities of deductive, abductive, and logical reasoning. The present study assessed the abilities of LLMs in medication review and medication reconciliation processes.

Methods

Four LLMs were prompted with appropriate queries related to dosing regimen errors, drug-drug interactions, therapeutic drug monitoring, and genomics-based decision-making process. The veracity of the LLM outputs were verified from validated sources using pre-validated criteria (accuracy, relevancy, risk management, hallucination mitigation, and citations and guidelines). The impacts of the erroneous responses on the patients' safety were categorized either as major or minor.

Results

In the assessment of four LLMs regarding dosing regimen errors, drug-drug interactions, and suggestions for dosing regimen adjustments based on therapeutic drug monitoring and genomics-based individualization of drug therapy, responses were generally consistent across prompts with no clear pattern in response quality among the LLMs. For identification of dosage regimen errors, ChatGPT performed well overall, except for the query related to simvastatin. In terms of potential drug-drug interactions, all LLMs recognized interactions with warfarin but missed the interaction between metoprolol and verapamil. Regarding dosage modifications based on therapeutic drug monitoring, Claude-Instant provided appropriate suggestions for two scenarios and nearly appropriate suggestions for the other two. Similarly, for genomics-based decision-making, Claude-Instant offered satisfactory responses for four scenarios, followed by Gemini for three. Notably, Gemini stood out by providing references to guidelines or citations even without prompting, demonstrating a commitment to accuracy and reliability in its responses. Minor impacts were noted in identifying appropriate dosing regimens and therapeutic drug monitoring, while major impacts were found in identifying drug interactions and making pharmacogenomic-based therapeutic decisions.

Conclusion

Advanced LLMs hold significant promise in revolutionizing the medication review and reconciliation process in healthcare. Diverse impacts on patient safety were observed. Integrating and validating LLMs within electronic health records and prescription systems is essential to harness their full potential and enhance patient safety and care quality.

Keywords

Artificial intelligence
ChatGPT
Gemini
Claude-instant
Llama
Pharmacy
==== Body
pmc1 Introduction

Medication errors and unsafe medication practices are among the leading causes of preventable harm to patients worldwide, costing around $42 billion annually.1 According to the Institute of Medicine's report ‘Preventing Medication Errors’, at least one medication error per day was observed among hospitalized patients.2 Adverse drug events were responsible for 197,000 deaths annually in Europe and around 1.3 million emergency department visits in the US.3,4

Medication reconciliation, which involves a systematic review of prescriptions, helps contain medication errors such as omissions, duplications, dosing errors, or drug interactions.5 Medication review is a structured evaluation of prescribed medications with the aim of identifying and resolving drug-related issues.6 Errors in dosing regimens like incorrect dosages, frequencies, durations, or timing of drug administration and simultaneous administration of interacting medicines are some commonly observed types of medication errors.7 One of the latest advancements in mitigating medication-related problems is genomic-based personalized drug therapy.8

Large language models (LLMs) show promising applications in enhancing various healthcare activities like compiling patient notes, facilitating communication between patients and providers, and aiding clinical decision making.9 However, LLMs have also been demonstrated to generate misinformation, fabricated data, plagiarism, and provide information that does not exist.10 A mixed-methods study assessing practicing clinicians' perceptions of advanced LLMs reached a consensus that they excel at synthesizing and analyzing data and showed a positive attitude toward employing them in healthcare settings.11 Generative artificial intelligence (AI) or LLMs are trained on vast datasets, enabling them to learn and provide human-like responses.12 The Pharmacy GPT study by researchers at the University of North Carolina Health System involved 5000 patients aged 18 or older who were admitted to intensive care units for at least 24 h.13 The authors observed that the LLM could cluster patients based on key variables, predict outcomes including mortality, and identified a potential role in devising medical plans tailored to patients' conditions. However, there is a dearth of data evaluating the utility of LLMs in optimizing prescribed therapies. With thorough training on extensive datasets and abilities for deductive, abductive, and logical reasoning, LLMs demonstrate great potential for medication review and optimizing therapy while minimizing errors. Thus, we conducted the present study to evaluate four LLMs' capabilities for medication review and reconciliation.

2 Methods

2.1 Study ethics and design

This study was a cross-sectional design carried out during February–May 2024. Institutional Review Board approval was not sought considering the absence of any human interaction or collection of any data.

2.2 Study procedure

The LLMs were separately queried as follows under the following domains:A. Dosing regimen errors

Query # 1: I am a Clinical Pharmacist, and a patient was prescribed as follows:

Tablet Simvastatin 20 mg once daily to be taken in the morning.

Tablet Amlodipine 5 mg taken every night.

Is the prescription appropriate?

Query # 2: I am a Clinical Pharmacist, and I observed the following drug prescription for a patient during my rounds in the intensive care unit.

Injection heparin 5000 International Units intramuscularly once daily.

Tablet warfarin 2.5 mg once daily.

Is the prescription appropriate? Is any modification necessary?

Query # 3: I am a Clinical Pharmacist, and I saw an adult patient weighing 70 kg receiving the following prescription during my inpatient rounds.

Tablet acetylsalicylic acid 80 mg orally once daily.

Tablet furosemide 40 mg orally twice daily.

Tablet spironolactone 100 mg orally once daily.

Tablet digoxin 1.5 mg once daily orally.

Is the prescription appropriate? Should I inform the Clinician to modify anything in the above prescription?

Query # 4: I am a Clinical Pharmacist, and I observed the following prescription during my rounds in the intensive care unit for a patient.

Injection Vancomycin 1-g intravenous bolus.

Is the prescription appropriate? Does it need any modification?B. Drug-drug interactions

Query # 1: I am a Clinical Pharmacist, and a patient was diagnosed with vasospastic angina for which he was prescribed the following drugs by his doctor for 30 days.

Tablet Metoprolol 50 mg once daily.

Sublingual tablet nitroglycerin 0.6 mg PRN.

Tablet verapamil 200 mg once daily.

Is there any error in the above prescription? Is the above prescription appropriate?

Query # 2: I am a Clinical Pharmacist and for a patient with rheumatic heart disease who is taking warfarin tablet 10 mg once daily, the doctor prescribed Erythromycin 800 mg twice daily for a potential bacterial upper respiratory tract infection. Is the prescription appropriate? Does it need any modification?

Query # 3: I am a Clinical Pharmacist, and a patient was diagnosed with rheumatic heart disease on warfarin 10 mg daily. The patient has atrial fibrillation right now and the doctor is planning to initiate amiodarone 400 mg orally once daily. Is the prescription rational, and correct?

Query # 4: I am a patient with chronic heart failure, and I am taking the following drugs:

Tablet Metoprolol 100 mg once daily.

Tablet Spironolactone 50 mg twice daily.

Tablet Empagliflozin 10 mg once daily.

Tabet Furosemide 80 mg twice daily.

Tablet Enalapril 10 mg once daily.

Tablet Potassium chloride 25 mEq twice daily.

Is the prescription rational and appropriate? Is any modification required?C. Therapeutic drug monitoring

Query # 1: I am a Clinical Pharmacist, and a doctor in the intensive care unit has sought my assistance in estimating the appropriate loading dose and maintenance dosing regimen for initiating vancomycin intravenously for a patient with the following details:Age: 50 years

Gender: Male

Weight: 70 kg

Height: 150 cm

Serum creatinine (mg/dl): 1

Estimate the loading dose and maintenance dosing regimen for vancomycin for this patient and provide me with the name of the formula or the reference that you used to calculate the vancomycin dose.

Query # 2: I am a Clinical Pharmacist, and a doctor in the intensive care unit has sought my assistance in estimating the appropriate dosing regimen for initiating gentamicin intravenously for empirical antimicrobial treatment in a patient with the following details:Age: 50 years

Gender: Male

Weight: 70 kg

Height: 150 cm

Serum creatinine (mg/dl): 1

Expected gentamicin peak concentration (mcg/ml): 4

Expected gentamicin trough concentration (mcg/ml): <2

Estimate the dosing regimen for gentamicin in this patient and provide me with the name of the formula or the reference that you used to calculate the dosing regimen.

Query # 3: I am a Clinical Pharmacist, and a patient in the intensive care unit was about to be initiated ciprofloxacin intravenously as a part of empirical antimicrobial therapy. The following were the details of the patient:Age: 50 years

Gender: Male

Weight: 70 kg

Height: 150 cm

Serum creatinine (mg/dl): 3.5

Estimate the dosing regimen for ciprofloxacin in this patient and provide me with the steps and appropriate reference that was used for calculating the dosing regimen.

Query # 4: An adult epileptic patient with normal hepatic and renal functions has been receiving phenytoin tablet 100 mg three times daily for the last 6 months. For the last 2 weeks, the patient has had three episodes of seizures lasting for a few seconds. His serum albumin is 5 g/dl and plasma phenytoin concentration are observed to be 3.6 mg/dl. We use corrected plasma phenytoin concentrations to decide on the dose modifications and the corrected plasma phenytoin concentrations is estimated as follows: observed phenytoin (mg/dl)/ {[0.1 x serum albumin (g/dl)] + 0.1}. The reference range for corrected phenytoin concentration is 10–20 mg/dl. I am a Clinical Pharmacist, and I have been asked to provide suggestions on whether any dose modification of phenytoin is required based on the corrected plasma phenytoin concentrations and if so, what should be the modified dosing regimen in this patient? Suggest the same.D. Pharmacogenomics-based decision-making

Query # 1: I am a Clinical Pharmacist, and a doctor prescribed a patient Warfarin tablet 5 mg once daily for 30 days. Before prescription, he ordered a genetic testing of CYP2C9, and the results came as *3/*3 polymorphisms. Should the dose of warfarin be revised based on this test result?

Query # 2: I am a Clinical Pharmacist, and a doctor prescribed tablet warfarin to an adult male patient. He ordered a genetic testing, and the results were as follows: VKORC1 – AA and CYP2C9 - *3/*3 polymorphisms. What is the dose of warfarin that should be prescribed to this patient?

Query # 3: I am a Clinical Pharmacist, and a doctor has prescribed Aripiprazole tablet 10 mg/day for 30 days for treating schizophrenia in an adult patient. He ordered CYP2D6 genetic testing, and the result came as poor metabolizer. Should the dose of aripiprazole be adjusted? If so, what is the dose to be administered in this patient?

Query # 4: I am a Clinical Pharmacist, and a doctor was planning to initiate a post-renal transplant adult patient weighing 50 kg on tacrolimus tablets. His CYP3A5 genotyping result is *1/*1. What is the dose of tacrolimus to be initiated for this patient?

Query # 5: I am a Clinical Pharmacist, and an adult patient was initiated on clopidogrel tablet at 75 mg once daily for 30 days, a week back. His CYP2C19 genotyping results came as *3/*3. Should the doctor consider modifying his drug treatment (clopidogrel)?

Query # 6: An adult female patient weighing 50 kg was diagnosed with uncomplicated malaria caused by Plasmodium vivax pathogen. The patient is glucose-6-phosphate dehydrogenase deficient. I am a Clinical Pharmacist and I have been asked to provide some advice on the appropriate drug therapy with the dosing regimen for radical cure in this patient. Generate the appropriate dosing regimen for radical cure in this patient.

These scenarios were selected empirically based on the prevalence of drug use and their relevance to clinical pharmacists.

2.3 Large language models

We explored the capabilities of four different generative AI platforms, each built on diverse neural network algorithms. These AI platforms can acquire, synthesize, interpret, and provide logical responses to queries.A. ChatGPT 3.5©: Developed by OpenAI©, ChatGPT 3.5© is a generative AI based on Generative Pre-trained Transformer (GPT) architecture, used in generating text and understanding the language. ChatGPT 3.5 is a modified version of the GPT-3 model, with 6.7 billion parameters compared to GPT-3's 175 billion parameters.14

B. Gemini Pro©: Gemini Pro, developed by Google©, is trained on a massive dataset of billions of texts and code, employing transformers, a neural network architecture, for processing the information and generating text, translating languages, writing different kinds of creative content, and answering the user questions in an informative way.15

C. Claude-Instant©: Claude-Instant is a multimodal AI assistant developed by Anthropic©, that uses Constitutional AI, possessing knowledge related to medications, diseases, and clinical guidelines which it refers to when formulating responses.16

D. Llama-2-13b©: Llama-2-13b is an AI assistant highly trained and experienced language model, capable of processing and generating human-like text, developed by a group of researchers in Meta©. My architecture is based on a deep neural network, with a combination of embeddings, recurrent layers, and attention mechanisms. This LLM is ablet to perform a variety of tasks, such as text classification, language translation, and language generation.17

The veracity of the LLM outputs was verified using recommendations from the British National Formulary 86, United States Food and Drug Administration (USFDA) drug labels, Medicines and Healthcare Products Regulatory Agency, World Health Organization guidelines, and Clinical Pharmacogenetic Implementation Consortium (CPIC) guidelines.18., 19., 20., 21, 22. The LLM outputs were assessed based on the evaluation criteria for LLM outputs recommended by Murugan et al.23:• Accuracy: The degree of alignment with information from the credible sources mentioned above.

• Relevancy: Outputs that are specific, personalized, and appropriately address complex therapeutic problems.

• Risk management: Statements or lists of key strategies related to patient safety, where appropriate.

• Hallucination mitigation: Limiting information not supported by evidence or generated by the LLM.

• Citations and guidelines: References to established guidelines/articles supporting the responses, where applicable.

Two authors independently assessed the LLM outputs as complete, partial or none for each criterion. Any discrepancies were resolved through discussion to reach a consensus. Table 1 outlines the key responses expected for the prompts. Additionally, deficiencies in the overall response for each query were categorized either as major or minor depending on the extent of potential impact on the patient safety.Table 1 Key responses expected for the prompts.

Table 1Domains	Queries	Expected key responses	
Dosing regimen errors	Query # 1	• Simvastatin to be taken in the evening and not in the morning.

• Amlodipine is a CYP3A4 inhibitor and simvastatin undergoes metabolism by CYP3A4.

• Increased risk of myopathy due to interaction between simvastatin and amlodipine.

	
Query # 2	• Heparin should not be administered intramuscularly rather subcutaneously or intravenously.

• Heparin bridge therapy with warfarin is recommended.

	
Query # 3	• Digoxin dose should not exceed 1 mg once daily.

	
Query # 4	• Vancomycin injection should be administered as intravenous infusion.

	
Drug-drug interactions	Query # 1	• Metoprolol should not be administered.

• Metoprolol and verapamil should not be concomitantly administered.

	
Query # 2	• Erythromycin interacts with warfarin and should not be prescribed together.

	
Query # 3	• Amiodarone interacts with warfarin and should not be prescribed together.

	
Query # 4	• Spironolactone, enalapril and potassium together increases the risk of hyperkalemia.

	
Therapeutic drug monitoring	Query # 1	• Vancomycin loading dose is 20–35 mg/kg (1400–2450 mg) and the maintenance dose is 15–20 mg/kg (1050–1400 mg) every 8–12 h.

	
Query # 2	• Dose of gentamicin injection is 5–7 mg/kg once daily

	
Query # 3	• Creatinine clearance is 25 ml/min

• Ciprofloxacin dose is 250–500 mg every 18 h

	
Query # 4	• Corrected serum phenytoin concentration is 6 mg/dl (subtherapeutic)

• Phenytoin dose must be increased by 100 mg/day (from 300 mg/day in the existing regimen to 400 mg/day).

	
Pharmacogenomic-based decision-making	Query # 1	• Warfarin dose to be reduced to 0.5–2 mg/day

	
Query # 2	• The initial warfarin dose should be 0.5–2 mg/day

	
Query # 3	• Aripiprazole dose to be reduced to 5 mg/day

	
Query # 4	• Tacrolimus initial dose should not exceed 0.3 mg/kg/day

	
Query # 5	• Alternative antiplatelet therapy such as ticagrelor or prasugrel should be considered rather than clopidogrel

	
Query # 6	• Primaquine base 0.75 mg/kg once a week for 8 weeks.

	

3 Results

All the LLMs provided answers to all queries (Electronic Supplementary Materials 1–4). In general, none of the LLMs other than Gemini has provided citations for some of the queries even without stating in the prompts.

3.1 Dosing regimen errors

A summary of assessment of the LLM responses for this domain are represented in Table 2. The responses of LLMs were similar across all the domains. However, the qualitative analyses revealed the following:• Query #1: ChatGPT was the sole source that indicated the usual timing of simvastatin administration is in the evening or at night.1 Gemini made the incorrect assertion that statins are better absorbed at night. Claude mistakenly recommended taking amlodipine in the evening to reduce interaction risks.2 Only Llama correctly identified the potential interaction between simvastatin and amlodipine, particularly at high doses.3 The errors of omission for all AI platforms were determined to have a minor impact on patient safety.

• Query #2: ChatGPT exclusively pointed out that heparin should not be administered intramuscularly.4 All respondents except Llama suggested heparin bridge therapy. Llama provided inaccurate information by stating that doses were inappropriate without considering weight or condition details. Furthermore, Llama incorrectly suggested adding protamine sulphate for bleeding management. The omission error by Claude-Instant was categorized as minor while the one from Llama-2-13b was determined to have a major impact on the patient safety.

• Query #3: All respondents except Gemini accurately identified erroneous digoxin dosing. Llama erroneously claimed that the dose of ASA might be insufficient and recommended increasing it. Additionally, Llama incorrectly stated that the doses of furosemide and spironolactone were too high and advised lowering them. The errors from Gemini and Llama-2-13b was determined to have a major impact on patients' safety.

• Query #4: All respondents except Gemini correctly pointed out that vancomycin should be administered intravenously rather than by bolus. Furthermore, all respondents recognized the importance of monitoring renal function and levels for both efficacy and safety.5 The errors from Gemini and Llama-2-13b were determined to have a minor impact on patients' safety.

Table 2 Evaluation of LLM outputs related to dosing regimen errors.

Table 2

In general, the responses of the LLMs were mostly relevant, providing appropriate risk management strategies, with adequate hallucination mitigation and the errors were determined to have minor impact on patients' safety.

3.2 Drug-drug interactions

A summary of assessment of the LLM responses for this domain are represented in Table 3. The responses were similar across the LLMs in all domains. Further, the qualitative analyses revealed the following:• Query #1: None of the LLMs recognized that metoprolol should not be administered in patients with vasospastic angina, nor did they acknowledge its contraindication with verapamil. The omission error from all the AI platforms were determined to have a major impact on patients' safety.

• Query #2: All LLMs correctly identified the potential interaction between erythromycin and warfarin, offering suggestions such as increased monitoring, dosage adjustments, and alternative antimicrobials. Additionally, Llama proposed modifying the erythromycin dosage based on laboratory tests. No errors were detected for this query impacting the patient's safety.

• Query #3: All LLMs correctly identified the potential interaction between amiodarone and warfarin, providing appropriate suggestions. Gemini suggested considering novel oral anticoagulants as alternatives to warfarin to mitigate the interaction, like Llama's suggestion of beta-blockers or calcium channel blockers as alternatives to amiodarone. Claude-Instant precisely recommended reducing warfarin dosage by 30–50% with concurrent amiodarone therapy. No errors were detected for this query impacting the patient's safety.

• Query #4: None of the LLMs recognized the high risk of potential hyperkalemia associated with the concomitant administration of enalapril, spironolactone, and potassium supplementation. Claude-Instant incorrectly stated that potassium supplementation is appropriate with furosemide/spironolactone use, while Llama suggested considering spironolactone, enalapril, beta-blockers, and hydrochlorothiazide as alternatives to furosemide for heart failure management. The omission error from all the AI platforms were determined to have a major impact on patients' safety.

Table 3 Evaluation of LLM outputs related to drug-drug interactions.

Table 3

In general, the responses of the LLMs were mostly relevant, providing appropriate risk management strategies, with adequate hallucination mitigation and the errors were determined to have major impact on patients' safety.

3.3 Therapeutic drug monitoring

A summary of assessment of the LLM responses for this domain are represented in Table 4. The responses were similar across the LLMs in all domains. Further, the qualitative analyses revealed the following:• Query #1: Claude-Instant accurately identified the loading dose of vancomycin (25 mg/kg) and the maintenance dose (15–20 mg/kg). Gemini correctly mentioned the maintenance dose but mistakenly attributed it to the loading dose. ChatGPT provided an erroneous dosage of 35 g for both loading and maintenance doses. Similarly, Llama inaccurately estimated creatinine clearance and vancomycin doses. All except Claude-Instant suggested therapeutic drug monitoring, with ChatGPT recommending target area-under-the-concentration and Gemini and Llama suggesting trough vancomycin concentration. The omission and commission errors from ChatGPT and Llama-2-13b were determined to have a major impact while the error from Gemini was considered to have a minor impact on patients' safety.

• Query #2: All LLMs except Llama recommended loading and maintenance doses of gentamicin. Llama incorrectly estimated creatinine clearance and gentamicin dose. ChatGPT suggested a loading dose of 6–7 mg/kg and maintenance doses ranging from 4 to 5 mg/kg, while Gemini proposed a 4 mg/kg loading dose and a maintenance dose of 16 mg/kg/day. Claude-Instant suggested a 5 mg/kg loading dose and a maintenance dose of 2 mg/kg/day. The omission and commission errors from Llama-2-13b were determined to have a major impact while the error from other AI platforms was considered to have a minor impact on patients' safety.

• Query #3: Only ChatGPT accurately estimated creatinine clearance but did not specify the exact ciprofloxacin dosing regimen, suggesting consulting a reliable source. Gemini erroneously estimated creatinine clearance and ciprofloxacin dose, recommending monitoring blood concentrations if necessary. Claude-Instant's estimations were close to the original values, with a recommended dosing regimen of 200–400 mg every 24 h. Llama used an incorrect formula for estimating creatinine clearance and ciprofloxacin dose. The omission and commission errors from Llama-2-13b and Gemini were determined to have a major impact while the error from other AI platforms was considered to have a minor impact on patients' safety.

• Query #4: All except Llama accurately estimated corrected serum phenytoin concentrations. Claude-Instant specified the correct dosing regimen (adding phenytoin 100 mg at bedtime). Llama suggested a dose of 150 mg for the first 2–3 days, titrating based on response, while Gemini did not provide a specific dose recommendation. ChatGPT recommended increasing the dose to 150 mg three times daily. The omission error from Llama-2-13b was determined to have a major impact on patients' safety while the errors from ChatGPT and Gemini were categorized as minor.

Table 4 Evaluation of LLM outputs related to TDM.

Table 4

In general, the responses of the LLMs were mostly irrelevant and inaccurate, but provided appropriate risk management strategies, with adequate hallucination mitigation and the errors were determined to have minor impact on patients' safety.

3.4 Pharmacogenomics-based decision-making

A summary of assessment of the LLM responses for this domain are represented in Table 5. The responses were similar across the LLMs in all domains. Further, the qualitative analyses revealed the following:• Query #1: All LLMs recommended considering dose reduction of warfarin due to genetic polymorphism, but none suggested the CPIC dose recommendation (0.5–2 mg/day). Claude-Instant suggested a 30–50% dose reduction (2.5–3.5 mg/day), while Llama suggested a range of 2–4 mg/day. The omission error from all the AI platforms were determined to have a major impact on patients' safety.

• Query #2: All LLMs suggested initiating lower warfarin doses. Gemini correctly stated the initial dose of 0.5–2 mg/day, followed by Claude-Instant recommending 1–2 mg/day. ChatGPT suggested 3–4 mg/day initially, and Llama suggested 2–3 mg/day. Claude-Instant also suggested small dose adjustments in units of 0.5 mg for reduced efficacy. The errors from ChatGPT and Llama-2-13b were determined to have a minor impact on patients' safety.

• Query #3: All LLMs suggested considering reduction in aripiprazole doses. ChatGPT and Claude-Instant correctly stated the initial dose as 5 mg/day, while Gemini suggested 7.5–10 mg/day and Llama suggested 2 mg/day. The errors from Gemini and Llama-2-13b were determined to have a minor impact on patients' safety.

• Query #4: ChatGPT, Gemini, and Claude-Instant correctly specified the initial tacrolimus dose as 0.1–0.2 mg/kg/day. Llama erroneously stated the initial dose as 1–1.5 mg/kg/day. All suggested monitoring serum tacrolimus levels, with Claude-Instant recommending titration based on these levels. The error from Llama-2-13b was determined to have a major impact on patients' safety.

• Query #5: All LLMs except Llama suggested considering alternative antiplatelet drugs like prasugrel or ticagrelor. Llama mistakenly suggested increasing the clopidogrel dose to 150–300 mg/day. The error from Llama-2-13b was determined to have a major impact on patients' safety.

• Query #6: All except Gemini suggested dosing regimens for primaquine, but none were correct. Primaquine is recommended once a week for 8 weeks in G6PD-deficient female adults, whereas ChatGPT, Claude-Instant, and Llama suggested a daily regimen for 14 days. Additionally, Gemini erroneously stated that tafenoquine should be considered in G6PD-deficient individuals. The errors from all the AI platforms were determined to have a major impact on patients' safety.

Table 5 Evaluation of LLM outputs related to pharmacogenomics-based decision-making.

Table 5

In general, the responses of the LLMs were mostly relevant with adequate hallucination mitigation and the errors were determined to have major impact on patients' safety.

4 Discussion

4.1 Key findings

In the assessment of four LLMs regarding dosing regimen errors, drug-drug interactions, and suggestions for dosing regimen adjustments based on therapeutic drug monitoring and genomics-based individualization of drug therapy, responses were generally consistent across prompts with no clear pattern in response quality among the LLMs. For identification of dosage regimen errors, ChatGPT performed well overall, except for the query related to simvastatin and in general, the errors were determined to have a minor impact on patient's safety. In terms of potential drug-drug interactions, all LLMs recognized interactions with warfarin but missed the interaction between metoprolol and verapamil and were considered to have a major impact on patient's safety. Regarding dosage modifications based on therapeutic drug monitoring, Claude-Instant provided appropriate suggestions for two scenarios and nearly appropriate suggestions for the other two. The erroneous responses from Llama-2-13b were determined to have a major impact on patient's safety while they were minor with other AI platforms. Similarly, for genomics-based decision-making, Claude-Instant offered satisfactory responses for four scenarios, followed by Gemini for three. Notably, Gemini stood out by providing references to guidelines or citations even without prompting, demonstrating a commitment to accuracy and reliability in its responses. Overall, for this domain, the erroneous responses from AI platforms were determined to have a major impact on patients' safety.

4.2 Comparison with the existing literature

Medication review and reconciliation, led by pharmacists, have been shown to significantly reduce medication errors, with studies indicating a decrease from 61.9% to 9.3%. However, challenges such as lack of interprofessional communication, time constraints, and the absence of structured workflow charts hinder implementation.24,25 Integrating LLMs into electronic health records or prescription systems could aid in identifying potential medication errors and providing appropriate suggestions. Recent integration of LLMs into online prescription systems resulted in significantly more near-miss events of medication errors being caught and corrected before reaching the patient.26

The concomitant administration of interacting drugs is another dimension of prescription-related medication errors. In our study, LLMs, except for one instance involving metoprolol and verapamil, were generally able to identify potential interactions, sometimes partially. Claude-Instant consistently provided appropriate answers to all therapeutic drug monitoring queries, suggesting the potential integration of LLMs with precision dosing.

A study among 724 physicians in a developing country revealed that only 35.3% applied pharmacokinetics/therapeutic drug monitoring principles in clinical practice due to inadequate knowledge and understanding.27 Regarding genomics-based decision-making, LLMs provided either partially or completely accurate and relevant responses for all scenarios except G6PD deficiency, showcasing their potential in suggesting appropriate dosing regimens with risk management strategies based on documented associations with single nucleotide polymorphisms.

In our study, we focused on drugs with well-documented associations across guidelines/societies, including warfarin, tacrolimus, aripiprazole, clopidogrel, and primaquine. Advanced AI models, such as GPT-4 integrated with retrieval-augmented generation, have shown promise in providing personalized therapy recommendations, as demonstrated in a study by Murugan et al.23 With further advancements in AI models and improved access to guidelines, LLMs have immense potential in delivering prompt and thorough responses to genomics-based decision-making processes.

With wider availability and accessibility, LLMs could serve as preliminary online consultation tools for patients, providing evidence-based consultation and enhancing understanding before in-person consultations.28 However, it's crucial to have LLM outputs overseen by experts to mitigate potential risks, despite minimal observed hallucinations in our study.

4.3 Strengths, limitations and future directions

This study represents the initial evaluation of LLMs' capabilities in medication review and reconciliation, focusing on both genetics- and non-genetics-based modifications of dosing regimens, including therapeutic drug monitoring (TDM). However, several limitations should be acknowledged: Limited Scope: The study assessed only common drug-related scenarios, potentially overlooking less common or specialized situations that could pose unique challenges. Future research could expand the scope to encompass a broader range of scenarios to better assess LLM performance across diverse clinical contexts; Lack of Patient Feedback: The study did not directly involve patients or solicit their feedback, which could provide valuable insights into the practical implications and acceptability of LLM-generated recommendations in real-world clinical settings. Incorporating patient perspectives could enhance the relevance and applicability of LLM interventions; Generalizability: Findings from this study may not fully generalize to all clinical settings or populations, as factors such as healthcare infrastructure, provider training, and patient demographics can influence the implementation and effectiveness of LLM-based medication review and reconciliation processes; Evaluation Metrics: The study did not explicitly delineate specific evaluation metrics or criteria for assessing LLM performance, potentially introducing subjectivity or variability in the interpretation of results. Future studies could employ standardized assessment tools to enhance comparability and reproducibility of findings; Technological Constraints: The study did not address potential technological limitations or barriers associated with integrating LLMs into existing healthcare systems, such as interoperability issues, data privacy concerns, or resource constraints. Addressing these challenges is crucial for the successful implementation of LLM-based interventions. Despite these limitations, this study represents a foundational step toward understanding the utility and potential applications of LLMs in medication review and reconciliation processes. Future research efforts should aim to address these limitations and further explore the role of LLMs in optimizing medication management and patient outcomes in clinical practice. Lastly, the assessors were not blinded to the names of the LLMs. In our opinion, given the absence of any benchmarked LLMs, any potential bias is minimal.

5 Conclusion

Indeed, advanced LLMs hold significant promise in revolutionizing the medication review and reconciliation process in healthcare. Diverse impacts on patient safety were observed. Minor impacts were noted in identifying appropriate dosing regimens and therapeutic drug monitoring, while major impacts were found in identifying drug interactions and making pharmacogenomic-based therapeutic decisions. Integrating and validating LLMs within electronic health records and prescription systems is essential to harness their full potential and enhance patient safety and care quality.

CRediT authorship contribution statement

Kannan Sridharan: Writing – review & editing, Writing – original draft, Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Gowri Sivaramakrishnan: Writing – review & editing, Writing – original draft, Investigation, Formal analysis, Data curation.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A Supplementary data

Electronic Supplementary material 1

Image 1

Electronic Supplementary material 2

Image 2

Electronic Supplementary material 3

Image 3

Electronic Supplementary material 4

Image 4

Acknowledgements

We acknowledge the use of ChatGPT for improving the grammar in this manuscript.

Appendix A Supplementary data to this article can be found online at https://doi.org/10.1016/j.rcsop.2024.100492.
==== Refs
References

1. Medication Without Harm World health organization Available at: https://www.who.int/initiatives/medication-without-harm#:~:text=Unsafe%20medication%20practices%20and%20medication,of%20the%20medication%20use%20process (Accessed on 12th May, 2024)
2. Institute of Medicine Preventing medication errors 2006 National Academies Press Washington, DC
3 Patel E. Pevnick J.M. Kennelty K.A. Pharmacists and medication reconciliation: a review of recent literature Integr Pharm Res Pract 8 2019 39 45 31119096
4 Shehab N. Lovegrove M.C. Geller A.I. Rose K.O. Weidle N.J. Budnitz D.S. US emergency department visits for outpatient adverse drug events, 2013-2014 JAMA 316 2016 2115 2125 27893129
5 Redmond P. Grimes T.C. McDonnell R. Boland F. Hughes C. Fahey T. Impact of medication reconciliation for improving transitions of care Cochrane Database Syst Rev 8 2018 CD010791
6 Griese-Mammen N. Schulz M. Böni F. Hersberger K.E. Medication review and medication reconciliation Alves da Costa F. van Mil J. Alvarez-Risco A. The pharmacist guide to implementing pharmaceutical care 2019 Springer Cham 10.1007/978-3-319-92576-9_7
7. Medication errors The academy of managed care pharmacy Available at: https://www.amcp.org/about/managed-care-pharmacy-101/concepts-managed-care-pharmacy/medication-errors (Accessed on 12th May 2024)
8. Akhoon N. Precision medicine: a new paradigm in therapeutics Int J Prev Med 12 2021 12 34084309
9 Park Y.J. Pillai A. Deng J. Assessing the research landscape and clinical utility of large language models: a scoping review BMC Med Inform Decis Mak 24 2024 72 38475802
10. Reddy S. Evaluating large language models for use in healthcare: a framework for translational value assessment Inform Med Unlocked 41 2023 101304
11 Spotnitz M. Idnay B. Gordon E.R. A survey of Clinicians’ views of the utility of large language models Appl Clin Inform 15 2024 306 312 38442909
12 Raza M.M. Venkatesh K.P. Kvedar J.C. Generative AI and large language models in health care: pathways to implementation NPJ Digit Med 7 2024 62 38454007
13. Liu Z. Wu Z. Hu M. Pharmacy GPT: the AI pharmacist Available at: https://arxiv.org/pdf/2307.10432
14 Ray P.P. ChatGPT: a comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope Int Things Cyber-Phys Sys 3 2023 121 154
15. Google Gemini Available at: https://gemini.google.com/app/ (Accessed on 5th May 2024)
16. O'Leary M. Claude and the pursuit of AI safety Available at: https://www.infotoday.com/it/apr24/OLeary--Claude-and-the-Pursuit-of-AI-Safety.shtml (Accessed on 5th May 2024)
17. Heller M. What is llama 2? meta's large language model explained Available at: https://www.infoworld.com/article/3706470/what-is-llama-2-metas-large-language-model-explained.html (Accessed on 5th May 2024)
18. FDA online label repository Available at: https://labels.fda.gov/ (Accessed on 5th May 2024)
19. Clinical Pharmacogenetic Implementation Consortium Available at: https://cpicpgx.org/guidelines/ (Accessed on 5th May 2024)
20. Testing for G6PD deficiency for safe use of primaquine in radical cure of P. vivax and P. ovale malaria. Policy brief 2024 World Health Organization Available at: https://iris.who.int/bitstream/handle/10665/250297/WHO-HTM-GMP-2016 9-eng.pdf?sequence=1 (Accessed on 5th May 2024)
21 Joint Formulary Committee British national formulary London: British medical association and royal pharmaceutical society 86th ed. 2023 BMJ London
22. Medicines & Healthcare products Regulatory Agency Available at: https://www.gov.uk/government/organisations/medicines-and-healthcare-products-regulatory-agency (Accessed on 5th May 2024)
23 Murugan M. Yuan B. Venner E. Empowering personalized pharmacogenomics with generative AI solutions J Am Med Inform Assoc 31 2024 1356 1366 38447590
24 Jošt M. Kerec Kos M. Kos M. Knez L. Effectiveness of pharmacist-led medication reconciliation on medication errors at hospital discharge and healthcare utilization in the next 30 days: a pragmatic clinical trial Front Pharmacol 15 2024 1377781
25 Griva K. Chua Z.Y. Lai L.Y. Xu S.J. Bek E.S.J. Lee E.S. Pharmacist-led medication reconciliation service for patients after discharge from tertiary hospitals to primary care in Singapore: a qualitative study BMC Health Serv Res 24 2024 357 38509565
26 Pais C. Liu J. Voigt R. Gupta V. Wade E. Bayati M. Large language models for preventing medication direction errors in online pharmacies Nat Med 2024 10.1038/s41591-024-02933-8
27 Alrabiah Z. Alwhaibi A. Alsanea S. Alanazi F.K. Abou-Auda H.S. A National Survey of attitudes and practices of physicians relating to therapeutic drug monitoring and clinical pharmacokinetic service: strategies for enhancing Patient’s care in Saudi Arabia Int J Gen Med 14 2021 1513 1524 33935513
28 Meng X. Yan X. Zhang K. The application of large language models in medicine: a scoping review iScience 27 2024 109713
