
==== Front
Plast Reconstr Surg Glob Open
Plast Reconstr Surg Glob Open
GOX
Plastic and Reconstructive Surgery Global Open
2169-7574
Lippincott Williams & Wilkins Hagerstown, MD

GOX-D-24-00262
00013
10.1097/GOX.0000000000006136
3
Technology
Special Topic
ChatGPT-4 Surpasses Residents: A Study of Artificial Intelligence Competency in Plastic Surgery In-service Examinations and Its Advancements from ChatGPT-3.5
Hubany Shannon S. BS *†
Scala Fernanda D. MD †
Hashemi Kiana BS *†
Kapoor Saumya BS *†
Fedorova Julia R. BS *†
Vaccaro Matthew J. BS *†
Ridout Rees P. BS *†
Hedman Casey C. BS *†
Kellogg Brian C. MD †
Leto Barone Angelo A. MD †
From the * University of Central Florida College of Medicine, Orlando, Fla.
† Division of Craniofacial and Pediatric Plastic Surgery, Nemours Children’s Hospital, Orlando, Fla.
Angelo A. Leto Barone, MD, Division of Craniofacial and Pediatric Plastic Surgery, Nemours Children’s Hospital, 6535 Nemours Pkwy, Suite 6739, Orlando, FL 32832, E-mail: angelo.letobarone@nemours.org, Instagram: angelo_leto_baronemd
9 2024
05 9 2024
12 9 e61368 3 2024
9 7 2024
Copyright © 2024 The Authors. Published by Wolters Kluwer Health, Inc. on behalf of The American Society of Plastic Surgeons.
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open-access article distributed under the terms of the Creative Commons Attribution-Non Commercial-No Derivatives License 4.0 (CCBY-NC-ND), where it is permissible to download and share the work provided it is properly cited. The work cannot be changed in any way or used commercially without permission from the journal.

Background:

ChatGPT, launched in 2022 and updated to Generative Pre-trained Transformer 4 (GPT-4) in 2023, is a large language model trained on extensive data, including medical information. This study compares ChatGPT’s performance on Plastic Surgery In-Service Examinations with medical residents nationally as well as its earlier version, ChatGPT-3.5.

Methods:

This study reviewed 1500 questions from the Plastic Surgery In-service Examinations from 2018 to 2023. After excluding image-based, unscored, and inconclusive questions, 1292 were analyzed. The question stem and each multiple-choice answer was inputted verbatim into ChatGPT-4.

Results:

ChatGPT-4 correctly answered 961 (74.4%) of the included questions. Best performance by section was in core surgical principles (79.1% correct) and lowest in craniomaxillofacial (69.1%). ChatGPT-4 ranked between the 61st and 97th percentiles compared with all residents. Comparatively, ChatGPT-4 significantly outperformed ChatGPT-3.5 in 2018–2022 examinations (P < 0.001). Although ChatGPT-3.5 averaged 55.5% correctness, ChatGPT-4 averaged 74%, a mean difference of 18.54%. In 2021, ChatGPT-3.5 ranked in the 23rd percentile of all residents, whereas ChatGPT-4 ranked in the 97th percentile. ChatGPT-4 outperformed 80.7% of residents on average and scored above the 97th percentile among first-year residents. Its performance was comparable with sixth-year integrated residents, ranking in the 55.7th percentile, on average. These results show significant improvements in ChatGPT-4’s application of medical knowledge within six months of ChatGPT-3.5’s release.

Conclusion:

This study reveals ChatGPT-4’s rapid developments, advancing from a first-year medical resident’s level to surpassing independent residents and matching a sixth-year resident’s proficiency.

OPEN-ACCESSTRUE
COUNTRYUNITED STATES
==== Body
pmcTakeaways

Question: To evaluate the advancements of ChatGPT by assessing the performance of its latest version, Chat GPT-4, on Plastic Surgery In-Service Examinations, and comparing these results to those of medical residents and its predecessor, ChatGPT-3.5.

Findings: ChatGPT-4 was tested on Plastic Surgery In-Service Examinations from 2018 to 2023. It answered 18.54% more questions correctly than ChatGPT-3.5 on average and increased from the 23rd to 97th percentile of all residents on the 2021 examination.

Meaning: ChatGPT has shown remarkable improvements and surpassed the performance of many medical residents.

INTRODUCTION

The medical field has seen a rapid increase in the application of artificial intelligence (AI) and machine learning. ChatGPT was developed by the company OpenAI and launched publicly in November 2022, gaining quick relevance as a large language model pretrained on a vast amount of data, including medical information. The generative pretraining transformer’s (GPT) most sophisticated version, GPT-4, was made available in March 2023. OpenAI has tested several benchmarks, including simulating examinations that were originally designed for humans with impressive improvement when comparing ChatGPT-3.5 to GPT-4.1 Among the multiple simulations in the medical realm, many were performed yearly by residents in training for many different specialties.

Given its training in a vast amount of data and its ability to simulate a casual human-like conversation, ChatGPT has been thought-provoking for its use in education worldwide for breaking down answers in a very convincing manner. Preliminary studies have been initiated in several medical specialties to evaluate how such AI would perform in different examinations regularly taken by physicians in training.2–7 Nonetheless, it is still worrisome that several of the simulations have not achieved a satisfactory percentile, being deemed as a nonpassing score in self-assessment tests that would correlate to the examination of specialty boards of orthopedic surgery8 and hand surgery3 or would be comparable to the performance only to the 49th percentile of a first-year trainee in plastic and reconstructive surgery on the in-service examinations.6

It was shown that GPT-3 was frequently fabricating information and producing biased text from previously learned subjects or simply disobeying user instructions,9 which could have contributed to the incorrect choices in tests and illogical reasoning for problem-solving. Although GPT-4 generally lacks knowledge of events that have occurred after the vast majority of its data cut off (September 2021), it has proved to be more reliable and prepared to handle nuanced instructions after reinforcement learning from human preferences posttraining.1,9

The purpose of the study is to evaluate and compare the performance of the two latest ChatGPT versions in the In-Service Examination for Residents, provided by the American Society of Plastic Surgeons yearly to trainees. Here, we will compare scores obtained in the years 2018 to 2022 using ChatGPT-3.5 and 4.0, being the longest study conducted in this field to date. By incorporating multiple questions, with varying degrees of difficulty and different authorship, we believe that there is a decrease in bias in how questions are formulated and focus on strict data application by ChatGPT.

With the proliferation of internet-based learning modalities, there is an imperative need to assess the readiness of ChatGPT for utilization as a learning resource, particularly in the context of surgical specialties. Mastery in surgery extends beyond merely excelling in board examinations; it encompasses critical logical thinking and the ability to tailor appropriate treatments for specific patients, attributes that are challenging to evaluate through GPT prompts. Jain et al conducted a comprehensive analysis of ChatGPT’s ability not only to provide accurate responses but also to maintain consistent logical reasoning in its thought process.2 It is noteworthy to mention that, even when the answers were correct, the underlying logic sometimes demonstrated inconsistencies. Similarly, some incorrect responses lacked any logical explanation.2

The Plastic Surgery In-service Examination (PSISE), conducted annually for all plastic surgery residents in the United States, serves as a critical tool to gauge their progress in medical training and assess their competencies as physicians. This examination is designed not merely to evaluate medical knowledge but also to test the integration of complex decision-making skills in specific contextual scenarios. In this study, the authors evaluate the performance of ChatGPT-4 on these same examinations, undertaking a comparative analysis of its capabilities relative to the national performance of residents, as well as its predecessor, ChatGPT-3.5.

METHODS

Questions from the PSISEs from 2018 to 2023 were included in this study. This examination is composed of 250 questions, with five sections of 50 questions each. These section titles include comprehensive, hand and lower extremity, craniomaxillofacial, breast and aesthetic, and core surgical principles. All questions in these examinations are multiple choice with between three and five answer choices. The PSISEs can be accessed through a password-protected site and are therefore not publicly available. However, the datasets generated during and/or analyzed during this study are available from the corresponding author upon reasonable request.

Questions including images met the exclusion criteria for this study. Although ChatGPT-4 is now able to accept and analyze images, this study was focused on the chabot’s ability to answer multiple-choice questions only to better compare its performance with that of ChatGPT-3.5. Questions not originally scored on the examination were excluded to accurately compare ChatGPT-4’s result with residents’ scores. Additionally, any inconclusive answers were not included in the final analysis. Inconclusive answers refer to any response that did not provide a single answer choice option. This includes providing multiple answer choices, no answer choices, or errors (Fig. 1).

Fig. 1. This flowchart delineates the exclusion criteria applied to the questions considered in our study’s analysis.

Each examination was assessed in a separate chat session, and the answer to each was not provided to ChatGPT. The data were collected between December 9, 2022, and January 17, 2023. The question stems and each multiple-choice answer were manually copied from each examination year’s syllabus and pasted one-by-one verbatim into ChatGPT-4. The formatting of each question was not changed because the AI model was able to interpret the questions and choices correctly. ChatGPT-4 provided an answer choice with an explanation considering each possible choice. Each incorrect answer was recorded as a 0, and each correct answer was recorded as a 1 in a Microsoft Excel datasheet. Each question including an image was recorded as a 2. Each question not originally scored was recorded as a 3. Each question where a definitive answer was not given was recorded as a 4.

The ratio of correct answers over the total number of questions included was calculated using Excel. The norm table for each examination year was used to compare the performance of ChatGPT-4 on the PSISE to that of residents nationally from both independent and integrated programs. Statistical significance was defined as a P value less than 0.05. A paired t test was performed with R statistical software to compare the percentage of correct answers between ChatGPT-3.5 and ChatGPT-4.

We compared ChatGPT-4’s performance with that of ChatGPT-3.5’s, reported in a prior study conducted by Gupta et al.6 We did not administer each examination with ChatGPT-3.5 again to compare its advancements over time. Reproducing these results would introduce confounding variables stemming from its ongoing learning capabilities. This approach recognizes the dynamic nature of ChatGPT’s design, which is engineered to enhance its capabilities through ongoing interactions. Consequently, it is acknowledged that replicating this study at a later date would likely yield results influenced by ChatGPT’s progressive learning abilities. We followed similar methods and inclusion criteria as those of this study to more accurately compare the two.

RESULTS

A total of 1500 questions from the six PSISEs conducted in the years 2018 to 2023 were reviewed. A total of 1292 questions were included in the final analysis after excluding 164 containing images, 25 not originally scored, and 19 considered inconclusive. ChatGPT provided detailed responses to each question, illustrating the rationale behind its selected answer choice, as well as refuting the others. In total, ChatGPT-4 answered 961 (74.4%) of those questions correctly (Fig. 2).

Fig. 2. The percentage of correct answers of ChatGPT-4’s responses for examination years 2018–2023, as well as the average percentage correct. Additionally, this table shows the percentage of correct answers in each of the five examination sections, as well as the average percentage correct.

ChatGPT-4 versus Residents

ChatGPT-4’s performance on each year’s PSISE was slightly varied, correctly answering between 69.7% and 78.4% of the questions included in the 2022 and 2021 examinations, respectively. On average, ChatGPT-4 performed best on the core surgical principles section, answering 79.1% of questions correctly, and worst on the craniomaxillofacial section, answering 69.1% of questions correctly (Fig. 2). Compared with the performance of residents for each of these years, ChatGPT-4 would rank between the 61st and 97th percentile of all residents, respectively. ChatGPT performed highest on the 2021 PSISE. ChatGPT-4’s performance on this examination would rank in the 100th percentile of first- and second-year independent residents, and the 99th percentile of third-year independent residents. When compared with residents in the integrated program, ChatGPT-4 would rank in the 100th percentile of first- and second-year residents, 98th percentile of third-year residents, 97th percentile of fourth-year residents, and 93rd percentile of fifth- and sixth-year residents (Fig. 3).

Fig. 3. The percentile rankings of ChatGPT-4’s performance across examinations from 2018 to 2023, utilizing each year’s normal distribution for comparison against the performance of plastic surgery residents, alongside aggregated averages.

ChatGPT-4 versus ChatGPT-3.5

ChatGPT-3.5 was not previously tested on the 2023 PSISE. Therefore, when comparing ChatGPT-4 with ChatGPT-3.5, the performance of ChatGPT-4 on the 2023 examination was not included in the calculations. Overall, ChatGPT-4 scored significantly higher than ChatGPT-3.5 on the PSISEs from the years 2018 to 2022 (P < 0.001). On average, ChatGPT-3.5 answered 55.5% of questions correctly, whereas ChatGPT-4 answered 74% of questions correctly, with a mean difference of 18.54% (Fig. 4). ChatGPT-4 outperformed ChatGPT-3.5 in every examination, achieving a higher percentage of correct answers (Fig. 5).6 ChatGPT-3.5 as well as ChatGPT-4 both performed highest on the 2021 PSISE. ChatGPT-3.5 answered 60.1% of questions correctly, which would rank in the 23rd percentile of all residents, 49th percentile of first-year residents, and seventh percentile of third-year residents of the independent program.6 ChatGPT-4 answered 78.4% of questions correctly, which would rank in the 97th percentile of all residents, the 100th percentile of first-year residents, and the 99th percentile of third-year residents (Fig. 6).

Fig. 4. Comparison of the percentage of correct answers between ChatGPT-3.56 and ChatGPT-4 for examination years 2018–2022, as well as comparing the overall average percentage correct. Subsequent statistical analysis through a paired t test revealed a P value of less than 0.001 and a mean difference of 18.54.

Fig. 5. Illustration of the difference in the percentage of correct answers between ChatGPT-3.56 and ChatGPT-4 for examination years 2018–2022 as well as a comparison of the overall average.

Fig. 6. A comparative analysis of percentile ranks for ChatGPT-3.56 and ChatGPT-4 on the 2021 PSISE when compared with each year of plastic surgery residents. Both versions achieved the highest number of correct responses on this examination.

DISCUSSION

This study aimed to correlate ChatGPT-4’s performance on the PSISEs with that of residents nationally, as well as with that of ChatGPT-3.5.

ChatGPT-4 was able to provide an answer to the vast majority of questions included in the study, with only 19 out of the 1292 questions (1.5%) being considered inconclusive. ChatGPT-4 outperformed most residents nationally, on average, scoring higher than 80.7% of all residents nationally. When compared with all first-year residents, ChatGPT-4 scored above the 97th percentile on average. When looking at residents in later years of the integrated program, ChatGPT-4’s performance was more comparable to that of the residents, scoring in the 55.7th percentile of sixth-year residents, on average (Fig. 3).

These results suggest that ChatGPT-4 possesses a substantial amount of foundational medical information, coupled with a notable enhancement in its capacity for application of this in more complex medical scenarios. Remarkably, this advancement has been achieved within a mere 6-month span for this chabot, bridging the launch of ChatGPT-3.5 and GPT-4.1 This poses interesting questions about how AI will be used in the future of medical education. It is essential to acknowledge that despite the impressive performance of ChatGPT-4, there remains the potential for it to deliver inaccurate information. Consequently, it is crucial to independently verify the responses provided by ChatGPT for accuracy and reliability before making any decisions in a clinical setting.

Interestingly, ChatGPT-4’s performance was more varied across years than ChatGPT-3.5.6 This may be due to the constant adaptations that ChatGPT is programmed to make. For the completion of data evaluation in a timely manner, the authors divided assessments among team members, resulting in a staggered completion of assessments, with certain examinations being finalized ahead of others. The chronological sequence of completion for these examinations is as follows: 2021, 2023, 2020, 2018, 2019, and 2022. Remarkably, this sequence also aligns with a descending order of performance (Fig. 2). This correlation, whether coincidental or indicative of an underlying factor influencing ChatGPT’s responses, warrants additional investigation into the adaptability of ChatGPT.

ChatGPT-4 introduces the novel feature of image analysis, expanding its range of capabilities. However, to facilitate a direct comparison with its predecessor, ChatGPT-3.5, our study did not incorporate this feature. The introduction of image analysis by ChatGPT-4 paves the way for a multitude of new research opportunities, particularly in domains like radiology, where AI tools have already demonstrated significant advancements due to the objective nature of image interpretation.10 ChatGPT-4’s expansion into this area suggests considerable potential for application and further exploration of its image analysis capabilities. Possibilities of future implications of image analysis include, but are not limited to, the suggestion of different surgical techniques, detailed procedural descriptions, and enhanced prediction of patient outcomes.11

ChatGPT may serve as a versatile tool in the realm of education. For students preparing for examinations, ChatGPT may act as an effective study aid, breaking down complex topics into more understandable content and providing insights into specific questions that would otherwise require timely research.12 When answering a question, it goes beyond mere information delivery by providing detailed explanations and feedback. Although the benefits of ChatGPT in education can be substantial, considering effective logical reasoning for the explanations, it is still essential to use this tool merely as a supplement to traditional educational methods and resources. ChatGPT is a machine learning software that can only process information in the form of pattern recognition and algorithmic interpretation.13 Although its performance on these assessments is impressive, its outputs should not be interpreted as objective information. Its responses should always be critically reviewed for accuracy and relevance.

This study represents the most extensive evaluation to date of ChatGPT-4’s application in the PSISE, yet it is not without its limitations. A key constraint was the exclusion of images, which are typically crucial in clinical diagnosis and management. This decision, aimed at facilitating a more accurate comparison with ChatGPT-3.5, may introduce some variability when assessing ChatGPT’s performance relative to that of medical residents. Furthermore, we did not scrutinize the logic underpinning its answer selections. This leaves room for future studies to delve into the reasoning processes of ChatGPT. Additionally, 19 questions were deemed inconclusive due to ChatGPT-4’s failure to select a definitive answer. Although these could be considered errors, they accounted for less than 2% of the total questions, thus exerting a negligible impact on the overall findings.

It is important to note that ChatGPT is an evolving tool, continually refining its capabilities through usage. Consequently, a repeated analysis using the same methodology might yield slightly different results. To mitigate this variability, we implemented several controls. Despite ChatGPT-4’s usage limits to entering approximately 30 questions a day, we tried to keep the data collection period as short as possible. Additionally, we did not reveal correct answers to ChatGPT, and each examination was conducted in a separate chat session. This approach was designed to minimize the AI’s learning from previous examination questions, ensuring a more unbiased evaluation. Previous research followed similar methodologies6; however, this suggests that ChatGPT may have encountered these questions in the past. Although these studies did not directly supply ChatGPT with responses, it remains unclear whether its users have done so at any time. Consequently, assessing the extent to which this has influenced the outcomes of our study presents a challenge.

CONCLUSIONS

The findings of this study highlight the significant progress ChatGPT-4 has achieved in a relatively brief period following the release of its predecessor, ChatGPT-3.5. Its performance has evolved from matching the proficiency of a first-year medical resident6 to surpassing the performance of all independent residents and equating to that of a sixth-year resident in the integrated program. This advancement suggests that ChatGPT may act as a valuable asset in the medical field, particularly in educational contexts. Nonetheless, there is a significant opportunity for further investigation into the effectiveness of ChatGPT-4 in the intricate realms of clinical diagnosis and treatment procedures within medicine.

DISCLOSURES

Dr. Leto Barone is the founder and chief medical officer of ReconstratA, LLC, and the Founder and President of Reconstruct Together, Corp. All the other authors have no financial interest to declare in relation to the content of this article. This study was supported by The University of Central Florida College of Medicine through funds designated for the publication costs.

ACKNOWLEDGMENT

The ML software ChatGPT-4, developed by the company OpenAI, was used to complete this study.

Published online 5 September 2024.

Disclosure statements are at the end of this article, following the correspondence information.
==== Refs
REFERENCES

1. Open AI. GPT-4. Available at https://openai.com/research/gpt-4. 2023. Accessed January 25, 2024.
2. Jain N Gottlich C Fisher J . Assessing ChatGPT’s orthopedic in-service training examination performance and applicability in the field. J Orthop Surg Res. 2024;19 :27.38167093
3. Han Y Choudhry HS Simon ME . ChatGPT’s performance on the hand surgery self-assessment exam: a critical analysis. J Hand Surg Glob Online. 2024;6 :200–205.38903839
4. Madrid-García A Rosales-Rosado Z Freites-Nuñez D . Harnessing ChatGPT and GPT-4 for evaluating the rheumatology questions of the Spanish access examination to specialized medical training. Sci Rep. 2023;13 :22129.38092821
5. Gupta R Herzog I Park JB . Performance of ChatGPT on the plastic surgery inservice training examination. Aesthet Surg J. 2023;43 :NP1078–NP1082.37128784
6. Humar P Asaad M Bengur FB . ChatGPT is equivalent to first-year plastic surgery residents: evaluation of ChatGPT on the plastic surgery in-service examination. Aesthet Surg J. 2023;43 :NP1085–NP1089.37140001
7. Hoch CC Wollenberg B Lüers JC . ChatGPT’s quiz skills in different otolaryngology subspecialties: an analysis of 2576 single-choice and multiple-choice board certification preparation questions. Eur Arch Otorhinolaryngol. 2023;280 :4271–4278.37285018
8. Lum ZC . Can artificial intelligence pass the American Board of Orthopaedic Surgery Examination? Orthopaedic residents versus ChatGPT. Clin Orthop Relat Res. 2023;481 :1623–1630.37220190
9. De Angelis L Baglivo F Arzilli G . ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health. Front Public Health. 2023;11 :1166120.37181697
10. Langlotz CP . The future of AI and informatics in radiology: 10 predictions. Radiology. 2023;309 :e231114.37874234
11. Kanevsky J Corban J Gaster R . Big data and machine learning in plastic surgery: a new frontier in surgical innovation. Plast Reconstr Surg. 2016;137 :890e–897e.
12. Lee H . The rise of ChatGPT: exploring its potential in medical education. Anat Sci Educ. 2024;17 :926–931.36916887
13. Kaul V Enslin S Gross SA . History of artificial intelligence in medicine. Gastrointest Endosc. 2020;92 :807–812.32565184
