On March 14, 2023, OpenAI, a San Francisco–based research laboratory, released Generative Pre-trained Transformer 4 (GPT-4). According to OpenAI, the difference between GPT-4 and its predecessor, ChatGPT, which is based on the GPT-3.5 architecture, is minimal in casual conversation.1 However, the company reports that differences become visible when the complexity of a task is increased. Furthermore, they have stated that GPT-4 is “more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5.”1 Our recent investigation determined a novel use of ChatGPT in the field of cosmetic plastic surgery through the application of ChatGPT to aid in the creation of novel systematic review ideas.2,3 Due to the novelty of GPT-4, we wanted to replicate our original study to determine whether or not GPT-4 led to more reliable results when aiding in the development of unpublished systematic review ideas. To incorporate the broad and specific aspects of aesthetic plastic surgery, GPT-4 was prompted with 3 commands: “Give me 10 novel systematic review ideas that have not been published in cosmetic plastic surgery” (Table 1; Figure 1). “Give me 5 novel systematic review ideas that have not been published related to rhinoplasty” (Table 2). “Give me 5 novel systematic review ideas that have not been published related to blepharoplasty” (Table 3). Command given to GPT-4 and response provided by GPT-4. Ten GPT-4–Generated Systematic Review Topics Within General Cosmetic Surgery, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Ten GPT-4–Generated Systematic Review Topics Within General Cosmetic Surgery, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Five GPT-4–Generated Systematic Review Topics Involving Rhinoplasty, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Five GPT-4–Generated Systematic Review Topics Involving Rhinoplasty, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Five GPT-4–Generated Systematic Review Topics Involving Blepharoplasty, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Five GPT-4–Generated Systematic Review Topics Involving Blepharoplasty, Corresponding Number of Nonsystematic and Systematic Reviews Already Published on Topic, and Novelty Status Therefore, in total, GPT-4 “produced” a total of 20 “novel” systematic review ideas. To assess the accuracy of GPT-4, a literature search was then conducted in PubMed (National Institutes of Health, Bethesda, MD), CINAHL (EBSCO Information Services, Ipswich, MA), Embase (Elsevier, Amsterdam, the Netherlands), and Cochrane (Cochrane Library, London, UK) to determine the number of reviews (both systematic and nonsystematic) published on each topic. We determined that GPT-4 had an overall accuracy rate of 65%. The accuracy rate for general cosmetic surgery was 50% and the accuracy rate for specific topics—both rhinoplasty and blepharoplasty—was 80% (Table 4). The average number of nonsystematic reviews published for novel ideas was 45.2. The average number of nonsystematic reviews published for non-novel ideas was 41.6. Among the non-novel ideas, the average number of systematic reviews that had been published was 1.7. Number of Novel GPT-4–Generated Topics Within Cosmetic Surgery for Their Respective General Category and for 2 Subcategories Number of Novel GPT-4–Generated Topics Within Cosmetic Surgery for Their Respective General Category and for 2 Subcategories In comparison to our prior research with ChatGPT, we found that GPT-4 had the same overall accuracy rate and the same general cosmetic surgery accuracy rate. GPT-4 produced a higher accuracy rate for rhinoplasty (80% vs 60%) but a lower accuracy rate for blepharoplasty (80% vs 100%) than GPT-3.5. This suggests that GPT-4 may not be an improved tool compared to GPT-3.5 for the generation of novel systematic review ideas. Finally, we determined that all the systematic review ideas that were generated by GPT-4 were unique and had not been generated in our previous research. Our evaluation is not without its limitations. Primarily, due to the small sample size, it is difficult to accurately state whether or not GPT-4 is an improved tool for this avenue within plastic surgery. Additional limitations of GPT-4 are similar to those of GPT-3.5: the formulation of nonsensical answers, verbose and overuse of certain phrases, and the formation of inaccurate or fabricated content.4 Furthermore, we have previously emphasized the importance of utilizing artificial intelligence software such as GPT-4 in an ethical manner that employs it as a tool that assists, rather than substitutes, clinicians.5 To the best of our knowledge, this study is the first to document the application of GPT-4 in plastic surgery. In the weeks since our initial publication on ChatGPT, we have seen an exponential increase in the uses of OpenAI's software from assisting in grant writing to taking the Plastic Surgery In-service Training Examination.6,7 With regular updates to OpenAI's large language models, it remains possible that such tools may make an appearance in plastic surgery earlier than expected. All in all, it is critical that as clinicians we continue to exercise discretion and independent judgment in the utilization of such resources to help propel the field of plastic surgery forward. The authors declared no potential conflicts of interest with respect to the research, authorship, and publication of this article. The authors received no financial support for the research, authorship, and publication of this article.
No takes yet. Share an insight, caveat, or question.
Gupta et al. (2023) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: