ABSTRACT Introduction With the increasing prevalence of inflammatory bowel disease (IBD) and its complex management, evaluating innovative approaches became essential. Large language models (LLM), such as ChatGPT‐4.0, have remarkably advanced, yet their application in IBD management remains unexplored. This study compares therapeutic recommendations generated by ChatGPT‐4.0 with those made by IBD specialists, evaluating its potential and limitations in clinical practice. Materials and Methods This prospective observational study, conducted between September 2024 and March 2025, included 47 cases of IBD patients requiring initiation or modification of biologic therapy. Standardized clinical vignettes were presented to both IBD specialists and ChatGPT‐4.0. For each case, three sequential questions were addressed: the choice of the therapeutic class of the treatment, the specific molecule in this class, and if the proposition is a monotherapy or combination therapy. Agreement was measured using the percentage concordance and Cohen's kappa coefficient ( κ ). Results The agreement between the IBD specialists and ChatGPT‐4.0 was 76.6% ( κ = 0.677) for the therapeutic class and 44.7% ( κ = 0.352) for specific molecule selection. Therapeutic class discordance occurred in 11 cases (23.4%); all were reviewed and deemed acceptable by the gastroenterologists. Notably, seven of ChatGPT‐4.0's second‐line suggestions matched with the IBD specialists' treatment decisions. Conclusion This study highlights a promising area of ChatGPT‐4.0 application in the therapeutic management of IBD. However, it is not suitable as a standalone or fully reliable resource for decision‐making in healthcare. Future development of tailored healthcare GPT integrating European guidelines and experts' feedback may further improve LLM performance.
Loncour et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: