Commentary highlights AI support's prevalence in academic writing, suggesting a need to revise ethical norms and publishing practices.
Commentary First and foremost, Callanan et al. need to be congratulated for their timely and important contribution to the literature. The authors answered questions that most might not even realize needed answering. Their data establishes the foundation for a framework for—or, at least, a conversation on—the use of AI support in academic writing that should challenge our perspective on the future of peer-reviewed literature, publishing practices, and ethical norms1. Prior to reading this study, one might not have expected that artificial intelligence (AI) detection software would overreport at a baseline rate of 11%. This finding underscores the authors' methodological wisdom in leveraging pre-AI-era comparative data to avoid overstating the prevalence of AI-based support and to provide a quantitative threshold of 33% (based on the mean AI detection percentage for pre-AI-era-manuscripts plus 2 standard deviations) for significant AI involvement in a manuscript. With this newly established threshold, 17% of publications were found to have leveraged considerable AI-based support. For argument's sake, ignore the limitations in the study. Accept that ZeroGPT accurately identified AI-based support no more often than the false positive rate of 11%. Accept that these observations are static and may not change overnight with a simple software update from a privately held company. The visceral reaction by readers upon learning that 17% of articles (38.3% in The Journal of Bone & Joint Surgery [JBJS] alone) exceeded the threshold of significant AI support is unlikely to be well met. Some may even call into question the morality of the authors of those articles. This reaction stems from the perceived circumvention of the deep study and heavy labor typically involved in the process of diving into background literature, communicating methodology clearly, and stringing together a cogent argument. This perspective is understandable albeit superficial, likely futile, and potentially hypocritical. The same people who traded in the Dewey Decimal System from the libraries they fled for the unprecedented immediacy of the Internet 20 years ago should withhold their outcry. The use of AI support for the purpose of scientific reporting feels wrong, but is it? It feels wrong because we were told by 4 Editors-in-Chief from Clinical Orthopaedics and Related Research, The Bone & Joint Journal, the Journal of Orthopaedic Research, and JBJS that misuse of large language models may "undermine the integrity of the scholarly record."2 We have bigger threats to the scholarly record (e.g., ghost authorship, inaccessible literature, undercompensated editorial services) than individuals who are honest enough to credit AI use. Moreover, the trite scoping and systematic reviews that litter the literature embody repackaged "legacy language models" that already water down the scholarly record—the key difference being that those are hardcoded and indoctrinated. The authorship guidelines from the International Committee of Medical Journal Editors are clear but have not matured to consider the value of AI assistance. This combined statement from these 4 journals unfortunately incentivizes authors to avoid disclosure altogether2. We say that we want global collaboration in science. We say that we want a diverse exchange of ideas. We know that communication is the foundational lynchpin to accomplish this. So why wouldn't we accept assistance from a tool that enables authors from across the globe to better communicate their imagination? Of all use cases for deploying large language models, scientific reporting is less concerning because methodologic communication in original research should be formulaic. A large language model does not provide ideation or inspiration; those exclusively come from the human imagination. AI is a tool, not a threat. A global scientific journal should welcome AI assistance in original research to level the playing field for researchers whose first language may not be English. However, we should have heightened wariness of misconduct with AI tools in review articles and editorials, where the scholarly record deserves increased depth and creativity, respectively. The context of AI assistance matters, which is why Callanan et al.'s finding of a 16.4% rate of AI involvement among original research concerns me far less than the 18.2% rate of AI involvement among review studies. As a final note on the subject of the inevitability of AI assistance, we need to recognize that society outwardly proclaims to value authenticity but that our actions repeatedly favor convenience over authenticity. If our indignation over AI support was sincere, we would not allow Netflix to nudge our Friday nights, permit Gmail (Google) to finish our missives, or engage with self-affirming content in our social media echo chambers. When it comes to clinical and scientific medicine, AI assistance should be welcomed as an adjunct—but not as an alternative—to human involvement3.
No takes yet. Share an insight, caveat, or question.
Prem N. Ramkumar (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: