Abstract While recent years have seen increasingly diverse modes of assessment in Higher Education English, the advent of generative artificial intelligence (AI) has prompted fresh concern that the essay remains the default practice, and one that – whether for administrative, workload, or pedagogical reasons – cannot suddenly be reimagined even though now heavily exposed to academic misconduct. In this paper, we try to unlock the ‘black box’ of generative AI by exploring how ChatGPT responds variously to different forms of wording commonly used in English Literature exam questions, taking as our case study a suite used in modules at a UK university. We know that, pre-AI, the precise wording of essay questions can significantly affect the learning outcomes, formal structures, and methods that students are expected to adopt in response. Drawing on computational and inductive analysis, we identify that slight changes in question framing can elicit small but identifiably different outputs from generative AI. This allows us to recommend how questions might be set in order either (potentially) to expose AI use where it has been employed, or (idealistically) to encourage students to engage in more independent writing and thinking to mitigate obvious AI limitations. It also informs a basic human-led criteria by which the most egregious AI use may, albeit unreliably and provisionally, be detected either at an individual or cohort level. Finally, it allows for a provisional outline of what an AI-assisted detection process might look like based jointly on the methodological framework of the study and the patterns identified.
Marshall et al. (Thu,) studied this question.