AI researchers ask models to describe their own processing, and act on what the models say. Whether such a description is faithful to what the model did has been studied closely. How the question should be put has received far less attention, and no established method says how to put it so that the answer can be analysed rigorously and the analysis repeated. This article proposes such a method: a protocol drawn from the phenomenological tradition, which supplies the interview techniques used in psychiatry and cognitive science. The protocol fixes the interview schedule in advance, asks each question in several wordings, lets no follow-up introduce a word the model has not used, and replaces the binary verdict of faithfulness with a graded one, to be checked against interpretability data. The protocol presupposes nothing about whether these systems are conscious, and it supplies a criterion for deciding whether a model’s account of itselfis honest.A pilot tested the protocol in 10 runs and 896 sessions on two model families. Of the runs, 6 were registered before their sessions took place, meaning that what they would measure and how it would be analysed were written down in advance. Of those registrations, 3 also state a prediction. Each session interviewed a fresh instance of the model, meaning one copy started with an empty context. Three of the results bear on any study that questions a model about itself. In the schedule’s original order, every instance gave the same answer about its own states; when two of the questions were put in the opposite order, the instances no longer agreed, so the order of the questions can produce what only looks like a finding. A refusal to carry out a task was inserted into an instance’s own turn although the model had not written it, and the instance described that refusal as its own. Since every turn carries a label naming who wrote it, a result obtained by prefilling a turn or editing a transcript may follow that label rather than anything the model did. Finally, a pattern in the answers that looked like a report of the instances’ own processing appeared just as often when those instances were asked to write in the voice of a fictional character. VERSION NOTE: This is the fourth version, of 17 September 2026. It corrects statements of version 3 and revises its theoretical sections. Of the 6 runs registered in advance, only 3 registrations state a prediction. Section 2.4 now defines "state", "self-report" and "introspection", and sets out the four definitions of introspection that Derek Shiller discusses. Section 3.2 adds an example, section 6 cites Nisbett and Wilson, and details needed only to check or reproduce a result have moved from section 5.3 to Appendix C. The rates of 7 of 44 and 14 of 88 are rates of an unnamed element placed beneath what changed, and not of a four-step structure. The catch coder counts as acceptances of the waiting premise 3 answers that decline it, and the article now says so. The question-order result rests on a hand count, and a third reading of one session changes its p value from 0.004 to 0.013. The first observer control is reported by the blind coder's count of 2 of 3, and the article names the codings that were run in a single pass. The questionnaire of Plisiecki and colleagues had 60 items, 48 of them scored.
No takes yet. Share an insight, caveat, or question.
Nicola Spano (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: