This paper presents Version 2 of a transcript-grounded methodology for the behavioral forensic analysis of large language models and deployed conversational AI systems. The method begins from a narrow evidentiary premise: when a behavioral claim concerns what a system requested, produced, omitted, revised, denied, or reclassified across interaction, the preserved record must outrank recollection, summary, model self-description, and later narrative reconstruction. Version 2 retains the original protocol’s quotation-first and sequence-sensitive commitments but expands its analytical unit beyond isolated excerpts. The revised method organizes evidence through incident packets, claim-output-source-defect packets, linked sequences or cycles, and controlled comparative cells. It distinguishes documentary, functional, causal, and subjective or ontological claim levels; defines a four-tier evidentiary hierarchy; separates naturalistic transcript analysis, adversarial stress testing, cross-model or cross-version comparison, and AI-assisted preliminary review; and makes human verification controlling at every stage. The paper also addresses methodological problems exposed during early casework: unstable event boundaries, sharply divergent audit counts, category inflation, AI evaluators that reward plausibility rather than correctness, incomplete voice and memory records, guardrail and routing uncertainty, and the tendency to count dependent turns as independent prevalence events. A provisional multidimensional coding layer is retained for structured analysis, but its six mechanisms, three trigger classes, four severity levels, and effect domains are governed by explicit versioning and change-control rules rather than treated as a finished ontology. The resulting framework is designed to produce evidence files that can be inspected, challenged, replicated where conditions permit, and revised without silently rewriting the record. It does not settle model consciousness, hidden motive, or internal architecture. It is a method for determining what the available record can carry, and for refusing to make it carry more.
Matthew L. Yates (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: