Large language models are now used to generate the items of psychological and organisational assessment instruments. The emerging literature on automatic item generation evaluates such items psychometrically, under laboratory conditions, on constructs chosen by the researcher. It does not tell us what happens when an LLM writes the questionnaire that a real company will actually send to its employees. We report a field audit of exactly that. Across two production pilots of an AI-native 360-degree feedback platform, we audited the items the system generated and the behaviour of the surrounding pipeline. At one site, the system generated four instruments (manager, peer, direct-report and self perspectives) from the organisation's own competency framework. We audited every generated item against established item-writing criteria. 37 of 48 items (77%) carried at least one item-writing defect. We organise the defects into eight classes: relevance failure (the item presumes a role or a customer-facing context the respondent does not have), double-barrelled items, content-validity imbalance, social-desirability exposure, leading items, format mismatch (an open-ended stem attached to a rating response), redundancy, and binary constructs forced onto a Likert scale. Relevance failure (15 items) and double-barrelling (12) were the most common. Beyond item writing, the other site surfaced a distinct and, we argue, more serious class of failure. A real-time assistant intended to help assessors write better feedback systematically steered them to rewrite legitimate negative feedback in positive terms. A multi-source feedback instrument whose assistant suppresses negative valence does not merely produce noisy data. It produces data that is wrong in a specific and undetectable direction, while appearing to work. We report additional pipeline-level failures (locale propagation across a shared session, a single model serving every pipeline stage, and central-tendency bias induced by an odd-numbered scale) and we set out mitigation directions at the level of design principle. We do not report our implementation. We are explicit about what this study is not. It is a retrospective single-rater audit of two pilots of one product. There is no inter-rater reliability statistic, no blinding, no control condition, and no comparison against human-written items from the same organisations. It establishes that these defects occur and are common. It does not establish their rate in general, nor that LLM-written items are worse than human-written ones. We set out the study that would.
Alessio Biancheri (Sun,) studied this question.