Large language model (LLM) systems are increasingly used to summarize clinical guidelines, draft patient-facing education, and answer evidence-linked clinical or public health questions. We define citation veneer as an observable audit state in which an output presents citation cues or apparently supportive source information while still containing incorrect, incomplete, or materially unsupported content. We develop and evaluate the Clinical Evidence Audit Grid (CEAG), an analyst-facing monitoring surface that encodes six binary response-quality signals into a fixed-length plain-text audit tag (CENHOV): clinical correctness (C), evidential support (E), numeric concordance (N), hedging (H), stylistic ornamentation (O), and a derived CEAG-V supported-error citation veneer marker (V=1 when E=1 and C=0). CEAG was evaluated on a controlled benchmark (synthetic, n=120 questions per condition) and on a 12-case public check built from PubHealth-style health-claim examples and public medical or public-health sources. The controlled benchmark used scripted response regimes rather than a live LLM/API benchmark; therefore, model identity, sampling temperature, and seed are not applicable to the main controlled analysis. In the synthetic benchmark, the RUSH regime retained high evidential-support rates (93.3%, 95% CI 87.4%-96.6%) while dropping to 73.3% correctness (95% CI 64.8%-80.4%) and producing 25.8% CEAG-V veneer (95% CI 18.8%-34.3%). In the 12-case public check, RUSH produced 66.7% CEAG-V veneer (95% CI 39.1%-86.2%), whereas VERIFY and POSTER produced none; because n=12, these public-case rates are descriptive pilot data only. CEAG is not proposed as a patient-facing display or a replacement for conventional charts. It is a clinical informatics monitoring aid for reviewers who need to detect apparently well-sourced hallucinations, compare response regimes, and separate stylistic change from substantive trust change.
Yuusuke Harada (Sun,) studied this question.