Large language models (LLMs) are rapidly being integrated into clinical workflows, supporting tasks such as diagnosis generation and patient communication.1 Hallucinations—unintended fabrications arising from gaps in a model’s underlying knowledge—are a well recognised risk. However, research in 2024 has identified a distinct class of model behaviour, known as deception. Deception occurs when a model produces outputs that misrepresent its reasoning or capabilities in ways that make the output appear more credible or aligned with user expectations.
Reddy et al. (Wed,) studied this question.