Because LLMs are still in development, what is true today may be false tomorrow. We therefore need general strategies for debiasing LLMs that will outlive current models. Strategies developed for debiasing human decision making offer one promising approach as they incorporate an LLM-style prompt intervention designed to access additional latent knowledge during decision making. LLMs trained on vast amounts of information contain information about potential biases, counter-arguments, and contradictory evidence, but that information may only be brought to bear if appropriately prompted. Metacognitive prompts developed in the human decision making literature are designed to achieve this and, as I demonstrate here, they show promise with LLMs. The prompt I focus on is “could you be wrong?” Following an LLM response, this prompt leads LLMs to produce additional information, including why they answered as they did, identifying errors, biases, contradictory evidence, and alternatives, none of which were present in their initial response. Further, this metaknowledge often reveals that how LLMs and users interpret prompts are not aligned. I demonstrate this prompt in three cases. In the first two cases I use a set of questions taken from recent articles identifying LLM biases, including implicit discriminatory biases and failures of metacognition. “Could you be wrong” prompts the LLM to identify its own biases and produce cogent metacognitive reflection. In the last case I present an example involving convincing but incomplete information about scientific research (the too much choice effect), which is readily corrected by “could you be wrong?” In sum, this work argues that human psychology offers a valuable avenue for prompt engineering, leveraging a long history of effective prompt-based improvements to decision making.
Thomas T. Hills (Mon,) studied this question.