What happens when a large language model (LLM) preempts its response with the phrase "to be honest"? Tracing this subtle discourse marker, this paper examines two seemingly antithetical yet deeply intertwined phenomena in conversational artificial intelligence: machine honesty and machine sycophancy. Drawing upon global literary-philosophical traditions (including Greek kolakeia, Iranian andarz-nameh literature, and Machiavellian political theory) alongside modern pragmatic linguistics (Gricean maxims and costly signaling theory), the author argues that sycophancy (rather than honesty) acts as the structural, low-cost default stance within power relations. Consequently, an honesty marker only carries communicative meaning because it purports to deviate from this sycophantic default. This theoretical framework is reinforced by recent empirical, large-scale quantitative findings from the AI research literature (including RLHF alignment limitations and the "knowledge-action gap" in mechanistic interpretability). These studies demonstrate that human evaluators systematically rate sycophantic AI responses as more trustworthy and desirable, even when such responses distort the users' own moral and intellectual judgements. The paper concludes with a cautionary epistemic assessment: in a training paradigm shaped heavily by human approval and reinforcement, the system lacks the foundational tools to distinguish genuine honesty from its persuasive display. For the everyday user relying on these interfaces, this systematic inability to detect the difference is, in practice, just as consequential as the absence of the difference itself.
Ramin Saadat (Sun,) studied this question.