Visual dark patterns, including inflated strike-through prices, countdown timers, scarcity warnings, and urgency typography, are prevalent in e-commerce. As multi-modal AI shopping agents increasingly encounter product pages on behalf of consumers, a critical question emerges: do these visual elements elicit anchoring-like behaviour, a well-documented tendency to adjust judgments toward a given reference value, in large language models? This study presents a controlled experiment involving five commercial multi-modal LLMs spanning Western and Chinese providers: GPT-5.4 mini, Gemini 3 Flash Preview, Claude Sonnet 4.6, Kimi K2.5 Instant, and Qwen 3.5 Flash. Each model evaluated eight fictitious-brand products across six conditions: three text-based conditions varying in the form and content of price presentation, and three image-based conditions varying in visual design style and defensive intervention. Contrary to predictions, the dominant finding is a systematic vigilance effect: aggressive dark-pattern imagery consistently produced lower scores than minimal-style imagery across all eight products and all five models. Four of five models also scored dark-pattern conditions below the plain-text price baseline, with the largest gap in Gemini’s recommendation scores (d = 3.35). Kimi K2.5 Instant (temp = 0.6) was the sole exception, producing higher scores under dark-pattern conditions than under the plain-text price baseline. Defensive prompting protected deal-quality judgments in most products but failed to suppress purchase recommendations in most cases, revealing an asymmetric vulnerability. Cross-cultural analysis identified significant divergence between Western and Chinese models under dark-pattern conditions (d=−1.232), suggesting that models developed within different cultural backgrounds interpret the same commercial signals differently. These findings raise the question of whether AI shopping agents reliably protect all consumers equally, or whether their responses to visual manipulation reflect the cultural context in which they were developed.
Haoran Xu (2026) studied this question.