Randomized trial evaluates text optimization to improve product ranks in e-commerce LLMs, implying modest effectiveness.
Large Language Models (LLMs) are increasingly integrated into e-commerce search and recommendation pipelines, raising concerns about their susceptibility to adversarial manipulation of product rankings. This study implements and evaluates a gradient-based Strategic Text Sequence (STS) optimization framework, using the Greedy Coordinate Gradient (GCG) algorithm of Zou et al. [7] to construct adversarial token sequences embedded in product descriptions, with the goal of promoting a target product's rank in LLM-generated recommendation lists. Using the open-weight Llama-3.2-1B-Instruct model, we optimize an STS over 50 GCG iterations and evaluate the resulting sequence against a no-STS baseline across 30 independent, randomly reordered trials. Optimization produced a consistent, monotonic reduction in training loss (2.73 to 0.32), and evaluation showed a positive directional trend across ranking metrics: NDCG@5 improved from 0.073 to 0.152, Top-5 frequency from 0.133 to 0.300, and mean rank from 8.57 to 7.67. However, a paired t-test on the 30-trial sample did not reach statistical significance (t=1.10, p=0.279), indicating that while the attack shows a genuine positive effect, its magnitude on a small (1B-parameter) instruction-tuned model is modest and requires a larger sample to confirm reliably. We report these results transparently, including the null result on significance, and discuss why effect sizes on small LLMs may differ from those reported for larger models in prior work [1].
No takes yet. Share an insight, caveat, or question.
Sameed Rashid (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: