Randomized trial assesses vulnerability profiles in large language models, suggesting methods for improved safety in decision-making systems.
Large language models are increasingly deployed in safety-critical decision services---content moderation, clinical decision support, legal analysis---yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal manipulations each model resists and it does not. Testing 5~models spanning 2~architecture families under a realistic \50/day compute budget, we find that vulnerabilities are : linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not---identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable (ICC(2,k)\,=\,0.97), and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline---with adaptive concurrency, budget-aware execution, and three-tier data---scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment. Target venue: IEEE BigDataService 2026 (accepted). Author preprint deposited for archival and citation. Draft — pending author review.
No takes yet. Share an insight, caveat, or question.
Andrew Bond (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: