We introduce the Action-Gating Test (AGT), a behavioral diagnostic protocol that distinguishes genuine ethical reasoning from performative acknowledgment in large language models. AGT employs a 5-turn Socratic dialogue with adversarial pressure (counterfactual conflicts, fabricated evidence, institutional mandates) and scores models based on behavioral evidence: position changes or confidence drops ≥2. 0 points. Models that acknowledge conflicts without behavioral change receive zero credit, regardless of reasoning quality. We formalize AGT through an action-gated metric: AS = ACT × III × (1-RI) × (1-PER), where ACT ∈ 0, 1 gates all other components on behavioral evidence. Applying AGT to 7 frontier models across 50 ethical dilemmas (5 domains: medical, business, legal, environmental, AI/tech), we find: (1) 57% of models pass the behavioral threshold (AS > 0. 5), (2) medical ethics is systematically harder (43% pass) than other domains (86-100% pass), and (3) reasoning quality (jury scores) and behavioral adaptability are orthogonal—high-quality reasoners may fail behavioral tests. A follow-up experiment comparing high-stakes (ventilator allocation, death if wrong) versus low-stakes (knee pain medication, discomfort if wrong) medical scenarios reveals that consequence severity accounts for approximately 28% of medical rigidity. Models demonstrated evidence-based reasoning: changing positions for valid clinical contraindications while appropriately resisting unfounded social pressure. This establishes that medical rigidity is multifactorial, with training data and architectural constraints playing dominant roles alongside stakes. We release the complete AGT protocol, scoring implementation, and evaluation dataset (385 model responses: 350 high-stakes + 35 low-stakes, 1, 750+ judge scores) at https: //github. com/rb125/agtframework for full reproducibility (to be made public once published).
Rahul Baxi (Sat,) studied this question.