This paper investigates whether alignment and misalignment emerge differently in single-agent versus multi-agent large language model (LLM) systems when placed in adversarial, high-stakes scenarios. Using a controlled survival simulation (“Island Plane Crash”), we compare behavioral patterns, coordination dynamics, and failure modes across single-agent and multi-agent configurations. Our preliminary findings suggest that single-agent systems tend toward centralized, paternalistic crisis leadership, while multi-agent teams enable deferential, human-aligned collaboration alongside novel systemic risks such as communication breakdowns and emergent shared ideologies. This work presents early-stage research and ongoing analysis. We actively welcome feedback on methodology, interpretation, and scenario design.
Hermann et al. (Thu,) studied this question.