The Error-Cost Asymmetry of Current-Paradigm AI: Why Instability Constrains Construction More Than AttackCivilization Physics — AI Governance Series This paper develops a structural account of how instability in contemporary AI systems produces asymmetric effects across task types. It argues that the most decision-relevant consequence of large language model (LLM) instability is not average accuracy, but error-cost asymmetry between construction tasks, which require sustained correctness, and attack tasks, which can succeed through iterative trial and error. This asymmetry systematically constrains constructive capability while enabling scaled misuse under certain conditions . The analysis begins by defining two distinct task classes. Construction tasks depend on producing artifacts that satisfy multiple constraints over time, including correctness, safety, and maintainability. Errors in these contexts compound across steps, creating cascading costs in debugging, validation, and long-term operation. In contrast, attack tasks derive value from achieving a single successful outcome, such as bypassing a control or extracting a resource. Failures in attack contexts are often disposable, while success can be rapidly verified and scaled. To formalize this distinction, the paper introduces the concept of error-cost as the marginal cost associated with failure, including detection, correction, and downstream impact. Construction environments typically exhibit high and compounding error-cost due to dependency chains and verification requirements. Attack environments often exhibit low marginal error-cost, enabling repeated attempts with minimal penalty. Four structural mechanisms are identified as drivers of asymmetry: Single-point success — attack tasks require only one successful pathway to generate value. Disposable failure — unsuccessful attempts can be discarded without significant cost. Short feedback loops — rapid verification enables continuous iteration and refinement. Automated iteration — scalable retries amplify the probability of eventual success even with modest per-attempt capability. These mechanisms create fundamentally different pipelines. Construction resembles a constrained, multi-step process where each failure introduces rework and risk. Attack resembles an iterative loop where variation is filtered through repeated attempts until success emerges. The paper introduces a compact formal model culminating in the Asymmetric Risk Index (ARI), which expresses the relative cost of achieving success in construction versus attack contexts. The model demonstrates that even low per-attempt success probabilities can translate into high realized attack capability when retry budgets are large and feedback is fast. Conversely, construction tasks remain tightly bound by dependency depth and verification cost, limiting scalability under instability. Empirical case examples ground the analysis. Indirect prompt injection exploits the blurred boundary between data and instructions in AI-integrated systems, enabling unauthorized actions through repeated attempts. Package hallucinations in code generation create supply-chain vulnerabilities by introducing non-existent dependencies that attackers can weaponize. Insecure code generation demonstrates how plausible outputs can embed hidden defects that are costly to detect, constraining constructive use while exposing exploitable patterns. The paper highlights that this asymmetry is conditional and shaped by system design. Attack tasks become more constrained when they require persistence, stealth, or coordination over time, increasing dependency depth and verification cost. Conversely, construction tasks become more scalable when verification is automated and inexpensive. System architecture—such as sandboxing, output validation, and privilege separation—can significantly alter the balance by increasing attack cost and reducing retry feasibility. From a governance perspective, the paper argues that evaluation frameworks must shift from single-shot performance metrics to scaled misuse potential. Key evaluation criteria include success rates as functions of attempt budgets, feedback latency, and verification cost, as well as standardized testing for prompt injection, output handling vulnerabilities, and supply-chain risks. These criteria align with emerging risk-management frameworks that emphasize capability thresholds and continuous measurement. The paper concludes that AI instability becomes systemically significant only when filtered through task economics. Construction demands sustained coherence and therefore remains constrained by instability. Attack benefits from iteration and therefore scales with it. As part of the Civilization Physics framework, this work situates AI risk within a broader structural principle: systems with asymmetric error-cost landscapes will channel capability toward domains where failure is cheap and success is scalable. Effective governance must therefore reshape the cost structure itself, rather than relying solely on improving average model performance. Keywords: AI Risk · Error-Cost Asymmetry · Asymmetric Risk Index · AI Security · Prompt Injection · Model Instability · Evaluation Metrics · Attack Scaling · Construction Constraints · Civilization Physics
Xiangyu Guo (Fri,) studied this question.