Taxonomy analysis delineates five distinct engineering layers in large language model applications, highlighting failure modes caused by missing iterative verification loops.
Large language model (LLM) Application Development has moved, in roughly five years, from hand-tuning a single input string to operating autonomous systems that plan, call tools, and revise their own output across multiple passes. Practitioners now use “prompt engineering,” “context engineering,” “agent engineering,” and, more recently, “harness engineering” and “loop engineering” almost interchangeably, even though each term names a distinct engineering discipline with its own failure modes and its own way of measuring success. This paper proposes a five-layer taxonomy— Prompt Engineering, Context Engineering, Harness Engineering, Agent Engineering, and Loop Engineering— that separates these disciplines by what is actually being engineered at each stage: the input string, the surrounding information, the tool-and-permission scaffold, the delegation and planning behavior, and the iterative actobserve-verify cycle, respectively. For each layer we define boundary conditions relative to its neighbors, propose a concrete evaluation metric, and ground the layer in one published system from the literature and one first-person practitioner artifact drawn from an active multi-agent deployment. The paper is a taxonomy and case-study contribution, not a controlled experiment: its claims about the practitioner artifacts are illustrative, not measured, and we say so explicitly rather than let the distinction blur. We argue the taxonomy’s principal practical use is diagnostic—identifying which layer a failing AI system is actually missing, particularly the common and under-diagnosed failure of building an agent without a loop, which produces confident but unverified output.
No takes yet. Share an insight, caveat, or question.
MFL Muhammad Faisal Laiq Siddiqui (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: