We describe the Global Network Model (GNM), an architecture that reduces the energy cost of large-scale inference by offloading operations from a high-energy probabilistic language model to *low-energy functor models and localized specialized models. The organizing principle is ontological: the semantic space is treated as a union of taxonomies, with the large language model (LLM) representing only the upper levels required for routing, and an ensemble of small language models (SLMs) and graph neural networks (GNNs) holding the deeper, domain-specific structure and executing at lower energy. Operations that can be performed deterministically retrieval, canonicalization, routing, and exact transforms are moved off the LLM entirely. The GNN ensemble supports governed cross-model message passing, coordinated by an orchestrator over a bounded, deterministic schedule, which carries cross-cutting structure between specialized models while preserving bounded cost and extending deterministic replay and rollback to the full ensemble. We position the locality and equivalence-class complexity arguments as assumptions built atop established sparse-computation results rather than as new complexity theorems, and we are explicit about the costs the offloading incurs: routing accuracy, the cost of computing equivalence classes, and orchestration overhead across the model fleet. The architecture’s advantage is the offloading and the ontological decomposition; its realized magnitude is workload-dependent and strongest where the ontology decomposes cleanly and structure is cheap to detect. Functor Models are patent pending but the GNM is open research.*Early functor model benchmarking has shown promising energy reduction but more thorough efforts are still required in this area
John Harby (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: