Protocol evaluation demonstrates problem-fragmented inference reduces coherence tax across context budgets in peer-to-peer networks, indicating viable distributed model execution.
Peer-to-peer language-model inference today splits the model: layers or tensors are spread across machines and activations cross the public internet on every generated token, against bandwidth and latency gaps of four to five orders of magnitude. This work specifies Swarmbly, a protocol that distributes the problem instead. A client-side orchestrator decomposes a request into a dependency graph of semantic micro-tasks, each dispatched once and asynchronously to a volunteer node running a complete 1–8B model; returned fragments are verified, selected and spliced locally, so the network is crossed once per fragment per session rather than once per layer per token. Substantive contributions. The context budget S — the shared context carried by each dispatched fragment — is identified as a single scalar on which assembly coherence, fragment verifiability, privacy by decontextualization and required worker capability all depend, with opposing signs. Viability reduces to whether some S satisfies all four thresholds at once, stated as a falsifiable proposition. A coverage model for semantic assembly, in which the pre-generation plan acts as the reference sequence and packet loss rather than sample placement is the stochastic element, yielding the design equation c ≥ ln(1/ε)/(1−p) and an operating range of 3–5 replicas per critical unit. Consensus by multiple alignment of replicas produced by nodes of deliberately different model families, which returns with every answer a per-unit agreement score and a map of low-confidence regions — a signal a provider running a single model has nothing to align, and therefore cannot produce. Disclosed as a mechanism: its correlation with correctness was measured and came back flat, so no reliability benefit is claimed for it. A fourth element, added in v1.3, is dynamic privacy tiering: a client-side classifier routes each request into an open volunteer mesh, a permissioned “trusted swarm” whose membership is a cryptographic public-key whitelist under mutual TLS, or purely local execution — the same protocol at every tier, with the replica count reducible inside a trusted swarm only under an explicit declaration that no confidence map was produced. First measurements. Against three model families served locally, the coherence tax falls monotonically in the context budget — 24.1 %, 20.4 %, 16.1 %, 13.7 % across the swept range — and the abandonment criterion fixed in advance is met in three task categories. The same run finds no relationship between inter-replica agreement and judged quality (r = −0.030 over 597 semantic units), which leaves the confidence map unsupported rather than refuted; both results are reported in full. The deposition also contains the wire protocol, the assembly and verification algorithms, and a reference harness that measures the coherence tax against a stated abandonment criterion. Scope of claims. No claim is made of latency parity, unlimited context, cryptographic confidentiality, or a demonstrated environmental benefit. Limitations and negative results are stated in full in Section 12 of the whitepaper. Contents: whitepaper v1.4 (English and Spanish), protocol specification v0.2 revision 2, and the V0 coherence-tax reference implementation (178 tests). Licensing: the software and specification are released under AGPL-3.0-or-later — clause 13 (remote network interaction) applies to hosted deployments, see the NOTICE file. The document text is CC BY 4.0.
No takes yet. Share an insight, caveat, or question.
Sebastian Espinoza‐Ulloa (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: