Controlled benchmarking demonstrates secure inference protocol performance in distributed AI architectures, indicating minimal latency overhead alongside verified governance.
Edition note: This record contains the authoritative English edition (DARI_Paper.pdf) and a complete Korean translation (DARI_Paper_KO.pdf). The English edition controls if a translation ambiguity affects technical meaning. 판본 안내: 이 레코드는 영문 원문과 한국어 완역본을 함께 수록한다. 번역상 모호성이 기술적 의미에 영향을 미치는 경우 영문판을 기준으로 한다. AI systems combine private context, model inference, tools, durable effects, multimodal streams, and distributed accelerators, yet their stacks divide endpoint authentication, policy, placement, execution, and evidence. A secure channel proves who connected, not whether context, models, workers, effects, and retained evidence were authorized. We present DARI (Delegated Authorization and Receipts for Inference), an authority-bearing protocol that binds peer identity, attenuated grants, context disclosure, policy state, model and endpoint identity, streaming, tool effects, and signed evidence in a transcript-bound exchange state machine. Decisions precede protected output consumption and external effects; receipts commit to the principals, grants, decisions, inputs, outputs, and outcomes that formed the exchange. DARI also treats governance as a placement constraint: signed worker capabilities and admission policy define the eligible set before a late-binding router considers load, latency risk, topology, adapter or media affinity, and tenant-scoped KV-cache overlap. A controlled localhost evaluation compared governed DARI, HTTP/JSON with server-sent events, and a persistent WebSocket under one deterministic 128-token workload. Across five 30-turn trials, median time to first token was 8.34 ms for DARI, 5.87 ms for HTTP/SSE, and 5.79 ms for WebSocket; median completion time was 309.98, 294.07, and 293.88 ms, respectively. Median inter-token latency was 2.28 ms in every arm. DARI used 5,688 application-layer bytes per turn, 13.1% fewer than HTTP/SSE and 26.6% fewer than WebSocket, while carrying mutual authentication, inline governance, and receipt delivery. In this workload, DARI's measured cost appears at first-token and completion boundaries while median steady-stream cadence is unchanged and application-layer wire volume is lower.
No takes yet. Share an insight, caveat, or question.
Siook (Patrick) Rho (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: