Randomized trial decomposes disagreement in machine-extracted argument structures, indicating unitization is not the failure point.
Machine-extracted argument structures are increasingly read as evidence about how fields move, so measurement error at extraction propagates into whatever is built on the extracted graph. A prior pre-registered run found two machine operators did not recover the dependency structure of five argumentative documents at a declared agreement level, and located the failure upstream of relation labelling, but could not say what the disagreement was made of. This paper fixed the unit inventory before extraction, supplied it identically to both operators, and repeated every extraction five times, decomposing disagreement into selection, typing and relation drawing against a within-operator baseline. Every threshold was derived from its own reference class and deposited before data collection. Across 150 extraction calls in three unitization conditions, edge-level agreement failed its .71 gate in all fifteen document-by-condition cells, including the ten where units were supplied identically, and separated below the within-operator baseline in every cell (p ≤ .001). On this corpus the predecessor’s closing recommendation is therefore overturned: sharing an inventory does not rescue edge agreement, so unitization was not where the failure lived. Segmenter fidelity failed on all five documents, which bounds every fixed-inventory claim here and leaves a shared inventory of established fidelity untested. Includes zharnikov-2026bl-unitization-before-extraction.yaml (Paper Spec v0.1.0) — a machine-readable specification of the paper’s claims, assumptions, and dependencies. The paper’s full machine-first bundle (the SPINE claim/dependency graph and the ONTOLOGY term module) lives in the public repository; see github.com/spectralbranding/paper-spec for the standard. This PDF is generated programmatically from that machine-first source under a research-as-repository model.
No takes yet. Share an insight, caveat, or question.
Dmitry Zharnikov (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: