Proof of format demonstrates the feasibility of standardized tags for academic papers, enhancing cross-disciplinary understanding.
Academic knowledge is written almost entirely in natural language, which makes it readable by humans but hard to query across disciplines that use different words for related things. We describe a tag layer that assigns every paper the same small set of field-neutral facets — claim type, phenomenon (a normalised join key), claim relations, replication status, and limitations — under one non-negotiable rule: every tag is grounded in an exact, verbatim sentence from the paper's own text. We report a proof-of-format on a curated corpus of 126 papers spanning 39 disciplines (1866–2023), of which 116 were fully faceted. The format is feasible and verifiable: 127 claim-relations and 87 limitations, each substring-checked against the source. It is also honest about what it did not do: a novelty scan surfaced no cross-disciplinary link that the literature had not already drawn, and the phenomenon facets, written as specific verbatim-anchored strings, rarely collide across fields — indicating that a canonical join vocabulary, not free text, is the missing ingredient for surfacing new links. We argue the value of the layer is as a compass for "far reading" that points where slow close reading may be worthwhile, not as an answer engine.
No takes yet. Share an insight, caveat, or question.
Shir Sivroni (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: