System evaluation reveals graph expansion alters retrieval for half of queries, highlighting trade-offs between thematic and entity-based searches.
Graph-augmented retrieval (Graph RAG) systems conventionally construct their graph by LLM-based entity and relation extraction. We study an alternative setting, common in personal knowledge management yet under-evaluated: the graph already exists as human-curated wikilinks accumulated in a linked note corpus. We describe a deployed retrieval system over such a corpus (dense retrieval with multilingual-e5, a bounded 1-hop wikilink expansion used strictly for candidate generation, and cross-encoder reranking; SQLite persistence; no graph database and no extraction pass), and report pilot A/B telemetry from production use: on 10 logged production queries compared vector-only versus vector-plus-graph, the graph hop changed the reranked top-12 in 5 cases (promoting 1-3 notes each, all arriving via the wikilink walk) and changed nothing in the other 5. The pilot also surfaced a directional failure mode: expansion appears to help thematic queries and hurt named-entity queries, whose hub-like cards flood the candidate pool. These observations motivate the paper's main contribution: a query-class-stratified evaluation protocol (entity, theme, bridge, compare, temporal strata) for measuring when graph expansion helps, together with ablations (hop depth, neighbour cap, entity gate) and a paired statistical analysis plan. The pilot sample is small (N=10 queries, one corpus, one user) and we state this plainly; the protocol, not the pilot, is the contribution. Code for the retrieval layer is public; the protocol is designed so that any owner of a linked corpus can replicate it on private data without disclosing that data.
No takes yet. Share an insight, caveat, or question.
Anton Dziatkovskii (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: