Systems evaluation demonstrates reliable automated claim triage with local language models, highlighting minimal argument generation as a key tool-calling design principle.
Insurance claim adjudication gets held up as a natural fit for LLM-based agents: chaining lookups, applying conditional coverage logic, and producing a rationale a human can audit, none of which calls for creativity so much as reliability. This report covers the design and evaluation of a two-tool ReAct agent, built with LangGraph, that triages synthetic insurance claims into "Approved" or "Route for Review" against policy coverage rules. The agent never issues a final denial by design; every failed check goes to a human, so the system triages instead of adjudicating. The main finding here is about tool-calling reliability, not reasoning quality. An early version of the agent asked the model to reproduce an entire patient record as a JSON string argument, and that failed unpredictably across several open-weight models served locally through Ollama (llama3.1, then qwen2.5): truncated JSON, tool calls narrated as text instead of invoked, and in the worst cases a single stray character as the entire argument. Replacing that tool with one that only takes a short patient identifier cut the model's actual generation task down to something close to copying a string it had already seen, and that alone eliminated the failures, with no change to model size and no new prompt engineering. On a 15-record synthetic validation set with a deterministic rule-checker reference, the resulting agent, run against a frontier model as its default backend but built to support local and other hosted backends as configurable options, matched the reference decision on all 15 records against a 10/15 floor for a policy that always defers to review. The practical lesson for anyone building something similar: the reliability gap between small local models and frontier models on structured tool-calling is a specific, trainable skill gap, separate from general reasoning ability, and the fix that worked here was cutting what the model had to generate down to nearly nothing, well before reaching for a bigger model or a cleverer prompt.
No takes yet. Share an insight, caveat, or question.
Lukmon Olayinka (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: