This study evaluated a tool-augmented large language model (LLM) agent system utilizing a traditional Asian medicine (TAM) metadata. A structured information comprising 4780 entries across five entity types (herbs, syndromes, TCM symptoms, modern medicine symptoms, and acupoints) was constructed. Four LLMs (GPT-5.2, GPT-5-mini, Claude Sonnet 4.6, and Claude Haiku 4.5) were evaluated under baseline, non-agentic retrieval (RAG), and tool-augmented (agent) conditions using two public benchmarks: TCMBench (1300 items) and TCMEval-SDT (600 items). McNemar’s test and bootstrap confidence intervals were applied to examine performance differences. Tool augmentation showed a nominally significant improvement for GPT-5-mini on TCMBench term-related items (+4.5 percentage points on paired items, McNemar p = 0.034, 95% CI +0.9, +8.5 pp), while Claude Sonnet 4.6 showed a nominally significant decline on SDT pathogenesis (−0.038, bootstrap p = 0.006). The agent condition outperformed the non-agentic RAG baseline for GPT-5-mini on TCMBench (agent 85.6% vs. RAG 77.7% vs. baseline 75.4%), suggesting that selective, autonomous tool invocation is more effective than fixed retrieval. Tool usage rates varied substantially across models (2–87%), with moderate usage (30–40%) associated with the most consistent gains. These findings provide empirical evidence on the potential and limitations of metadata-based tool augmentation for LLMs in the TAM domain.
Lee et al. (Tue,) studied this question.