PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

Diagnostics of cognitive failures in multi-agent expert systems using dynamic evaluation protocols and subsequent mutation of the processing context

View Full Paper
ASAndrejs SorstkinsJBJ. William BaileyDBDaniel Barón

Key Points

  • The framework reveals latent cognitive failures in multi-agent systems' decision processes, improving their performance.
  • It utilizes curated golden datasets alongside silver datasets for controlled behavioral mutation to enhance evaluation.
  • An LLM-based Agent Judge scores agent performance, providing targeted prescriptions in a vectorized recommendation map.
  • This approach promotes reproducible expert behavior transfer for refined system performance through active interventions.

Abstract

The rapid evolution of neural architectures - from multilayer perceptrons to large-scale Transformer-based models - has enabled language models (LLMs) to exhibit emergent agentic behaviours when equipped with memory, planning, and external tool use. However, their inherent stochasticity and multi-step decision processes render classical evaluation methods inadequate for diagnosing agentic performance. This work introduces a diagnostic framework for expert systems that not only evaluates but also facilitates the transfer of expert behaviour into LLM-powered agents. The framework integrates (i) curated golden datasets of expert annotations, (ii) silver datasets generated through controlled behavioural mutation, and (iii) an LLM-based Agent Judge that scores and prescribes targeted improvements. These prescriptions are embedded into a vectorized recommendation map, allowing expert interventions to propagate as reusable improvement trajectories across multiple system instances. We demonstrate the framework on a multi-agent recruiter-assistant system, showing that it uncovers latent cognitive failures - such as biased phrasing, extraction drift, and tool misrouting - while simultaneously steering agents toward expert-level reasoning and style. The results establish a foundation for standardized, reproducible expert behaviour transfer in stochastic, tool-augmented LLM agents, moving beyond static evaluation to active expert system refinement.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sorstkins et al. (2025) studied this question.

synapsesocial.com/papers/68de5da283cbc991d0a2077bhttps://doi.org/10.48550/arxiv.2509.15366
Ask AI
Helpful
Bookmark
Share
View Full Paper