PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 7, 2026International Journal of Advanced Computer Science and Applications0 citationsOpen Access

Agentic Accountability: “The Buck Stops Where?” Ethical Frameworks for Human Oversight of Autonomous AI Systems

OBOrnella Bahidika

Key Points

  • This study investigates the ethical and governance aspects of accountability in autonomous AI systems, questioning who is responsible when harm occurs.
  • Developed an oversight framework mapping agent actions to four human involvement tiers.
  • Analyzed 177 incidents from the AI Incident Database for tier assignability and impact.
  • Assessed inter-rater reliability among independent annotators for tier classification.
  • Found 23% of tier-assignable incidents fell within T4 (prohibited) under the framework.
  • Identified that of 41 tangible-harm events, 63% involved Medium or High autonomy relevant to tier-T4 enforcement.
  • Measured inter-rater reliability across annotators ranged from 0.58 to 0.77, indicating moderate-to-substantial agreement.

Abstract

The rapid evolution from generative Artificial Intel-ligence (AI) toward agentic AI, systems capable of autonomously planning and executing multi-step actions, has introduced an unprecedented accountability gap in modern computing. Unlike traditional AI tools that respond to discrete prompts, agentic sys-tems pursue goals across extended time horizons, invoke external services, and produce cascading real-world consequences. This shift raises a fundamental ethical question: When an autonomous agent causes harm, who bears responsibility? This study examines the ethical and governance dimensions of agentic accountability, drawing on recent literature in AI ethics, regulatory studies, and human-computer interaction. Building on prior tiered ap-proaches to automation oversight in the human-factors literature, in regulatory risk classification, and in recent agent-autonomy frameworks, we present an action-level oversight framework that maps individual agent actions to four tiers of human-in-the-loop involvement, ranging from full automation to mandatory prohibition, calibrated by stakes and reversibility. We further analyze design patterns for “emergency brakes” (circuit breakers, action budgets, reversibility constraints, audit trails, kill switches), and propose a composition-aware extension that detects tier laundering, where individually low-tier actions compose into a higher-tier outcome. We then conduct an empirical pilot applying the framework to 177 incidents from an April 2026 snapshot of the AI Incident Database with full CSET classification, finding that 23% of tier-assignable incidents fall in T4 (prohibited under the framework) and that of 41 tangible-harm events, 26 (63%) involved Medium or High autonomy where tier-T4 enforcement would have been most directly applicable. Inter-rater reliability across the database’s three independent annotators ranges from ?=0.58 to 0.77 at the tier level (moderate-to-substantial agreement). The contribution of this work is an action-level operationalization of tiered oversight, anchored in real-world incident data, with explicit identification of composition-aware detection as the highest-leverage methodological direction for follow-up research.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ornella Bahidika (2026) studied this question.

synapsesocial.com/papers/6a250c507def13d035e1c6a0https://doi.org/10.14569/ijacsa.2026.0170502
Ask AI
Helpful
Bookmark
Share
View Full Paper