In this paper, we present an agent-based model of a controlled detonation system for dynamic sandbox analysis of suspicious software. Instead of treating the sandbox as a passive observer, the model places an AI operator inside the analysis loop and allows it to perform adaptive GUI interactions in a plausible, isolated execution environment. The controlled detonation process is formulated as a partially observable Markov decision process (POMDP), while the proposed proof-of-concept architecture combines initial profiling, VM preparation, multi-layer telemetry, and an RL policy with visual perception and temporal memory. Evaluation in a controlled emulation setting on 180 malware samples from three threat classes shows higher Activity Rates and Coverage, and shorter Time-to-Reveal than passive and fixed scripted baselines. These results support the feasibility of adaptive interactions as a promising direction for sandbox analysis, while broader external validation, matched comparisons with prior systems, and component-wise ablation remain future work.
Ivanchenko et al. (2026) studied this question.