Abstract This study investigates the application of deep reinforcement learning (DRL) to the control of smart base-isolated buildings exhibiting highly nonlinear isolation behavior. A five-story structure with lead–rubber bearings (LRBs) is modeled using the Bouc–Wen formulation, and a DRL-based controller is trained via proximal policy optimization (PPO) within a partially observable Markov decision process (POMDP) framework. The agent relies solely on measurable signals such as acceleration and displacement, making the approach suitable for realistic implementation. To promote generalization and robustness, training is conducted using synthetic ground motions generated by the Kanai–Tajimi filter with randomly varied spectral and temporal parameters. Numerical simulations confirm that the trained controller significantly reduces both peak and root-mean-square (RMS) displacements and accelerations compared to passive isolation, especially under strong seismic inputs. In particular, the controller substantially suppresses peak base displacement, with reductions exceeding 40% in certain cases, which addresses a critical limitation of passive systems and accelerates the decay of residual vibrations. Importantly, the controller generalizes well to previously unseen earthquake records, maintaining high performance across a wide range of scenarios. These results highlight DRL as a promising data-driven strategy for robust and adaptive control of nonlinear structural systems under partial observability.
Takehiko Asai (Mon,) studied this question.