PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 17, 2026Journal of Vibration and Control2 citations

A proximal policy optimization framework for bearing condition monitoring using low-dimensional time-domain features

View Full Paper
WAWaqar AhmadYuan Ze UniversitySSSyed Humayoon ShahYuan Ze UniversityNINaeem Ul IslamYuan Ze University

Key Points

  • This research aims to explore the effectiveness of Proximal Policy Optimization in developing maintenance strategies for bearing conditions.
  • Developed a custom OpenAI Gym environment to simulate decision-making for bearing maintenance.
  • Utilized experimental vibration data from bearings with various conditions: normal, ball, inner-race, and outer-race faults.
  • Trained an RL agent to establish maintenance policies like inspection, repair, and replacement.
  • Evaluated training performance through cumulative rewards, loss metrics, and sensitivity analysis.
  • Achieved 94.2% accuracy in decision-making over 10 epochs.
  • Observed limited improvement in performance with additional training iterations.
  • Identified instability in policy updates and sensitivity to reward structures.

Abstract

Bearings are essential elements of rotating machinery and their malfunction may result in considerable operational interruptions and financial detriment. This paper investigates Proximal Policy Optimization (PPO), a reinforcement learning (RL) technique, to formulate data-driven policies for bearing maintenance. A custom OpenAI Gym environment was developed to replicate the decision-making process employing experimental vibration data from normal bearings, as well as bearings with ball, inner-race, and outer-race faults. The RL agent was trained to determine the appropriate maintenance policies such as inspection, repair, and replacement to reduce total costs and prevent breakdowns. In addition, training performance was evaluated using essential measures such as cumulative rewards, loss, KL divergence, and value loss. The experimental findings demonstrate that the PPO agent achieved 94.2% accuracy in decision making in 10 epochs with limited improvement from additional training. Furthermore, the method shows instability in policy updates, value loss, and sensitivity to sparse-reward structure. These findings demonstrate that PPO holds considerable potential for vibration-based CBM; however, its performance in real-world operational environments remains highly dependent on reward design and hyperparameter tuning. This research showcases a balanced evaluation of PPO’s strengths and limitations in bearing maintenance and provides a foundation for future studies on hybrid and alternative reinforcement learning strategies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ahmad et al. (2026) studied this question.

synapsesocial.com/papers/69e1ce605cdc762e9d857690https://doi.org/10.1177/10775463261443503
Ask AI
Helpful
Bookmark
Share
View Full Paper