PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 19, 2026Discover Networks1 citationsOpen Access

A reinforcement learning framework for self-healing fault recovery in intent-based SDNs

EME. H. MakiyahMRM. N. RasoolFRF. A. Rawdhan

Key Points

  • The research aims to develop a reinforcement learning framework for autonomous fault recovery in software-defined networking environments.
  • Developed a self-healing framework using a Mininet testbed with the Faucet SDN controller.
  • Collected multi-source telemetry data with Prometheus and simulated network faults to train a PPO agent.
  • Evaluated the performance of the PPO agent in selecting state-conditional recovery actions across various fault types.
  • The PPO agent achieved state-conditional recovery actions in over 85% of fault-state timesteps.
  • Demonstrated stable optimization behavior with significant improvement in episodic rewards during training.
  • Showcased ability to handle diverse fault types beyond just link failures, improving overall SDN operational reliability.

Abstract

This paper presents a reinforcement learning-based self-healing framework for Software-Defined Networking (SDN) that autonomously manages diverse network faults in a realistically emulated environment. A Mininet testbed controlled by the Faucet SDN controller is instrumented with Prometheus to collect multi-source telemetry, while an automated fault injector and congestion generator produce link, port, flow and controller events alongside UDP-induced bottlenecks to create rich training data. Network features–including controller CPU and memory usage, OpenFlow statistics, port status and explicit fault labels–are periodically scraped and aggregated into a structured dataset that forms the state space of a custom Gym-compatible environment. A Proximal Policy Optimisation (PPO) agent with a multilayer perceptron policy learns discrete self-healing actions such as no-op, port resets, switch restarts and bespoke recovery procedures, guided by a reward function that penalises persistent faults and unnecessary interventions while rewarding timely and appropriate recovery. Experimental evaluation over multiple PPO training runs shows stable optimisation behaviour and high episodic rewards with long episode lengths. Policy output analysis stratified by fault state confirms that the agent has learned state-conditional recovery behaviour, selecting fault-type-appropriate actions in over 85% of fault-state timesteps, thereby providing direct evidence that the agent successfully distinguishes healthy from faulty conditions and among different fault types at the level of individual recovery decisions. Compared with existing RL-based approaches that focus primarily on link failure recovery or service function chain reconfiguration, the proposed framework handles a broader spectrum of SDN fault types and integrates control-plane, data-plane and congestion indicators, thereby offering a more general and robust self-healing capability for operational SDN environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Makiyah et al. (2026) studied this question.

synapsesocial.com/papers/6a0bfde8166b51b53d379383https://doi.org/10.1007/s44354-026-00028-z
Ask AI
Helpful
Bookmark
Share
View Full Paper